arXiv paper (Sept. 2026) introducing SkillGym, a framework that turns human-written agent skills into executable, verifiable training environments so language model agents learn them as built-in capabilities.
A recommendation page from Artificial Analysis that suggests AI models for a given use case, drawing on the site's benchmark data across intelligence, speed, and cost. Free to use.
Why I recommend it: A quick way to turn "we need a model for X" into a shortlist with cost and speed tradeoffs attached. Useful in the first meeting; verify with your own tests before committing.
A leaderboard comparing more than 250 AI language models across intelligence, price, output speed, latency, and context window, with per-model provider analysis. Free to read.
Why I recommend it: The single table to open when someone claims one model is "the best." Sort by cost per task or speed and the answer often changes — a useful reality check in vendor conversations.
A free way to use several well-known chat models — including ones from OpenAI and Anthropic alongside open models like Llama and Mistral — without an account and without the model provider seeing who you are. DuckDuckGo strips your identity and passes the request on, and says the providers agree not to train on what goes through it. Supports image and PDF uploads, image generation and voice chat, with daily limits on the free tier.
Why I recommend it: The most frictionless privacy win on this list: no sign-up, nothing to cancel, and it takes about four seconds to start. Good for the everyday questions you do not want attached to a profile. Be clear about what it does and does not do — DuckDuckGo hides who you are from the model provider, but the words you type still travel to that provider's servers, so it is not the same as encryption or running a model on your own machine. There is a paid upgrade for higher limits.
Research and product writing from the team behind the Arena model leaderboards: how coding-agent harnesses change cost and success rates, how the agent leaderboards are built, and their academic partnership calls.
Why I recommend it: Free to read. The clearest writing anywhere on why two people using the same model get very different results — the tool wrapped around the model changes the cost and the outcome. They run the leaderboards they write about, so treat their rankings as one measurement, not the verdict.
Xiaomi's open-weight AI model family — text, image, video, and audio understanding — released under the MIT license, free to download and self-host.
Why I recommend it: Weights are free (MIT license, commercial use allowed) on Hugging Face. The hosted API is paid per token. If you can run models locally, this is a free frontier-tier option; if you want a chat interface, use the API and expect a bill.
A practical engineering comparison of decision models (Jev AI) and generative large language models: when to use each, how they can work together, and a five-dimension framework for choosing.
Why I recommend it: Community article on Hugging Face by sora-2, published 21 September 2026. It explains that Jev AI turns state into a choice, score or yes/no judgment inside a defined answer space, while generative LLMs handle open-ended writing, explanation and reasoning. Rule of thumb: if the answer space can be defined in advance and the result will be reused, ranked, routed or blocked by code, evaluate Jev AI first.
#Jev AI#decision model#LLM#large language model#AI architecture#agent#Hugging Face#chat model
An open-source drop-in replacement for the OpenAI API that runs models on your own hardware, including CPU-only machines. MIT licensed.
From the site: LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required. - mudler/LocalAI
Why I recommend it: Point existing code at your own server instead of a paid API — no code changes beyond the address.
A self-hosted, open-source chat interface for local or hosted models — chat history, documents, multiple users. Works on top of Ollama or any OpenAI-compatible API.
From the site: User-friendly AI Interface (Supports Ollama, OpenAI API, ...) - open-webui/open-webui
Why I recommend it: If you like the ChatGPT window but not the subscription, this is that window running on your own machine.
Open-source software that downloads and runs open AI models on your own computer, with a single command and an OpenAI-compatible local API. MIT licensed.
From the site: Ollama is the easiest way to automate your work using open models, while keeping your data safe.
Why I recommend it: The easiest honest way to use AI privately — nothing you type leaves your machine. Start with a small model before you judge the speed.
A hands-on write-up of wiring Google's open Gemma 4 model into the Codex command-line coding agent so it runs locally instead of calling a hosted API.
From the site: I wanted to know whether Gemma 4 could replace a cloud model for my day-to-day agentic coding. Not in theory, in practice. I use Codex CLI…
Why I recommend it: Useful if you want to try coding agents without paying per token — local models are slower, but free and private.
A free newsletter explaining ideas in AI in plain English — mostly what and why, a little how.
From the site: Ideas in AI, preferably in English. Mostly what and why, a little how. Click to read Very Sane AI Newsletter, by SE Gyges, a Substack publication with thousands of subscribers.
Why I recommend it: A good weekly read if you want to understand AI without the hype or the math.
A plain-language explainer on RLCD, a way of aligning language models by learning from contrasting outputs rather than human ratings alone.
From the site: RLCD is a method developed to adjust language models to human preferences without using human feedback data. This approach aims to address…
Why I recommend it: Good background reading if you want to understand how the models you use are actually steered.
A free ebook walking through reinforcement learning from the basics to RLHF, written for practitioners rather than researchers.
From the site: Reinforcement learning (RL) is transforming how reliable AI agents are trained and deployed. Discover real-world use cases, efficiency techniques like LoRA, and practical patterns you can apply today.
Why I recommend it: Free download in exchange for an email address. Solid grounding if you keep seeing "RLHF" and nodding along.
A free, open-source AI coding agent that runs in your terminal, works with multiple model providers and can be installed with a single command.
From the site: OpenCode - The open source coding agent.
Why I recommend it: Free and open source, so you can point it at whichever model you already have access to instead of paying for another subscription.
A free, open-source tool for running and monitoring multiple AI coding agents from one place, with a plugin ecosystem and documentation for the agent CLIs it supports.
From the site: Run them anywhere. Leave them running. Herdr holds real terminals open so your agents keep working when you close the laptop, and gets you back in from any tty.
Why I recommend it: Open source with an active plugin community — useful if you are experimenting with more than one AI coding tool.
Independent, free benchmarks testing leading AI models on real-world finance, software, science and safety tasks, with cost and latency alongside accuracy.
From the site: Private, domain-specific benchmarks in legal, tax, and finance.
Why I recommend it: When someone claims a model is "the best," check here — these are independent evaluations, not vendor marketing.
Free OpenRouter developer guide walking through building a terminal-based agent harness — tool calls, loops and model routing — with working code you can adapt.
Why I recommend it: Building a small agent harness yourself is one of the clearest portfolio projects for AI-adjacent roles right now.
TypeSafe AI announcement from founder Diogo Almeida (formerly OpenAI) introducing System One models and Jev, aimed at cheaper automation rather than better chat.
Why I recommend it: Worth skimming to track where new AI labs are placing bets — useful context for interviews at AI companies.
A 2026 research paper from Google's Paradigms of Intelligence team and the University of Chicago showing that safety fine-tuning meant to stop models claiming consciousness also suppresses how they represent minds in animals and people, shifting their answers on values, religiosity and well-being.
Why I recommend it: Useful if you want to speak credibly about AI alignment trade-offs in an interview or a policy conversation.
Official OpenAI guide for designing prompts for Realtime voice models, including gpt-realtime-2 and gpt-realtime-1.5. Covers role definition, guardrails, tool delegation and iterative testing.
Why I recommend it: Start here if you are building voice agents or want cleaner spoken-AI interactions. The guide recommends starting minimal and adding rules only for behaviors that fail in testing.
Analysis based on interviews with 56 experts across 24 countries on how AI language models are used differently in the Global South and the human rights risks that follow.
Why I recommend it: A rare look at AI harms and benefits outside the US and Europe.
An OpenAI-compatible API for unrestricted language models aimed at red teaming, security research, evaluations, and synthetic data, paired with a policy gateway for per-project keys, audit logs, and no data retention.
Why I recommend it: I keep this in the ethics shelf on purpose. Seeing how guardrails get removed for testing is the clearest way to understand why they matter in the tools you actually use at work.