UK AI Security Institute
The UK government's AI Security Institute: research and evaluations of advanced AI systems' capabilities and risks, with public reports and tools.
18 free resources on this topic. Everything here is free and hand-picked. You can also search within this topic.
The UK government's AI Security Institute: research and evaluations of advanced AI systems' capabilities and risks, with public reports and tools.
A platform from Artificial Analysis for building custom benchmarks from your own files, agent traces, or coding environment, then running them across leading models to compare quality, cost per task, and time per task. Benchmarks can be graded against objective rubrics or pairwise judging. Optima is a commercial product; the public announcement and product overview are free to read.
Why I recommend it: Standard benchmarks tell you which model is best in general; they cannot tell you which is best for your workload. If you are choosing a model for a real product, a custom benchmark on your own tasks is the right move — this is one way to do it without building the harness yourself.
Bespoke Labs' open-source toolkit for local typed decisions, contrastive data curation and model evaluation.
Why I recommend it: Free and open source, from a company that also sells services — the repo is usable on its own.
Technical newsletter by Vinoth Govindarajan taking AI agents apart: control loops, memory, orchestration, evaluation and what breaks in production. Recent series walks through the OpenCode architecture end to end.
Why I recommend it: Free to subscribe, over 2,000 readers; written for people who build software, not for beginners. Substack publications can add paid-only posts at any time, so check the top of a post before counting on it.
Free service that gathers public expert evaluations and reviews of preprints in one place, so you can see what other researchers said about a study before it was formally published.
Why I recommend it: Use it as a sanity check on a preprint someone is quoting at you. No evaluation on Sciety means nobody independent has publicly assessed it yet — which is not the same as the study being wrong.
The free research library and blog of Snorkel AI, the company spun out of Stanford's Snorkel project on programmatic labelling. The papers and posts explain how training data for AI models is actually built — labelling, evaluation sets, and the 'environments' used to train agents. Useful if you want to understand the unglamorous data work behind model quality, which is where a lot of the real jobs are.
Why I recommend it: Research and blog posts are free to read with no signup. Read it for the how, not the verdict: Snorkel sells data services to frontier AI labs, so posts arguing that better data beats bigger models are also a sales case. Everything else on the site is a paid enterprise product — 'request dataset samples' means a sales call.
An AI safety lab focused on "scheming" — models that pursue their own goals while appearing aligned. Publishes research on detecting deception, evaluations and governance advice.
From the site: Apollo Research is focused on reducing risks from scheming frontier AI. Our goal is to secure frontier AI systems across development, deployment, and governance.
Why I recommend it: Their research and blog are free to read. Apollo also sells a monitoring product, so read claims about their own tool as company claims.
Personal site of Adam Gleave, CEO and co-founder of the AI safety research lab FAR.AI, with his papers and writing on making models robust and evaluable.
From the site: Adam Gleave is the CEO of FAR.AI, an alignment research non-profit. His research interests include adversarial robustness and value learning.
Why I recommend it: Useful if you want the research side of AI safety rather than the commentary side. Papers first, opinions second.
Non-profit AI safety research lab publishing technical work on model robustness and evaluation, plus events and a fellowship pipeline for researchers entering the field.
From the site: FAR.AI is an AI safety nonprofit advancing technical research across robustness, deception, and red-teaming to ensure AI systems remain safe and beneficial.
Why I recommend it: Look at their fellowships and events pages, not just the papers — that is where the actual entry points are.
An open-source framework from the UK's AI Security Institute for evaluating models — writing tests, scoring answers and logging what happened. Free.
From the site: Open-source framework for large language model evaluations
Why I recommend it: What a government safety institute actually uses to test models. Technical, but the docs explain the thinking behind each kind of test.
An open-source tool for testing and red-teaming prompts and AI apps — run the same prompts across models, compare answers, and catch regressions. Free and self-hosted.
From the site: The AI Security Platform that catches vulnerabilities in development. Trusted by 156 of the Fortune 500 and 300,000+ developers worldwide.
Why I recommend it: The practical one: if you have built anything on top of a model, this is how you check a prompt change did not quietly make it worse.
Research-grade tracking of what AI models can do and how that has changed over time, with the data and methods published. Free.
From the site: Our hub for benchmark results, featuring the performance of leading AI models on challenging tasks. It includes results from benchmarks administered internally by Epoch AI as well as data collected from external sources. Explore trends in AI capabilities across time, by benchmark, or by model.
Why I recommend it: For the longer view rather than this week's launch — they show their working, which most leaderboards do not.
Independent benchmarking of the major AI models on speed, price and quality, with the numbers side by side. Free to read.
From the site: Comparison and analysis of AI models and API hosting providers. Independent benchmarks across key performance metrics including quality, price, output speed & latency.
Why I recommend it: Where I check what a model actually costs per million words before believing a "cheap" claim.
Head-to-head model comparisons voted on by the public: you see two anonymous answers to the same prompt and pick the better one, and the rankings come from those votes. Free.
From the site: Chat, compare, vote for the world's best AI models. Join the community shaping the public leaderboard for LLMs, image, and code models through real-world evaluation.
Why I recommend it: The closest thing to a fair fight between models on ordinary prompts, instead of marketing claims. Votes are taste as much as accuracy, so read it as popularity with a purpose.
A free, regularly updated leaderboard benchmarking how well leading AI models actually search the web, with the methodology and benchmarks published alongside.
Why I recommend it: Check this before assuming your favorite chatbot is the best one for research. The rankings move month to month.
A fellowship with Vals AI, the team that benchmarks how AI models actually perform, for people who want to work on model evaluation.
Why I recommend it: Evaluation work is where a lot of AI hiring is heading — careful, skeptical people do well here even without a PhD.
An OpenAI-compatible API for unrestricted language models aimed at red teaming, security research, evaluations, and synthetic data, paired with a policy gateway for per-project keys, audit logs, and no data retention.
Why I recommend it: I keep this in the ethics shelf on purpose. Seeing how guardrails get removed for testing is the clearest way to understand why they matter in the tools you actually use at work.
Public leaderboard and open-source benchmark that drops AI agents into realistic business environments with 47 real tools across sales, marketing, operations, support, finance, and HR. Scores are based on final environment state, not an LLM-as-judge.
Why I recommend it: The leaderboard and the benchmark code are free; running it yourself means paying the model APIs at the costs shown. The test design is based on Zapier's own task data, so it's a realistic lens on agent work, but Zapier also sells automation tools — treat the benchmark as a useful public dataset, not a neutral referee.