Skip to content
Launchpad Library logo

ai evaluation

3 free resources on this topic. Everything here is free and hand-picked. You can also search within this topic.

FreePublication
AI & Assistive Tools

Arena Blog (LLM and Agent Evaluation Research)

Research and product writing from the team behind the Arena model leaderboards: how coding-agent harnesses change cost and success rates, how the agent leaderboards are built, and their academic partnership calls.

Why I recommend it: Free to read. The clearest writing anywhere on why two people using the same model get very different results — the tool wrapped around the model changes the cost and the outcome. They run the leaderboards they write about, so treat their rankings as one measurement, not the verdict.

#ai-evaluation#benchmarks#llm#ai-agents#research#free#ai-and-assistive-tools
arena.aiAdded Sep 22, 20260 opens
FreeReport
Technology & Ethics

A Call for Control of Frontier AI Models

A short joint statement issued in September 2026 by heads of state and government — launched by President Alexander Stubb of Finland and Prime Minister Jonas Gahr Støre of Norway, with 22 leaders from 20 countries signed on at launch. It asks for three things: mandatory pre-deployment testing and independent evaluation by qualified evaluators with real access; coordinated common standards and shared reporting of serious safety incidents, with scientific capacity available to countries in every region; and for UN member states to explore an international institution that could set standards, verify compliance and convene states when capability thresholds are crossed.

Why I recommend it: Read this as the primary source rather than someone's summary of it — it is one page, in plain language, and you can read the whole thing in three minutes. Two things worth noticing: it is a political appeal, not a law or a treaty, so nothing in it binds any company today; and the United States and China are not among the signatories, which matters given where the frontier labs are. Useful if you are writing or interviewing about AI policy and want to quote what governments actually asked for, dated September 2026.

#ai ethics#ai-evaluation#ai-governance#ai-policy#ai-regulation#ai-safety#frontier-models#international-law#oversight#regulation#research#united-nations
presidentti.fiAdded Sep 21, 20260 opens
FreeOrganization
Technology & Ethics

AI Evaluator Forum

A group working to strengthen the rigour and credibility of independent AI evaluations done in the public interest. Free to read.

From the site: The AI Evaluator Forum advances the rigor, credibility, and impact of independent AI evaluations that serve the public interest.

Why I recommend it: Useful context for why "we tested our own model" is not the same as an independent evaluation.

#ai ethics#ai-evaluation#ai-governance#ai-safety#auditing#regulation
AI Evaluator ForumAdded Sep 21, 20260 opens