Research and product writing from the team behind the Arena model leaderboards: how coding-agent harnesses change cost and success rates, how the agent leaderboards are built, and their academic partnership calls.
Why I recommend it: Free to read. The clearest writing anywhere on why two people using the same model get very different results — the tool wrapped around the model changes the cost and the outcome. They run the leaderboards they write about, so treat their rankings as one measurement, not the verdict.
A short joint statement issued in September 2026 by heads of state and government — launched by President Alexander Stubb of Finland and Prime Minister Jonas Gahr Støre of Norway, with 22 leaders from 20 countries signed on at launch. It asks for three things: mandatory pre-deployment testing and independent evaluation by qualified evaluators with real access; coordinated common standards and shared reporting of serious safety incidents, with scientific capacity available to countries in every region; and for UN member states to explore an international institution that could set standards, verify compliance and convene states when capability thresholds are crossed.
Why I recommend it: Read this as the primary source rather than someone's summary of it — it is one page, in plain language, and you can read the whole thing in three minutes. Two things worth noticing: it is a political appeal, not a law or a treaty, so nothing in it binds any company today; and the United States and China are not among the signatories, which matters given where the frontier labs are. Useful if you are writing or interviewing about AI policy and want to quote what governments actually asked for, dated September 2026.