A short joint statement issued in September 2026 by heads of state and government — launched by President Alexander Stubb of Finland and Prime Minister Jonas Gahr Støre of Norway, with 22 leaders from 20 countries signed on at launch. It asks for three things: mandatory pre-deployment testing and independent evaluation by qualified evaluators with real access; coordinated common standards and shared reporting of serious safety incidents, with scientific capacity available to countries in every region; and for UN member states to explore an international institution that could set standards, verify compliance and convene states when capability thresholds are crossed.
Why I recommend it: Read this as the primary source rather than someone's summary of it — it is one page, in plain language, and you can read the whole thing in three minutes. Two things worth noticing: it is a political appeal, not a law or a treaty, so nothing in it binds any company today; and the United States and China are not among the signatories, which matters given where the frontier labs are. Useful if you are writing or interviewing about AI policy and want to quote what governments actually asked for, dated September 2026.
Non-profit research organisation working on theoretical alignment and on evaluations that test what frontier models are capable of, with public reports.
From the site: ARC is a non-profit research organization whose mission is to align future machine learning systems with human interests.
Why I recommend it: Their evaluations work is why "dangerous capability testing" is now a normal phrase. Read the reports, they are short.
On-device AI with cloud fallback for smartphones, laptops, and edge devices, designed to cut inference costs by knowing when to hand off to frontier cloud models.
Why I recommend it: I am watching on-device AI closely because it could make powerful tools accessible at lower cost and with more privacy. Cactus is a useful example of the "know when to hand off" design pattern.
OpenAI's announcement of GPT-6 Astra, its most capable model, with reported results on computer use, browsing, software engineering, cybersecurity, and professional work.
Why I recommend it: Read the capability list as a job-task list. Whatever a model does well this year reshapes entry-level work the next.