Elements of AI
Free, no-math introduction to artificial intelligence from the University of Helsinki.
Why I recommend it: The clearest starting point if AI still feels abstract. You will be able to talk about it accurately in an interview.
AI safety and evaluation
Six steps built from resources already in this library, in the order that actually builds on itself — the basics first, then the primary papers and the arguments between them, then the governance language, then two tools you run yourself against a real model, then a structured course to tie it together.
This is the path I'd follow if I wanted AI safety and evaluation work on my resume rather than a certificate on a shelf. Each step produces something you can show: a small project, notes on the papers everyone cites, a filled-in risk register, a scan report, a red-team write-up.
Every step is free. Nothing here needs an employer to sponsor you, and nothing here asks for a card. Tick the steps off as you go — your progress is saved in this browser only, and never leaves your device.
Roughly six to nine months alongside a job
0 of 6 steps done
Loading your saved progress…
Six to ten weeks, a few hours a week
Everything later assumes you know what a model is, what training does, and enough Python to run somebody else's script. Elements of AI is the gentlest honest start with no math panic; Harvard's CS50 AI course is the step up with real code; Microsoft's beginners' course is the shortest route to building something small.
Worth knowing: None of these three is a credential anybody hires on. They're the groundwork that makes the later steps possible, and skipping them is why most self-taught attempts stall in month two.
Free, no-math introduction to artificial intelligence from the University of Helsinki.
Why I recommend it: The clearest starting point if AI still feels abstract. You will be able to talk about it accurately in an interview.
Harvard course covering graph search algorithms, optimization, machine learning, and natural language processing with hands-on Python projects.
Why I recommend it: The most rigorous free option here. Finish the projects — a CS50 project portfolio carries real weight with hiring managers.
21 structured lessons on prompt engineering and building generative AI applications, with practical exercises in Python and TypeScript.
Why I recommend it: The best free course for actually building something. Work one lesson at a time and keep the code you write — that is your proof of skill.
About 20 hours of reading, over six to eight weeks
Every safety argument you'll meet at work is a secondhand version of one of these papers. Read them in this order and you get the vocabulary first, then the mainstream research consensus, then the strongest worried case, then the people who think that case is wrong, then what a lab says it actually does about it. This step is the papers and the arguments only — the tools come later.
Worth knowing: These are arguments, not settled findings, and several serve the interests of whoever published them — I've said which, above. Reading the worried case without the dissent, or the other way round, leaves you with half the field.
The 2016 paper that framed AI safety as a set of specific engineering problems — side effects, reward hacking, unsafe exploration — rather than a philosophical worry. Free on arXiv.
From the site: Rapid progress in machine learning and artificial intelligence (AI) has brought increasing attention to the potential impacts of AI technologies on society. In this paper we discuss one such potential impact: the problem of accidents in machine learning systems, defined as unintended and harmful behavior that may emer…
Why I recommend it: Start here if the safety conversation sounds abstract. It is plain about what can go wrong and why, and almost everything since cites it.
A short consensus paper from Geoffrey Hinton, Yoshua Bengio and two dozen other researchers on the risks they consider serious and the governance they think is needed. Free on arXiv.
From the site: Artificial Intelligence (AI) is progressing rapidly, and companies are shifting their focus to developing generalist AI systems that can autonomously act and pursue goals. Increases in capabilities and autonomy may soon massively amplify AI's impact, with risks that include large-scale social harms, malicious uses, an…
Why I recommend it: The clearest statement of what the safety-concerned researchers actually agree on, signed rather than paraphrased.
A structured survey of the risks — malicious use, competitive pressure, organizational failure, and systems pursuing goals of their own — with the evidence for each. Free to read.
From the site: There are many potential risks from AI. CAIS focusses on mitigating risks that could lead to catastrophic outcomes for society, such as bioterrorism or loss of control over military AI systems.
Why I recommend it: The best single map of the different worries, which are usually mashed together into one. Written by a safety organization, so read it as advocacy with citations.
A detailed scenario for how AI might develop through 2027, written by former OpenAI researcher Daniel Kokotajlo and colleagues, with the reasoning and uncertainties spelled out. Free to read in full.
From the site: A research-backed AI scenario forecast.
Why I recommend it: The forecast everyone in this field argued about. Read it as one carefully argued scenario, not a prediction — the authors say as much themselves.
Yann LeCun's position paper arguing that today's language models are the wrong architecture, and sketching what he thinks should replace them. Free to read.
Why I recommend it: The serious technical case against scaling language models further. Dense, but it is the argument itself rather than a summary of it.
Anthropic's paper describing how Claude is trained against a written set of principles instead of relying only on human ratings. Free on arXiv.
From the site: As AI systems become more capable, we would like to enlist their help to supervise other AIs. We experiment with methods for training a harmless AI assistant through self-improvement, without any human labels identifying harmful outputs. The only human oversight is provided through a list of rules or principles, and s…
Why I recommend it: Worth reading to see what "aligned" means in practice at one lab — and note it comes from the company selling the model.
Two to three weeks of evenings
Every safety job description asks for a risk framework by name. NIST's is the one US employers cite, it's free, and it's short enough to read properly.
Worth knowing: It's voluntary guidance, not law, and it deliberately tells you what to consider rather than what to do. Don't expect thresholds or pass marks.
The US government's voluntary framework for identifying and managing AI risk, plus its playbook of concrete practices. Free.
Why I recommend it: The one your employer's legal team is most likely already citing. Useful vocabulary if you want to raise AI risk at work and be taken seriously.
A weekend to install, a month to read the results well
garak probes a model for the failures people worry about — prompt injection, leaked data, toxic output, jailbreaks — and hands you a report. It's the fastest way to stop treating safety as an abstraction.
Worth knowing: Scans call a model, so they cost money if you point them at a paid API. Cap your spend first, and never scan a system you don't have permission to test.
An open-source scanner that probes a language model for weaknesses — prompt injection, data leakage, jailbreaks, toxic output — and reports what it found. Free.
From the site: the LLM vulnerability scanner. Contribute to NVIDIA/garak development by creating an account on GitHub.
Why I recommend it: Point it at a model you are about to rely on and see how it fails before your users do.
Six to eight weeks alongside step 4
PyRIT is the step up from scanning: you script multi-turn attacks, score the responses automatically, and repeat them as the model changes. This is what an evaluation job looks like day to day.
Worth knowing: PyRIT is a framework, not a verdict — it generates and scores attacks, it doesn't tell you whether a system is safe to ship. That judgment stays yours.
Microsoft's open-source toolkit for red-teaming AI systems: automated attack prompts, scoring of the responses, and repeatable runs. Free.
From the site: The Python Risk Identification Tool for generative AI (PyRIT) is an open source framework built to empower security professionals and engineers to proactively identify risks in generative AI system...
Why I recommend it: Built by the team that red-teams Microsoft's own AI products, and released as-is. Best paired with a written idea of what you are testing for.
Around 12 weeks, a few hours a week
BlueDot's AI Safety Fundamentals is the free course safety teams actually recognize: a reading order, weekly facilitated discussion and a final project. Taken last rather than first, the discussions land properly — you arrive with scan findings and a risk register instead of opinions.
Worth knowing: Cohorts are competitive and run on a schedule, so you may wait for an intake — start the papers in step 2 meanwhile. It teaches you the arguments, not the engineering; the engineering was steps 4 and 5.
A free structured course in AI alignment and AI governance — readings, exercises and facilitated cohorts. Self-paced version free to anyone.
From the site: Free online courses, grants, and intensive in-person programs from the leading talent accelerator for beneficial AI and societal resilience. Join 10,000+ alumni and start today.
Why I recommend it: The usual route in for people trying to move into safety work. The reading list alone is worth the visit even if you never join a cohort.
Once the five steps above are done and written up, these are the natural extensions: the UK safety institute's evaluation framework, a lighter prompt-testing tool for everyday work, and the NIST framework again — reread it after you have real scan findings and it lands differently.
An open-source framework from the UK's AI Security Institute for evaluating models — writing tests, scoring answers and logging what happened. Free.
From the site: Open-source framework for large language model evaluations
Why I recommend it: What a government safety institute actually uses to test models. Technical, but the docs explain the thinking behind each kind of test.
An open-source tool for testing and red-teaming prompts and AI apps — run the same prompts across models, compare answers, and catch regressions. Free and self-hosted.
From the site: The AI Security Platform that catches vulnerabilities in development. Trusted by 156 of the Fortune 500 and 300,000+ developers worldwide.
Why I recommend it: The practical one: if you have built anything on top of a model, this is how you check a prompt change did not quietly make it worse.
The US government's voluntary framework for identifying and managing AI risk, plus its playbook of concrete practices. Free.
Why I recommend it: The one your employer's legal team is most likely already citing. Useful vocabulary if you want to raise AI risk at work and be taken seriously.