Skip to content
Launchpad Library logo

AI safety and evaluation

An AI upskilling plan: from beginner to red-teaming your own models

Six steps built from resources already in this library, in the order that actually builds on itself — the basics first, then the primary papers and the arguments between them, then the governance language, then two tools you run yourself against a real model, then a structured course to tie it together.

This is the path I'd follow if I wanted AI safety and evaluation work on my resume rather than a certificate on a shelf. Each step produces something you can show: a small project, notes on the papers everyone cites, a filled-in risk register, a scan report, a red-team write-up.

Every step is free. Nothing here needs an employer to sponsor you, and nothing here asks for a card. Tick the steps off as you go — your progress is saved in this browser only, and never leaves your device.

Roughly six to nine months alongside a job

Who this is for

  • You're starting from little or no AI background, or you know the tools but not the theory.
  • You want evaluation, red-teaming or AI governance work — not model training research.
  • You'd rather finish five things properly than start twelve.

Your progress

0 of 6 steps done

Loading your saved progress…

1. Get the basics down first

Six to ten weeks, a few hours a week

Everything later assumes you know what a model is, what training does, and enough Python to run somebody else's script. Elements of AI is the gentlest honest start with no math panic; Harvard's CS50 AI course is the step up with real code; Microsoft's beginners' course is the shortest route to building something small.

  1. Work through Elements of AI end to end first. It's free, non-technical, and it gives you the words for everything that follows.
  2. Then take CS50's Introduction to AI with Python, doing the problem sets rather than watching the lectures. This is where Python stops being scary.
  3. Alongside it, build one tiny thing from Microsoft's Generative AI for Beginners — a prompt-based app, a small script — so you've handled a model API before step 5 asks you to attack one.
  4. Get comfortable with basic probability and enough linear algebra to read a matrix multiply without flinching. Every later step assumes both.

Worth knowing: None of these three is a credential anybody hires on. They're the groundwork that makes the later steps possible, and skipping them is why most self-taught attempts stall in month two.

FreeTraining Program
AI & Assistive Tools

Elements of AI

Free, no-math introduction to artificial intelligence from the University of Helsinki.

Why I recommend it: The clearest starting point if AI still feels abstract. You will be able to talk about it accurately in an interview.

#ai#ai-literacy#ai-tools#artificial-intelligence#course#free#free-training#fundamentals#learning#program#students#training#university
elementsofai.comAdded Aug 31, 20260 opens
FreeTraining Program
AI & Assistive Tools

CS50s Introduction to AI with Python (Harvard)

Harvard course covering graph search algorithms, optimization, machine learning, and natural language processing with hands-on Python projects.

Why I recommend it: The most rigorous free option here. Finish the projects — a CS50 project portfolio carries real weight with hiring managers.

#ai#ai-tools#course#cs50#free#free-course#harvard#learning#machine-learning#nlp#program#python#students#training
cs50.harvard.eduAdded Sep 1, 20260 opens
FreeTraining Program
AI & Assistive Tools

Generative AI for Beginners (Microsoft)

21 structured lessons on prompt engineering and building generative AI applications, with practical exercises in Python and TypeScript.

Why I recommend it: The best free course for actually building something. Work one lesson at a time and keep the code you write — that is your proof of skill.

#ai#ai-tools#course#free#free-course#generative-ai#learning#microsoft#program#prompt-engineering#python#training#typescript
github.comAdded Sep 1, 20260 opens

2. Read the primary sources

About 20 hours of reading, over six to eight weeks

Every safety argument you'll meet at work is a secondhand version of one of these papers. Read them in this order and you get the vocabulary first, then the mainstream research consensus, then the strongest worried case, then the people who think that case is wrong, then what a lab says it actually does about it. This step is the papers and the arguments only — the tools come later.

  1. Concrete Problems in AI Safety (two evenings) — the paper that gave the field its words: side effects, reward hacking, scalable oversight, safe exploration. It predates language models, so read it for the categories, not the examples.
  2. Managing Extreme AI Risks Amid Rapid Progress (one evening) — short, signed by senior researchers including Hinton and Bengio. A negotiated consensus statement says what a large group could all sign, not what any one author believes most strongly.
  3. An Overview of Catastrophic AI Risks (three or four evenings) — the risk taxonomy with citations: misuse, race dynamics, organizational failure, rogue systems. The Center for AI Safety exists to make this argument; read it as the best case for worry, argued by people who hold it.
  4. AI 2027 (one evening) — Daniel Kokotajlo's month-by-month scenario forecast, the most specific published attempt to say what the next few years look like. It's a forecast, not a finding, and its authors say so and publish their reasoning.
  5. A Path Towards Autonomous Machine Intelligence (two evenings) — Yann LeCun's dissent: today's language models are the wrong architecture to reach human-level intelligence at all, which changes what's worth worrying about. It's a research agenda by someone building the alternative, so it argues for his own direction.
  6. Constitutional AI (one evening) — Anthropic's published method for training a model against written principles. A company describing its own safety work on a model it sells; the method is real and reproducible, the framing is theirs.
  7. Write a page of your own on where you land and why. Naming which argument convinced you, and what would change your mind, is the thing an interview actually probes.

Worth knowing: These are arguments, not settled findings, and several serve the interests of whoever published them — I've said which, above. Reading the worried case without the dissent, or the other way round, leaves you with half the field.

FreeDocument
AI & Assistive Tools

Concrete Problems in AI Safety

The 2016 paper that framed AI safety as a set of specific engineering problems — side effects, reward hacking, unsafe exploration — rather than a philosophical worry. Free on arXiv.

From the site: Rapid progress in machine learning and artificial intelligence (AI) has brought increasing attention to the potential impacts of AI technologies on society. In this paper we discuss one such potential impact: the problem of accidents in machine learning systems, defined as unintended and harmful behavior that may emer…

Why I recommend it: Start here if the safety conversation sounds abstract. It is plain about what can go wrong and why, and almost everything since cites it.

#ai ethics#ai-risk#ai-safety#alignment#arxiv#foundational#free#machine-learning#reading#research#technology-and-ethics
arXiv.orgAdded Sep 17, 20260 opens
FreeDocument
AI & Assistive Tools

Managing Extreme AI Risks Amid Rapid Progress

A short consensus paper from Geoffrey Hinton, Yoshua Bengio and two dozen other researchers on the risks they consider serious and the governance they think is needed. Free on arXiv.

From the site: Artificial Intelligence (AI) is progressing rapidly, and companies are shifting their focus to developing generalist AI systems that can autonomously act and pursue goals. Increases in capabilities and autonomy may soon massively amplify AI's impact, with risks that include large-scale social harms, malicious uses, an…

Why I recommend it: The clearest statement of what the safety-concerned researchers actually agree on, signed rather than paraphrased.

#ai ethics#ai-policy#ai-risk#ai-safety#alignment#arxiv#free#governance#reading#regulation#research#technology-and-ethics
arXiv.orgAdded Sep 17, 20260 opens
FreeReport
AI & Assistive Tools

An Overview of Catastrophic AI Risks

A structured survey of the risks — malicious use, competitive pressure, organizational failure, and systems pursuing goals of their own — with the evidence for each. Free to read.

From the site: There are many potential risks from AI. CAIS focusses on mitigating risks that could lead to catastrophic outcomes for society, such as bioterrorism or loss of control over military AI systems.

Why I recommend it: The best single map of the different worries, which are usually mashed together into one. Written by a safety organization, so read it as advocacy with citations.

#ai ethics#ai-policy#ai-risk#ai-safety#alignment#free#governance#overview#reading#regulation#research#technology-and-ethics
Center for AI SafetyAdded Sep 17, 20260 opens
FreeReport
AI & Assistive Tools

AI 2027

A detailed scenario for how AI might develop through 2027, written by former OpenAI researcher Daniel Kokotajlo and colleagues, with the reasoning and uncertainties spelled out. Free to read in full.

From the site: A research-backed AI scenario forecast.

Why I recommend it: The forecast everyone in this field argued about. Read it as one carefully argued scenario, not a prediction — the authors say as much themselves.

#ai ethics#ai-policy#ai-risk#ai-safety#alignment#forecasting#free#reading#regulation#research#scenario#technology-and-ethics
ai-2027.comAdded Sep 17, 20260 opens
FreeDocument
AI & Assistive Tools

A Path Towards Autonomous Machine Intelligence

Yann LeCun's position paper arguing that today's language models are the wrong architecture, and sketching what he thinks should replace them. Free to read.

Why I recommend it: The serious technical case against scaling language models further. Dense, but it is the argument itself rather than a summary of it.

#ai ethics#ai-research#ai-safety#alignment#architecture#free#machine-learning#reading#research#technology-and-ethics#world-models
openreview.netAdded Sep 17, 20260 opens
FreeDocument
AI & Assistive Tools

Constitutional AI: Harmlessness from AI Feedback

Anthropic's paper describing how Claude is trained against a written set of principles instead of relying only on human ratings. Free on arXiv.

From the site: As AI systems become more capable, we would like to enlist their help to supervise other AIs. We experiment with methods for training a harmless AI assistant through self-improvement, without any human labels identifying harmful outputs. The only human oversight is provided through a list of rules or principles, and s…

Why I recommend it: Worth reading to see what "aligned" means in practice at one lab — and note it comes from the company selling the model.

#ai ethics#ai-risk#ai-safety#alignment#anthropic#arxiv#free#reading#research#technology-and-ethics#training
arXiv.orgAdded Sep 17, 20260 opens

3. Learn the language governance people use

Two to three weeks of evenings

Every safety job description asks for a risk framework by name. NIST's is the one US employers cite, it's free, and it's short enough to read properly.

  1. Read the core framework end to end once — Govern, Map, Measure, Manage. It's a vocabulary, not a rulebook.
  2. Then do the useful part: pick one AI system you've actually used at work and write its risk register against the Map and Measure functions. Two pages is plenty.
  3. Keep that document. It's the single best thing to bring to an interview for a governance or assurance role, and almost nobody brings one.

Worth knowing: It's voluntary guidance, not law, and it deliberately tells you what to consider rather than what to do. Don't expect thresholds or pass marks.

FreeDocument
AI & Assistive Tools

NIST AI Risk Management Framework

The US government's voluntary framework for identifying and managing AI risk, plus its playbook of concrete practices. Free.

Why I recommend it: The one your employer's legal team is most likely already citing. Useful vocabulary if you want to raise AI risk at work and be taken seriously.

#ai ethics#ai-policy#ai-safety#compliance#free#governance#government#regulation#risk-management#standards#technology-and-ethics#workplace
NISTAdded Sep 17, 20260 opens

4. Break a model on purpose, with a scanner

A weekend to install, a month to read the results well

garak probes a model for the failures people worry about — prompt injection, leaked data, toxic output, jailbreaks — and hands you a report. It's the fastest way to stop treating safety as an abstraction.

  1. Install it: `python -m pip install -U garak` (Python 3.10+; a virtual environment saves you pain later).
  2. Run a first scan against a model you can reach: `garak --model_type openai --model_name gpt-4o-mini --probes encoding` — start with one probe family, not all of them.
  3. Read the failure lines, then go one level down: pick a single failing probe and work out why that phrasing got through. That explanation is the skill, not the run.
  4. Write up one scan as a short report: what you probed, what failed, how often, and what you'd change. That's the deliverable.

Worth knowing: Scans call a model, so they cost money if you point them at a paid API. Cap your spend first, and never scan a system you don't have permission to test.

FreeTool
AI & Assistive Tools

garak

An open-source scanner that probes a language model for weaknesses — prompt injection, data leakage, jailbreaks, toxic output — and reports what it found. Free.

From the site: the LLM vulnerability scanner. Contribute to NVIDIA/garak development by creating an account on GitHub.

Why I recommend it: Point it at a model you are about to rely on and see how it fails before your users do.

#ai ethics#ai-risk#ai-safety#developer-tools#free#open-source#prompt-injection#red-teaming#research#security#technology-and-ethics#testing
GitHubAdded Sep 17, 20260 opens

5. Run structured red-team exercises

Six to eight weeks alongside step 4

PyRIT is the step up from scanning: you script multi-turn attacks, score the responses automatically, and repeat them as the model changes. This is what an evaluation job looks like day to day.

  1. Install it: `pip install pyrit` (Python 3.10–3.12), then work through the first notebook in their docs before writing anything of your own.
  2. Rebuild one of your garak findings as a PyRIT orchestrator — same weakness, now as a repeatable multi-turn attack with a scorer attached.
  3. Add a scoring rule you wrote yourself for a harm that matters in your target industry. Generic scorers are where everyone stops; a domain-specific one is what gets you hired.
  4. Map each finding back to the NIST Measure function from step 3, so your report speaks to engineers and to governance in the same document.

Worth knowing: PyRIT is a framework, not a verdict — it generates and scores attacks, it doesn't tell you whether a system is safe to ship. That judgment stays yours.

FreeTool
AI & Assistive Tools

PyRIT

Microsoft's open-source toolkit for red-teaming AI systems: automated attack prompts, scoring of the responses, and repeatable runs. Free.

From the site: The Python Risk Identification Tool for generative AI (PyRIT) is an open source framework built to empower security professionals and engineers to proactively identify risks in generative AI system...

Why I recommend it: Built by the team that red-teams Microsoft's own AI products, and released as-is. Best paired with a written idea of what you are testing for.

#ai ethics#ai-risk#ai-safety#developer-tools#free#microsoft#open-source#red-teaming#security#technology-and-ethics#testing
GitHubAdded Sep 17, 20260 opens

6. Finish with a structured course and a final project

Around 12 weeks, a few hours a week

BlueDot's AI Safety Fundamentals is the free course safety teams actually recognize: a reading order, weekly facilitated discussion and a final project. Taken last rather than first, the discussions land properly — you arrive with scan findings and a risk register instead of opinions.

  1. Apply for a facilitated cohort rather than reading the curriculum alone. The discussion is the part that makes it stick, and the curriculum is published free either way if you miss the intake.
  2. Use your garak and PyRIT work from steps 4 and 5 as the basis of the final project, so it's evidence rather than an essay.
  3. Pick a problem in the industry you're leaving — you know its data and its edge cases better than any classmate will.

Worth knowing: Cohorts are competitive and run on a schedule, so you may wait for an intake — start the papers in step 2 meanwhile. It teaches you the arguments, not the engineering; the engineering was steps 4 and 5.

FreeTraining Program
AI & Assistive Tools

AI Safety Fundamentals

A free structured course in AI alignment and AI governance — readings, exercises and facilitated cohorts. Self-paced version free to anyone.

From the site: Free online courses, grants, and intensive in-person programs from the leading talent accelerator for beneficial AI and societal resilience. Join 10,000+ alumni and start today.

Why I recommend it: The usual route in for people trying to move into safety work. The reading list alone is worth the visit even if you never join a cohort.

#ai ethics#ai-risk#ai-safety#alignment#career-change#course#free#governance#regulation#research#study#technology-and-ethics#training
BlueDot ImpactAdded Sep 17, 20260 opens

Where to go after this

Once the five steps above are done and written up, these are the natural extensions: the UK safety institute's evaluation framework, a lighter prompt-testing tool for everyday work, and the NIST framework again — reread it after you have real scan findings and it lands differently.

FreeTool
AI & Assistive Tools

Inspect

An open-source framework from the UK's AI Security Institute for evaluating models — writing tests, scoring answers and logging what happened. Free.

From the site: Open-source framework for large language model evaluations

Why I recommend it: What a government safety institute actually uses to test models. Technical, but the docs explain the thinking behind each kind of test.

#ai ethics#ai-risk#ai-safety#alignment#benchmarks#developer-tools#evaluation#free#open-source#research#technology-and-ethics#testing
InspectAdded Sep 17, 20260 opens
FreeTool
AI & Assistive Tools

promptfoo

An open-source tool for testing and red-teaming prompts and AI apps — run the same prompts across models, compare answers, and catch regressions. Free and self-hosted.

From the site: The AI Security Platform that catches vulnerabilities in development. Trusted by 156 of the Fortune 500 and 300,000+ developers worldwide.

Why I recommend it: The practical one: if you have built anything on top of a model, this is how you check a prompt change did not quietly make it worse.

#ai ethics#ai-risk#ai-safety#developer-tools#evaluation#free#open-source#prompts#red-teaming#technology-and-ethics#testing
promptfoo.devAdded Sep 17, 20260 opens
FreeDocument
AI & Assistive Tools

NIST AI Risk Management Framework

The US government's voluntary framework for identifying and managing AI risk, plus its playbook of concrete practices. Free.

Why I recommend it: The one your employer's legal team is most likely already citing. Useful vocabulary if you want to raise AI risk at work and be taken seriously.

#ai ethics#ai-policy#ai-safety#compliance#free#governance#government#regulation#risk-management#standards#technology-and-ethics#workplace
NISTAdded Sep 17, 20260 opens