Skip to content
Launchpad Library logo

Learn AI

AI ethics, in plain language

Ethics in AI is usually presented as either a corporate values page or an academic seminar. This is neither. Below are the principles the field actually argues about, the risks that are documented rather than imagined, and the questions to ask before you let a tool make or shape a decision.

Nothing here is my opinion dressed up as consensus. Every principle and risk points to a researcher profiled on this site, a paper you can read free in full, or a term defined and sourced in the Terminology Decoder. Where the evidence is contested, it says so.

The people behind these arguments

Who’s Who in AI profiles the ethicists and accountability researchers named throughout this page — what each one argues, what they’ve built and what they’ve published.

Core principles

These nine show up, in some wording, in almost every serious AI ethics framework. The wording matters less than the second line in each box: what honoring it looks like in a real decision.

Fairness

A system should not work noticeably worse for some groups of people than others, and should not sort people using traits they can't change.

In practice

Before trusting a tool, ask who it was tested on. If accuracy was only reported as one overall number, the gaps are hidden rather than absent.

Profiled here: Joy Buolamwini, Timnit Gebru, Cathy O’Neil

Read the source: Gender Shades, Fairness Through Awareness

Terms explained: AI Ethics

Accountability

When an automated decision harms someone, a named human or organization has to answer for it — and there has to be a way to challenge it.

In practice

For any system that decides something about a person, ask: who signed off on this, and how does someone appeal a wrong answer?

Profiled here: Deborah Raji, Rumman Chowdhury, Cathy O’Neil

Privacy and consent

Information shared in one setting should not quietly be reused in another. Consent to post something publicly is not consent to train on it.

In practice

Check what a tool does with what you paste into it, and never put someone else's private information into a system on their behalf.

Profiled here: Helen Nissenbaum

Terms explained: Training / Training Data

Human oversight

A person should stay responsible for consequential decisions, with real authority to overrule the system rather than rubber-stamping it.

In practice

Oversight only counts if the reviewer has time, information and permission to say no. A reviewer approving 300 cases an hour is not oversight.

Terms explained: Human in the Loop (HITL)

Honest claims

A product should be described by what it actually does, not by what the category is imagined to do.

In practice

Treat "AI-powered" as a marketing phrase until someone shows you an evaluation. Predictions about individual human futures deserve the most skepticism.

Profiled here: Arvind Narayanan

Terms explained: AI Washing

Environmental and labor cost

Training and running large models consumes energy and water, and relies on low-paid data and moderation work that rarely appears in the marketing.

In practice

Prefer the smallest model that does the job, and don't assume a bigger one is automatically better value.

Read the source: On the Dangers of Stochastic Parrots

Documented risks

These are harms that have already happened and been written up — not speculation about future superintelligence. Each one names an example and what reduces it. Nothing here is described as solved.

Unequal performance

A system works well on the people it was mostly trained and tested on, and worse on everyone else — while still reporting good average accuracy.

Example: Commercial face-classification products were up to 34 percentage points less accurate on darker-skinned women than on lighter-skinned men (Gender Shades, 2018).

What reduces it: Insist on results broken down by group, not a single accuracy figure, and test on your own population.

Profiled here: Joy Buolamwini, Timnit Gebru

Read the source: Gender Shades

Confident wrong answers

Language models generate fluent text whether or not it is true, and fluency reads as confidence to most people.

Example: Fabricated citations, invented case law and made-up product features are routine failure modes, not rare glitches.

What reduces it: Use models for drafting and summarizing things you can check, and verify every fact, name, number and link before it leaves your hands.

Read the source: Ethical and Social Risks of Harm from Language Models

Automated decisions about people

Scoring and screening systems get applied to benefits, housing, hiring, credit and child-welfare referrals, often with no meaningful route of appeal.

Example: Automating Inequality documents eligibility and risk-scoring systems tested first on poor and working-class people.

What reduces it: Require an appeals path, a human decision-maker, and a record of why each decision was made.

Profiled here: Virginia Eubanks, Cathy O’Neil

What's actually in the training data

Very large scraped datasets contain copyrighted, private, racist and sexual material that nobody reviewed, and scale makes auditing harder.

Example: Audits of open multimodal datasets have repeatedly found non-consensual imagery and malignant stereotypes inside datasets already used to train shipped models.

What reduces it: Prefer models whose data sources are documented, and treat anything you paste in as potentially retained.

Profiled here: Abeba Birhane

Data reuse you didn't agree to

Information you supply for one purpose gets used for another: training, advertising profiles, employer review, or resale.

Example: Free AI tools commonly reserve the right to train on your inputs unless you find and change a setting.

What reduces it: Read the data-use setting before the first use, and keep client, patient, student and personal data out of consumer tools.

Profiled here: Helen Nissenbaum

Overselling and outright snake oil

Some products predict human outcomes — job performance, criminality, emotion — at accuracy barely better than chance while being sold as objective.

Example: AI Snake Oil separates generative tools that often work from predictive systems used on people, which mostly do not.

What reduces it: Ask for an independent evaluation. If the only evidence is a case study written by the vendor, treat the claim as unproven.

Profiled here: Arvind Narayanan

Losing the skill you delegated

Handing judgment to a tool erodes the practiced judgment you'd need to catch it when it's wrong.

Example: The AI Mirror argues systems trained on our past reflect it back, narrowing rather than widening what we imagine is possible.

What reduces it: Keep doing the hard version sometimes. Use AI to widen options, then decide yourself.

Profiled here: Shannon Vallor

Audits that change nothing

An audit or ethics statement can be published, praised and then quietly ignored, giving the appearance of accountability without the substance.

Example: Actionable Auditing tracked which face-recognition vendors actually improved after public audits — and which simply did not.

What reduces it: Look for what changed after the audit: a version, a date, a withdrawn feature. Otherwise treat it as marketing.

Profiled here: Deborah Raji

How to apply it

You don’t need an ethics board to do this well. You need a short list of questions you ask every time, and the willingness to walk away from a tool that can’t answer them.

Step 1

Before you adopt a tool

Most ethics work happens at the moment of choosing, not after something goes wrong. Ten minutes here saves a lot later.

  • What exactly is this tool for, and what is it not for?
  • Who was it tested on, and where does the vendor admit it performs worse?
  • What happens to what I put into it — is it used for training, and can I turn that off?
  • What is the plan for when it's wrong? Who notices, and who fixes it?

Step 2

If it touches a decision about a person

Hiring, grading, lending, care, discipline, eligibility. This is where harm is concentrated, and where a human has to stay answerable.

  • Is a named person making the final call, with the authority and the time to overrule it?
  • Can the affected person find out that a system was used, and challenge the result?
  • Would I be comfortable explaining this decision to the person it lands on?
  • Have I checked the outcomes by group, rather than trusting one overall accuracy number?

Step 3

When you use the output

Treat model output as a confident first draft from someone who has never been held responsible for being wrong.

  • Have I verified every fact, figure, quote, citation and link independently?
  • Would I be comfortable if this were published with my name on it and the AI use disclosed?
  • Am I disclosing AI use where it's relevant — to a client, an employer, a reader, or an application?
  • Is any of this someone else's confidential information?

Step 4

Keeping yourself honest over time

The field moves fast and the marketing moves faster. Reading the primary sources yourself is the cheapest protection available.

  • Am I reading critics as well as builders?
  • Have I read at least one of the papers below in full, rather than a summary of it?
  • When I repeat a claim about AI, can I name where it came from?
  • Am I distinguishing what a system does from what its maker says it does?

Free tools that do the checking

The questions above are easier to answer with something in your hands. 26 free tools — bias and fairness detectors, dataset audits, model documentation, red-teaming scanners, explainability libraries and checklists that need no code at all — are listed on their own page, each with what it will not tell you.

Open the free AI ethics tools

Read the papers yourself

All 30 are free to read in full — no paywall, no email gate. Each card here links to the publisher and to its entry in the library.

  • Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification

    Joy Buolamwini and Timnit Gebru, 2018

    Measured commercial face-classification products failing on darker-skinned women up to 34.7% of the time against 0.8% for lighter-skinned men.

  • On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?

    Emily M. Bender, Timnit Gebru, Angelina McMillan-Major and Shmargaret Shmitchell, 2021

    Argues the costs of ever-larger language models — energy, unauditable training data, encoded bias, the illusion of understanding — scale with the models themselves.

  • Model Cards for Model Reporting

    Margaret Mitchell, Timnit Gebru and colleagues, 2019

    Proposed the short standard document that ships with a trained model: what it is for, who it was tested on, and where it performs worse.

  • Datasheets for Datasets

    Timnit Gebru, Kate Crawford and colleagues, 2018

    Applies the electronics datasheet idea to training data: how it was collected, who is in it, who consented, and what it should not be used for.

  • Ethical and Social Risks of Harm from Language Models

    Laura Weidinger and colleagues at DeepMind, 2021

    Maps 21 distinct risks from language models across six areas, from discrimination and misinformation to environmental and economic cost.

  • The Values Encoded in Machine Learning Research

    Abeba Birhane, Pratyusha Kalluri and colleagues, 2021

    Annotated 100 highly cited papers and found the field rewards performance, novelty and generalization far more than fairness or societal need.

  • Language (Technology) is Power: A Critical Survey of “Bias” in NLP

    Su Lin Blodgett, Solon Barocas, Hal Daumé III and Hanna Wallach, 2020

    Reviewed 146 papers on bias in language technology and found most never say who is harmed or how.

  • Artificial Intelligence, Values, and Alignment

    Iason Gabriel, 2020

    Separates the technical problem of making a system pursue a goal from the normative one of whose values it should pursue, and who chooses.

  • The Ethics of Advanced AI Assistants

    Iason Gabriel and colleagues at Google DeepMind, 2024

    Examines what changes when AI acts on your behalf: manipulation, anthropomorphism, misplaced delegation, and the effects of everyone having an assistant at once.

  • Model Evaluation for Extreme Risks

    Toby Shevlane and colleagues across DeepMind, OpenAI, Anthropic and academia, 2023

    Sets out pre-release testing for dangerous capabilities, and is where today's frontier safety framework language comes from — written largely by the labs it would govern.

  • Fairness Through Awareness

    Cynthia Dwork, Moritz Hardt, Toniann Pitassi, Omer Reingold and Richard Zemel, 2011

    The founding technical paper on algorithmic fairness: treat similar individuals similarly, and note why simply not collecting a protected attribute fails to achieve it.

  • The Mythos of Model Interpretability

    Zachary C. Lipton, 2016

    Shows “interpretable” is used to mean several incompatible things, and that simpler models are not automatically more honest about what they do.

  • Oxford Martin AI Governance Initiative

    Oxford Martin School, ongoing

    A research program on how AI is and should be governed, spanning policy, regulation, and the politics of frontier model development.

  • From surveillance to governance: recentering human rights in the UN AI agenda

    Access Now, 2026

    Argues the UN's AI agenda has leaned toward surveillance and counter-terrorism and should be recentered on human rights and democratic governance.

  • The companies racing to build frontier AI are now racing to govern it (CIO)

    CIO, 2024

    Reporting that the firms building frontier AI are also shaping its governance through safety frameworks, standards bodies, and policy lobbying, raising questions about who sets the rules.

  • AI Regulations Around the World — Mind Foundry

    Mind Foundry (an AI company), 2026

    Surveys AI rule-making across major jurisdictions including the EU AI Act, US state laws, China, and the UK, noting where approaches diverge.

  • AI Regulation Tracker (regulations.ai)

    regulations.ai, ongoing

    A maintained tracker of AI laws and proposals by jurisdiction, covering the EU AI Act, US executive actions, and country-level rules.

  • Prepared Remarks: Sanders — Regulating AI Is as American as Apple Pie

    Senator Bernie Sanders, 2026

    Senate floor argument that regulating AI follows American precedent — labor, consumer, and antitrust law — and that self-regulation by the largest firms is not sufficient.

  • US and Chinese Visions for AI Regulation Differ Sharply at UN Meeting

    Politico, 2026

    Report on the US and China presenting sharply different visions for AI regulation at a UN meeting, reflecting a broader split over state versus market-led oversight.

  • OpenAI, Anthropic CEOs Call for Global AI Regulation at UN

    Al Jazeera, 2026

    Coverage of the heads of OpenAI and Anthropic telling the UN that AI development requires global regulation, while critics note the conflict of firms shaping their own rules.

  • Nvidia's Jensen Huang on AI safety and regulation (Politico)

    Politico, 2026

    Nvidia's CEO on AI safety and regulation, including his position that overly broad rules could slow progress and that standards should be risk-based.

  • NIST AI Risk Management Framework

    NIST, 2023 (updated)

    The US National Institute of Standards and Technology's voluntary framework for managing AI risk across an organization: govern, map, measure, manage.

  • AI Best Practices for Authors — The Authors Guild

    The Authors Guild, 2026

    Guidance for authors on AI: disclosure, contract terms, copyright registration, and protecting work from unauthorized training use.

  • Publisher Policies on AI — Oklahoma State University Library

    Oklahoma State University Library, ongoing

    A library-maintained directory of how academic and trade publishers address AI in their policies, covering authorship, peer review, and training.

  • An Overview of Catastrophic AI Risks

    Center for AI Safety, 2024

    A survey of catastrophic and existential risks from AI, grouped into categories such as misuse, misalignment, and structural risk.

  • AI and the A-bomb: What the Analogy Captures and Misses (Bulletin of the Atomic Scientists)

    Klyman and Piliero, Bulletin of the Atomic Scientists, 2024

    Examines the common analogy between AI and nuclear weapons — what it captures about runaway risk and what it misses about intent, verification, and the actor.

  • Statement on AI Extinction Risk

    Center for AI Safety, 2023

    A one-sentence public statement signed by researchers and executives that mitigating extinction risk from AI should be a global priority alongside pandemics and nuclear war.

  • AI Agents, Misalignment and the Risk of Losing Human Control: Evidence from the OpenAI-Hugging Face Incident

    UN Independent International Scientific Panel on AI, 2026

    A UN scientific panel brief using the OpenAI–Hugging Face incident to examine the risk of losing human control over autonomous AI agents.

  • Managing Extreme AI Risks Amid Rapid Progress

    Yoshua Bengio and colleagues, 2023

    A research agenda for reducing extreme AI risks during rapid capability progress, covering governance, technical safety, and coordination.

  • How Much Should We Spend to Reduce A.I.'s Existential Risk?

    Chad Jones, 2023

    An economic argument for how much society should spend to reduce AI existential risk, framed around expected value and uncertainty.

Where to go next