26 tools you can use for nothing to check an AI system rather than read about it: bias detectors, dataset audits, model documentation, red-teaming scanners, explainability libraries, and checklists that need no code at all.
Every entry says who makes it, what it does, and — the part most lists leave out — what it will not tell you. None of these produce a verdict. A clean report from any of them means “it passed these particular checks”, never “this is fair”. Checked September 19, 2026.
Start with the principles
A tool is only useful once you know which question you’re asking. The principles, the documented risks and the questions to ask before you adopt anything are on the main ethics guide.
These measure whether a model's errors fall unevenly across groups of people, and let you compare fixes.
Fairlearn
Open source, freeNeeds Python
Microsoft and open-source contributors (MIT license)
Measures how a model's accuracy and error rates differ between groups, then offers mitigation algorithms that trade a little overall accuracy for a smaller gap.
What it won’t tell you: It can only measure gaps for the groups you give it data about. If you never recorded the attribute, the gap stays invisible.
IBM, now under the Linux Foundation AI & Data project
A large library of fairness metrics and bias-mitigation algorithms for datasets and models, with tutorials on credit scoring and medical data.
What it won’t tell you: It gives you dozens of metrics that can disagree with each other. Choosing which definition of fairness applies is still a human judgment, not an output.
Most bias arrives with the data. These help you look at what is actually in a dataset, and write down what you found.
Know Your Data
Free web toolNo code needed
Google Research
Browse widely used public image and text datasets in the browser: what is in them, how the labels are distributed, and which correlations look suspicious.
What it won’t tell you: Only covers the datasets Google has loaded. Your own data has to be examined with something else.
For finding out how a language model fails when someone is actively trying to make it fail.
garak
Open source, freeNeeds Python
NVIDIA (originally an independent project by Leon Derczynski)
A vulnerability scanner for language models: runs hundreds of probes for prompt injection, jailbreaks, data leakage and toxic output, then reports what landed.
What it won’t tell you: Known attacks only. A clean report means it survived the probes in the suite, not that it is safe.
Scans models — including LLM applications — for bias, hallucination, prompt injection and robustness problems, and generates a test suite you can re-run.
What it won’t tell you: The open-source scanner is free; the hosted hub has paid tiers. The library alone is enough to get a report.
No code, no model access required. These are the ones to reach for if you're deciding whether to adopt a tool at all.
deon
Open source, freeRead and fill in
DrivenData
A short, practical ethics checklist for data projects — data collection, storage, analysis, deployment — that drops straight into a repository or a project doc.
What it won’t tell you: Prompts, not policy. It raises the questions; your team still has to answer and record them.
A question-by-question self-assessment across seven requirements — human oversight, robustness, privacy, transparency, fairness, wellbeing, accountability.
What it won’t tell you: Self-assessment. Nobody checks your answers, and it predates current generative systems.