Skip to content
Launchpad Library logo

jailbreaks

3 free resources on this topic. Everything here is free and hand-picked. You can also search within this topic.

FreeTool
AI Models & Benchmarks

JailbreakBench

An open benchmark for testing how well large language models resist jailbreak attacks — prompts designed to get a model to produce harmful or unwanted content. It provides a public repository of jailbreak prompts, a standardized evaluation library with a defined threat model and scoring, and a leaderboard of attacks and defenses.

#ai safety#benchmarks#jailbreaks#red teaming#open source
jailbreakbench.github.ioAdded Oct 1, 20260 opens
FreeTool
AI Models & Benchmarks

HarmBench

A standardized evaluation framework for automated red teaming of large language models, measuring how often attacks get models to comply with harmful requests and how reliably models refuse them.

#ai safety#benchmarks#red teaming#jailbreaks
harmbench.orgAdded Oct 1, 20260 opens
FreeTool
AI Models & Benchmarks

StrongREJECT

Documentation for StrongREJECT, an open-source benchmark and Python package for evaluating LLM jailbreaks. It includes rubric-based and fine-tuned evaluators, several dozen baseline jailbreaks, and a dataset of prompts across six categories of harmful behavior, from disinformation to violence.

#ai safety#benchmarks#jailbreaks#open source#python
strong-reject.readthedocs.ioAdded Oct 1, 20260 opens