An open benchmark for testing how well large language models resist jailbreak attacks — prompts designed to get a model to produce harmful or unwanted content. It provides a public repository of jailbreak prompts, a standardized evaluation library with a defined threat model and scoring, and a leaderboard of attacks and defenses.
A standardized evaluation framework for automated red teaming of large language models, measuring how often attacks get models to comply with harmful requests and how reliably models refuse them.
Documentation for StrongREJECT, an open-source benchmark and Python package for evaluating LLM jailbreaks. It includes rubric-based and fine-tuned evaluators, several dozen baseline jailbreaks, and a dataset of prompts across six categories of harmful behavior, from disinformation to violence.