2665 hand-picked resources, updated every week. Search it, filter it, or just browse a collection and see what catches your eye. Want today’s headlines instead? Read the free AI news feed.
Type
Platform
Topics
Cost
Ask the library about workforce, startup, and technology trends
Answers come only from resources in this hub, with the sources listed underneath.
An open benchmark for testing how well large language models resist jailbreak attacks — prompts designed to get a model to produce harmful or unwanted content. It provides a public repository of jailbreak prompts, a standardized evaluation library with a defined threat model and scoring, and a leaderboard of attacks and defenses.
A standardized evaluation framework for automated red teaming of large language models, measuring how often attacks get models to comply with harmful requests and how reliably models refuse them.
Documentation for StrongREJECT, an open-source benchmark and Python package for evaluating LLM jailbreaks. It includes rubric-based and fine-tuned evaluators, several dozen baseline jailbreaks, and a dataset of prompts across six categories of harmful behavior, from disinformation to violence.
Independent benchmarks comparing AI models and API providers on quality, speed, price and context length, with free charts and leaderboards.
Why I recommend it: Free to browse; the company sells API access to its data. Rankings depend on which tests they choose — compare with a second benchmark source before deciding.
Independent, domain-specific AI benchmarks in legal, tax and finance, testing how models actually perform on professional work rather than academic tests.
Why I recommend it: Free leaderboards; the company sells private benchmarking to enterprises. Domain-specific results are more useful than general scores if you work in these fields.
Artificial Analysis on Solar Mini 4, a proprietary reasoning model from Korean lab Upstage. It scores 24 on the Artificial Analysis Intelligence Index, with Upstage reporting 35B total and 3B active parameters, and is priced at $0.10/$0.40 per million input/output tokens — though it costs about 5x as much per task as GPT-6 Luna (max).
A plain-language guide from OpenK3 explaining what open-weight and closed-weight AI models are, how licenses still restrict open-weight models, real examples including Kimi K3, and how teams can choose between self-hosting and using a hosted API.
Simon Willison's annotated slides and notes from his closing keynote at WeAreDevelopers World Congress North America (Sept. 2026): a chronological tour of the year's LLM developments, starting with the November 2025 models that made coding agents start working.
A public scoreboard that tracks predictions about AI made by public figures — executives, investors and commentators — and scores each person's resolved predictions for accuracy, so readers can see who has a track record worth weighing.
METR's November 2024 release of RE-Bench, a benchmark comparing frontier model agents with human experts on seven machine-learning research-engineering tasks. It includes data from 71 human expert attempts and results for Claude 3.5 Sonnet and o1-preview.
A page listing benchmark scores (HLE, GPQA Diamond, MMLU-Pro, AI BENCHY) for a model called space-bunny-alpha, with confidence intervals.
Why I recommend it: Unverified: I couldn't find who runs this site or who made the model, and several scores are on subsets of the tests. Treat the numbers as unconfirmed until an independent lab reproduces them.