The free research library and blog of Snorkel AI, the company spun out of Stanford's Snorkel project on programmatic labelling. The papers and posts explain how training data for AI models is actually built — labelling, evaluation sets, and the 'environments' used to train agents. Useful if you want to understand the unglamorous data work behind model quality, which is where a lot of the real jobs are.
Why I recommend it: Research and blog posts are free to read with no signup. Read it for the how, not the verdict: Snorkel sells data services to frontier AI labs, so posts arguing that better data beats bigger models are also a sales case. Everything else on the site is a paid enterprise product — 'request dataset samples' means a sales call.
Meera Desai, Sang T. Truong, Hanna Wallach, Alex Chouldechova, A. Feder Cooper, Jean Garcia-Gathright, Daniel E. Ho, Abigail Z. Jacobs, Sanmi Koyejo, Nicholas Pangakis, Angelina Wang
Venue
COLM 2026 (Third Annual Conference on Language Modeling)
Published
2026-07-09
ID
OpenReview forum 889XnQKyhM
COLM 2026 paper from Stanford researchers and collaborators that turns the yardsticks on themselves: applying convergent and discriminant validity from the social sciences to 56 capability and safety benchmarks run across 53 models. Rankings on safety benchmarks sharing the same assigned concept often correlate only weakly, capability concepts overlap so much they barely discriminate from one another, and some benchmarks appear to measure something other than what they claim - for example, the bias benchmark BBQ-accuracy tracks reasoning benchmarks more closely than other bias benchmarks. The authors release their full dataset of item- and benchmark-level model outputs and scores for future validity research.
An NCDA Career Convergence article by Safaa Amer (Oct. 2026) on how AI tools are reshaping career coaching — expanding access and automating tasks — and the AI literacy and ethical practice coaches need to use them responsibly.
METR's November 2024 release of RE-Bench, a benchmark comparing frontier model agents with human experts on seven machine-learning research-engineering tasks. It includes data from 71 human expert attempts and results for Claude 3.5 Sonnet and o1-preview.
AIES (AAAI/ACM Conference on AI, Ethics, and Society)
Published
August 2025
The methodology paper behind Data Workers' Inquiry, by Milagros Miceli, Alex Hanna, Timnit Gebru and colleagues at the Distributed AI Research Institute. It lays out a participatory framework for centering workers' own questions and epistemic authority in AI research, turning hidden and precarious labor into a site for knowledge-making and change.
A global participatory research initiative spanning nine countries where data workers themselves become community researchers, documenting labor conditions in the AI industry through zines, documentaries, comics, essays, podcasts and animations. Inquiries cover workers at Sama, CloudFactory and Remotasks in Kenya, Venezuela, Syria, Brazil, France and Germany.
The personal site of Noah Shinn, a programmer and researcher who was one of the first employees and a researcher at Sierra. His papers include tau-bench, a benchmark for tool-agent-user interaction (ICLR 2025), work on code-editing instruction following, type prediction, and the Reflexion paper on language agents with verbal reinforcement learning.