What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
- Authors
- Meera Desai, Sang T. Truong, Hanna Wallach, Alex Chouldechova, A. Feder Cooper, Jean Garcia-Gathright, Daniel E. Ho, Abigail Z. Jacobs, Sanmi Koyejo, Nicholas Pangakis, Angelina Wang
- Venue
- COLM 2026 (Third Annual Conference on Language Modeling)
- Published
- 2026-07-09
- ID
- OpenReview forum 889XnQKyhM
COLM 2026 paper from Stanford researchers and collaborators that turns the yardsticks on themselves: applying convergent and discriminant validity from the social sciences to 56 capability and safety benchmarks run across 53 models. Rankings on safety benchmarks sharing the same assigned concept often correlate only weakly, capability concepts overlap so much they barely discriminate from one another, and some benchmarks appear to measure something other than what they claim - for example, the bias benchmark BBQ-accuracy tracks reasoning benchmarks more closely than other bias benchmarks. The authors release their full dataset of item- and benchmark-level model outputs and scores for future validity research.
