Skip to content
Launchpad Library logo

validity

1 free resource on this topic. Everything here is free and hand-picked. You can also search within this topic.

FreeResearch Paper
Research & Papers

What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks

Authors
Meera Desai, Sang T. Truong, Hanna Wallach, Alex Chouldechova, A. Feder Cooper, Jean Garcia-Gathright, Daniel E. Ho, Abigail Z. Jacobs, Sanmi Koyejo, Nicholas Pangakis, Angelina Wang
Venue
COLM 2026 (Third Annual Conference on Language Modeling)
Published
2026-07-09
ID
OpenReview forum 889XnQKyhM

COLM 2026 paper from Stanford researchers and collaborators that turns the yardsticks on themselves: applying convergent and discriminant validity from the social sciences to 56 capability and safety benchmarks run across 53 models. Rankings on safety benchmarks sharing the same assigned concept often correlate only weakly, capability concepts overlap so much they barely discriminate from one another, and some benchmarks appear to measure something other than what they claim - for example, the bias benchmark BBQ-accuracy tracks reasoning benchmarks more closely than other bias benchmarks. The authors release their full dataset of item- and benchmark-level model outputs and scores for future validity research.

#benchmarks#evaluation#validity#measurement#safety#COLM 2026
openreview.netAdded Oct 5, 20260 opens