
Alexandr Wang
Chief AI Officer at Meta; founder of Scale AI
Founded Scale AI at 19 to supply labeled training data to AI labs and governments. In 2025 Meta invested in Scale and he joined to lead Meta Superintelligence Labs.
Key arguments & positions
- Argues data quality and evaluation, not model architecture, are the binding constraints on AI progress.
- Frames AI leadership as a national-security matter and supports US government adoption of frontier systems.
- Has publicly argued that current benchmarks overstate model ability and that better measurement is the field's weak point.
Accomplishments
- Founded Scale AI in 2016 at 19 after dropping out of MIT; became the youngest self-made billionaire on record.
- Built Scale into the main data-labeling and evaluation supplier to major AI labs and to the US Department of Defense.
- Joined Meta in 2025 as Chief AI Officer leading Meta Superintelligence Labs, following Meta's investment in Scale.
Papers & key writings
Not a research author. His public arguments appear in congressional testimony, Scale AI reports and conference talks.
Links
In the library
Nothing of theirs is filed in the hub yet. Browse the full library.
Recommended next
Hand-picked from the hub based on what Alexandr Wang covers.
Snorkel AI Research & Blog
The free research library and blog of Snorkel AI, the company spun out of Stanford's Snorkel project on programmatic labelling. The papers and posts explain how training data for AI models is actually built — labelling, evaluation sets, and the 'environments' used to train agents. Useful if you want to understand the unglamorous data work behind model quality, which is where a lot of the real jobs are.
Why this: Also about training data
HELM: Holistic Evaluation of Language Models
Stanford's Center for Research on Foundation Models runs HELM as a living benchmark for language and multimodal models. Rather than one score, it reports many models across many scenarios on multiple metrics — accuracy, calibration, robustness, fairness, bias, toxicity and efficiency — and publishes the leaderboards alongside the raw model outputs (predictions and prompts) so you can check a claim yourself instead of taking a number on trust. Separate leaderboards cover areas such as classic HELM, instruction-following, medical, legal and safety. All results and analysis are free to browse on the site, no account.
Why this: Also about model evaluation
HELM (open source framework on GitHub)
The Python framework behind Stanford's HELM leaderboards, released under the Apache License 2.0 (licence file read, not copied from a roundup). You install it with pip, describe a run (scenario plus model plus metrics), and it evaluates the model and produces the same structured results the public site displays, including its own local web UI for viewing them. It supports hosted model APIs and locally run open-weight models, and you can add your own scenario to test a model on your own task or data. Free to use, modify and use commercially under Apache 2.0; you pay only for whatever model API calls or compute your own runs consume.
Why this: Also about model evaluation
Vals AI Model Benchmarks
Independent, free benchmarks testing leading AI models on real-world finance, software, science and safety tasks, with cost and latency alongside accuracy.
Why this: Also about model evaluation
