Cognitive scientist; founder of the AI Accountability Lab at Trinity College Dublin
Abeba Birhane audits the enormous datasets frontier models are trained on, and has repeatedly found racist, misogynistic and non-consensual content inside widely used open datasets. She founded and leads the AI Accountability Lab at Trinity College Dublin.
Research labaial.ie
#dataset audits#AI accountability#AI ethics
Read them in their own words
Published elsewhere, not by me — so you get real context beyond the name and the title.
ProfileAP News·July 2024
AI industry is influencing the world. Mozilla adviser Abeba Birhane is challenging its core values
What she found when she audited the giant image datasets that picture generators were trained on.
Argues claims about model capability mean little without audits of the training data behind them.
Has shown repeatedly that widely used open datasets contain racist, misogynistic and non-consensual material, and that scaling makes this worse rather than better.
Holds that the values of machine-learning research are visible in what its papers reward — performance and novelty over fairness or consent.
Accomplishments
Founder and principal investigator of the AI Accountability Lab at Trinity College Dublin, and an assistant professor there.
Lead author of "The Values Encoded in Machine Learning Research" (2021), which read 100 influential papers for the values they actually reward.
Audited LAION-scale multimodal datasets and documented what dataset scaling does to racial classification.
Vinay Prabhu and Abeba Birhane examine widely used image datasets and find non-consensual photos of real people, offensive labels and no realistic route to consent.
From the site: In this paper we investigate problematic practices and consequences of large scale vision datasets. We examine broad issues such as the question of consent and justice as well as specific concerns such as the inclusion of verifiably pornographic images in datasets. Taking the ImageNet-ILSVRC-2012 dataset as an example…
Why I recommend it: This is the audit that got a major benchmark dataset withdrawn. Short, readable, and free.
Birhane and colleagues show that training on a larger scrape makes hateful content and racist misclassification worse, not better — direct evidence against "more data fixes it".
From the site: `Scale the model, scale the data, scale the GPU-farms' is the reigning sentiment in the world of generative AI today. While model scaling has been extensively studied, data scaling and its downstream impacts remain under explored. This is especially of critical importance in the context of visio-linguistic datasets wh…
Why I recommend it: Useful whenever someone argues scale solves bias. Free in full on arXiv.
Abeba Birhane, Vinay Prabhu and Emmanuel Kahembwe audit the LAION-400M dataset used to train popular image models and document the racist, misogynistic and non-consensual material inside it.
From the site: We have now entered the era of trillion parameter machine learning models trained on billion-sized datasets scraped from the internet. The rise of these gargantuan datasets has given rise to formidable bodies of critical work that has called for caution while generating these large datasets. These address concerns sur…
Why I recommend it: The paper to read before anyone tells you a model is fine because the data was "publicly available". Free in full on arXiv.