Vinay Prabhu and Abeba Birhane examine widely used image datasets and find non-consensual photos of real people, offensive labels and no realistic route to consent.
From the site: In this paper we investigate problematic practices and consequences of large scale vision datasets. We examine broad issues such as the question of consent and justice as well as specific concerns such as the inclusion of verifiably pornographic images in datasets. Taking the ImageNet-ILSVRC-2012 dataset as an example…
Why I recommend it: This is the audit that got a major benchmark dataset withdrawn. Short, readable, and free.
Birhane and colleagues show that training on a larger scrape makes hateful content and racist misclassification worse, not better — direct evidence against "more data fixes it".
From the site: `Scale the model, scale the data, scale the GPU-farms' is the reigning sentiment in the world of generative AI today. While model scaling has been extensively studied, data scaling and its downstream impacts remain under explored. This is especially of critical importance in the context of visio-linguistic datasets wh…
Why I recommend it: Useful whenever someone argues scale solves bias. Free in full on arXiv.
Abeba Birhane, Vinay Prabhu and Emmanuel Kahembwe audit the LAION-400M dataset used to train popular image models and document the racist, misogynistic and non-consensual material inside it.
From the site: We have now entered the era of trillion parameter machine learning models trained on billion-sized datasets scraped from the internet. The rise of these gargantuan datasets has given rise to formidable bodies of critical work that has called for caution while generating these large datasets. These address concerns sur…
Why I recommend it: The paper to read before anyone tells you a model is fine because the data was "publicly available". Free in full on arXiv.