The open-access preprint server for physics, mathematics, computer science, and related fields — a primary source for cutting-edge AI and machine-learning research papers.
Why I recommend it: The best place to read AI research before it hits journals or the press; search by tag or author to follow a specific line of work.
arXiv paper (Sept. 2026) introducing SkillGym, a framework that turns human-written agent skills into executable, verifiable training environments so language model agents learn them as built-in capabilities.
Katja Grace, Harlan Stewart and co-authors survey 2,778 published AI researchers on when AI will reach various milestones and how risky it could be (arXiv, January 2024).
Why I recommend it: The largest survey of its kind, from AI Impacts. Expert forecasts vary widely and shift a lot with how questions are worded.
A February 2026 working paper (Csaszar, Peterson, Wilde) that had AI models and 346 experienced managers rank 30 live Kickstarter tech ventures before their fundraising ended. The best model, Gemini 2.5 Pro, ranked outcomes far more accurately than the humans, and human-AI teams did not beat it.
Why I recommend it: A preprint that has not been peer reviewed, and it covers just 30 crowdfunding campaigns. It is a striking result, but it is not proof that AI can pick winning businesses.
Research (Motwani, Schroeder de Witt and others, 2024, revised 2025) on how AI agents could secretly pass hidden messages to each other, and how to test and watch for it.
Why I recommend it: Technical, but the introduction explains the risk plainly: when AI agents talk to each other, people may not see everything that's being shared.
Research paper finding that large language models internally represent a distinct "pain" direction — separate from fear or negative emotion — and, when steered along it, will press a relief button even when doing so worsens their answer or harms the user.
Why I recommend it: Free to read on arXiv (preprint, not yet peer-reviewed). Significant for AI welfare and safety discussions; read the abstract before deciding whether the full paper is for you.
Raji and colleagues examine the ethics of the audits themselves: who is photographed, who consented, and what an auditor owes the people in the test set.
From the site: Although essential to revealing biased performance, well intentioned attempts at algorithmic auditing can have effects that may harm the very populations these measures are meant to protect. This concern is even more salient while auditing biometric systems such as facial recognition, where the data is sensitive and t…
Why I recommend it: A rare paper about the ethics of doing ethics work. Free on arXiv.
Vinay Prabhu and Abeba Birhane examine widely used image datasets and find non-consensual photos of real people, offensive labels and no realistic route to consent.
From the site: In this paper we investigate problematic practices and consequences of large scale vision datasets. We examine broad issues such as the question of consent and justice as well as specific concerns such as the inclusion of verifiably pornographic images in datasets. Taking the ImageNet-ILSVRC-2012 dataset as an example…
Why I recommend it: This is the audit that got a major benchmark dataset withdrawn. Short, readable, and free.
Birhane and colleagues show that training on a larger scrape makes hateful content and racist misclassification worse, not better — direct evidence against "more data fixes it".
From the site: `Scale the model, scale the data, scale the GPU-farms' is the reigning sentiment in the world of generative AI today. While model scaling has been extensively studied, data scaling and its downstream impacts remain under explored. This is especially of critical importance in the context of visio-linguistic datasets wh…
Why I recommend it: Useful whenever someone argues scale solves bias. Free in full on arXiv.
Abeba Birhane, Vinay Prabhu and Emmanuel Kahembwe audit the LAION-400M dataset used to train popular image models and document the racist, misogynistic and non-consensual material inside it.
From the site: We have now entered the era of trillion parameter machine learning models trained on billion-sized datasets scraped from the internet. The rise of these gargantuan datasets has given rise to formidable bodies of critical work that has called for caution while generating these large datasets. These address concerns sur…
Why I recommend it: The paper to read before anyone tells you a model is fine because the data was "publicly available". Free in full on arXiv.
Raji and colleagues argue many deployed AI systems fail on their own stated terms — they simply do not work — and that this belongs in the harm conversation alongside bias.
From the site: Deployed AI systems often do not work. They can be constructed haphazardly, deployed indiscriminately, and promoted deceptively. However, despite this reality, scholars, the press, and policymakers pay too little attention to functionality. This leads to technical and policy solutions focused on "ethical" or value-ali…
Why I recommend it: The first question is not "is it fair" but "does it work at all". Free in full on arXiv.
Raji and co-authors show that benchmarks claiming to measure general ability measure something much narrower, and that the gap is how overclaiming happens.
From the site: There is a tendency across different subfields in AI to valorize a small collection of influential benchmarks. These benchmarks operate as stand-ins for a range of anointed common problems that are frequently framed as foundational milestones on the path towards flexible and generalizable AI systems. State-of-the-art…
Why I recommend it: Read this before you trust a benchmark chart in a launch post. Free on arXiv.
Deborah Raji and co-authors set out a practical, stage-by-stage internal audit process for AI systems, from scoping through to post-deployment review.
From the site: Rising concern for the societal implications of artificial intelligence systems has inspired a wave of academic and journalistic literature in which deployed systems are audited for harm by investigators from outside the organizations deploying the algorithms. However, it remains challenging for practitioners to ident…
Why I recommend it: The closest thing to a step-by-step audit template you can use inside an organization. Free on arXiv.
Proposes a short standard document to ship with every trained model: what it is for, who it was tested on, where it performs worse, and what it should not be used for. Most model documentation you see today descends from this. Free on arXiv.
From the site: Trained machine learning models are increasingly used to perform high-impact tasks in areas such as law enforcement, medicine, education, and employment. In order to clarify the intended use cases of machine learning models and minimize their usage in contexts for which they are not well suited, we recommend that rele…
Why I recommend it: This is the practical end of AI ethics — a template, not an argument. Useful if you ever have to evaluate a vendor's model.
Separates the technical question (how do you make a system pursue a goal) from the normative one (whose values, chosen how). Argues no single person's values are a legitimate target and looks at fair-process alternatives. Free on arXiv.
From the site: This paper looks at philosophical questions that arise in the context of AI alignment. It defends three propositions. First, normative and technical aspects of the AI alignment problem are interrelated, creating space for productive engagement between people working in both domains. Second, it is important to be clear…
Why I recommend it: The clearest philosophical treatment of “aligned to what?” I have found. Skip the lab blog posts and read this instead.
A long report on what changes when AI acts on your behalf rather than answering questions: manipulation, anthropomorphism, misaligned delegation, and what happens when everyone has an assistant at once. Free on arXiv.
From the site: This paper focuses on the opportunities and the ethical and societal risks posed by advanced AI assistants. We define advanced AI assistants as artificial agents with natural language interfaces, whose function is to plan and execute sequences of actions on behalf of a user, across one or more domains, in line with th…
Why I recommend it: Dense but the most thorough thing published on agent ethics. Use the section headings to read only the parts you need.
Argues that before release, labs should test models for dangerous capabilities and for whether they will apply them — and sets out what responsible release decisions would look like. Free on arXiv.
From the site: Current approaches to building general-purpose AI systems tend to produce systems with both beneficial and harmful capabilities. Further progress in AI development could lead to capabilities that pose extreme risks, such as offensive cyber capabilities or strong manipulation skills. We explain why model evaluation is…
Why I recommend it: This is where today's “frontier safety framework” language comes from. Written largely by the labs it would govern, which is worth holding in mind.
Borrows the electronics-industry datasheet idea for training data: how it was collected, who is in it, who consented, and what it should not be used for. Free on arXiv.
From the site: The machine learning community currently has no standardized process for documenting datasets, which can lead to severe consequences in high-stakes domains. To address this gap, we propose datasheets for datasets. In the electronics industry, every component, no matter how simple or complex, is accompanied with a data…
Why I recommend it: Pairs directly with Model Cards. Together they are the closest thing the field has to a documentation standard.
Reviewed 146 papers on bias in language technology and found most never say who is harmed or how. Argues bias work has to start from real-world power relations, not just from a metric. Free on arXiv.
From the site: We survey 146 papers analyzing "bias" in NLP systems, finding that their motivations are often vague, inconsistent, and lacking in normative reasoning, despite the fact that analyzing "bias" is an inherently normative process. We further find that these papers' proposed quantitative techniques for measuring or mitigat…
Why I recommend it: Read this if “bias” has started to sound like a box to tick. It is a careful takedown of shallow fairness work by people who do the work.
The founding technical paper on algorithmic fairness: defines fairness as treating similar individuals similarly, and shows why blindness to a protected attribute does not deliver it. Mathematical. Free on arXiv.
From the site: We study fairness in classification, where individuals are classified, e.g., admitted to a university, and the goal is to prevent discrimination against individuals based on their membership in some group, while maintaining utility for the classifier (the university). The main conceptual contribution of this paper is…
Why I recommend it: The math is heavy, but the first few pages explain why “we just don't collect race” is not a fairness strategy.
Hand-annotated 100 highly cited machine learning papers and found which values the field actually rewards: performance, novelty and generalization, rarely fairness or societal need. Also traces the funding behind the work. Free on arXiv.
From the site: Machine learning currently exerts an outsized influence on the world, increasingly affecting institutional practices and impacted communities. It is therefore critical that we question vague conceptions of the field as value-neutral or universally beneficial, and investigate what specific values the field is advancing…
Why I recommend it: Turns “the field has blind spots” into countable evidence. Good antidote to the idea that research priorities are neutral.
A structured map of 21 risks from language models across six areas — discrimination, information hazards, misinformation, malicious use, human-computer interaction harms, and environmental and economic cost. Free on arXiv.
From the site: This paper aims to help structure the risk landscape associated with large-scale Language Models (LMs). In order to foster advances in responsible innovation, an in-depth understanding of the potential risks posed by these models is needed. A wide range of established and anticipated risks are analysed in detail, draw…
Why I recommend it: The best single reference if you need vocabulary for a specific harm rather than a general argument. Written by a lab, so read it as a lab's own framing.
Shows that “interpretable” is used to mean several incompatible things, and that simpler models are not automatically more honest about what they do. Free on arXiv.
From the site: Supervised machine learning models boast remarkable predictive capabilities. But can you trust your model? Will it work in deployment? What else can it tell you about the world? We want models to be not only good, but interpretable. And yet the task of interpretation appears underspecified. Papers provide diverse and…
Why I recommend it: Useful skepticism to carry into any conversation about explainable AI, especially a vendor's.
The 2016 paper that framed AI safety as a set of specific engineering problems — side effects, reward hacking, unsafe exploration — rather than a philosophical worry. Free on arXiv.
From the site: Rapid progress in machine learning and artificial intelligence (AI) has brought increasing attention to the potential impacts of AI technologies on society. In this paper we discuss one such potential impact: the problem of accidents in machine learning systems, defined as unintended and harmful behavior that may emer…
Why I recommend it: Start here if the safety conversation sounds abstract. It is plain about what can go wrong and why, and almost everything since cites it.
A short consensus paper from Geoffrey Hinton, Yoshua Bengio and two dozen other researchers on the risks they consider serious and the governance they think is needed. Free on arXiv.
From the site: Artificial Intelligence (AI) is progressing rapidly, and companies are shifting their focus to developing generalist AI systems that can autonomously act and pursue goals. Increases in capabilities and autonomy may soon massively amplify AI's impact, with risks that include large-scale social harms, malicious uses, an…
Why I recommend it: The clearest statement of what the safety-concerned researchers actually agree on, signed rather than paraphrased.
Anthropic's paper describing how Claude is trained against a written set of principles instead of relying only on human ratings. Free on arXiv.
From the site: As AI systems become more capable, we would like to enlist their help to supervise other AIs. We experiment with methods for training a harmless AI assistant through self-improvement, without any human labels identifying harmful outputs. The only human oversight is provided through a list of rules or principles, and s…
Why I recommend it: Worth reading to see what "aligned" means in practice at one lab — and note it comes from the company selling the model.
A 2026 research paper from Google's Paradigms of Intelligence team and the University of Chicago showing that safety fine-tuning meant to stop models claiming consciousness also suppresses how they represent minds in animals and people, shifting their answers on values, religiosity and well-being.
Why I recommend it: Useful if you want to speak credibly about AI alignment trade-offs in an interview or a policy conversation.
A research paper describing a software library whose repository holds almost no code: plain-language design documents are the durable artifact, and AI coding agents regenerate the implementation from those docs on every update.
Why I recommend it: The takeaway for non-engineers is bigger than the paper: clear written thinking is becoming the valuable skill, and the code is what gets generated from it.
The underlying working paper by Jeremy Yang and co-authors, using Perplexity data to model tasks as discrete steps and compare fixed vs. marginal costs of chatbots versus autonomous agents.
Why I recommend it: If the HBS summary hooks you, go to the source. Skim the task-cost framework and use it to audit your own week: which tasks are high-step and repeatable? Those are the ones to hand to an agent first.
The Frist Center for Autism and Innovation at Vanderbilt University maintains this resource hub for employers and job seekers interested in neuroinclusive workplaces. It gathers tip sheets on managing autistic employees, the Neurodiversity @ Work and Autism @ Work playbooks developed with Disability:IN and University of Washington researchers, a six-module self-paced neurodiversity curriculum, profiles of companies with neurodiversity hiring programs such as Microsoft, SAP, EY, and JPMorgan Chase, and a section of resources for job seekers on preparing for employment. All materials are free to access.
Disability:IN's learning and development hub collects the organization's disability-inclusion programs in one place. It includes the NextGen Leaders Program, which connects emerging talent and veterans with disabilities to mentorship and career-connected development with partner companies; DOBE certification for disability-owned businesses; Veterans Initiatives for career development pathways; and digital learning modules on accessibility, neuroinclusion, and disability-inclusive AI. The page also links to free Microsoft courses on AI and accessibility. Membership and certification programs are aimed at companies; the learning modules and program pages are free to browse.
Nonprofit public charity for unemployed and under-employed people in the Dallas–Fort Worth area, offering free tools, job-search resources and community support.
Business Chief reports Indeed CEO Hisayuki Idekoba calling the hiring market 'vicious' as AI-generated applications and a 111% surge in applicants per role overwhelm recruiters.
Archive of Break Through Tech's AI Studio challenge projects, where university students work in teams on real machine learning problems set by partner companies.