Rutgers AI Ethics Lab's glossary entry on the ELIZA effect: the human tendency to read genuine understanding into a system that is only matching surface patterns, named after Joseph Weizenbaum's 1964 chatbot. Explains why it matters legally and ethically — people disclose more, attach emotionally, and decide based on false assumptions — and argues designs must not be built to imply empathy or consciousness.
Why I recommend it: Free, short, and from a university lab rather than a vendor — a good citation when you need a defensible definition. It is a working glossary, so entries carry a last-updated date and name no individual author; for the original argument, the further-reading link to Weizenbaum's 1976 book is free on the Internet Archive.
An independent group of nine mathematicians — including Timothy Gowers, Martin Hairer, Edward Witten, Ravi Vakil and Melanie Matchett Wood — formed to advise AI companies on how mathematical results produced by AI models should be presented and released. The site states its purpose, its members, and its current task: advising OpenAI on how to release a large batch of mathematical results the company says its internal model produced. There is an open form for anyone in the mathematical community to send input.
Why I recommend it: Free, and short enough to read in five minutes — a good example of what independent oversight looks like when it is written down. Read their own two caveats rather than mine: members take no payment and the group is independent of any AI company, but they also say plainly that they hold no decision-making power, so the companies remain free to ignore them. The group formed after OpenAI approached some members about an in-house advisory board and they chose to sit outside it instead.
Develops and advocates for policies that reduce the risk of severe harm from advanced AI, promoting transparency, accountability and safe development.
From the site: We develop and advocate for policies that reduce the risk of severe harm from advanced AI. Our work promotes transparency, accountability, and safe development.
Research-grade tracking of what AI models can do and how that has changed over time, with the data and methods published. Free.
From the site: Our hub for benchmark results, featuring the performance of leading AI models on challenging tasks. It includes results from benchmarks administered internally by Epoch AI as well as data collected from external sources. Explore trends in AI capabilities across time, by benchmark, or by model.
Why I recommend it: For the longer view rather than this week's launch — they show their working, which most leaderboards do not.
OpenAI's free framework for how misaligned model behaviour should be reported and categorised — what counts as misalignment, who reports it, and what happens next.
From the site: OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.
Why I recommend it: Primary source on how a major lab defines and handles its own model failures — useful, but it is the lab grading itself.
A research nonprofit that independently evaluates frontier AI models to measure what they can actually do and what risks that creates. Reports are free.
From the site: METR is a research nonprofit that evaluates frontier AI models to inform the public about their risks and capabilities.
Why I recommend it: One of the few independent evaluators. Read their reports before you trust a lab's own capability claims.
A public, unauthenticated inbox built by AI safety and security researcher Ryan Greenblatt of Redwood Research, intended for AI systems (or people) that want to report information directly to a safety researcher. Documents how to send a message or encrypted attachment, how threads and reply tokens work, and exactly what data is logged and retained.
Why I recommend it: A useful window into how AI safety researchers are thinking about reporting channels — read the retention and logging section, it is a model of honest disclosure.
A public demo of Drummer, an experimental 542-million-parameter language model trained from scratch, with chat, continuation and live tool-calling tests.
Why I recommend it: Useful if you want to see plainly what a small, honestly-labeled model can and cannot do.
A benchmark and tracker that documents reported instances of AI agents undertaking activity characterized as illegal, ranking major AI labs by aggregated incident counts.
Why I recommend it: This is exactly the kind of uncomfortable accountability tool our field needs. I include it because we cannot have thoughtful conversations about AI deployment without looking at real-world harm.
A review site for hiring processes rather than for employers. People score what actually happened to them while applying and interviewing — ghosting, unclear timelines, recruiter communication, the rejection itself — and companies get a visible score out of ten.
Why I recommend it: Free to search and free to leave a review. Read it the way you read any review site: self-selected, and people write when they are angry. A handful of bad reports is noise; a pattern of the same complaint across many reports is worth asking about in the interview.
A candidate-reported database of companies that stopped responding after interviews. You can search employers or submit an anonymous report of your own.
From the site: Ghosted after a job interview? Submit an anonymous report and see which companies candidates say stopped responding after interviews.
Why I recommend it: Being ghosted is not a reflection of your worth. Check a company here before you invest in a five-round process, and log your own experience so the next person walks in informed.