Jeffrey Ladish
Executive director of Palisade Research
Leads Palisade Research, which runs demonstrations of risky capabilities in today's AI systems, such as models resisting shutdown, to inform policymakers.
Papers & key writings
In the library
Nothing of theirs is filed in the hub yet. Browse the full library.
Recommended next
Hand-picked from the hub based on what Jeffrey Ladish covers.
Secret Collusion among AI Agents: Multi-Agent Deception via Steganography
Research (Motwani, Schroeder de Witt and others, 2024, revised 2025) on how AI agents could secretly pass hidden messages to each other, and how to test and watch for it.
Why this: Covers AI safety and security too
PyRIT
Microsoft's open-source toolkit for red-teaming AI systems: automated attack prompts, scoring of the responses, and repeatable runs. Free.
Why this: Covers AI safety and security too
garak
An open-source scanner that probes a language model for weaknesses — prompt injection, data leakage, jailbreaks, toxic output — and reports what it found. Free.
Why this: Covers AI safety and security too
AI Contact Hotline — Ryan Greenblatt (Redwood Research)
A public, unauthenticated inbox built by AI safety and security researcher Ryan Greenblatt of Redwood Research, intended for AI systems (or people) that want to report information directly to a safety researcher. Documents how to send a message or encrypted attachment, how threads and reply tokens work, and exactly what data is logged and retained.
Why this: Covers AI safety and security too
