Philosophies & Movements
Friendly AI / the Friendliness Theorem
The early name for the project of building advanced AI that reliably wants what is good for humanity, and keeps wanting it as it becomes more capable. “Friendliness theorem” refers to the hoped-for mathematical result that would let builders prove a system stays benevolent under self-improvement — an aspiration discussed in this literature, not an established theorem anyone has proved.
Origin · no single documented coiner
“Friendly AI” was coined by Eliezer Yudkowsky (profiled on our Who’s Who page) in his 2001 document “Creating Friendly AI,” published by the Singularity Institute; the phrasing “friendliness theorem” circulated in that same community rather than from one named paper, and no such theorem exists today. The vocabulary was later largely replaced by “AI alignment.”
