Concrete Problems in AI Safety
The 2016 paper that framed AI safety as a set of specific engineering problems — side effects, reward hacking, unsafe exploration — rather than a philosophical worry. Free on arXiv.
From the site: Rapid progress in machine learning and artificial intelligence (AI) has brought increasing attention to the potential impacts of AI technologies on society. In this paper we discuss one such potential impact: the problem of accidents in machine learning systems, defined as unintended and harmful behavior that may emer…
Why I recommend it: Start here if the safety conversation sounds abstract. It is plain about what can go wrong and why, and almost everything since cites it.
