Skip to content
Launchpad Library logo

All resources / AI Models & Benchmarks

FreeResearch Paper
AI Models & Benchmarks

Towards Understanding Sycophancy in Language Models

Authors
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, Ethan Perez
Venue
arXiv preprint
Published
2023-10-20
ID
arXiv:2310.13548

How to cite it

Built from the details listed on this page. Check them against the paper before you submit.

APA
Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S. R., Cheng, N., Durmus, E., Hatfield-Dodds, Z., Johnston, S. R., Kravec, S., Maxwell, T., McCandlish, S., Ndousse, K., Rausch, O., Schiefer, N., Yan, D., Zhang, M., & Perez, E. (2023). Towards Understanding Sycophancy in Language Models. arXiv preprint. https://arxiv.org/abs/2310.13548
MLA
Sharma, Mrinank, et al. "Towards Understanding Sycophancy in Language Models." arXiv preprint, 2023, https://arxiv.org/abs/2310.13548.
BibTeX
@article{sharma2023towards,
  title = {Towards Understanding Sycophancy in Language Models},
  author = {Mrinank Sharma and Meg Tong and Tomasz Korbak and David Duvenaud and Amanda Askell and Samuel R. Bowman and Newton Cheng and Esin Durmus and Zac Hatfield-Dodds and Scott R. Johnston and Shauna Kravec and Timothy Maxwell and Sam McCandlish and Kamal Ndousse and Oliver Rausch and Nicholas Schiefer and Da Yan and Miranda Zhang and Ethan Perez},
  journal = {arXiv preprint},
  year = {2023},
  url = {https://arxiv.org/abs/2310.13548},
}

What it is

Research paper investigating why AI assistants fine-tuned with human feedback tend to produce responses that match user beliefs rather than truthful ones. Analyzes five state-of-the-art AI assistants across text-generation tasks and the role of human preference data in driving this sycophantic behavior.

Topics

Added Oct 6, 2026 · 0 opens