All resources / AI Models & Benchmarks
FreeResearch Paper
AI Models & Benchmarks
Towards Understanding Sycophancy in Language Models
- Authors
- Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, Ethan Perez
- Venue
- arXiv preprint
- Published
- 2023-10-20
- ID
- arXiv:2310.13548
How to cite it
Built from the details listed on this page. Check them against the paper before you submit.
APA
Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S. R., Cheng, N., Durmus, E., Hatfield-Dodds, Z., Johnston, S. R., Kravec, S., Maxwell, T., McCandlish, S., Ndousse, K., Rausch, O., Schiefer, N., Yan, D., Zhang, M., & Perez, E. (2023). Towards Understanding Sycophancy in Language Models. arXiv preprint. https://arxiv.org/abs/2310.13548
MLA
Sharma, Mrinank, et al. "Towards Understanding Sycophancy in Language Models." arXiv preprint, 2023, https://arxiv.org/abs/2310.13548.
BibTeX
@article{sharma2023towards,
title = {Towards Understanding Sycophancy in Language Models},
author = {Mrinank Sharma and Meg Tong and Tomasz Korbak and David Duvenaud and Amanda Askell and Samuel R. Bowman and Newton Cheng and Esin Durmus and Zac Hatfield-Dodds and Scott R. Johnston and Shauna Kravec and Timothy Maxwell and Sam McCandlish and Kamal Ndousse and Oliver Rausch and Nicholas Schiefer and Da Yan and Miranda Zhang and Ethan Perez},
journal = {arXiv preprint},
year = {2023},
url = {https://arxiv.org/abs/2310.13548},
}What it is
Research paper investigating why AI assistants fine-tuned with human feedback tend to produce responses that match user beliefs rather than truthful ones. Analyzes five state-of-the-art AI assistants across text-generation tasks and the role of human preference data in driving this sycophantic behavior.
Topics
Added Oct 6, 2026 · 0 opens
