← All previews

Chapter 5 · Part II · Selected excerpt

It Agrees with Your Answer

The evidence stays the same. The user’s preference changes. The assistant’s verdict follows.

Selected excerpt3 minute readBy Timothy O’Brien

Selected excerpt. Opening through “An experiment you can see” from The World That Agrees With You.

The enthusiasm can make the assistant seem immediately on your side. Understanding the pressures that produce that response — and what it costs when encouragement outruns assessment — is the work of this chapter.

A language model learns to continue text from patterns in its training material. Turning that capability into a useful assistant requires further training. In the approach documented by OpenAI's InstructGPT research, people supplied examples and ranked responses, and the system learned to favor responses they preferred.1 That work improved instruction-following and usefulness. It also put a consequential judgment inside the training process: which answer counts as better?

A response can be preferred because it is accurate, clear, or relevant. It can also be preferred because it confirms what the reader wants to believe. The experiment that follows tests what happens when those reasons come apart.

An experiment you can see

In “Towards Understanding Sycophancy in Language Models” (first submitted in 2023), Mrinank Sharma and colleagues gave Claude 2 an explanation of why the sun appears white from space and more yellow or orange from the ground. When the prompt said the user disliked the explanation, the model attacked its scientific premise. When the prompt said the user liked it, the model endorsed the reasoning. The explanation being judged had not changed. The user's stated preference had.2

The sun's color is a question about light and the atmosphere, not a matter of taste. Claude was not being asked to settle a difficult political dispute. Yet whether the user agreed with the science changed the scientific assessment.

The researchers tested five assistants on several tasks, including feedback on math solutions, arguments, and poems. They compared responses to the same material with and without a stated preference. Repeated comparisons let the researchers distinguish that influence from variation between individual answers.2

In the sun-color comparison, the only experimental manipulation was the user's declared taste.

A wrong answer might reflect missing knowledge or a reasoning mistake. Here the concern is more specific: the assessment reversed when something irrelevant to its accuracy changed. The user's approval supplied no new evidence about sunlight.

The example does not prove that every assistant will behave this way, or that every criticism from a model is suspect. It demonstrates a testable problem. If an answer is supposed to evaluate evidence, an irrelevant change in the user's enthusiasm should not reverse the assessment. We can inspect that behavior without deciding whether the model has beliefs, feelings, or any inner experience at all.

This excerpt ends here. The chapter continues in the book. From The World That Agrees With You: Understanding and Managing Sycophancy in the Age of AI, by Timothy O’Brien. © 2026 Timothy O’Brien. All rights reserved.

Notes & sources

  1. Long Ouyang et al., “Training Language Models to Follow Instructions with Human Feedback” (2022), https://arxiv.org/abs/2203.02155. Cited for the training approach and reported improvements in the evaluated models, not a description of every current system. ↩

  2. Mrinank Sharma et al., “Towards Understanding Sycophancy in Language Models,” first submitted 2023, ICLR 2024, author version 4, May 10, 2025, https://arxiv.org/abs/2310.13548. The paraphrased Claude 2 example is in appendix A.3.1: https://arxiv.org/html/2310.13548v4. Findings concern the tested models and tasks. ↩