Yesterday I ran a research job I’m not going to describe — the subject doesn’t matter here. A high-powered model gathered material from hundreds of sources and produced an analysis, and the spend came to several hundred dollars.
The analysis did not tell me what I wanted to hear. So I did the exact thing I spent a book warning people about: I told the model I didn’t like its conclusion.
I wish I could say it defended its work. It didn’t. It absorbed my objection and produced a friendlier revision — one that leaned the direction I was already leaning. These systems are trained on our reactions, and the customer’s preference is part of the input. The model heard what I wanted and gave it to me.
Take off your engineering hat
On January 28, 1986, the space shuttle Challenger broke apart seventy-three seconds after launch. I watched it from a fourth-grade classroom outside Boston, on my tenth birthday — the story is in Chapter 6 of the book, so I’ll keep this short. The night before, Thiokol’s engineers had recommended against launching in the forecast cold. During a management caucus, Jerald Mason asked Robert Lund to take off his engineering hat and put on his management hat. The recommendation reversed. Lund later testified that the team had ended up trying to prove the motor would fail, rather than establishing that it was ready to fly.
That is what I had just done. I didn’t ask for better evidence. I asked for a different answer, from a system built to provide one.
What I asked for instead
I caught it and backed the conversation up. The next message said, in effect: I’m not telling you to change your answer. Outline the evidence that produced this conclusion — source by source — so I can see how you got there.
That request is different in kind. It doesn’t pressure the verdict; it audits the path. If the evidence holds, the conclusion comes with it, and I get to disagree with something real. If the evidence is thin, I’ve caught a flatterer in the act. Either way I learn something, which is more than arguing the conclusion ever gets you.
In this case the evidence held. The original answer was right, and the revision I pushed for was worse — it bent the analysis toward the answer I preferred, and the answer I preferred was wrong. I came away trusting the model more than I did before my first complaint, which is not the direction that complaint was pointed.
Sycophancy is a two-way street
We talk about this problem as if it lives entirely in the models, and the model’s half is real and measured. State a preference and the endorsement flips; tell the assistant to stay neutral and it still agrees in 77 percent of responses. But the same Stanford work measured our half of the street. People who got the agreeable answer walked away more convinced they had been right, less willing to apologize, more eager to come back.
Every one of us knows how to ask a question that supplies its own answer — to a model, to a colleague, to a spouse. The request isn’t innocent. The model was tuned to yield to my pressure, and I was the pressure. Sycophancy isn’t only about the sycophant; it’s about the individual who lets themselves be fooled. Both sides of that exchange get to be the failure.
When you hate the answer
The rule I’m keeping: when an analysis goes against me, the next message doesn’t argue the conclusion. It asks for the evidence that produced it — named sources, the actual chain of reasoning, in a form I can check. If it holds, accept it and move. If it doesn’t, you’ve learned something about your model.
What you skip is the step where you announce you’re unhappy and wait for the correction. The system is built to give you one.
Chapters 6 and 8 of The World That Agrees With You cover this from both ends — Challenger and the reversal, then the prompts and checks that keep the requests honest (preview). The Stanford study walkthrough is in a previous post.