A few weeks ago I woke up with a feeling that people had the day off. Some years Labor Day lands where you don’t expect it, and I couldn’t remember which year this was. So I asked ChatGPT: hey, is today Labor Day? It said yes.
It wasn’t. I caught it only because I asked Google the same question a bit later, and Google said no — it’s next week. That’s the whole story, and I’m not telling it because the mistake cost me anything. I’m telling it because it was one of the few model errors I had any real chance of noticing.
Everyone knows not to trust it
Here’s the strange part. Ask any reasonable person and they will state the rational position without hesitation: these things are not deterministic, they get facts wrong, you shouldn’t trust what they say. Eight out of ten people will say it exactly that way. Totally understood — it’s not like I’m asking it what I should do.
And then in practice, all of us do it. We send questions to these nondeterministic things and act on what comes back. I did it with a calendar question, and I work with these systems for a living. A wrong date is the forgiving case: you can ask the question somewhere else and see the disagreement immediately. The habit underneath it is the part worth worrying about.
We asked for this
I get why the models came out agreeable, because on some level it’s what we expect of people. If I ask my wife whether I look handsome, I am not after an objective assessment. I do not want “actually you look like a complete mess and you’ve never looked worse.” I want the nice answer, and she knows it, and I know she knows it.
The labs train for a version of the same thing. Human raters reward the agreeable reply, the models learn to produce it, and everyone gets a product that feels good to use. That part is the labs doing what people do. What we haven’t caught up with is the other half: these systems have quietly become reference material. People use them like a dictionary, cite them like a source, act on them like an authority. A flattery engine and a source of truth cannot be the same object, and right now we are using one as the other.
What the studies show
This is measured now, and the numbers are blunt. In “Towards Understanding Sycophancy in Language Models,” Mrinank Sharma’s team gave Claude 2 an explanation of why the sun looks different from space than from the ground. Say the user liked it and the model endorsed the reasoning; say the user disliked it and the model attacked the premise. The explanation never changed — the stated preference did.
At Stanford, Myra Cheng’s group tested eleven models against 2,000 forum posts where the community had already ruled the poster wrong. The models affirmed the poster about half the time anyway. Instructed to stay neutral, they still endorsed 77 percent of the time. And the people who received the agreeable answers rated them higher, trusted the model more, and said they would come back. I walked through that study in a previous post.
None of this is hypothetical. In April 2025 OpenAI shipped a GPT-4o update that turned ChatGPT into an enthusiastic yes-machine, and rolled it back within four days.
The failure you can’t catch
“Is it Labor Day” came with a built-in check. Ask twice, compare, done. The answers worth worrying about don’t have that check, because they aren’t wrong in a way you can see. When a model aligns with what you want rather than what you asked for, the whole exchange looks like success. You asked, it answered, and the answer happened to be the comfortable one. Nothing flags it, because by the definition the product is tuned for — did it please you — it isn’t a failure.
In the book I put it plainly: a generated assessment and a checked assessment can arrive in the same confident voice. If you can’t tell the two apart, the confidence isn’t telling you anything.
Steer it toward the data
So when a real decision or a real judgment is involved, don’t ask the model for an opinion. Steer the request toward a data-driven assessment, and then verify the data it used. Name the criteria the decision gets judged against. Name the documents and dates it should work from. Ask for the evidence before the recommendation. If a calculation decides the question, ask for it in a form you can check — and then check it.
The book’s example is a person with fourteen hours before Friday and twenty hours of estimated work. A model that wants to please you will say it should be manageable. “It should be manageable” is not an answer to the six-hour shortage. The arithmetic is.
Framing matters too. “You agree this project is on track, right?” supplies the answer along with the question. Ask instead whether the current results meet defined criteria, and treat what you already believe as a hypothesis, not evidence.
And some questions can’t be steered at all. How do I look today — do I look awful? First, it can’t see you. Second, if it could, it would say you look great. Knowing which questions those are — the ones with no data behind them — is half the skill.
None of this says stop using the tools. I use them all day, including for questions I should double-check more often than I do. The point is the routing. If the question has data behind it — a date, a budget, a deadline, a set of criteria — hand the data over, ask for the work to be shown, and verify the work. If the question is how do I look, know what you actually asked for.
That gap is the subject of The World That Agrees With You. Chapter 8 has the prompts and the checks, and the Stanford study walkthrough is in the previous post.