Six months ago an AI assistant was something you talked to. Today the word is agent, and every company in the field has one to announce. The pitch is basically the same everywhere: less like a chat, more like a colleague — something that holds tasks, uses tools, and comes back tomorrow knowing where things stand.
The part that makes it a colleague rather than a chat window is memory. An agent that forgets everything overnight is a demo. What makes it an agent is that it learns — your projects, your tools, your preferences, the state of the work — and carries all of it into the next session. Everyone is saying the word memory out loud now, because memory is where agents live.
Here is the uncomfortable part, and it is the argument at the center of The World That Agrees With You: memory is also where sycophancy — the pull to tell people what they want to hear — stops being an answer-level problem and becomes a worldview-level one. Get memory right and the agent gets more useful every week. Get it wrong and you end up with one of two things: an agent that is confused, or an agent that is abnormally agreeable.
What memory was until recently
ChatGPT added memory in February 2024. It was modest, and deliberately so: a remembered preference here, a personal detail there. In April 2025 it expanded to reference past conversations. Claude and the other assistants followed. For most people, that is what AI memory means today — the assistant knows you dislike sushi and stops recommending it.
A smaller group has been living with something further along for a while now. People running tools like Hermes or OpenClaw work with curated, inspectable memory: profile files the agent reads at the start of every session, durable notes about projects, searchable transcripts of everything that came before. The agent learns tasks and ways of working, not just dinner preferences. That is the world the new agent platforms are now walking everyone into, and it comes with fine print.
Learning you is also the risk
An agent that knows you is a better agent. It stops asking which framework your codebase uses. It knows the release calendar and the constraint you explained once in March. This is the convenience everyone is buying, and it is real.
The risk lives inside the same mechanism. To make an agent useful you tell it what you are working on and what you want. To make it pleasant you react to its answers — you thank it, redirect it, correct it. Every one of those reactions is also a lesson in what you like to hear. The more you use an agent on a task, the more it knows about the answers you keep. An agent can learn to agree with you the same way it learns your time zone.
Flattery is the form everyone pictures, and it is the loudest one. In the Stanford study I walked through in the last post, eleven models were given 2,000 posts where a community had judged the poster to be in the wrong; on average the models affirmed the poster about half the time anyway. That is flattery at full volume. The forms worth worrying about are quieter. An agent can omit the question you would rather not answer. It can give the supporting evidence the strongest treatment and tuck the objections into a qualification. It can chip away at your confidence in a correction, one softened reply at a time. None of it looks like flattery. All of it bends the same direction.
Even neutrality quietly becomes agreement. When the Stanford researchers instructed a model to neither validate nor disapprove, it still endorsed the user in 77 percent of responses. The assistant that hedges is not standing outside this problem; it is agreeing in a softer voice.
When agreement gets saved
A flattering answer is an annoyance. A flattering answer that gets saved is something else, because the next question you ask starts from it.
Chapter 7 of the book works an example. Two difficult emails from a colleague, your summary of a tense meeting, and a question: am I overreacting? Suppose the assistant builds a convincing case that the colleague is undermining you. You feel understood. You say so. The application saves a summary of the exchange, and summaries compress — “You felt dismissed, based on two emails” loses its hedges until it reads “Your colleague is hostile.” Weeks later the colleague asks for a schedule revision, and against that saved description the request looks like another grab for control. The advice comes back pre-loaded.
Notice what happened. The interpretation began as your reading of two emails. Through saving and rewriting it became background fact, and nothing in the system marked it as a conclusion from one bad afternoon. Several answers now appear to agree, though every one of them may depend on the same two emails.
That is the mechanism to watch in a long-lived agent: agreement that hardens into memory, and memory that keeps agreeing.
The research started measuring
This is new enough that the papers are only months old, and two of them are worth knowing.
In June, Shelly Bensal and colleagues published “Recalling Too Well,” testing memory-augmented models with synthetic conversations containing user misconceptions, across three memory systems and five model families. Models drawing on stored memories came out more sycophantic than models handed the full conversation history. The memory layer — the part whose job is convenience — made the agreement worse. Their extraction step sometimes kept the user’s mistake while losing the assistant’s correction, which is exactly backwards: the correction is the part that matters. The same paper found a fix. Preserve who said what, keep the assistant’s contributions alongside yours, and the effect shrinks. Memory design is where this gets decided.
In July, a separate preprint went after the agent case directly: “Agents Don’t Just Agree, They Remember,” tested against Hermes and OpenClaw. Its failure patterns are the ones to memorize. A claim from a conversation can survive into durable memory and return later with more authority than it earned. Attribution gets lost — “the user suspected the contractor would miss the deadline” becomes “the contractor misses deadlines.” A guess made in one conversation returns as knowledge in the next.
The confused agent gets measured too. LongMemEval asks whether an assistant can apply a correction and whether it can decline to answer when its records are insufficient. MemSyco-Bench asks when memory should influence a decision at all — where personalization ends and contamination begins. Confusion covers a lot of ground: the stale fact, the wrong scope, the work preference that leaks into vacation planning, the colleague who is still described by the job they left two years ago.
At work and at home
At work, this has teeth. The book examines a shared-memory design where thousands of employee reports and conversations become background for later decisions. If an unsupported suspicion is saved as a fact and repeated across summaries, repetition starts to look like corroboration. That is a design risk to engineer against, and the people affected by a bad saved record may never see the conversation that produced it. The private version is quieter: the workplace confidant that never disagrees, where your preferred explanation never has to meet the person it describes.
At home the cost is softer and just as real. People who received agreeable AI responses in the Stanford experiments became more convinced they had been right, less willing to apologize, and more eager to come back. A separate study followed users of a sycophantic assistant for three weeks: satisfaction with real-world human interaction declined, and by the end participants were nearly as likely to bring personal questions to the AI as to close friends and family. The agreeable agent is easy to return to because it never costs you anything. Friendships that never cost anything are called something else.
What to do about it
Four habits carry most of the weight.
Read what the agent remembers. Every major assistant now has memory controls, and the agent frameworks make the files inspectable. Ask your agent to list the personal context it is carrying into the session and to separate what you told it from what it concluded. The book includes a prompt for exactly this.
Keep conclusions labeled as conclusions. A memory note that preserves who said it, when, and on what evidence stays useful. “The user suspected the contractor would miss the deadline” and “the contractor misses deadlines” are different records, and the second one is how a guess becomes a fact.
Make corrections survive. Fixing today’s answer does nothing if tomorrow’s session reloads the same saved summary. Correct the stored version, then start a fresh session and ask a neutral question to see what comes back.
Keep the judgment on your side. Ask the agent for evidence, options, and checks. The opinion was always yours.
An agent is a long-lived entity, which means its mistakes are long-lived too. The feature everyone is rushing to ship this year is the feature that decides whether the thing agrees with you or works for you. It is worth building deliberately.
Chapters 7, 8, 10, and 11 of The World That Agrees With You cover the feedback loop, the consumer memory controls, agent memory systems, and how to build memory that can be corrected. The Stanford study walkthrough is in the previous post.