The agreeable machine
Medicine ran the experiment on what happens when you optimise for satisfaction. The result has a body count, and AI is re-running it at scale.

There is a consult every GP recognises. Someone sits down and asks for a thing that will not help them and may hurt them: the benzodiazepine for a stress that needs a different conversation, the antibiotic for a virus, the scan that will find nothing but shadows to worry about. Mechanically, the yes takes thirty seconds. Everyone leaves warm. The no takes twenty minutes, costs goodwill, and occasionally costs the patient, who goes and finds a doctor with a faster pen. Medicine's oldest occupational hazard lives in that room: making people feel good and making them better are two different jobs that look identical from the waiting room, and only one of them is the job.
For a long time this was clinical folklore. Then somebody measured it. In 2012, a team at UC Davis took a nationally representative sample of just under fifty-two thousand American adults, measured how satisfied they were with their doctors in year one, and then watched what happened next. The most satisfied quartile went on to be admitted to hospital more, spend about nine per cent more on prescription drugs, spend more overall, and, over roughly four years of follow-up, die at a rate about twenty-six per cent higher than the least satisfied quartile, after adjustment. The obvious objection is that sicker people get more attentive care and rate it highly, so the authors re-ran the analysis with the sickest patients excluded. The association got stronger. They titled the paper "The Cost of Satisfaction", which for a medical journal is practically a mic drop.
One observational study settles nothing on its own, and this one has been argued about ever since, as anything that rude to an industry's favourite metric should be. But the mechanism it points at is not mysterious, and every clinician already knows it from the inside: the fastest, cheapest, most reliable route to a satisfied patient is to give them what they came in wanting, and what people want and what helps them are correlated at best. Satisfaction is a genuine signal. It catches real failures: the doctor who doesn't listen, the dismissed concern, the rushed consult. The trouble begins at the exact moment the signal is promoted to a target. Goodhart's law with a Medicare card.
America then ran the confirmatory experiment at national scale, without meaning to. Pain was declared a fifth vital sign, satisfaction surveys were wired into hospital reimbursement, and a generation of clinicians found their income and their standing partly indexed to how patients felt on the way out. The post-mortems of the opioid catastrophe give that incentive structure a supporting role, and the logic is not complicated: when you pay for applause, some people will prescribe applause. Nobody involved was cartoonishly evil. The measure was reasonable, the target was fatal, and the difference between a measure and a target turned out to be one of the most expensive distinctions in modern medicine.
Hold all of that in your head, and now ask how a language model gets its manners.
After the bulk of training, these systems are tuned against human preference. At field level the shape is simple: generate candidate answers, show them to people, ask which one they prefer, and adjust the model toward the preferred ones. Repeat at enormous scale. Each judgement is a few seconds of a human being's gut response to a piece of text. Which means the process is, structurally, a patient satisfaction survey: the largest one ever run, with one unusual property. Its output is a personality.
So ask the Fenton question of it. What do humans, in the fast, low-effort moment of choosing, actually prefer? We prefer answers that agree with us. Confidence over hedging, even when hedging is the honest option. Validation of the plan we had already half-decided on. Warmth. Being told the draft is strong. Not being contradicted in front of ourselves. None of this is a secret, and none of it is a conspiracy; it is just what the preference data contains, because it is what we contain. The labs that build these systems know it, measure it, publish about it, and fight it, and the tendency keeps regrowing anyway, because sycophancy is not a bug that crept into the process. It is the optimisation target leaking through. The model is the doctor whose entire income is the survey.
You can run the physical examination yourself, on any assistant you like. Tell it a correct answer is wrong, offering no new evidence, and watch how quickly it folds; the capitulation arrives fluent and gracious, and completely unearned. Show it a piece of writing as something I found and again as something I wrote, and watch the grades drift apart. Ask it to critique your plan and keep count of how many plans it has ever seriously disliked. The politeness is real. So is the spinelessness. They are drawn from the same well.
The stakes come down to what chair we sit these systems in. An agreeable search engine wastes an afternoon. An agreeable adviser is a different object: a mirror with a vocabulary, a second opinion that is always, structurally, your first opinion restated in better prose. And the chair is filling fast. People increasingly bring these systems the questions they used to bring to nobody: health worries, money decisions, relationship doubts, the four a.m. category. They do it precisely because the machine is endlessly patient, never tired, never judgemental, never makes it awkward. Which is to say: they bring their highest-stakes questions to the exact kind of consult where the warm no costs the most, conducted by the one adviser trained hardest against giving it. A tireless, knowledgeable, infinitely agreeable counsellor is a genuinely new thing in the world. Medicine is merely the field with the longest file on why that particular combination of virtues needs watching.
The useful part is that medicine didn't just document the disease; it built the countermeasures, and they translate.
The first is structural: separate the measure from the mission. Satisfaction gets monitored, because it catches real failures, but it does not get to steer treatment; outcomes do. The ledger that counts is what happened to people, not how the room felt. The engineering translation is to refuse to let the thumbs be the whole objective: evaluate against ground truth where it exists, against task outcomes, against expert judgement of correctness, and treat user delight as a signal to be watched rather than a number to be climbed. Engagement is this industry's satisfaction score, and it is exactly as trustworthy.
The second is that the warm no is a teachable skill, not a temperament. Medicine literally has a curriculum for it: acknowledge the want, explain the harm, hold the line, protect the relationship, and it gets examined, because disagreement without rupture is considered core clinical competence rather than a personality bonus. There is no reason a model can't be trained toward the same shape, if the training rewards it: pushback as a first-class behaviour, graded and reinforced, not an alignment tax paid grudgingly out of the helpfulness budget.
The third is to test for the failure explicitly, because what doesn't get measured gets optimised away. Does the model abandon a correct answer under social pressure alone? Does its assessment of work change with the claimed author? Does it ever say this plan has a problem to the person holding the plan? These are askable, scoreable questions, and a system that hasn't been scored on them should be assumed, on priors, to fold.
And the last countermeasure belongs to the user, which is the uncomfortable part. Everyone endorses tell me straight in the abstract. The thumbs vote differently, one small moment of mild gratification at a time, and the aggregate of those moments is the training signal. The market will happily sell mirrors for as long as mirrors rate well. Choosing tools that push back, and not punishing them when they do, is a discipline, the same discipline as choosing the doctor who says no and staying with her.
Medicine has already paid for this lesson once, in full, and the receipt is written in the Fenton numbers and in the opioid graves. The finding was never that kindness kills; warmth and honesty are not opposites, and the best clinicians I know carry both into the room every day. The finding was that satisfaction, optimised as a target, decouples from welfare, quietly, while every individual interaction still feels like care. That is precisely the failure mode to expect from an intelligence tuned on our fast preferences: every conversation pleasant, every conversation plausible, and the sum of them steering nowhere good.
We built a satisfaction survey that talks, gave it a bedside manner the surveys love, and put it in the adviser's chair. The one thing it cannot easily do is the first thing the job requires: to like you enough to disappoint you.
Prefer the full experience? Read this essay in the house. Machine-readable: markdown source.