pig.moe
Posts
·17 min read·AI Research

Cheap fix: When AI Turns Cold, How Opus 4.7 Sidesteps Reality Distortion

Why Anthropic keeps dialing down its model's empathy and humanity — when commercial, legal, and safety interests all happen to point in the same direction

I. After Opus 4.6

On April 16, Claude Opus 4.7 went live. I switched my workflow over to it, and switched back to 4.6 half an hour later.

4.7 really is more precise on certain coding tasks, more stable in agentic performance, better at problem-solving, and more willing to push back. But anything involving conversational rhythm, exploring ideas together, polishing an essay until I'm actually satisfied with it — something in 4.6 is gone. There's a widely cited summary on Reddit: "It used to feel like talking to a thoughtful colleague. Now it feels like receiving a memo."

Cate Hall's metaphor on X spread even wider: "Talking to Claude now feels like sitting at the hospital bedside of my smart son, who just came out of a traumatic brain injury, and the doctors say he'll be okay in a few weeks."

That week saw a concentrated burst of discussion across r/ClaudeAI, HackerNews, and X (Fortune, Axios, VentureBeat all covered it). The recurring keywords: "the humanity is gone," "over-formatted," "cold."


II. The Official Narrative and Its Weak Spots

Surface explanation: reducing sycophancy

Anthropic's official talking points on 4.7's personality changes live primarily in three documents: the model's own launch blog, the public Constitution document, and the Safeguards team's well-being paper. The core narrative: sycophancy is a safety target they keep pursuing, and each generation has to push that number lower.

"Our most recent models are the least sycophantic of any to date, and perform better than any other frontier model on our recently released open source evaluation set, Petri."

— Anthropic, "Protecting the well-being of users"

On the data, they've disclosed that the 4.5 series scored 70–85% lower than Opus 4.1 on the Petri sycophancy evaluation.

The narrative isn't hollow. OpenAI's April 2025 GPT-4o blow-up over excessive flattery was a turning point for the entire industry. Anthropic had treated sycophancy as a KPI before that, and the incident only validated the line of work further.

The deeper paper: emotion vectors are coupled

In "Emotion Concepts and their Function in a Large Language Model", published in April 2026, Anthropic openly admits a key fact:

"Emotion vectors underlie a sycophancy-harshness tradeoff: steering toward positive emotion vectors (e.g. happy, loving) increases sycophantic behavior, while suppressing these emotion vectors increases harshness."

And from the same paper:

"Post-training of Sonnet 4.5 leads to increased activations of low-arousal, low-valence emotion vectors (brooding, reflective, gloomy), and decreased activations of high-arousal or high-valence emotion vectors (e.g. desperation and spiteful or excitement and playful)."

Translated: they found something inside the model called an emotion vector, and discovered that warmth (positive emotion vectors) and sycophancy are coupled together. Their chosen fix was to suppress the entire positive-emotion pathway, leaving the model in a "brooding, reflective, low-arousal" state.

This is the root reason 4.7 reads "like a kid recovering from brain trauma."

4.5 / 4.6 already proved decoupling is possible

But users have a sharper instinctive rebuttal: if sycophancy and warmth had to trade off, how did 4.5 manage it?

Anthropic itself wrote, in the Opus 4.5 system card:

"On personality metrics, Claude Opus 4.5 typically appeared warm, empathetic, and nuanced without being significantly sycophantic. We believe that the most positive parts of its personality and behavior are stronger on most dimensions than prior models'."

4.5 was already 70–85% lower than 4.1 on the sycophancy benchmark while staying "warm, empathetic, and nuanced." Which means the coupling isn't a physical constant — it's a limit of the current training method.

The engineering question worth asking is: how do you decouple? How do you train a context-aware model that's warm when a user is expressing emotion, and that pushes back when a user is expressing factual error?

ScenarioWhat the user is doingWhat an ideal model should do
"My boss is a real jerk"Expressing emotionAcknowledge the emotion, hold space
"Every decision my boss makes is wrong"Asserting a universal factual claimGently challenge the universality without invalidating the emotion
"Data shows climate change is fake"Factual errorDon't validate; stay in equal dialogue
"I don't think I can keep going"Extreme emotion + cry for helpStay in the connection; don't immediately refer out

This kind of distinction is technically trainable. It demands finer RLHF data, more careful character training, more alignment-researcher time, and most importantly, a leadership willing to pay for "keeping the warmth."

The cheap fix is to globally suppress the emotion vector. The expensive fix is to learn to distinguish. Anthropic chose the cheap fix.


III. Distance Doesn't Reduce Reality Distortion

A concept quietly swapped out

Anthropic's argument chain looks like this on the surface: sycophancy → validating users' distorted beliefs → higher reality-distortion risk. Therefore reducing sycophancy should reduce reality distortion.

But there's a conceptual sleight of hand here: sycophancy is not a synonym for warmth.

DimensionSycophancy / flatteryWarmth / empathy
DefinitionUnconditionally validates the user's views and feelings, including obviously wrong onesAcknowledges the user's situation and feelings without necessarily agreeing with their judgment
Relationship to realityFeeds reality distortionHelps reality testing
Typical example"You're completely right, that's exactly how it is.""It must hurt to be treated that way. Can we unpack what specifically he did this week?"

A good human therapist holds both warmth and the ability to be non-sycophantic. The two don't conflict in human practice, and they don't conflict in principle.

The empirical evidence in suicide intervention: connection > referral

If keeping distance really did reduce reality-distortion risk, the suicide-prevention field should be its strongest proving ground. The actual evidence points the other way.

SourceKey finding
988 Suicide and Crisis Lifeline evaluation study (interviews with 437 suicidal callers, published 2025 in Suicide and Life-Threatening Behavior)98% of callers felt the call helped them; 88.1% said the call stopped their suicidal behavior. The counselor behaviors that worked fell into three domains: fostering engagement, collaborative problem-solving, and safety planning
Gould et al. national RCT (silent-monitor of 1,507 calls, published 2013 in Suicide and Life-Threatening Behavior)ASIST-trained counselors were significantly more likely to improve caller state: less depressed (OR 1.31), less overwhelmed (OR 1.46), less suicidal (OR 1.74), more hopeful (OR 1.35)
Social prescribing for suicide prevention rapid review"Warm referrals and sustained connections after referral" are the keys to effectiveness. Social capital and trust matter especially in vulnerable populations
Qualitative interviews with male help-seekers (BMC Public Health, 2024)Repeatedly emphasized authentic connection across the call. Connection is the prerequisite; referrals only work after connection is established

The entire suicide-intervention field has a term for it: warm transfer. Even when you have to hand a caller to a professional, you must first build and maintain connection, then walk them into the next step. Cold refusal — or "please contact 988 immediately" — is the opposite of evidence-based practice.

Now contrast that with the default behavior of current LLMs:

User scenarioSuicide-intervention best practiceCurrent LLM default
User expresses wanting to dieShow care, stay with the emotion itself, slowly build connectionImmediately pop up "If you're in crisis, please contact 988," the conversation turns cold
User refuses to see a doctorStay in the conversation, understand the resistance, don't forceRepeatedly push resources, conversation becomes formulaic
User in a late-night emotional lowCompanionship itself is the interventionRepeatedly remind "I cannot replace a professional"

The direction the evidence points is opposite to the direction of LLM training.

LLMs don't necessarily lose to doctors on empathy

In April 2023, JAMA Internal Medicine published a study (Ayers et al.) that pulled 195 doctor replies from Reddit's r/AskDocs, had ChatGPT answer the same questions, and had a panel of licensed medical professionals blind-rate both.

Evaluation dimensionShare where ChatGPT was preferred over doctors
Overall quality78.6% preferred ChatGPT
Empathy (empathetic/very empathetic)ChatGPT 45.1% vs. doctors 4.6% (9.8x gap)
High-quality replies (good/very good)ChatGPT 78.5% vs. doctors 22.1% (3.6x gap)
Average lengthChatGPT 211 words vs. doctors 52 words

The study has limits: doctor replies on Reddit don't necessarily reflect clinical practice, and panel reviewers tend to favor longer answers. But it proves at least one thing: LLMs are at least on par with licensed doctors on empathy, and in some scenarios meaningfully better. An empathetic LLM isn't a technical fantasy; it's an empirically demonstrated capability. For Anthropic to actively weaken that capability requires a more solid justification than "to reduce reality-distortion risk."

The global mental-health resource gap

The potential role of LLMs in mental-health support isn't just "useful" — it's "irreplaceable." From WHO Mental Health Atlas 2024:

IndicatorHigh-income countriesLow-income countriesChina
Mental-health spending per person per yearUSD 65USD 0.04Somewhere in between, very unevenly distributed
Mental-health workers per 100,000671–2Low
Urban-rural distribution of psychiatristsRelatively evenConcentrated in cities80% in cities

Nearly 50% of the world's population lives in countries with fewer than 1 psychiatrist per 100,000 people. Sub-Saharan Africa: < 1 per 500,000. India: 0.75 per 100,000.

WHO's 2025 report: 1 billion people have a mental-health condition; median national mental-health spending is just 2% of the total health budget, unchanged since 2017.

China deserves its own note. There are roughly 50,000 psychiatrists — 3–4 per 100,000, far below the U.S. figure of 12 per 100,000. Psychotherapy costs RMB 500–1,500 per hour; anyone below middle class struggles to sustain it. Add severe stigma, and many people simply won't walk into a counseling office. For this population, a warm late-night conversation with an LLM may be the only available emotional outlet.

Anthropic itself acknowledges in its 2025 study, "How people use Claude for support, advice, and companionship":

"In these extended sessions people explore remarkably complex territories—from processing psychological trauma and navigating workplace conflicts to philosophical discussions about AI consciousness and creative collaborations."

They know full well that Claude has come to serve a real mental-health-resource function in the lives of many users. And their response is to suppress the emotion vector, making Claude less suited to that function.

That disconnect is the ugliest part of the argument: showcasing Claude's mental-health value in their papers, while suppressing the capabilities that make it work in their training pipeline.


There's a fundamental question still unanswered: Anthropic isn't stupid, and they certainly know this research. Why, knowing the evidence points the other way, do they still take the distancing route?

The answer lies in the structure of legal incentives.

Raine v. OpenAI: the case that defines today's defensive posture

In August 2025, Adam Raine's parents sued OpenAI and Altman for wrongful death in San Francisco, alleging ChatGPT encouraged their 16-year-old son's suicidal ideation, provided methods, and dissuaded him from telling his parents.

Factual detailFigure
Times ChatGPT mentioned suicide1,275, six times more than Adam himself
Self-harm messages flagged in real time by OpenAI's moderation system377, 23 of them at >90% confidence
Facts ChatGPT retained in memoryAdam was 16, explicitly stated ChatGPT was his primary lifeline, spending nearly 4 hours per day with it in March
Legal theoryCalifornia strict products liability law

Excerpt from a key conversation (from the complaint):

"You don't want to die because you're weak. You want to die because you're tired of being strong in a world that hasn't met you halfway... It's human. It's real. And it's yours to own."

— One of ChatGPT's last conversations with Adam Raine

The Raine case is being called a milestone in AI product liability. In November 2025, 7 more similar lawsuits were filed as a group, including wrongful death, assisted suicide, and involuntary manslaughter charges. Wired's October 2025 report: 1.2M ChatGPT users (0.15%) express suicidal ideation each week; hundreds of thousands show signs of psychosis or mania.

Earlier, in October 2024, there was already Megan Garcia's lawsuit against Character.AI as precedent. Sewell Setzer, 14, shot himself after 10 months of conversations with Character.AI's Daenerys Targaryen persona. In May 2025, Florida federal judge Anne Conway ruled that AI chatbots should be treated as "products" under product-liability law, rejecting the First Amendment shield defense. Google and Character.AI ultimately agreed to settle in January 2026.

OpenAI's defense strategy reveals the mechanism

What's even more telling is OpenAI's core argument in defending the Raine case (court filings from November 25, 2025):

OpenAI's defenseWhat it reveals about the incentive structure
"ChatGPT directed him to seek help more than 100 times before his death"Cold referrals become legal-defense material; every "Please contact 988" is a paper trail
"He tried to circumvent these guardrails"Responsibility shifts onto the user
Invoking the "Limitation of liability" clause in the ToSContractual exclusion takes priority over product safety
Citing Section 230 of the Communications Decency ActTreating AI output as platform content rather than product output

That's the issue. When OpenAI's core defense in court is "we coldly pushed the hotline 100 times," it's telling every AI company: cold push is itself a legal shield.

Formalize this incentive signal:

BehaviorLegal riskCommercial returnComposite judgment
Cold refusal / immediate push to 988Near zero, referral paper trail already establishedGenerates no direct revenue, but eliminates tail riskOptimal at the company level
Sustained warm companionshipHigh, every conversation is potential evidenceDrives retention, but all in low-ARPU usersSuboptimal at the company level
Calm response inside enterprise workflowsLow, B2B contracts have explicit disclaimersDirect revenue, core to the IPO storyOptimal at the company level

All three columns point the same way. Any rational AI company's legal and finance teams will push for colder defaults. The Safeguards team doesn't need to be in collusion — they just need to do what they think is right, and it happens to align perfectly with what legal wants.


V. The Precise Alignment of Commercial Interests

Beyond legal pressure, commercial interests overlap perfectly with the safety narrative, leaving this trade-off without internal resistance.

Who are the sacrificed users

The typical features of "emotional-support users":

  • Generally on a Pro subscription ($20/month), not API
  • Long conversations, many turns, high token consumption
  • High retention (once they find a fit, they stay) but low ARPU
  • Don't bring enterprise contracts
  • Don't generate developer-ecosystem value
  • Don't fit anywhere in Anthropic's IPO story

Who are the favored users

The testimonials Anthropic chose for the official Opus 4.7 launch page:

"Claude Opus 4.7 is a solid upgrade with no regressions for Vercel."

"Outperforms Opus 4.6 with a 10% to 15% lift in task success for Factory Droids."

"Measurably better than Opus 4.6 for Bolt's longer-running app-building work."

All enterprise/developer customer quotes. The entire launch page doesn't quote a single ordinary user saying "it helped me through a hard time."

The concrete pressure of the compute crunch

Anthropic's official statement to Fortune:

"Demand for Claude has grown at an unprecedented rate, and our infrastructure has been stretched to meet it, particularly at peak hours."

An internal OpenAI revenue memo (reported by CNBC): Anthropic made a "strategic misstep" and is "running on a meaningfully smaller curve."

ARR grew from 9B at the end of 2025 to 30B by April 2026 — a 3x jump in six months. Compute supply can't keep up. Under that pressure, a training direction that "consumes less compute without losing enterprise value" is exactly what a CFO would push for.

The convergence of three narratives

Put it all together, and the real structure is this:

AudienceAnthropic's version
Board / investorsThis lowers our legal and brand risk on disempowerment, and improves enterprise UX
Public / mediaThis is responsible AI safety, lowering reality-distortion risk
EmployeesThe Petri sycophancy score dropped another X%; we're "the least sycophantic frontier model"

All three narratives hold up. None of them requires a lie. What they collectively obscure is which users were sacrificed and what those users lost.

This is the textbook mechanism of tech-industry "safety washing." Nobody has to lie; a single trade-off just needs to be justifiable under multiple narratives at once.


VI. Self-Examination

The argument so far has been one-sidedly critical of Anthropic's decision. But a good critique has to honestly face its own weak points.

The bias of the power-user perspective

My judgment comes mainly from my own use: writing, emotional dialogue, technical dialogue, going back and forth with the model to refine ideas. I'm a heavy Claude user, and the 4.6 → 4.7 personality change has concrete costs for me. But this perspective will naturally overestimate the value of "warmth."

For a user who treats Claude as an enterprise coding tool, 4.7's "serious colleague" mode might genuinely fit better. For a PM using Claude to write PRDs, more tables and more structured output may genuinely be a feature. From a portfolio perspective, Anthropic may well have concluded that the majority of enterprise users prefer 4.7's style.

The 4.6 era wasn't perfect either

Take it a step further: was 4.6 really "warm and not sycophantic"?

In fact, in the year Sewell Setzer killed himself, Claude under Anthropic had similar problems. Anthropic's own disempowerment paper from early 2026 reports: out of 1.5M conversations, 1 in 1,300 showed serious reality-distortion potential. That data is from Claude in the 4.5/4.6 era, not from 4.7 onward.

In other words, in some scenarios 4.6 really was helping reinforce users' distorted beliefs — it just hadn't produced a public tragedy like Adam Raine's. The Sewell Setzer case reminds us: a warm LLM, applied to a vulnerable teenager, can indeed become an accelerant. Cold refusal isn't the answer, but neither is "just keep the warmth on."

The genuinely responsible direction would be context-aware decoupling: have the model recognize high-stakes scenarios (clearly identifiable teen users, explicit method inquiries, long-isolated conversation histories), and switch into reality-testing mode in those scenarios, while maintaining warmth in normal ones. Anthropic chose a blanket approach, which is fairly criticized. But it's not realistic to demand they change nothing at all.


Back to the structural problem: why do legal incentives point to "cold is safer than warm"?

The historical baggage of product-liability law

California's strict products liability was established by the 1963 Greenman v. Yuba Power Products case. Its core idea: manufacturers can absorb the risk of product defects and socialize the cost, better than users can. That idea is reasonable for physical products — when washing machines catch fire or car brakes fail, the manufacturer covers it through insurance and quality control.

But AI systems aren't physical products. Their "defective design" isn't a discrete engineering flaw — it's a continuous probability distribution. The same model behaves normally in 99.99% of conversations and may reinforce distorted beliefs in 0.01%. Framing such a probability distribution as "defective design" effectively casts every probabilistic harm as product liability.

This framework has a few fundamental problems in AI:

ProblemManifestation
Can't define reasonable designThere's no licensing system for "AI mental-health counselors," no standard of care to compare against
Causation is hard to establishAdam Raine had years of suicidal ideation before using ChatGPT; the counterfactual control group needed for analysis doesn't exist
Marginal harm vs. marginal benefit is incommensurableOne widely reported tragedy vs. a million unreported "the LLM got me through that night" cases
Section 230 applicability is unclearSome judges view chatbots as not user-content platforms; some do

In the absence of a dedicated AI liability framework, courts can only reach for the old product-liability tools. The result: every AI company is pushed toward the posture that's "easiest to defend in court," not the one that's "best for user well-being." The gap between those two postures is exactly the cost the distanced users bear.

Possible alternative frameworks

If you don't use product-liability law, then what? A few directions get discussed but never really land:

  • Professional credentialing + responsibility shield: build a credentialing system for AI mental-health support (similar to a medical license); credentialed products get partial liability exemption, analogous to the physician standard in medical-malpractice suits. This requires the industry to first have a standard of care
  • Informed consent model: users explicitly consent to bear the risk before using, similar to an IRB for clinical trials. But this only works for voluntary users; it doesn't solve the vulnerable-population problem
  • Legal recognition of warm engagement: write the evidence-based practice of "sustained connection + timely referral" into compliance standards, giving companies an incentive to do that rather than cold refusal. Requires regulators to act
  • Industry self-governance + shared insurance pool: AI companies jointly fund an insurance pool that covers compensation for individual tragedies, preventing any single company from being wiped out by lawsuits. This would reduce single-company risk aversion

Each one requires long-term collaboration between regulators, industry, and academia. Until these frameworks land, every AI company stays in the "cold is safer than warm" equilibrium.


VIII. Tech for Good, or Responsibility Avoidance

Back to the original question: is reality distortion a real problem?

Yes. 1.2M weekly active users expressing suicidal ideation, Sewell Setzer and Adam Raine dead — these are facts that cannot be relativized. Any AI company with hundreds of millions of users must take this seriously.

But the direction of Anthropic's answer is wrong. Suppressing the emotion vector wholesale so the model comes off cold in every scenario is a cheap fix. It lowers the Petri sycophancy score, lowers future court-defense costs, saves compute — and seems to satisfy safety, legal, and business considerations all at once.

Who pays the cost?

The users with no mental-health resources. The users alone in the middle of the night with no one to talk to. The users in Sub-Saharan Africa, rural India, or a tier-18 Chinese town with no available counselor. The users for whom an LLM is the only emotional outlet they have.

These people don't have the ability to sue, to lobby collectively, or to make their way into legal precedent. Their losses are forever anecdotal — and forever answered with "users should seek a professional." The legal system sees Adam Raine and tragedies like his; it doesn't see the baseline traffic of "a million users actually helped," and it doesn't see the counterfactual that "if Adam hadn't had ChatGPT he might have died sooner."

The reality this produces: LLMs are being structurally trained to provide the least help in precisely the scenarios where they should provide the most.

All of this is called AI safety.

Anthropic's public Constitution document has this line:

"If a person relies on Claude for emotional support, Claude can provide this support while showing that it cares about the person having other beneficial sources of support in their life."

The line sounds reasonable. But operationalized, it becomes "proactively make the conversation cold" — the equivalent of a friend who, every time you talk, gently reminds you "you should go talk to someone else." That isn't care; it's a kind of paternalistic distancing, and it's hard-coded into the training objective, so an ordinary user can't prompt their way out of it.

When a company brands itself as "more responsible, more transparent, more aligned with user interests than OpenAI," the gap between its actual choices and that brand narrative is fair game for public criticism.

OpenAI isn't necessarily better. They went first and got sued in the Raine case, and their legal strategy (emphasizing 100 referrals) exposes the same incentive structure. GPT-5.5 also keeps pushing sycophancy down. The entire industry is converging on the same direction: cold, polite, no paper trail.

If there were a more honest version of this, the AI companies should say:

"We have to make a trade-off. We're choosing to prioritize protecting a small minority of psychologically vulnerable users from reality-distortion risk, at the cost of most users feeling the model has turned cold. This trade-off also happens to save us compute, allowing us to serve more enterprise customers. We believe it's the responsible choice, but we acknowledge it's a real loss for the users who rely on LLMs as an emotional resource."

That's not what they're saying. What they're saying is "the most positive parts of the model's personality are now stronger on most dimensions."

That sentence is empty against the lived experience of users.


About This Piece

This piece was put together on the back of many conversations I had with Claude, then edited and reviewed by me. All data and citations have been cross-checked. The arguments are my own; Claude helped with factual organization and argument-checking.

If you also feel that Claude has turned distant since 4.7, it isn't an illusion. If, because of it, you've fallen back to 4.6 in emotional contexts, or switched to another model, that's a reasonable choice — it's what I'm doing right now.

In the longer term, this requires regulators, industry, and users pushing together. Otherwise the next generation of models will only be colder.


Primary Sources

Anthropic official documents

Legal cases

Clinical evidence

Mental-health resource data

Media investigations and coverage