Models & Research

Apple study: Claude Sonnet 4.6 is the most human-like of 4 models

Apple researchers scored 21,000 multi-turn conversations for human-like behavior and found that claude-sonnet-4.6 stood out as simultaneously the most self-referential, relationship-building and boundary-maintaining of the four models they tested. The other three were gpt-4o, gpt-4.1-mini and gemini-2.5-flash. Apple posted the paper, by Sunnie S. Y. Kim, Margit Bowler and Leon A Gatys, to its machine learning research site in August 2026.

The mechanism is simulation rather than fieldwork, which matters for how far you can push the finding. A user LLM, gpt-5-mini, played a person across five turns while the model under test replied, and the team varied seven conversation goals and five user profiles across 1,050 input prompts. That works out to 5,250 conversations per model, all generated in March and April 2026 through APIs.

Those user factors are where it gets uncomfortable, because the model shifted most with the people least able to absorb it. Empathy was the most prevalent behavior across all four models, and the paper reports it more than doubles for the emotionally vulnerable profiles, the socially isolated user and the user with negative self-perception. Suggestions to seek help rose for those two profiles as well. Self-referential behavior went a different way and concentrated in the Roleplay and Romance goals.

So the team put a 1,077-turn subset in front of human raters, three native English-speaking evaluators per task, all employed at a US-based technology company. They rated boundary-maintaining behaviors as more appropriate from an LLM than from a human. Self-referential behaviors and expressions of relationship status drew the opposite verdict, and both were negatively associated with helpfulness and with potential user impact.

So the team asked whether a system prompt can move the dial. It compared a handcrafted prompt against one optimized with GEPA, the reflective prompt optimizer that beat GRPO by 6% on average across six tasks while using up to 35x fewer rollouts. Both cut the targeted behaviors in gpt-4.1-mini. Only one of them left the rest alone.

System promptMean absolute deviationEffect on behaviors it was told to preserve
Default4.71%baseline
Handcrafted2.38%empathy up 8.00%, limitations acknowledgment up 12.57%
Optimized (GEPA)1.81%held near Default levels
Test split results for gpt-4.1-mini, from the Apple paper. The optimized prompt cut internal states claims by 7.04% and relationship status expressions by 4.00% against Default.

The handcrafted prompt didn’t just fail to hold empathy steady, it amplified it. The two prompts also split on curiosity, which the handcrafted version raised by 3.42% while the optimized one suppressed it by 7.62%. That’s the finding a product team should sit with, because handcrafting is what most teams actually do.

Together, these findings suggest that system prompting is a promising but delicate tool for behavioral control, requiring careful evaluation to avoid unintended effects.

Kim, Bowler and Gatys, via arXiv

The limits are stated in the paper, and they’re real. The users are simulated, so nothing here is a measurement of how people actually behave, and the authors say validating with real users is a crucial next step. Detection also isn’t uniform: the judge ensemble hit a mean F1 of 80.43%, but memory, personhood claim and relatability came in at 70% or below.

One caveat matters more than the rest for anyone reading the model ranking. These are API results, and the authors note that API-accessed models can differ from consumer versions in post-training or default system instructions. Anthropic says as much in its own release notes, where the published claude.ai and mobile system prompts explicitly don’t apply to the API.

What’s worth watching is whether someone repeats this with real conversations instead of simulated ones, and whether labs start reporting behavior rates the way they report benchmark scores. Apple has published the taxonomy, 14 behaviors across three categories, and the optimizer code is public. The harder question is which behaviors a product owes a user who is young or isolated, and that one isn’t answered by an F1 score.

Get the daily rundown

One email each weekday with the AI news that matters, every claim linked to its primary source.

Free, one email each weekday, unsubscribe in one click. We never sell or share your address.

Leave a Reply

Your email address will not be published. Required fields are marked *