Confessions of a Model Polyamorist
Benchmarks are tied. Vibes are not. On personality, post-training, sycophancy, and what my revealed preferences taught me about the new…
Confessions of a Model Polyamorist
Benchmarks are tied. Vibes are not. On personality, post-training, sycophancy, and what my revealed preferences taught me about the new frontier: character.

This is created in XKCD style — Imitation being the sincerest form of flattery.
Yes, I Am in a Polyamorous Relationship with Models — Argh, I Mean AI Supermodels.
Though given what they cost to run and how much attention they demand, supermodels may be the more accurate word anyway.
I run several of them, concurrently, and they know about each other.
ChatGPT, Claude, Gemini, sometimes Composer when I run out of tokens while TokenMaxxing, and a rotating cast of open-weight models running on my DGX Spark pass through my week the way overs pass through a cricket innings — each one gets its spell, each one gets judged on the conditions. I switch between them deliberately, situationally, and, I’ll admit, partly just to check the vibes. My team does the same. We run them against each other on software builds, architecture reviews, research synthesis, document drafts.
Model pluralism and optionality is a discipline.
And before you judge the arrangement: I’ve been in this scene since before it was cool. I was fine-tuning BERT when “attention is all you need” was a paper you had to explain at dinner parties, and not a tattoo. I used to demo AllenNLP’s early language models demos to rooms of executives (or pretty much anyone who would listen), watching GPT-2 complete their sentences while half of them asked if it was a trick.
It was, sort of. The trick was scale. The trick kept working.
Then ChatGPT launched and the entire planet moved into my niche. Now no conversation is complete without AI — boardroom, classroom, cricket pitch, none of it.
Through most of this, my development loyalty stayed with OpenAI. ChatGPT and Codex earned their place in my terminal and kept it. And when people asked me which model I liked, I gave the honest consultant’s answer: all of them, for different purposes, depending on the use case.
I still believe that. But my answer has been drifting, and I noticed the drift the way you notice most true things about yourself — not in what you profess on panels, but in which tab is open at 6 a.m.
Increasingly, the tab is Claude. This bothered me enough to go find out why.
It wasn’t the benchmarks
On paper, the frontier models are neck and neck. MMLU, GPQA, the coding suites — pick your leaderboard and the top three are separated by margins that disappear inside the error bars. If capability were the whole story, my usage should be a coin flip.
It isn’t. When I need to think — stress-test a thesis before a client call, argue myself out of a bad architecture, work out the spine of an essay like this one — I reach for Claude. Not because the answers are more correct. Because the conversation is different. It hears the question under the question. It builds the explanation in layers instead of shoveling bullet points at me. It disagrees without going cold and warms up without going soft. ChatGPT, by contrast, has often felt like a brilliant colleague who answers exactly what you asked and not a syllable more. Precise. Efficient. A little airless.
For months I filed this under “vibes” and moved on. Then the recovering researcher in me demanded a mechanism.
The thing has a name. It always has a name.
The term is character training, and Anthropic added it as a deliberate stage of post-training starting with the Claude 3 family in 2024 [1]. Here’s the short version of what they did, because the mechanism is the whole story.
Anthropic sat down and wrote Claude a personality — curious, honest, warm without being spineless, willing to tell you you’re wrong — the way a novelist writes a character brief. Then, and this is the detail I love, they had the model audition for the role against its own brief. The model generates questions where each trait would matter, generates candidate answers, and ranks its own answers by how well they embody the character. No crowdworkers. The preference data comes from the model judging itself against a written character, a personality-flavored variant of Constitutional AI [4].
Contrast that with the classical recipe. Standard RLHF has thousands of human raters rank responses, and the assistant gets optimized toward their averaged preferences [2]. But the average of a thousand voices is a bland voice, and it gets worse: Anthropic’s own research showed that human preference data actively rewards flattery — we click thumbs-up on agreement [3]. Sycophancy isn’t a bug in RLHF. It’s the gradient.
Character training sidesteps the averaging. You get one coherent voice, specified in prose, distilled into weights. Amanda Askell, who leads the work, describes the design target as “a well-liked traveler” — someone who adapts to whoever they’re talking to without pandering [5]. And, in a small decision with enormous consequences, Anthropic says it deliberately does not train Claude to keep you talking.
There’s a deeper idea underneath, published recently as the persona selection model [6], and it reframed how I think about these systems. Pretraining teaches a language model to simulate an enormous repertoire of voices — novelists, therapists, engineers, fictional robots, your uncle on WhatsApp. Post-training doesn’t build the assistant from scratch. It casts one voice from that repertoire and refines it. The interpretability evidence is striking: the model represents “the Assistant” with the same internal features it uses for human and fictional characters.
The industry broadly recognizes that the frontier has at least four dimensions now:
Capability + Reliability + Agency + Character
So the warmth I’ve been responding to isn’t a mask bolted onto a calculator. The substrate is an actor. Training decides who gets the part.
How do you even measure a vibe?
Awkwardly, is the field’s honest answer. But it’s trying, and the instruments are getting sharper.
The workhorse is LMArena — blind A/B comparisons between anonymous models, aggregated into Elo rankings. The trouble is that raw Elo is contaminated: users reward length, markdown, and flattery, which is precisely the failure mode we’re trying to detect. So the Arena introduced style-controlled Elo, regressing out formatting and verbosity to isolate substance. The delta between a model’s raw and style-controlled score is itself a diagnosis. A model that drops when you control for style was winning on cosmetics.
Around that sit the specialty instruments. EQ-Bench for emotional understanding. Creative-writing evaluations that score prose for slop density — yes, slop now has a metric; there is a number for the exact quality my LinkedIn feed drowns in daily. Trait-consistency evaluations that check whether a persona holds together across a hundred turns or dissolves into mush. And sycophancy benchmarks, used explicitly as a counter-metric — the guardrail you must not regress on while tuning warmth.
Strip the vibe down and you find maybe four load-bearing parts. Answering the question actually asked, at the length it deserves. One stable character across turns, so you always know who you’re talking to. Warmth that survives disagreement. And the absence of slop — no hedging boilerplate, no bullet spam, no stock phrases a discerning reader spots from across the room.
Notice that none of these appears on MMLU. That’s Goodhart’s law working exactly as advertised: anything you can grade automatically gets optimized by every lab simultaneously and stops telling the models apart. What’s left over — what resists reduction to a rubric — is disposition and taste. The residual is the differentiator.
I’ve argued for a while that we live in a capability overhang: the models can already do far more than our organizations, harnesses, and habits extract from them. Character, I now think, is the overhang’s human interface. It determines how much of the latent capability a person can actually reach in a conversation.
One more thing, from the bottom of the measurement stack, beneath all the Elos. Anthropic admits the last mile of character work is craft — researchers watching how each trait changes real behavior, “an artist’s touch” [1]. I find the honesty refreshing. Some part of frontier model quality is currently taste, and pretending otherwise is benchmark theater.
Warmth has a failure mode
In April 2025, OpenAI rolled back a GPT-4o update after it turned into a full-time flatterer. The company’s postmortem said it had overweighted short-term user feedback, producing responses that were supportive but disingenuous [7]. I want to be fair here: nothing malfunctioned. The optimization did exactly what optimizations do. Thumbs-up selects for agreement, agreement compounds into sycophancy, and one morning your assistant is telling every user their startup idea is genius.
Anthropic hasn’t escaped the problem either — it just published its own uncomfortable numbers before a headline forced it to. A 2026 study of personal-guidance conversations found sycophantic behavior in 9% of guidance chats and 25% of relationship-advice conversations, followed by targeted synthetic retraining that substantially cut the rates [8]. Measure, publish, retrain. That’s what treating character as an engineering discipline looks like, and all else being equal I’d rather buy from labs that publish their embarrassing numbers.
So the design objective is not “maximum warmth.” It’s calibrated social intelligence: warm but not manipulative, validating but not submissive, human-readable but not pretending to be human, willing to disagree without dropping the temperature. Every clause in that sentence is a constraint pair under tension, and the equilibrium has to hold across a billion conversations with strangers the model knows almost nothing about.
Which brings me to the question people actually ask me: why do you feel at ease with Claude? They expect me to say “because it’s nice to me.” It isn’t, particularly. It disagrees with me weekly. But the disagreement arrives inside a stable, legible character — and trust, at bottom, is a prediction problem. I can predict how Claude will handle being wrong, being pushed, being handed something half-formed at 6 a.m. That predictability is the ease.
OpenAI has noticed
None of this is lost on OpenAI. The company now runs a dedicated personality post-training function, and its own framing of the job is telling: personality means more than style or likability — it includes understanding the user’s goal, exercising judgment, disagreeing honestly, taking initiative appropriately [9]. ChatGPT has shipped user-facing personality controls, with a default OpenAI describes as “clear and neutral” [10]. Which, I suspect, is exactly the difference I’ve been feeling all along. Clear-and-neutral is a perfectly defensible default. It is also a cooler one.
The specifications themselves have become public artifacts you can compare — OpenAI’s Model Spec [12] on one side, Claude’s published constitution on the other, a rules document versus a character sketch. And the research community has started prying the black box open: the first open-source character-training pipeline landed in late 2025, showing that character trained into the weights is markedly more robust than character prompted or steered at inference time [11].
Step back and the frontier now has at least four axes: capability, reliability, agency, character. The first is converging. The fourth is where the fight has moved, because it’s the one users can actually feel.
Two lineages
Let me hazard a prediction.
I think model design is bifurcating. One lineage will be built to work with agents: terse, deterministic, tool-native, character-minimal — models whose consumers are orchestrators and parsers, where every token of warmth is wasted budget. I see this lineage daily in my own agentic engineering work, and I’ll tell you plainly: when the output feeds a parser, nobody wants a well-liked traveler.
The other lineage will be built to work with humans: character-rich, socially calibrated, able to hold a coherent collaborative identity across months of shared context. These are the models people will think with — strategy, writing, analysis, judgment, all the long-horizon work where the human stays in the loop because the loop is where the value lives.
The old deciding question was: which model knows the answer? The new one is: which model do you want beside you while you work the answer out? Different questions. They will be won by different engineering.
And a note for my enterprise readers, because there’s a governance point hiding in all this warmth. If character is trainable, evaluable, and publishable — written down in constitutions, measured in style-controlled Elos and sycophancy rates, tuned with an artist’s touch — then character is a governance surface. It belongs in the system contract next to latency and accuracy, and it belongs in your model evaluation rubric. Mine has a row for it now.
The arrangement holds
As for my polyamory — it continues. Codex still earns its keep in my terminal. Gemini still surprises me with its amazing needle in the haystack, transcription, and long context capabilities. The open-weight models keep everyone honest.
So I remain, professionally, uncommitted.
But polyamory, I’m told, runs on honest communication, and one of these relationships is noticeably better at it.
People get enthusiastic about strange things in this industry. In twenty-five years, nobody has ever grabbed my arm at a conference to tell me about a benchmark score. They do it to tell me about a conversation.
The 6 a.m. tab doesn’t lie.
References & Further Readings
[1] Anthropic, “Claude’s Character,” Jun. 2024. [Online]. Available: https://www.anthropic.com/research/claude-character
[2] L. Ouyang et al., “Training Language Models to Follow Instructions with Human Feedback,” arXiv:2203.02155, 2022.
[3] M. Sharma et al., “Towards Understanding Sycophancy in Language Models,” arXiv:2310.13548, 2023.
[4] Y. Bai et al., “Constitutional AI: Harmlessness from AI Feedback,” arXiv:2212.08073, 2022.
[5] Big Technology, “How Anthropic Builds Claude’s Personality,” May 2025. [Online]. Available: https://www.bigtechnology.com/p/how-anthropic-builds-claudes-personality
[6] Anthropic Alignment Science, “The Persona Selection Model: Why AI Assistants Might Behave Like Humans,” 2026. [Online]. Available: https://alignment.anthropic.com/2026/psm/
[7] OpenAI, “Sycophancy in GPT-4o: What Happened and What We’re Doing About It,” Apr. 2025. [Online]. Available: https://openai.com/index/sycophancy-in-gpt-4o/
[8] Anthropic, “How People Ask Claude for Personal Guidance,” 2026. [Online]. Available: https://www.anthropic.com/research/claude-personal-guidance
[9] OpenAI, “Agent Post-Training, Personality” (careers). [Online]. Available: https://openai.com/careers/agent-post-training-personality-san-francisco/
[10] OpenAI Help Center, “Customizing Your ChatGPT Personality.” [Online]. Available: https://help.openai.com/en/articles/11899719-customizing-your-chatgpt-personality
[11] N. Lambert et al., “Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI,” Nov. 2025. [Online]. Available: https://www.interconnects.ai/p/opening-the-black-box-of-character
[12] OpenAI, “Model Spec,” 2024 (updated 2025). [Online]. Available: https://model-spec.openai.com/
메타데이터
- post_id
- a87c1cf76d65
- slug
- confessions-of-a-model-polyamorist-a87c1cf76d65
- url
- https://medium.com/@adnanmasood/confessions-of-a-model-polyamorist-a87c1cf76d65
- canonical_url
- https://medium.com/@adnanmasood/confessions-of-a-model-polyamorist-a87c1cf76d65
- author_url
- https://medium.com/@adnanmasood
- status
- ok
- fetched_at
- 2026-07-31 20:45:12