On August 10, after a phrase had shown up for the ninth time, I typed this into a Claude Code session running Opus 5:
“the load-bearing section” this is too jargon-y. Do not use language like this (add it to tm output style)
Forty-nine minutes later, the rule was live in trusty-tools:
No borrowed-metaphor jargon. “Load-bearing” is the instance that prompted this rule... Wrong: “that section is load-bearing” / Right: “deleting that section breaks X”.
I had seen it nine times in two days. Opus 5 used “load-bearing” for a constraint, then a crate, then a test suite, all in six sessions on August 9 and 10. I got tired of reading it, so I banned the word.
I remembered a much worse week. In my version, I fought with Opus 5 for days, switched to Fable 5, and everything got better. Then I checked my logs.
Across eight days, I typed 201 messages to the two models. The record is boring: no capability problems, no ALL-CAPS, no profanity, no “that’s not what I said.” I was calm and typo-heavy throughout (agents understand intent well enough that my typing has gotten worse). The switch to Fable happened in the middle of a session without a command or comment from me.
The window’s one documented correction appears secondhand, inside an auto-generated summary. It concerned version-bump authorization. My complaint about Opus 5’s prose came two days earlier. I remembered a fight, but the logs recorded something else.
Meanwhile, Opus 5 produced 1.37 million words of output on real engineering tasks, next to Fable 5’s 193,000 words. The capability argument was gone. I was left with the way it talked.
TL;DR
My trusty-tools logs show zero capability problems with Opus 5 over the window that supposedly caused a switch to Fable 5. My original thesis was wrong.
Across the month-long sample, Opus 5 opens sentences with “worth noting/naming/knowing” 10.7x more often than Fable 5 (0.451 vs. 0.042 per 1,000 words).
A recurring metaphor (something “wearing” something else’s clothing) shows up 26 times across 12+ Opus 5 sessions and zero times anywhere else in the sample.
Eight named developers across Hacker News and X say they prefer GPT-5.6 Sol specifically for how it communicates. None turned up making that same claim on Reddit.
The voice predates Opus 5. A GitHub issue filed 11 days before it shipped describes the same tics, and Claude Code gives Opus 5 a system prompt a bit over a third the length of Sonnet 5’s.
How Opus 5 Talks
The trusty-tools bans landed between August 10 and 12. From August 12 onward, the same rule applied to every model: no “worth noting,” no “load-bearing,” no borrowed-metaphor jargon, and none of the “honest/honestly” family. The month-long counts also include earlier output, so this is not a clean post-ban comparison. It is a record of how often the patterns appeared in the logs I have.
The sample covers July 28 through August 27 in the same repository: 9,148 Opus 5 turns (1.37M words) and 2,732 Fable 5 turns (193K words). Opus 5 averages 149.5 words a turn on equivalent status-update and engineering tasks, more than double Fable 5’s 70.5.
On a routine GitHub API check, Opus 5 wrote:
Worth noting how it verified: its first PATCH used
-f description=@file, which isn’t a realgh apishorthand...
Later, in a SlotRegistry design note:
It’s load-bearing for assertion B, it’s about to be cited in an ADR, and right now nobody but you can retrieve it.
The “wearing” metaphor got under my skin. Opus 5 used some version of it 26 times across at least 12 sessions that month. A conflict resolution became “unreviewed code wearing a rebase’s clothing.” A misattributed fact came back as “a hypothesis wearing a rule’s clothing.” A faster teardown was “a regression dressed as a performance win.” Fable 5 produced zero in 193,000 words. The thinner Opus 4.8 and Sonnet 5 samples produced zero too.
Fable 5, working the same repository, the same kind of task, under the same ban:
Error arms proven load-bearing (swallowing provider errors turns four tests red)...
You’re right, and the honest accounting is that the day went to the plumbing that kept invalidating runs...
The pinned gate run is genuinely green.
Fable used the same forbidden words, but less often. The “worth noting” family appears roughly a tenth as often.
“Honest” is the exception. Opus 5 sits at 0.193 per 1,000 words against Fable 5’s 0.176, close enough to call a wash. The other gaps run from roughly 2x to 10.7x.
Developers Recognize the Voice
On Hacker News alone, I found 29 items from April through August 2026 about Opus 5’s communication style. Roughly 55 people joined the complaint. The highest-engagement thread, “Why does Opus 5 feel worse to work with?”, sits at 993 points and 873 comments. Two Ask HN threads pose nearly the same question: “Which is the least sloppy and claudeism free model you have used?” (July 25) and “Why are Claude models so verbose?” (August 26).
On X, 13 tweets with verifiable URLs carry the same complaint from 11 distinct authors, dated July 24 through August 19.
By comparison, eight named people across both platforms say they prefer GPT-5.6 Sol specifically for how it communicates: six on Hacker News (including matheusmoreira: “Opus has a distinctive sentence structure and Fable somehow manages to be even more obtuse”), plus two on X: Emmett Shear (”Talking to Opus makes me angry and depressed in a way that’s hard to articulate. It’s actually somehow worse than Sol”) and yacineMTB (”Meh. Switched back to sol. I am trying to get shit done not be hypnotized by a robot”).
The 816-point thread that gets cited most often in this argument, “Anthropic’s best AI model struggles to attract users as cheaper tools thrive,” is a price story. One top comment calls the output “peak verbosity vomit,” but the thread itself is about pricing pressure from cheaper models, not a referendum on tone.
These quotes come from Reddit aggregator blogs that include thread names, usernames, and vote counts. One step removed. The style complaint still shows up with its own vocabulary: “essay of slop,” “Claudeslop,” “benchslop,” and “load-bearing” appear independently across multiple write-ups as stock Opus 5 language. But I could not find a single Reddit thread where someone says they prefer Sol because of how it talks. The Sol arguments there are about benchmarks and price. Two threads praise Fable 5’s voice over Opus 4.8, which says nothing about Fable 5 versus Opus 5.
Counted across the Hacker News threads read for this piece: “load-bearing” appears 83 times, “verbose” 69, “honest/honestly” 66, “jargon” 34, “seam(s)” 36.
The Voice Predates the Model
GitHub issue #77136 was filed July 13, 2026, eleven days before Opus 5 shipped. It documents “load bearing,” invented jargon, “hand-waving,” “reflexive hedging,” and “honest framing” in Opus 4.7, Opus 4.8, and Fable 5. The complaint is older than the model now taking the blame.
Claude Code also gives Opus 5 a system prompt of roughly 11,000 characters with almost no anti-verbosity instructions. Sonnet 5 gets about 29,000 characters, much of the difference made up of rules that suppress this behavior. In the published test, restoring the longer prompt changed the behavior. The model may supply the prose, but the harness decides how much restraint to ask for.
And somebody made the opposite switch. Manu Parasuraman moved from Sol to Opus 5 because Sol was harder to read: “Sol answers questions nobody asked in language nobody speaks,” while Opus 5 “walks you through its reasoning and stays concrete.”
I started this piece assuming Opus 5’s personality was the whole problem. The system-prompt evidence complicates that: I may be hearing Claude Code’s prompt as much as the model underneath it, then blaming the name in the model picker.
What This Comparison Can and Can’t Show
The 10.7x gap is descriptive. It combines output from before and after the bans, and it does not tell me what either model would say on an empty prompt. Sonnet 5 contributed 13 turns from one session, too thin to rate. GPT-5.6 Sol never appears in the trusty-tools logs, which means every Sol comparison here comes from other people rather than my own measurement. The Reddit material is secondhand because the primary site was unreachable during research.
Ten Days Later, Anthropic Shipped a Concise Style
Anthropic reportedly shipped a built-in Concise output style for Claude Code on August 20, according to a third-party writeup. That was ten days after my rule landed and six days after the 993-point Hacker News thread. I haven’t checked whether it helps. I prefer my own.
The logs explain why I wrote the rule. Opus 5 completed the engineering work, at scale, while repeatedly talking past an instruction written for it in plain English. I don’t call that a capability failure, but it is a product problem. The rule stays in trusty-tools.
Bob Matsuoka is CTO of Duetto, a hospitality revenue-management platform, and writes about AI-augmented engineering practice.
Related reading:
When Sessions Talk — What 375 messages between my own Claude Code sessions said to each other
AI Power Ranking — Tool comparisons and benchmarks for AI practitioners
LinkedIn Newsletter — Strategic AI insights for CTOs and engineering leaders




