Anthropic released Sonnet 5.5 on September 28. My writing agents name Claude Code’s sonnet alias, so when the alias moved from Sonnet 5 to Sonnet 5.5 that day, they moved with it.
I ran both Sonnets through the same writer agent on two tasks from my September 24 article, “Why Now?”: a copyedit of lines 48 to 71, and a Substack Note of 120 to 150 words. Opus 5.5 ran the same tasks as a reference.
The Test and Its Limits
Each Sonnet ran at its Claude Code default. That’s Medium for Sonnet 5.5, per the model configuration docs, and High for Sonnet 5, per the Sonnet 5 overview. So this tests what the alias delivered, not equal effort. Opus ran at High, set in my config.
Sonnet 5 ran two days later through headless Claude Code, pinned by model ID, so its wall times include CLI startup. There was one run per cell. The judge, an Opus 5.5 agent, wasn’t blind.
Artificial Analysis measured the same pairing independently. Sonnet 5.5 at Medium scored 41 on its Intelligence Index at $0.59 a task, against 32 at $1.79 for Sonnet 5 at High.
The Numbers
TaskModelEffortTokensTool callsWall timeCopyeditSonnet 5Highabout 109,3007139 s*CopyeditSonnet 5.5Medium119,686638 sCopyeditOpus 5.5High145,97815119 sNoteSonnet 5Highabout 109,900791 s*NoteSonnet 5.5Medium98,651831 sNoteOpus 5.5High145,0811092 s
* Includes headless CLI startup.
Sonnet 5.5’s token counts came within about 10% of Sonnet 5’s. Anthropic says testers saw Sonnet 5.5 batch tool calls more than Sonnet 5. These runs don’t show that.
The Copyedit
Sonnet 5.5 made two required fixes to Sonnet 5’s one. Both removed the trailing space at line 58. Only Sonnet 5.5 turned the semicolon at line 70 into a period, per my voice rules’ budget of “essentially never.” Sonnet 5 left it and reported that the section “already matches Bob’s HyperDev voice.”
The NSA sentence at line 54 split them. “Including the reduction in their research burden” treats a reduction as a kind of distillation. Sonnet 5 rewrote it without asking, to “noting it reduces their research burden,” which presents the reduction as something the NSA stated. Sonnet 5.5 flagged it as a query. My editorial-craft.md says “when in doubt, query rather than silently resolve,” and this sentence paraphrases a cited source.
Neither Sonnet split the comma-and at line 62, which rule 9 in my learned-from-edits-2026-09.md says to split with a period. Opus did.
The Note
Both Sonnet Notes ran 135 words with nine claims, and neither invents a fact. Both drop the qualifier on the index figure, so “three points” has no unit.
The judge counted about seven rule issues in Sonnet 5’s Note against three in Sonnet 5.5’s. Sonnet 5.5 kept the article’s “my hypothesis” framing and wrote “Anthropic alleges.” Sonnet 5 stated the hypothesis as fact (”there’s a second goal”), wrote “Anthropic says,” and added an em dash and three contractions the article doesn’t use.
Sonnet 5 had the better-prepared ending. “Someone has to pay for it” follows from the $800 billion compute figure before it, while nothing in Sonnet 5.5’s Note sets up its closing “pricing pressure.” Still, Sonnet 5’s line is a closing button, which rule 5 in the same file advises against.
Opus’s 150-word Note kept the index qualifier and stated the article’s thesis, which neither Sonnet did. It is closer to publishable than either.
Neither Sonnet ran the linter, and Opus ran it both times. Sonnet 5 didn’t load writing-bob-voice, the skill that holds the numbered rules above, on either task. Sonnet 5.5 skipped it on the Note.
Where Sonnet 5.5 Beat Opus 5.5
Speed and tokens. Sonnet 5.5 took 38 seconds to Opus’s 119 on the copyedit, and 31 to 92 on the Note. It used 18% and 32% fewer tokens, and 6 tool calls to 15 on the copyedit. It ran at Medium and Opus at High, so some of that gap may come from the effort setting. I didn’t measure dollar cost against Opus, so the token gap isn’t a cost figure.
The sourced sentence. Opus rewrote the NSA sentence silently, as Sonnet 5 did. Sonnet 5.5 made the query my rules ask for.
Nothing to undo. Opus made one edit I’d reverse: “attributed to DeepSeek, Moonshot, and MiniMax campaigns involving...” reads “MiniMax campaigns” as a compound noun. None of Sonnet 5.5’s three changes needed reversing.
Terminal work, on benchmarks. Anthropic reports Sonnet 5.5 at 70.6% on Terminal-Bench 4.0 against 66.4% for Opus 5.5 at Xhigh. Artificial Analysis measured 64 against 60, on a pre-release build. My runs didn’t test this.
Opus stayed ahead on the rule 9 fix, the linter runs and the Note. On the same benchmark tables it leads on FrontierCode 1.1, 54.4% to 46.2% for Sonnet 5.5 at Max, and on Artificial Analysis’s AA-Omniscience factual-accuracy test, 66 to 54. GDPval-AA is close to even, 1846 to 1844.
Recommendation
Just use it. If your agents name the sonnet alias, they already do.
Bob Matsuoka is CTO of Duetto, a hospitality profit and revenue-management platform, and writes about AI-augmented engineering practice. Previously, Bob has been CTO of Tripadvisor, Citymaps, and Runtime Technologies.
Related Reading:
What Is Harness Engineering? (And Do You Need to Learn It?). The distinction between model assistance and the surrounding systems that control a workflow.
AI Power Ranking. Tool comparisons and benchmarks for AI practitioners.
LinkedIn Newsletter. Strategic AI insights for CTOs and engineering leaders.



