<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Hyperdev]]></title><description><![CDATA[HyperDev is a technical publication exploring practical agentic AI development and AI-powered coding tools. As a veteran technology executive with 25+ years of experience, I provide honest, hands-on reviews and strategic insights about which AI coding too]]></description><link>https://hyperdev.matsuoka.com</link><image><url>https://substackcdn.com/image/fetch/$s_!j9a7!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab665959-5546-4469-9e93-9e1518976e2b_1024x1024.png</url><title>Hyperdev</title><link>https://hyperdev.matsuoka.com</link></image><generator>Substack</generator><lastBuildDate>Mon, 03 Aug 2026 12:30:59 GMT</lastBuildDate><atom:link href="https://hyperdev.matsuoka.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Robert Matsuoka]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[hyperdev@matsuoka.com]]></webMaster><itunes:owner><itunes:email><![CDATA[hyperdev@matsuoka.com]]></itunes:email><itunes:name><![CDATA[Robert Matsuoka]]></itunes:name></itunes:owner><itunes:author><![CDATA[Robert Matsuoka]]></itunes:author><googleplay:owner><![CDATA[hyperdev@matsuoka.com]]></googleplay:owner><googleplay:email><![CDATA[hyperdev@matsuoka.com]]></googleplay:email><googleplay:author><![CDATA[Robert Matsuoka]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[OpenAI's "Comeback"]]></title><description><![CDATA[Can it really be a comeback with they have so much money, so much exposure, and so little revenue?]]></description><link>https://hyperdev.matsuoka.com/p/openais-comeback</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/openais-comeback</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Fri, 31 Jul 2026 11:30:03 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!hQ4A!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!hQ4A!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!hQ4A!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png 424w, https://substackcdn.com/image/fetch/$s_!hQ4A!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png 848w, https://substackcdn.com/image/fetch/$s_!hQ4A!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png 1272w, https://substackcdn.com/image/fetch/$s_!hQ4A!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!hQ4A!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png" width="1456" height="874" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:874,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1378063,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/209163913?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!hQ4A!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png 424w, https://substackcdn.com/image/fetch/$s_!hQ4A!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png 848w, https://substackcdn.com/image/fetch/$s_!hQ4A!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png 1272w, https://substackcdn.com/image/fetch/$s_!hQ4A!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>GPT-5.6 Sol scores 80 on the <a href="https://artificialanalysis.ai/articles/gpt-5-6-has-landed">Artificial Analysis Coding Agent Index</a> against 77 for Claude Fable 5, at a lower cost per task (about $1.04). That post went up on July 9.</p><p>I called model parity back in April, in <a href="https://hyperdev.matsuoka.com/p/its-the-harness-stupid">a piece</a> with a section headed &#8220;GPT-5.4 Caught Up.&#8221; What I missed was everything around the model: that OpenAI would close the product gap, and buy a decade of compute while doing it.</p><p>Over the last few months I&#8217;ve been using ChatGPT more and more for work-like tasks. Not Duetto work, because we don&#8217;t have an enterprise plan and Duetto material stays out of it. Work-like things. Both the capabilities and the UX have improved to the point where I&#8217;d put the product at least on par with claude.ai, to my surprise. There are still gaps. It&#8217;s still good.</p><h2>TL;DR</h2><ul><li><p>GPT-5.6 Sol leads the Artificial Analysis Coding Agent Index 80 to 77 over Claude Fable 5 (July 9) and Terminal-Bench 2.1 85.77% to 84.64% over Claude Opus 5 (July 22), the second one with a methodology caveat that widens the gap.</p></li><li><p>Claude Opus 5 still leads the broader Artificial Analysis Intelligence Index at 60.7 against GPT-5.6 Sol&#8217;s 58.9 (July 29). Anthropic&#8217;s remaining model lead is general capability, not agentic coding.</p></li><li><p>Anthropic leads on revenue ($47B run-rate, May 28, 2026) and valuation ($965B), and leads enterprise API share 40% to 27% (Menlo, December 2025). OpenAI closed the capability gap and most of the product gap while trailing commercially.</p></li><li><p>OpenAI is far ahead on planned compute (Stargate at nearly 7GW toward a stated 10GW commitment, the Broadcom &#8220;Jalape&#241;o&#8221; inference ASIC). In July it moved its usage limits three times in seventeen days, in both directions, around a near-global outage on the 25th. Anthropic has been raising its limits since May.</p></li><li><p>No open-weight provider I found fields an app ecosystem within reach of ChatGPT or Claude Code, which keeps this a two-horse race even as the weights converge. Chinese domestic app numbers could complicate that.</p></li></ul><h2>Where the coding lead sits</h2><p>Two current measures put GPT ahead on agentic coding. The Coding Agent Index lists GPT-5.6&#8217;s three effort tiers at Sol 80, Terra 77.4 and Luna 74.6, with Claude Fable 5 at 77.2 (Terra and Fable 5 land 0.2 apart). Sol&#8217;s per-task cost runs below both Fable 5 and Opus 4.8. On <a href="https://www.vals.ai/benchmarks/terminal-bench-2-1">Terminal-Bench 2.1</a> (89 tasks, Terminus 2 harness, snapshot dated July 22), GPT-5.6 Sol leads Claude Opus 5 by about a point, 85.77% to 84.64%.</p><p>The Terminal-Bench result comes with a caveat from the publisher. Opus 5&#8217;s run used Claude Opus 4.8 as a refusal fallback. Count the nine affected passing results as failures instead and Opus 5 drops to 81.27%, which widens the GPT lead to roughly four and a half points. The lead is real. Its size depends on how you count.</p><p>Now the other direction. On the broader Artificial Analysis Intelligence Index, <a href="https://benchlm.ai/benchmarks/artificialAnalysis">as of July 29</a>, Claude Opus 5 sits first at 60.7 and Claude Fable 5 second at 59.9, with GPT-5.6 Sol third at 58.9. Anthropic holds the top two spots on general capability while losing the coding-agent tables. My &#8220;diminishing edge&#8221; framing was directionally right if imprecise. The edge moved. It didn&#8217;t disappear.</p><p>SWE-bench Pro, the number everybody reached for six months ago, has no usable current standing. Scale&#8217;s <a href="https://labs.scale.com/leaderboard/swe_bench_pro_public">public-set leaderboard</a> doesn&#8217;t list GPT-5.6, Claude Opus 5 or Claude Fable 5 at all, and the top of it reshuffled again this month. The aggregator sites quoting newer figures disagree with each other and cite at least one model name I can&#8217;t confirm exists. Skip it.</p><h2>The product story is a two-way race now</h2><p>This is the first half of what I missed. Codex spans CLI, IDE, web, and a hosted cloud mode for long-running tasks. The <a href="https://learn.chatgpt.com/docs/changelog?type=codex-cli">Codex CLI changelog</a> for the last two weeks of July includes resumable sessions with paginated thread history, session naming and branching, configurable sub-agents, audio input, and an <code>/import</code> command that migrates project-scoped memories out of Cursor and Claude Code. That last one reads like a decision made by somebody thinking hard about switching costs. On <a href="https://9to5mac.com/2026/07/09/openai-announcing-the-next-chapter-for-chatgpt-today-watch-here/">July 9</a> OpenAI also folded the standalone Codex desktop app into a single ChatGPT desktop app with Chat, Work, and Codex modes. Existing Codex users were moved across automatically, and Codex remains a full mode inside the merged app.</p><p>What I notice in use isn&#8217;t a benchmark, it&#8217;s flow. Codex and GPT live in the same product, and when I go looking for something I worked on weeks ago, ChatGPT finds it wherever I left it. Claude&#8217;s Projects still behave like sealed folders for me. That isolation was an advantage once, back when context was scarce and a hard wall around what the model could see was the safer design. With current context budgets it mostly means I go hunting. GPT is also the far better image renderer, which makes for a short comparison, since Claude doesn&#8217;t ship a first-party image model at all.</p><p>Anthropic didn&#8217;t stand still while any of this happened. <a href="https://www.anthropic.com/news/claude-design-anthropic-labs">Claude Design</a> shipped out of Anthropic Labs on April 17, and it&#8217;s excellent: you talk to it and get slides, wireframes, one-pagers and pitch decks, with export to PPTX or PDF, inline comments, adjustment sliders, and handoff into Claude Code. Codex has no equivalent that I&#8217;m aware of (and I&#8217;d bet real money they&#8217;re building one). Claude Cowork expanded to web and mobile on <a href="https://techcrunch.com/2026/07/07/the-coding-agent-wars-are-spilling-into-the-rest-of-the-office-claude-cowork/">July 7</a>. Both labs shipped serious B2B product work inside the same four months.</p><h2>Capacity: a much bigger future, a strained present</h2><p>The second half is compute. OpenAI is buying a much bigger future than Anthropic is, and is visibly straining in the present.</p><p>On committed buildout it isn&#8217;t close. OpenAI, Oracle and SoftBank put combined planned Stargate capacity at <a href="https://openai.com/index/five-new-stargate-sites/">nearly 7 gigawatts</a> across the Abilene flagship plus five more sites, with over $400B invested and the full commitment still stated as $500B and 10GW. Tomasz Tunguz&#8217;s aggregation of the announced vendor deals puts OpenAI&#8217;s 2025&#8211;2035 infrastructure commitment at <a href="https://tomtunguz.com/openai-hardware-spending-2025-2035/">roughly $1.15 trillion</a> across seven suppliers, which is derived rather than OpenAI-confirmed. On June 24, OpenAI and Broadcom unveiled <a href="https://openai.com/index/openai-broadcom-jalapeno-inference-chip/">Jalape&#241;o</a>, OpenAI&#8217;s first custom inference ASIC, designed to tape-out in about nine months, with engineering samples already running Codex workloads in the lab and first production deployment targeted at gigawatt scale in late 2026 (<a href="https://www.cnbc.com/2026/06/24/openai-and-broadcom-reveal-jalapeno-first-ai-chip-in-partnership.html">CNBC&#8217;s coverage</a> has the same details).</p><p>Anthropic&#8217;s disclosed commitments are significant but much smaller: <a href="https://www.anthropic.com/news/anthropic-amazon-compute">up to 5GW with AWS</a> announced April 20 (with nearly 1GW of Trainium online by year end and $100B+ committed over ten years), <a href="https://www.anthropic.com/news/google-broadcom-partnership-compute">well over 1GW of Google TPU capacity in 2026</a> with the Broadcom-built expansion arriving from 2027, and <a href="https://www.anthropic.com/news/higher-limits-spacex">more than 300MW from SpaceX&#8217;s Colossus 1</a>. Its custom silicon is at the conversation stage, <a href="https://techcrunch.com/2026/07/02/anthropic-is-discussing-a-new-custom-chip-with-samsung/">reportedly early talks with Samsung</a> on a 2nm accelerator with nothing shipping before late 2027.</p><p>The strain shows up in the usage limits, which OpenAI moved three times in seventeen days, in both directions. Sol&#8217;s launch week doubled traffic inside 48 hours, so on July 12 and 13 it lifted the five-hour cap on Plus, Pro and Business tiers to absorb the load. On July 25 came a <a href="https://status.openai.com/incidents/01KYC921K145JTR1JK7DYKGWH1">near-global outage</a>, 9:17 to 11:08 AM ET, elevated errors across ChatGPT, the API and Codex, resolved the same morning. On July 29 it reset banked limits for ChatGPT Work and Codex users, whose quota Sol was burning faster than the math anticipated, then put the five-hour cap back the next day. Anthropic has been moving the other way since spring. On May 6 it permanently doubled Claude Code&#8217;s five-hour limits for Pro, Max, Team and seat-Enterprise users and dropped peak-hour throttling for Pro and Max, on the back of that SpaceX capacity.</p><h2>Where Anthropic still leads</h2><p>Anthropic&#8217;s remaining, potentially durable, advantages are enterprise adoption and commercial scale. Anthropic is ahead on total revenue and on valuation, and well ahead on enterprise API share. Its <a href="https://www.anthropic.com/news/series-h">Series H</a>, a $65B round closed May 28, 2026, disclosed a $47B run-rate and a $965B post-money valuation. OpenAI&#8217;s most recent confirmed figures are roughly $25B annualized (about $2B a month) at the <a href="https://www.cnbc.com/2026/03/31/openai-funding-round-ipo.html">March 31, 2026 close</a> of its $122B round, at an <a href="https://www.bloomberg.com/news/articles/2026-03-31/openai-valued-at-852-billion-after-completing-122-billion-round">$852B post-money valuation</a>. Menlo Ventures&#8217; <a href="https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/">December 2025 enterprise report</a> put Anthropic at 40% of enterprise LLM API usage against OpenAI&#8217;s 27%, and at 54% in AI coding specifically. No 2026 Menlo report exists yet, so December is the current number.</p><p>So OpenAI has closed the capability gap and most of the product gap while trailing on money and on enterprise accounts. That&#8217;s the actual shape of the comeback. The lead Anthropic still holds on the Intelligence Index is general capability, and general capability is the harder thing to put in front of a CTO who&#8217;s buying a coding agent this quarter.</p><h2>Apple versus Microsoft, again?</h2><p>I reached for that analogy myself in May 2025, when OpenAI spent close to $10 billion in three weeks on Jony Ive&#8217;s io Products and on Windsurf, and I called it <a href="https://hyperdev.matsuoka.com/p/openais-apple-moment-building-a-walled">OpenAI&#8217;s Apple moment</a>: control the stack from silicon to screen, consumer-first, design-led. Anthropic in that mapping is Microsoft, enterprise and developer-led, winning the accounts while nobody writes headlines about it.</p><p>I don&#8217;t fully trust it. Apple versus Microsoft was never settled by who had the better quarter. It ran on which company owned the layer everybody else had to build on top of, and on switching costs measured in years. Nobody owns that layer here. Both companies rent it, from Amazon, Google, Microsoft, Nvidia, Broadcom and now SpaceX. For a developer, changing frontier models is a config change and an eval run. What costs real time is changing the tool you work in, which is a much smaller lock than owning the platform. The Apple/Microsoft dynamic assumed you couldn&#8217;t leave.</p><p>The mapping also keeps slipping. Anthropic is the enterprise player in it, which is right, but it&#8217;s also the one with the bigger revenue base and the higher valuation, which is not the role the analogy assigns it.</p><h2>Can either of them make money</h2><p>Neither is profitable and both are burning billions, so anything past that is projection.</p><p>OpenAI&#8217;s advertising business is <a href="https://247wallst.com/investing/2026/07/21/openai-is-on-pace-to-miss-its-own-ad-revenue-forecast-by-90-heres-what-it-means-for-the-ai-trade/">on pace to miss its own five-year forecast by about 90%</a>, OpenAI projected $2.5B in ad revenue this year, scaling to $100B by 2030. eMarketer puts the entire US chatbot-ad market, not just OpenAI, under $1B this year and around $5.4B by 2030, per 24/7 Wall St. on July 21. The widely circulated 2026 cash-burn estimates in the $25B range are analyst syntheses rather than disclosures, because OpenAI is private and doesn&#8217;t publish the figures that would settle it. Same caution applies to the gross-margin numbers people quote for both companies.</p><p>Anthropic&#8217;s topline looks better but is contested. Ed Zitron argues in <a href="https://www.wheresyoured.at/anthropics-profitability-swindle/">&#8221;Anthropic&#8217;s &#8216;Profitability&#8217; Swindle&#8221;</a> that the near-term EBITDA profitability claim is a one-quarter accounting artifact of temporarily discounted SpaceX compute during exactly the months profitability was claimed, and that costs still rise linearly with revenue. He also flags CFO Krishna Rao&#8217;s March 9 sworn filing citing cumulative revenue &#8220;exceeding $5 billion to date&#8221; against the $19B run-rate announced six days earlier. That second point is weaker than it reads, since cumulative revenue and annualized run-rate aren&#8217;t the same measure. The first point is the one I&#8217;d want answered, and it hasn&#8217;t been.</p><p>Hardware doesn&#8217;t resolve it. Chris Lehane said at Davos in January that OpenAI&#8217;s first consumer device would debut in the <a href="https://9to5mac.com/2026/01/19/openai-teases-hardware-unveil-this-year-as-jony-ives-team-hires-more-apple-alumni/">&#8221;latter part&#8221; of 2026</a>, which is an unveiling and explicitly not an on-sale date. MacRumors reported in <a href="https://www.macrumors.com/2026/02/20/jony-ive-openai-smart-speaker-2027/">February</a> that the thing is a smart speaker with a camera launching in 2027. Everything else circulating about it (screenless, voice-first, 360-degree camera, Foxconn in Vietnam) is unconfirmed. As of today nothing has shipped and no date is confirmed, so I&#8217;d file the device as a 2027 question and stop considering it this year.</p><h2>Open weights are a side argument</h2><p>The open-weight families are close to the frontier on raw capability. Kimi K3, the highest-scoring open model on the Artificial Analysis Intelligence Index, sits at 57, about four points behind Claude Opus 5&#8217;s 60.7 and under two behind GPT-5.6 Sol&#8217;s 58.9, the same index cited above. On <a href="https://arena.ai/leaderboard">LMArena</a> the best open-weight entry, Kimi K3-max at 1491, trails Claude Fable 5 at the top by 17 points.</p><p>What none of them have is the app ecosystem. No open-weight provider fields anything within reach of ChatGPT or Claude Code on integration depth or developer reach, at least none I found, though I didn&#8217;t dig into DeepSeek&#8217;s or Qwen&#8217;s domestic app numbers in China, which could complicate that. Weights converging matters much less than it sounds like it should when the surface people work in belongs to two companies.</p><p>Which is where I end up. The model tables are the smallest part of it. OpenAI put Codex and ChatGPT in one app with an import command pointed straight at Claude Code&#8217;s users, and it has nearly 7 gigawatts of planned capacity standing behind that. Claude Opus 5 still holds the top of the Intelligence Index at 60.7, and Anthropic still held 40% of enterprise API spend as of last December. OpenAI is knocking on the door. A year ago I&#8217;d have told you that door was closed.</p><p>A coda. Could we imagine a world in which Anthropic and OpenAI run products on other models? Seems very unlikely now. But the cost of developing new models is astronomical, the cost of derived models much less (even though both companies are <a href="https://www.completeskeptic.com/p/is-it-even-possible-for-the-chinese">fighting distillation</a> tooth and nail). So there is a world in which they use the model building work they&#8217;ve done as an audience and tool building exercise, and focus on expanding that share by expanding options, let the models fight for their own justification. Farfetched, as I mentioned, but if they start losing to application clones that use open weight models, they&#8217;re essentially inviting someone to compete. The existing assumption that the app ecosystem was just to drive captured inference falls apart if that inference is cost prohibitive. We shall see.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and also writes about AI business at <a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496">AI Power Ranking</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/p/openais-apple-moment-building-a-walled">OpenAI&#8217;s &#8216;Apple Moment&#8217;: Building a Walled-Garden AI Stack</a> &#8212; Where I first made the Apple analogy, three weeks and $10 billion into OpenAI&#8217;s acquisition run.</p></li><li><p><a href="https://hyperdev.matsuoka.com/p/its-the-harness-stupid">It&#8217;s the Harness, Stupid</a> &#8212; The April piece where I already conceded model parity, under a section header that says so. Everything I got wrong after that is in this article.</p></li><li><p><a href="https://hyperdev.matsuoka.com/p/i-switched-to-claudeai-from-chatgpt">I Switched to Claude.AI from ChatGPT As My Main AI Assistant</a> &#8212; The May 2025 call I&#8217;m revisiting here, with the context-management and stability reasons that drove it.</p></li><li><p><a href="https://aipowerranking.com">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners.</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders.</p></li></ul>]]></content:encoded></item><item><title><![CDATA[What A Difference A Year Makes]]></title><description><![CDATA[claude-mpm to trusty-tools/mpm]]></description><link>https://hyperdev.matsuoka.com/p/what-a-difference-a-year-makes</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/what-a-difference-a-year-makes</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Fri, 24 Jul 2026 12:30:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Ko1F!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3035844-db97-4357-a169-309d987dc3b0_1619x971.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Ko1F!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3035844-db97-4357-a169-309d987dc3b0_1619x971.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Ko1F!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3035844-db97-4357-a169-309d987dc3b0_1619x971.png 424w, https://substackcdn.com/image/fetch/$s_!Ko1F!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3035844-db97-4357-a169-309d987dc3b0_1619x971.png 848w, https://substackcdn.com/image/fetch/$s_!Ko1F!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3035844-db97-4357-a169-309d987dc3b0_1619x971.png 1272w, https://substackcdn.com/image/fetch/$s_!Ko1F!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3035844-db97-4357-a169-309d987dc3b0_1619x971.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Ko1F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3035844-db97-4357-a169-309d987dc3b0_1619x971.png" width="1456" height="873" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e3035844-db97-4357-a169-309d987dc3b0_1619x971.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:873,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2305672,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/208282371?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3035844-db97-4357-a169-309d987dc3b0_1619x971.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Ko1F!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3035844-db97-4357-a169-309d987dc3b0_1619x971.png 424w, https://substackcdn.com/image/fetch/$s_!Ko1F!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3035844-db97-4357-a169-309d987dc3b0_1619x971.png 848w, https://substackcdn.com/image/fetch/$s_!Ko1F!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3035844-db97-4357-a169-309d987dc3b0_1619x971.png 1272w, https://substackcdn.com/image/fetch/$s_!Ko1F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3035844-db97-4357-a169-309d987dc3b0_1619x971.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I made the first commit to my project, called <a href="https://github.com/bobmatnyc/claude-mpm">claude-mpm</a>, on July 24, 2025. It was a multi-agent project manager built on top of Claude Code: a Python package that gave you specialist agents, a ticket workflow, a thin memory layer, and a way to route work between them. Building the agents was most of the work, because at that time Claude Code had no notion of a subagent. You wrote your agents as Markdown, loaded them yourself, and orchestrated the handoffs by hand.</p><p><a href="https://code.claude.com/docs/en/sub-agents">Custom subagents</a> shipped in Claude Code on July 25, 2025, the day after that first commit.</p><p>I&#8217;m not telling this story because of that coincidence. I&#8217;m telling it because it marks a moment. In July 2025, multi-agent orchestration was something novel. A year later it&#8217;s commonplace. And that single move (capability migrating out of the code you write and into the tool you run) is a line connecting the past twelve months in agentic coding, measurable in two of my own repositories.</p><h2>TL;DR</h2><ul><li><p><strong>Same author, same subscription, two eras.</strong> claude-mpm (Python) took roughly twelve months and 4,552 commits, with memory and search either hand-rolled or bolted on from outside. <a href="https://github.com/bobmatnyc/trusty-tools">trusty-tools</a>, a 25-crate Rust monorepo with first-class memory, semantic search, worktree orchestration, and PR review, came together in about two months.</p></li><li><p><strong>The budget held. The developer improved, but the harness and the models leapt.</strong> I spent the year getting better at driving a team of agents. Even so, most of the difference sits with the tooling, not with me. Both projects ran on the same Claude Max subscription. What changed underneath was the frontier model and everything Claude Code learned to do on its own.</p></li><li><p><strong>A year ago you hand-built the PM layer.</strong> Agents-as-Markdown, an external vector-search dependency, a memory subsystem you maintained yourself. Working in tickets felt like an edge.</p></li><li><p><strong>Now the harness carries it.</strong> Subagents nest and run in the background, worktrees are a native flag, skills and memory and search are first-class. Ticket-driven development is table stakes. Running many worktrees against one repo is my working model now.</p></li><li><p><strong>The safe extrapolation:</strong> The observed path is tickets &#8594; worktrees &#8594; many concurrent sessions per repo. If the next year rhymes with the last, the harness manages more parallelism than a person can hold in their head.</p></li></ul><h2>A year ago: you built the orchestration yourself</h2><p>claude-mpm was shaped by what the harness couldn&#8217;t do at that time.</p><p>The frontier models in late July 2025 were Claude Opus 4 and Sonnet 4, which had <a href="https://www.anthropic.com/news/claude-4">reached general availability</a> on May 22, 2025. Claude Code itself was about two months into GA, capable but young. Subagents didn&#8217;t exist until the day after I started. <a href="https://www.anthropic.com/news/agent-skills">Agent Skills</a> wouldn&#8217;t launch until October 2025. There was no native worktree support. If you wanted isolated parallel work, you ran <code>git worktree add</code> yourself and wired it up by hand. MCP was maybe eight months old. Memory and semantic search were not primitives the harness handed you.</p><p>So we (claude-mpm and me) built all of it. The agents were Markdown templates the package loaded and orchestrated. The ticket workflow lived in the CLI. Memory was thin by necessity, a couple of files under <code>src/claude_mpm/memory/</code>, and even that was on its way out. The changelog shows the memory hooks being removed and handed off to an external successor rather than maintained in-tree. Semantic search wasn&#8217;t native either. It came in as an outside dependency, <code>mcp-vector-search</code>, referenced across dozens of files. Worktree awareness existed at the level of a consumer that knew the concept, not a subsystem that owned it.</p><p>That was the state of the art then, and it worked. Over about twelve months claude-mpm accumulated 4,552 commits (roughly 380 a month) and grew into a real system: multiple MCP channel servers built into the Python package, a plugin path exposing 56-odd skills, agent templates for a spread of specialist roles. Ticket-driven development, routing discrete units of work through a queue instead of narrating one long conversation, felt like an advance at the time. You had to construct the scaffolding to get there, and the scaffolding was the hard part of the project.</p><h2>Now: the harness carries it</h2><p>trusty-tools took its first commit on May 19, 2026. It was still under active development the day I pulled these numbers. In roughly two months it reached 1,726 commits on <code>main</code>, about 860 a month. Same author, same Claude Max subscription, a different era.</p><p>That first commit wasn&#8217;t a blank slate. It already carried claude-mpm&#8217;s PM scaffolding, and one absorbed component holds claude-mpm session logs dated May 11, a week before the new repo existed. The Python predecessor was building its Rust successor. Then the successor began building itself: on July 6 the commit attribution switched to &#8220;generated with trusty-mpm,&#8221; in a commit that was fixing trusty-mpm&#8217;s own guard for its own subagents. Three days later the instructions moved out of CLAUDE.md into trusty-mpm&#8217;s own convention. About seven weeks were built by its predecessor before it took over its own development.</p><p>A lot shipped in that window. The frontier moved to <a href="https://www.anthropic.com/news/claude-opus-4-8">Claude Opus 4.8</a> (late May 2026) and <a href="https://www.anthropic.com/news/claude-sonnet-5">Claude Sonnet 5</a> (end of June 2026), with a Mythos-class model, <a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">Fable 5</a>, arriving in June at a million tokens of context. The million-token context window had <a href="https://claude.com/blog/1m-context-ga">gone GA at standard pricing</a> in early 2026. Inside Claude Code, subagents now nest several deep and run in the background by default. <code>/fork</code><a href="https://code.claude.com/docs/en/changelog"> and </a><code>/subtask</code><a href="https://code.claude.com/docs/en/changelog"> landed</a> in mid-2026, and a <a href="https://code.claude.com/docs/en/changelog">native </a><code>--worktree</code><a href="https://code.claude.com/docs/en/changelog"> flag</a> had arrived in 2026. The harness I was building on top of in mid-2026 was a different animal from the one I started claude-mpm against.</p><p>And trusty-tools reflects that, because it didn&#8217;t have to build the parts the harness now provides. It could spend that effort building capabilities further out. The result is a 25-crate workspace consisting of about 600K lines of Rust. The pieces claude-mpm hand-rolled or imported are now first-class crates of their own:</p><ul><li><p><strong>Memory</strong> is <a href="https://github.com/bobmatnyc/trusty-tools">trusty-memory</a>, a dedicated crate with knowledge-graph operations, &#8220;dream&#8221; consolidation that compacts and reorganizes stored facts, multiple namespaced memory &#8220;palaces,&#8221; and chat-session persistence. claude-mpm had no equivalent. Its two-file memory layer was being handed off precisely because maintaining that by hand no longer made sense.</p></li><li><p><strong>Search</strong> is a first-class crate (semantic, lexical, and knowledge-graph search, call-chain lookup, typeahead, indexing) rather than an external MCP dependency referenced across the codebase.</p></li><li><p><strong>Worktree orchestration</strong> is heavy and native to the design: <code>EnterWorktree</code>/<code>ExitWorktree</code> as real operations, and fifteen-plus simultaneous live worktrees as the ordinary way of working, not a party trick.</p></li><li><p><strong>Ticketing</strong> is a dedicated subsystem spanning GitHub issues and JIRA, not an example command.</p></li><li><p>And then the crates with no claude-mpm analogue at all: PR and diff review, git analytics, a code intelligence layer, a TUI, an embedding daemon.</p></li><li><p>It also has an original (not meta) harness called trusty-code (still a work in progress but early indications are that it will perform similarly to open-code), and a personal agents harness called trusty-agents. Both leverage many of the same orchestration tools that trusty-mpm uses, which is why I&#8217;m including them in the package.</p></li></ul><p>On top of that sits a catalog of 37 specialist agents over five foundation layers, and a live skill catalog surfacing more than 190 skill names, triple claude-mpm&#8217;s 56. The prompting style changed too. A year ago you spent your prompt budget teaching the model how to be an agent. Now you spend it telling a competent agent what you want, because the harness supplies the how (in the form of memory, search, specs, and tickets).</p><p>Speaking of, ticket-driven development, the edge a year ago, is now the assumed baseline. The frontier moved up a level: many worktrees against a single repo or monorepo, several sessions of work in flight at once.</p><h2>The measured contrast</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wZXl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa668e6d3-c3a0-4473-934c-49b405983fd7_1600x688.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wZXl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa668e6d3-c3a0-4473-934c-49b405983fd7_1600x688.png 424w, https://substackcdn.com/image/fetch/$s_!wZXl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa668e6d3-c3a0-4473-934c-49b405983fd7_1600x688.png 848w, https://substackcdn.com/image/fetch/$s_!wZXl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa668e6d3-c3a0-4473-934c-49b405983fd7_1600x688.png 1272w, https://substackcdn.com/image/fetch/$s_!wZXl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa668e6d3-c3a0-4473-934c-49b405983fd7_1600x688.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wZXl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa668e6d3-c3a0-4473-934c-49b405983fd7_1600x688.png" width="1456" height="626" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a668e6d3-c3a0-4473-934c-49b405983fd7_1600x688.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:626,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:91029,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/208282371?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa668e6d3-c3a0-4473-934c-49b405983fd7_1600x688.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!wZXl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa668e6d3-c3a0-4473-934c-49b405983fd7_1600x688.png 424w, https://substackcdn.com/image/fetch/$s_!wZXl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa668e6d3-c3a0-4473-934c-49b405983fd7_1600x688.png 848w, https://substackcdn.com/image/fetch/$s_!wZXl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa668e6d3-c3a0-4473-934c-49b405983fd7_1600x688.png 1272w, https://substackcdn.com/image/fetch/$s_!wZXl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa668e6d3-c3a0-4473-934c-49b405983fd7_1600x688.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Two projects, one author, one subscription. Here&#8217;s the comparison. Effort and wall-clock framing are estimates. Commit counts and dates are exact. The productivity story they imply is inference.</p><p><em>Time to build a comparable system: claude-mpm ~12 months versus trusty-tools ~2 months, roughly a sixth of the wall-clock time.</em></p><p>A more capable system (25 crates with first-class memory, search, worktree orchestration, PR review, and git analytics) came together in about a sixth of the calendar time.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!mZkp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffffec223-9ac6-409c-8a8e-eaa391cea2b4_1600x736.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!mZkp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffffec223-9ac6-409c-8a8e-eaa391cea2b4_1600x736.png 424w, https://substackcdn.com/image/fetch/$s_!mZkp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffffec223-9ac6-409c-8a8e-eaa391cea2b4_1600x736.png 848w, https://substackcdn.com/image/fetch/$s_!mZkp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffffec223-9ac6-409c-8a8e-eaa391cea2b4_1600x736.png 1272w, https://substackcdn.com/image/fetch/$s_!mZkp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffffec223-9ac6-409c-8a8e-eaa391cea2b4_1600x736.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!mZkp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffffec223-9ac6-409c-8a8e-eaa391cea2b4_1600x736.png" width="1456" height="670" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fffec223-9ac6-409c-8a8e-eaa391cea2b4_1600x736.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:670,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:97055,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/208282371?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffffec223-9ac6-409c-8a8e-eaa391cea2b4_1600x736.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!mZkp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffffec223-9ac6-409c-8a8e-eaa391cea2b4_1600x736.png 424w, https://substackcdn.com/image/fetch/$s_!mZkp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffffec223-9ac6-409c-8a8e-eaa391cea2b4_1600x736.png 848w, https://substackcdn.com/image/fetch/$s_!mZkp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffffec223-9ac6-409c-8a8e-eaa391cea2b4_1600x736.png 1272w, https://substackcdn.com/image/fetch/$s_!mZkp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffffec223-9ac6-409c-8a8e-eaa391cea2b4_1600x736.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Commit velocity: claude-mpm ~380/month versus trusty-tools ~860/month, roughly 2.3x, though squash-merges across many worktrees likely undercount trusty-tools&#8217; true activity.</em></p><p>Velocity roughly 2.3x per month. A year of tech-lead work sharpened the exact skills this way of working rewards, driving a team, now a team of agents, and holding SDLC discipline as the branches multiply. And it shows in the git record, not just in my own say-so. Three signals I can actually measure. Commit messages that name a driving issue or decision rose from roughly one commit in ten to about nine in ten. The code now logs <em>why</em>, not only <em>what</em>. The eight near-identical agent files I copy-pasted early in claude-mpm collapsed into one composed base definition, a fix I started mid-project and carried further into trusty-tools. And architecture decision records went from essentially none to eighteen numbered ADRs plus a per-crate decisions taxonomy, a habit that matured across both projects rather than one trusty-tools invented. Developer growth and harness capability compound. They push the same direction. But even with that, a person going from good to better doesn&#8217;t buy you a sixth of the calendar time. The impressive part is still the tooling.</p><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tj7D!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19050f4c-efa9-4513-9c5a-0c97c3860b5b_1650x1110.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tj7D!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19050f4c-efa9-4513-9c5a-0c97c3860b5b_1650x1110.png 424w, https://substackcdn.com/image/fetch/$s_!tj7D!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19050f4c-efa9-4513-9c5a-0c97c3860b5b_1650x1110.png 848w, https://substackcdn.com/image/fetch/$s_!tj7D!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19050f4c-efa9-4513-9c5a-0c97c3860b5b_1650x1110.png 1272w, https://substackcdn.com/image/fetch/$s_!tj7D!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19050f4c-efa9-4513-9c5a-0c97c3860b5b_1650x1110.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tj7D!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19050f4c-efa9-4513-9c5a-0c97c3860b5b_1650x1110.png" width="1456" height="979" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/19050f4c-efa9-4513-9c5a-0c97c3860b5b_1650x1110.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:979,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:180011,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/208282371?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19050f4c-efa9-4513-9c5a-0c97c3860b5b_1650x1110.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!tj7D!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19050f4c-efa9-4513-9c5a-0c97c3860b5b_1650x1110.png 424w, https://substackcdn.com/image/fetch/$s_!tj7D!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19050f4c-efa9-4513-9c5a-0c97c3860b5b_1650x1110.png 848w, https://substackcdn.com/image/fetch/$s_!tj7D!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19050f4c-efa9-4513-9c5a-0c97c3860b5b_1650x1110.png 1272w, https://substackcdn.com/image/fetch/$s_!tj7D!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19050f4c-efa9-4513-9c5a-0c97c3860b5b_1650x1110.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Capability checklist across three columns, a year ago, now, and a speculative year from now, for orchestration, memory, search, worktrees, skills, ticketing, and PR review.</em></p><p>Every row that read &#8220;you build this&#8221; a year ago reads &#8220;the harness provides this&#8221; now. Multi-agent orchestration: hand-built then, native flag now. Memory: thin and hand-maintained then, a knowledge-graph crate now. Search: external dependency then, first-class now. Worktrees: DIY shell commands then, a <code>--worktree</code> flag and fifteen live trees now.</p><p>The harness now enforces the workflow, not just supplies the parts. Spec-linked documentation was a prompt hint in claude-mpm, off by default. In trusty-tools it&#8217;s a build-blocking lint gate (<code>trusty-sld-lint</code>) wired into CI and pre-commit. Documentation that fails the build, not documentation you&#8217;re reminded to write. Ticket discipline hardened the same way: that rise in issue-referenced commits now sits inside a codified chain (spec &#8594; issue &#8594; PR-linked-to-issue &#8594; review gate &#8594; squash-merge), process-enforced, not yet a hard check that rejects an unlinked PR. And claude-mpm&#8217;s main branch required zero approvals and no passing checks, protected in name only. trusty-tools&#8217; main requires a review approval and six passing CI checks, with an LLM review pass (<code>trusty-review</code>) as a named gate. A year ago a disciplined developer chose these. Now the harness refuses to merge without a review and passing checks.</p><h2>A word on the subscription plan</h2><p>Claude Max was announced in April 2025 at <a href="https://support.claude.com/en/articles/11049741-what-is-the-max-plan">$100/month (Max 5x) and $200/month (Max 20x)</a>, with usage shared across chat and Code. Those price points appear to have held from then through mid-2026. The most visible change over that whole span was the five-hour rate limits for Claude Code, which Max subscribers draw on, <a href="https://www.anthropic.com/news/higher-limits-spacex">being doubled</a> in mid-2026. More headroom, not a different product tier.</p><p>So the input that stayed roughly constant was the money and the person. The input that changed was the model quality and the harness capability. When you hold the developer and the budget still and the output grows by this much, the variable that moved most would be the model and the harness (mostly the model, though in my experience good memory and search make a huge difference), even granting the developer improved and the language changed.</p><h2>A year from now</h2><p>The trajectory has a clear shape: tickets &#8594; worktrees &#8594; multi-session parallelism. Working in tickets was novel in mid-2025 and is common sense now. Native worktrees arrived in early 2026 and, within months, running many of them at once became my working model rather than a demo. Each step took a workflow that used to live in the developer&#8217;s head (tracking the units of work, isolating the parallel branches) and moved it into the tool. There are no hard adoption statistics for this progression. I&#8217;m describing a tooling timeline and a direction, not a measured majority practice. The tooling timeline itself is real, though: worktree support arriving across editors and agents through late 2025 and into 2026 follows the same curve.</p><p>The next frontier isn&#8217;t hard to name. If tickets became common sense and worktrees became my working model, the thing after worktrees is many concurrent sessions against one repo or monorepo. Enough parallel work in flight so that no human is tracking all of it, and the harness keeps the branches, the memory, and the review gates coherent. The <code>/fork</code> and <code>/subtask</code> primitives that landed in mid-2026, and subagents running in the background by default, are early moves in exactly that direction. The developer&#8217;s job shifts further from writing the steps toward specifying the outcome and reviewing the merge.</p><p>That&#8217;s not a prediction of artificial general anything. It&#8217;s the same migration, run one more turn: complexity leaving the code you write by hand and entering the tool you run. A year ago the hard part of the project was the scaffolding. This year it&#8217;s the harness. Next year it&#8217;s the coordination of more parallel work than one person can follow. So that one person gets pushed further up into the spec and the design.</p><h2>What a difference a year makes</h2><p>I built claude-mpm, and I&#8217;m very proud of it. For its moment it was a good answer to a real constraint. That&#8217;s the ordinary fate of scaffolding once the platform grows the feature underneath it, the design working as intended.</p><p>The measure of the year isn&#8217;t that I got better, though I hope I did. It&#8217;s that the same person, on the same max plan, with the same working habits, could build a materially more capable system in a fraction of the time, because the models got sharper and the harness absorbed the work that used to be mine to do. Twelve months of hand-built PM layer on one side, two months of composing first-class parts on the other. Same author. Same subscription.</p><p>What a difference a year makes.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and also writes about AI business at <a href="https://aipowerranking.com/">AI Power Ranking</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/p/its-the-harness-stupid">It&#8217;s the Harness, Stupid</a> &#8212; Why orchestration, not raw model quality, drives the spread in outcomes &#8212; the mechanism underneath this whole comparison.</p></li><li><p><a href="https://hyperdev.matsuoka.com/p/what-is-harness-engineering">What Is Harness Engineering? (And Do You Need to Learn It?)</a> &#8212; The durable skill under the tooling: designing the loop, not just the scaffolding.</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners.</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders.</p></li></ul>]]></content:encoded></item><item><title><![CDATA[If You’re Not Writing Specs, You’re Vibe Coding]]></title><description><![CDATA[And your specs need to get better.]]></description><link>https://hyperdev.matsuoka.com/p/if-youre-not-writing-specs-youre</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/if-youre-not-writing-specs-youre</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Thu, 16 Jul 2026 11:30:08 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!sThl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24749677-8226-49c5-b65d-26f7f8abb916_1600x1067.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sThl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24749677-8226-49c5-b65d-26f7f8abb916_1600x1067.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sThl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24749677-8226-49c5-b65d-26f7f8abb916_1600x1067.png 424w, https://substackcdn.com/image/fetch/$s_!sThl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24749677-8226-49c5-b65d-26f7f8abb916_1600x1067.png 848w, https://substackcdn.com/image/fetch/$s_!sThl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24749677-8226-49c5-b65d-26f7f8abb916_1600x1067.png 1272w, https://substackcdn.com/image/fetch/$s_!sThl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24749677-8226-49c5-b65d-26f7f8abb916_1600x1067.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sThl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24749677-8226-49c5-b65d-26f7f8abb916_1600x1067.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/24749677-8226-49c5-b65d-26f7f8abb916_1600x1067.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2368071,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/207192723?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24749677-8226-49c5-b65d-26f7f8abb916_1600x1067.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!sThl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24749677-8226-49c5-b65d-26f7f8abb916_1600x1067.png 424w, https://substackcdn.com/image/fetch/$s_!sThl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24749677-8226-49c5-b65d-26f7f8abb916_1600x1067.png 848w, https://substackcdn.com/image/fetch/$s_!sThl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24749677-8226-49c5-b65d-26f7f8abb916_1600x1067.png 1272w, https://substackcdn.com/image/fetch/$s_!sThl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F24749677-8226-49c5-b65d-26f7f8abb916_1600x1067.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Every engineering org I&#8217;ve worked in has a folder like this somewhere. A <code>docs/</code> directory, a Confluence space, a Google Doc linked from a wiki page nobody visits. Inside it: the spec. Written carefully before a feature shipped. Reviewed by three people. Approved. And then, the moment the code merged, functionally dead &#8212; a fossil of what the team intended, drifting from what the system does.</p><p>I don&#8217;t say that to complain about lazy documentation habits. Nothing in the workflow ties the spec to what ships &#8212; nothing breaks when the code changes and the doc doesn&#8217;t, and nothing in CI cares. The traceability matrix that maps requirements to implementation gets updated once, at launch, because keeping it current isn&#8217;t anyone&#8217;s job. Six months later an engineer greps the spec for the retry policy, finds three paragraphs of prose, and can&#8217;t tell whether they match the code or a decision quietly overridden in a hotfix a year back.</p><p>Spec-driven development, in its traditional form, fixes this with process: write the spec first, get sign-off, then build to it. That helps at the start of a project and does nothing for the middle, where the spec is still a separate artifact checked by separate tooling &#8212; if it&#8217;s checked at all. The argument here is narrower than &#8220;write better specs.&#8221; The fix isn&#8217;t a better document; it&#8217;s making the document part of the same system that builds and tests the code, so it can&#8217;t drift without something noticing.</p><h2>TL;DR</h2><ul><li><p><strong>Nothing enforces the spec-to-code link, so it rots.</strong> A standalone spec has nothing tying it to what ships; nothing breaks when it drifts, so it does.</p></li><li><p><strong>The reframe:</strong> the spec becomes part of the workflow &#8212; referenced inline from code comments, parsed as data by the build, versioned in the same commits as the implementation.</p></li><li><p><strong>trusty-mpm is a working instance</strong>, not a proposal: a spec requirement runs verbatim inside the prompt the model reads each session, code comments cite spec sections by ID, and a Rust module parses spec markdown at build time to check code against it &#8212; flagging code that cites a spec revision that&#8217;s since changed.</p></li><li><p><strong>This isn&#8217;t new</strong> &#8212; literate programming, Gherkin&#8217;s living documentation, rustdoc doctests, and design-by-contract all tried versions of it &#8212; but AI-generated code gives it new urgency: the volume of change makes a hand-maintained spec even less credible than before.</p></li><li><p><strong>The industry is converging on the same shape</strong>: GitHub Spec Kit, AWS Kiro, OpenSpec, and Tessl all bet that spec and code must co-evolve as tracked artifacts, not a document and its distant cousin.</p></li><li><p><strong>A minimal frontmatter schema</strong> &#8212; stable ID, version, status, <code>applies_to</code> globs &#8212; is the concrete thing to adopt this quarter, whether or not you touch any of the tools above.</p></li></ul><h2>The reframe: the spec as workflow, not deliverable</h2><p>Here&#8217;s the alternative. Instead of describing the system from outside, the spec becomes a component the system consumes. Code references it by a stable ID; the build (or a linter, or a review bot) parses the spec and checks the referenced section still says what the code assumes. Editing the spec means touching something with known consumers, not filing a document into a drawer.</p><p>None of this is new. Knuth&#8217;s <a href="https://en.wikipedia.org/wiki/Literate_programming">literate programming</a> interleaved explanation and code decades ago; <a href="https://cucumber.io/docs/gherkin/">Gherkin</a> made acceptance criteria executable; Rust&#8217;s <a href="https://doc.rust-lang.org/rustdoc/write-documentation/documentation-tests.html">doctests</a> and <code>mdbook test</code> run documentation&#8217;s code examples in the test suite, so a doc that lies about the API fails CI; design-by-contract (Eiffel, Ada/SPARK) puts pre- and postconditions in the interface; <a href="https://datatracker.ietf.org/doc/html/rfc2119">RFC 2119</a> gave the field MUST/SHOULD/MAY, precise enough for a machine to check.</p><p>What&#8217;s changed is the pace, and who&#8217;s writing the code. When a team ships every two weeks, a stale spec is an annoyance you catch next planning cycle. When an agent generates hundreds of lines and opens a PR before you&#8217;ve finished coffee, an unchecked spec is stale before the reviewer finishes the diff. The old techniques weren&#8217;t wrong; they just didn&#8217;t have to survive this throughput.</p><h2>Why code alone isn&#8217;t enough</h2><p>Open a function you didn&#8217;t write and you can read exactly what it does. What you can&#8217;t read is whether that&#8217;s what it was supposed to do. Code is a precise record of behavior and a silent one about intent &#8212; in the source tree a bug and a feature look identical, and the only thing that tells them apart is what someone wanted, which lives outside the code.</p><p>Tests don&#8217;t fill that gap, even though we reach for them as if they do. A test locks in behavior &#8212; it keeps the code doing what it does now &#8212; but it can&#8217;t tell you that behavior was the one you wanted. A test that a webhook retries three times passes just as happily whether &#8220;three&#8221; was a deliberate choice or a number someone typed on a Friday. The spec carries the intent, separate from the code meant to deliver it.</p><p>The deadline hits, so you take the shortcut. The hotfix goes in at 2 a.m. and nobody circles back. Every codebase is a running negotiation with expediency, and each shortcut gets absorbed into the code, where it stops looking like a shortcut and starts looking like a decision. A year later nobody can tell which lines were chosen and which were just the fastest way through. The spec is the one term of that negotiation you agree not to move &#8212; a survey marker: the code shifts around it, and because it stays put, you can measure how far. Take it away and a shortcut is just the code, indistinguishable from what&#8217;s near it.</p><p>That&#8217;s what trusty-mpm&#8217;s drift check does, later in this piece: when a spec section moves to a new revision and the code still points at the old one, the check flags it. The marker moved; the tool says so.</p><p>That agreement is the line between formal engineering and vibe coding. Vibe coding keeps the output and throws away the intent behind it. It works until the output has to change and neither the humans nor the model can reconstruct what it was for. Sean Grove, whose argument I reach below, calls this version-controlling the binary and shredding the source. AI sharpens the line rather than softening it: ask an agent to fix the case in front of it and it will often comply by quietly narrowing a behavior &#8212; solving the immediate problem, dropping an edge case nobody mentioned &#8212; and without a spec, that narrowing is just the new code, shipped with passing tests. The difference between formal engineering and vibe coding is the spec: the idea that doesn&#8217;t drift for expediency.</p><h2>An exemplar: trusty-mpm</h2><p><a href="https://github.com/bobmatnyc/trusty-tools">trusty-mpm</a> &#8212; the <code>tm</code> binary &#8212; is a Rust-based multi-agent orchestration harness I&#8217;ve been building: the layer that composes the prompt assets, agent definitions, and skills a Claude Code session consumes. (Not the unrelated Python project of a similar name.) The spec-as-workflow pattern here wasn&#8217;t designed up front; it emerged from a failure. What follows is what it does and why it matters &#8212; the file paths and function names are collected in the appendix.</p><h3>A spec the running system reads</h3><p>A live trusty-mpm session once misidentified its own harness &#8212; told the user it was running under something it wasn&#8217;t. The fix wasn&#8217;t a prompt patch and a changelog note. It was a spec: a short document stating what the system must know about itself, and when.</p><p>What makes it more than a postmortem is where that spec ended up. Its central requirement doesn&#8217;t just sit in a docs folder &#8212; it runs, close to verbatim, inside the prompt the model reads at the start of every session. And the live instruction cites its own source, naming the exact spec section that specified it. Change that section without updating the prompt and the citation is still pointing at it, flagging the mismatch. The requirement can&#8217;t quietly fall out of the running system, because the running system reads it every time.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OCiI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fe0a313-4d07-474d-a5ff-c1ddea8757ca_1600x960.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OCiI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fe0a313-4d07-474d-a5ff-c1ddea8757ca_1600x960.png 424w, https://substackcdn.com/image/fetch/$s_!OCiI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fe0a313-4d07-474d-a5ff-c1ddea8757ca_1600x960.png 848w, https://substackcdn.com/image/fetch/$s_!OCiI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fe0a313-4d07-474d-a5ff-c1ddea8757ca_1600x960.png 1272w, https://substackcdn.com/image/fetch/$s_!OCiI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fe0a313-4d07-474d-a5ff-c1ddea8757ca_1600x960.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OCiI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fe0a313-4d07-474d-a5ff-c1ddea8757ca_1600x960.png" width="1456" height="874" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6fe0a313-4d07-474d-a5ff-c1ddea8757ca_1600x960.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:874,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1718660,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/207192723?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fe0a313-4d07-474d-a5ff-c1ddea8757ca_1600x960.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!OCiI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fe0a313-4d07-474d-a5ff-c1ddea8757ca_1600x960.png 424w, https://substackcdn.com/image/fetch/$s_!OCiI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fe0a313-4d07-474d-a5ff-c1ddea8757ca_1600x960.png 848w, https://substackcdn.com/image/fetch/$s_!OCiI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fe0a313-4d07-474d-a5ff-c1ddea8757ca_1600x960.png 1272w, https://substackcdn.com/image/fetch/$s_!OCiI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6fe0a313-4d07-474d-a5ff-c1ddea8757ca_1600x960.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Links that run both ways</h3><p>For that to hold up across a whole codebase, specs need addresses. Each spec section has a stable ID, and the code that implements it carries a comment naming the section it satisfies. The spec points back the other way: each requirement lists the modules that implement it. Either side can find the other &#8212; which means either side drifting from the other gets caught, rather than surfacing months later as a mystery.</p><h3>A check that catches an outdated pointer</h3><p>Those section IDs carry a revision number. When a spec section changes in a way that matters, its revision moves forward. Code that still points at the old revision was written against a spec that has since moved &#8212; so a check flags it instead of trusting the stale link (the specifics are in the appendix). It&#8217;s the traceability-matrix problem from the opening, turned into something a machine checks on every change instead of a spreadsheet cell a human forgets to update.</p><h3>A test that fails when the shipped artifact drifts</h3><p>That self-awareness spec has acceptance criteria: things the shipped prompt must contain. A CI test asserts exactly those. If a later edit strips a required section out of the prompt, the test fails on the pull request that caused it, not in a bug report three sprints later. The spec&#8217;s requirement and the test&#8217;s pass condition are the same statement, written once.</p><h3>The smallest version of the same idea</h3><p>The discipline scales all the way down. Inside individual functions, a short comment states why a behavior exists and names the tests that enforce it, right beside the code. And the house rule is explicit: a spec is a behavior contract &#8212; it says what, not how &#8212; and code must link to its spec section from the first pull request that implements it, even while the spec is still a draft. Load-bearing before it&#8217;s finished.</p><h2>The industry is formalizing the same idea</h2><p>trusty-mpm isn&#8217;t the only place this is showing up; the convergence from different directions is itself evidence the constraint is real.</p><p>The clearest version of the thesis comes from Sean Grove at OpenAI, in <a href="https://lawwu.github.io/transcripts/8rABwKRsec4.html">a talk transcribed as &#8220;The New Code&#8221;</a>. Code, he argues, captures maybe 10-20% of a piece of software&#8217;s value; the rest is the structured communication &#8212; intent, constraints, tradeoffs &#8212; that produced it. Most teams version-control the binary and shred the source: they keep the generated code and throw away the spec. He cites OpenAI&#8217;s own <a href="https://model-spec.openai.com/">Model Spec</a> as one meant to stay living and cross-functional rather than one team&#8217;s document the rest of the org ignores.</p><p>Several tools are building toward the same shape. <a href="https://github.com/github/spec-kit">GitHub Spec Kit</a> runs spec &#8594; plan &#8594; tasks &#8594; implement with artifacts colocated in the repo (<code>specs/[branch]/spec.md</code>), <code>FR-001</code>-style IDs, and a project <code>constitution.md</code> of non-negotiables &#8212; though it doesn&#8217;t use YAML frontmatter, a gap I address below. <a href="https://kiro.dev/docs/specs/">AWS Kiro</a> structures each feature as <code>requirements.md</code>, <code>design.md</code>, and <code>tasks.md</code>, plus persistent &#8220;steering files&#8221; for conventions. <a href="https://github.com/Fission-AI/OpenSpec">OpenSpec</a> is the closest to spec and code co-evolving as tracked diffs: a propose &#8594; apply &#8594; archive cycle built on delta specs (ADDED/MODIFIED/REMOVED/RENAMED, in GIVEN/WHEN/THEN form). And <a href="https://tessl.io/">Tessl</a> <a href="https://techcrunch.com/2024/11/14/tessl-raises-125m-at-at-500m-valuation-to-build-ai-that-writes-and-maintains-code/">raised $125M</a> on the strongest version of the bet &#8212; the spec is the source, the code a regenerable output marked <code>// GENERATED FROM SPEC - DO NOT EDIT</code> &#8212; though as of mid-2026 still a thesis, not a practice with a track record; I cite it as direction, not proof.</p><p>None of these do exactly what trusty-mpm&#8217;s drift check does &#8212; parse the spec at build time, check it against the linked code, and flag an outdated pointer automatically, in CI rather than argued for as a future state.</p><h2>The spec stops being about the system</h2><p>The through-line is narrower than &#8220;documentation matters&#8221; &#8212; most teams agree it matters, and stale docs keep happening anyway. The claim is mechanical: a spec with a stable ID, cited from code, parsed by a build step, checked for drift, behaves differently than one in a wiki &#8212; not because anyone tries harder, but because something other than good intentions is now checking.</p><p>That&#8217;s the demarcation from the opening, made concrete: the difference between formal engineering and vibe coding is the spec &#8212; the idea a team agrees not to drift for expediency. The mechanisms in this piece exist to keep that agreement from decaying back into a good intention.</p><p>I should hold trusty-mpm to that standard, and it half meets it. The conventions are there: document numbers, stable section IDs, links from code to the spec sections it implements, the drift check, a status / owner / last-updated header on each spec. What&#8217;s missing is the last mile I&#8217;m prescribing to other teams: that header is semi-prose, not a machine-readable schema &#8212; no formal frontmatter, no published schema, no CI gate validating it. trusty-mpm parses spec <em>prose</em> today, not spec <em>metadata</em>, because the metadata isn&#8217;t structured enough to parse.</p><p>So <a href="https://github.com/bobmatnyc/trusty-tools/issues/2679">the next step</a> is to close that gap in trusty-tools: a formal frontmatter block on each spec, a schema behind it, and a CI check that fails when a spec&#8217;s metadata drifts from the code it governs. That&#8217;s what adopting the standard means here &#8212; making the build enforce conventions the project only half-follows, instead of prescribing to others a discipline I&#8217;ve left to good intentions in my own repo. The spec becomes a component of the system at the point where the build won&#8217;t let it drift. That&#8217;s the part still to build.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and also writes about AI business at <a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">AI Power Ranking</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/p/its-the-harness-stupid">It&#8217;s the Harness, Stupid</a> &#8212; What a harness is and why orchestration, not model quality, drives the spread in outcomes &#8212; the same infrastructure that makes spec-as-workflow possible.</p></li><li><p><a href="https://hyperdev.matsuoka.com/p/what-is-harness-engineering">What Is Harness Engineering? (And Do You Need to Learn It?)</a> &#8212; The durable skill underneath the harness: designing the loop, not just the scaffolding &#8212; spec resolution is one piece of that loop.</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners.</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders.</p></li></ul><div><hr></div><h2>Appendix: how it actually works</h2><h4>The practical takeaway: a frontmatter schema that plugs into the build</h4><p>If you want a piece of this without a whole tool, start with a frontmatter schema for your specs &#8212; the structured metadata block at the top of a markdown file &#8212; that your build can validate. A precedent to copy: <a href="https://smadr.dev/reference/specification/overview/">Structured MADR</a> ships a CI-validated YAML schema (JSON Schema plus a GitHub Action) with fields like <code>title</code>, <code>status</code>, <code>created</code>/<code>updated</code>, <code>author</code>, and <code>related</code>. Kubernetes&#8217; KEP (<code>kep.yaml</code>) and Python&#8217;s PEP headers are older versions of the same idea: metadata structured enough for tooling to check, not just a reviewer.</p><p>Here&#8217;s a schema to start from, adapted for code-adjacent specs rather than pure design docs:</p><p>Field Example Purpose <code>id</code> <code>SPEC-042</code> Stable anchor code comments point at (e.g. <code>// impl of SPEC-042#retry-policy</code>) <code>title</code> <code>"Retry policy for outbound webhook delivery"</code> Human-readable name <code>status</code> <code>accepted</code> One of draft / review / accepted / deprecated / superseded <code>version</code> <code>1.2.0</code> Bump on any semantically meaningful change &#8594; enables revision-drift detection <code>owners</code> <code>["@handle"]</code> Who owns the spec <code>applies_to</code> <code>["src/webhooks/**"]</code> Globs this spec governs; CI flags code changed under a glob with no spec touch (and vice versa) <code>supersedes</code> <code>[]</code> IDs of specs this one replaces <code>requirement_ids</code> <code>["FR-014", "FR-015"]</code> Requirement IDs this spec governs <code>last_verified</code> <code>2026-06-01</code> Last time someone confirmed spec still matches code</p><p>Each field does something a build can act on, not just something a human can read. Two do the heavy lifting: <code>version</code> makes drift checkable &#8212; without it, &#8220;the code cites an old section&#8221; is a judgment call, not a fact a machine can determine &#8212; and <code>applies_to</code> lets CI flag a PR that touches a governed glob with no spec commit, and the reverse. Together they turn the staleness signal a traceability matrix only promised into something a linter checks, not a column a human forgets.</p><p>trusty-mpm grew its own version of this organically, out of that self-identification failure. Before any schema existed, its section IDs were already doing the job <code>id</code> does above, and its revision-tagged sections plus the drift check were doing the job <code>version</code> does. Adopting this from scratch, you don&#8217;t need to rediscover that path &#8212; the frontmatter above is the same mechanism, explicit from day one.</p><h3>How it works</h3><p>The mechanics behind each mechanism in the exemplar section &#8212; file paths, identifiers, and the code that backs them.</p><ul><li><p><strong>The self-awareness spec</strong> &#8212; <code>docs/specs/trusty-mpm-self-awareness.md</code>, tracked as DOC-28. Its requirement R2 runs close to verbatim in the session prompt, <code>BASE_SM.md</code>, and in the <a href="https://github.com/bobmatnyc/trusty-tools/blob/main/crates/trusty-mpm/src/assets/output-styles/trusty-mpm.md">output-style file</a>. The live instruction cites it inline: <em>&#8220;&#8230;the active palace carries an </em><code>is_fact</code><em> triple identifying this framework (see docs/specs/trusty-mpm-self-awareness.md &#167;5).&#8221;</em></p></li><li><p><strong>The ID grammar</strong> &#8212; each spec carries a <code>DOC-N</code> document number; each governed section a stable ID in the <code>SPEC-{SUBSYSTEM}-{NN}~{rev}</code> grammar (e.g. <code>SPEC-CONFORMANCE-02~draft</code>), where <code>~{rev}</code> is the revision the drift check watches. A <code>{#SPEC-&#8230;}</code> heading marker anchors each section &#8212; like an HTML <code>id</code> &#8212; so code links to the section, not the whole file.</p></li><li><p><strong>Code-to-spec links</strong> &#8212; a <code># Spec References</code> block in a module&#8217;s doc comment (the rustdoc Rust renders into API docs). From <code>front_gate.rs</code>:</p></li></ul><pre><code><code>  //! # Spec References
  //! - [`SPEC-CONFORMANCE-02~draft`](docs/specs/intent-conformance.md#SPEC-CONFORMANCE-02~draft) (&#167;5.1 FRONT gate)
  //! - [`SPEC-CONFORMANCE-01~draft`](docs/specs/intent-conformance.md#SPEC-CONFORMANCE-01~draft) (&#167;4 decision matrix)</code></code></pre><p>The reverse direction: each spec requirement ends with an &#8220;Implementing Modules&#8221; table plus inline <code>file:line</code> citations.</p><ul><li><p><strong>The parser and drift check</strong> &#8212; <code>spec_resolve.rs</code>. <code>parse_spec_refs()</code> scans Rust source for <code># Spec References</code> blocks; <code>resolve_spec_section()</code> reads the spec markdown, finds the anchored section, and extracts its &#8220;Behavior Contract&#8221; and &#8220;Rationale.&#8221; Two gates run it through the same path &#8212; <code>front_gate.rs</code> before work starts, <code>conformance.rs</code> (in trusty-review) after &#8212; so they can&#8217;t diverge on what a section requires. If code links <code>~v1</code> of a section that&#8217;s since become <code>~v2</code>, it flags <code>revision_drift = true</code> rather than trusting the citation.</p></li><li><p><strong>The CI test</strong> &#8212; <code>bundle_tests.rs</code>, test <code>output_styles_carry_identity_protocol_and_load_marker</code>, asserts every bundled output style contains the non-overridable Identity protocol section and the marker <code>&lt;!-- trusty-mpm-instructions-loaded: v1 --&gt;</code>. Drift from DOC-28&#8217;s acceptance criteria fails the PR that introduced it.</p></li><li><p><strong>The doc-comment convention</strong> &#8212; <code>harness_doc.rs</code> pairs a Why / What / Test triad inside <code>///</code> comments: the rationale plus the exact test names that enforce it. The house rule is stated in <code>docs/specs/README.md</code> &#8212; &#8220;a behavior contract&#8230; without prescribing the implementation&#8221; &#8212; and requires code to link to spec sections from the first implementation PR, even while the spec is <code>~draft</code>.</p></li></ul>]]></content:encoded></item><item><title><![CDATA[What Is Harness Engineering?]]></title><description><![CDATA[And Do You Need to Learn It?]]></description><link>https://hyperdev.matsuoka.com/p/what-is-harness-engineering</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/what-is-harness-engineering</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Fri, 10 Jul 2026 11:30:51 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!yCIE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!yCIE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yCIE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png 424w, https://substackcdn.com/image/fetch/$s_!yCIE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png 848w, https://substackcdn.com/image/fetch/$s_!yCIE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png 1272w, https://substackcdn.com/image/fetch/$s_!yCIE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yCIE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png" width="1456" height="1092" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1092,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2396213,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/206397665?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!yCIE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png 424w, https://substackcdn.com/image/fetch/$s_!yCIE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png 848w, https://substackcdn.com/image/fetch/$s_!yCIE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png 1272w, https://substackcdn.com/image/fetch/$s_!yCIE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Enhancing the Golden Egg</figcaption></figure></div><p>An engineer at work asked me a good question last week. &#8220;What&#8217;s harness engineering, and do I actually need to learn it &#8212; or is it going to be obsolete by the time I do?&#8221; Fair question. The term is about five months old, half the people using it mean different things by it, and the people who build the most capable coding agents around keep going on record to say the thing you&#8217;d build will get absorbed into the next model.</p><p>So here is my answer, stated plainly, because I think most engineers are getting it wrong: harness engineering is the most important skill you can build right now beyond coding and architecture themselves. And the strongest evidence that it matters is that most engineers don&#8217;t yet believe they need it.</p><p>I&#8217;ve been writing harnesses for over a year, so treat that as a disclosed bias rather than a neutral survey. What follows argues against the smartest version of the other side &#8212; the case that harness work is disposable scaffolding the models will eat for breakfast &#8212; because that case is largely correct, and it still doesn&#8217;t touch the skill I&#8217;m talking about.</p><h2>TL;DR</h2><ul><li><p><strong>Harness engineering is real but young.</strong> The phrase traces to <a href="https://mitchellh.com/writing/my-ai-adoption-journey">Mitchell Hashimoto in February 2026</a> and an <a href="https://openai.com/index/harness-engineering/">OpenAI Codex case study days later</a>; <a href="https://martinfowler.com/articles/harness-engineering.html">Fowler and B&#246;ckeler</a> formalized it as &#8220;Agent = Model + Harness.&#8221; Report it as an emerging frame, not a settled discipline &#8212; and no &#8220;harness engineer&#8221; job title exists yet.</p></li><li><p><strong>The labs are right that the crutch layer shrinks.</strong> Anthropic&#8217;s Cat Wu says <a href="https://www.lennysnewsletter.com/p/how-anthropics-product-team-moves">&#8220;the models will eat your harness for breakfast&#8221;</a>; Boris Cherny says scaffolding gets &#8220;pushed into the model itself.&#8221; No one at Anthropic said &#8220;don&#8217;t learn it.&#8221;</p></li><li><p><strong>The benchmark fight lives at the crutch layer.</strong> Same model, different scaffold moves scores by <a href="https://arxiv.org/abs/2606.08529">up to 28 points on GAIA</a> and <a href="https://www.tbench.ai/leaderboard/terminal-bench/2.0">18.6 points on Terminal-Bench 2.0</a> &#8212; but <a href="https://agents-last-exam.org/blogs/harness-matters">Agents&#8217; Last Exam</a> shows model choice drives roughly 3x the spread of harness choice. Scores were beside the point.</p></li><li><p><strong>The skill lives at a layer the benchmarks don&#8217;t measure.</strong> Not a better crutch &#8212; a different unit of work: research &#8594; spec &#8594; ticket &#8594; build &#8594; PR &#8594; review &#8594; merge &#8594; deploy, driven by the harness. <a href="https://openai.com/index/harness-engineering/">OpenAI ran that loop with 3 engineers to ~1M lines and 1,500 PRs</a> with zero human-written code.</p></li><li><p><strong>Harness engineering is not IDE engineering.</strong> If you treat the harness as an extension of your old editor workflow, you cap your ceiling. The durable move is letting it drive and jumping in when needed.</p></li><li><p><strong>Yes, you need to learn it.</strong> The specific syntax is throwaway. That&#8217;s exactly why the skill matters &#8212; you&#8217;re learning to operate a new unit of work, not one vendor&#8217;s config file.</p></li></ul><h2>The question, stated bluntly</h2><p>Start with what &#8220;harness engineering&#8221; is even asking you to do, because the word carries two arguments at once and people talk past each other constantly.</p><p>I already made the case that orchestration beats model quality in <a href="https://hyperdev.matsuoka.com/p/its-the-harness-stupid">It&#8217;s the Harness, Stupid</a> back in April &#8212; same model, large spread in outcomes, the competitive edge moving from model superiority to ecosystem superiority. That piece defined what a harness <em>is</em> and showed that it dominates results. I&#8217;m not going to re-argue it. This is the follow-up question that piece left open: if the harness matters that much, is <em>building</em> one a skill worth learning &#8212; or a treadmill that resets every model release?</p><p>That distinction is the whole article. Because the answer the evidence points to is: parts of it reset every release, and the part that doesn&#8217;t is the part almost nobody is naming. Most of the public argument is being had about the parts that reset.</p><h2>What a harness actually is</h2><p>The cleanest definition comes from Martin Fowler and <a href="https://martinfowler.com/articles/harness-engineering.html">Birgitta B&#246;ckeler</a>: &#8220;the harness&#8221; is everything in an AI agent except the model itself. Agent = Model + Harness. They split it into guides that push instructions forward and sensors that feed results back, and they frame the whole practice as a specific form of context engineering. <a href="https://simonwillison.net/guides/agentic-engineering-patterns/how-coding-agents-work/">Simon Willison</a> puts it the same way from the other direction: a coding agent is software that acts as a harness for an LLM.</p><p>When I say &#8220;using a harness,&#8221; here&#8217;s the concrete inventory I mean:</p><ul><li><p><strong>Agents</strong> &#8212; the loop that calls the model and routes its tool calls, plus any sub-agents you dispatch work to.</p></li><li><p><strong>Skills</strong> &#8212; reusable capabilities you can invoke by name instead of re-explaining every time.</p></li><li><p><strong>Hooks</strong> &#8212; deterministic gates that fire on events: run the tests, block a commit, reformat on save.</p></li><li><p><strong>Workflow</strong> &#8212; the orchestrated path from research to deploy, and who (or what) drives each step.</p></li><li><p><strong>Harness-specific instructions</strong> &#8212; how <em>this</em> harness should behave, kept distinct from project-specific instructions about <em>this</em> codebase. Conflating those two is one of the most common configuration mistakes I see.</p></li><li><p><strong>Memory and search</strong> &#8212; what persists across sessions, and how the agent retrieves it.</p></li></ul><p>Internalize that inventory before you touch any tool, because every product arranges these pieces differently and calls them different things. <a href="https://huggingface.co/blog/agent-glossary">Hugging Face</a> is the one source I&#8217;ve found that formally separates the <em>scaffolding</em> (the behavior layer &#8212; prompts, tool descriptions, memory) from the <em>harness</em> proper (the execution layer that calls the model and decides when to stop), then notes that most products just call the whole bundle a harness. That ambiguity is real, not something to paper over: the term is contested, <a href="https://haverin.substack.com/p/what-is-harness-engineering-ai-hype">some practitioners call it an old idea in new packaging</a>, and Latent Space literally ran a piece titled <a href="https://www.latent.space/p/ainews-is-harness-engineering-real">&#8220;Is Harness Engineering Real?&#8221;</a>. I don&#8217;t want to oversell a five-month-old buzzword. I want to separate the durable part from the disposable part, and to do that I have to give the skeptics their strongest swing first.</p><h2>&#8220;The models will eat your harness for breakfast&#8221;</h2><p>Here&#8217;s the skeptical case in its own words, and it&#8217;s a good case.</p><p>Cat Wu, who heads product for Claude Code, has a line for it: <a href="https://www.lennysnewsletter.com/p/how-anthropics-product-team-moves">the models will eat your harness for breakfast</a>. Her team does a system-prompt audit on every new model and deletes the reminders the model no longer needs. Her example: a to-do enforcement tool built to stop Claude Code from overclaiming that a refactor was finished became dead weight once newer models completed multi-step refactors on their own.</p><p>Boris Cherny, who created Claude Code, says the same thing from the architecture side. In an <a href="https://every.to/podcast/transcript-how-to-use-claude-code-like-the-people-who-built-it">Every.to interview</a>, he described the tool as &#8220;the thinnest possible wrapper over the model&#8221; and said that as models advance, &#8220;stuff that used to be scaffolding... gets pushed into the model itself.&#8221; His team builds harness features they expect to delete: &#8220;we build most things... even if that means we&#8217;ll have to get rid of it in three months. If anything, we hope that we will get rid of it in three months.&#8221; Disposable on purpose. And at Sequoia&#8217;s AI Ascent this spring he extended it forward &#8212; prompt-injection defenses, static command verification, permission modes, human-in-the-loop gates would all become less critical, he argued, &#8220;because models will do the right thing themselves.&#8221;</p><p>This isn&#8217;t only an Anthropic view. It&#8217;s the <a href="http://www.incompleteideas.net/IncIdeas/BitterLesson.html">bitter lesson</a> applied to agents: building in how we think the work should be structured tends to lose, over time, to raw capability and scale. Han Lee makes the practitioner version <a href="https://leehanchung.github.io/blogs/2026/05/08/hidden-technical-debt-agent-harness/">bluntly</a>: &#8220;Almost all of it is going to dissolve into the next generation of models... build each production harness like you mean to replace it.&#8221; Tool wrappers dissolve because models read OpenAPI specs directly now. Elaborate memory layers collapse into &#8220;plain text in progress.md plus git log.&#8221;</p><p>Two concessions I&#8217;ll make up front, because the fair version of this argument requires them. First: nobody at Anthropic said &#8220;don&#8217;t learn it.&#8221; The claim that they did is a paraphrase &#8212; a real cluster of &#8220;the harness shrinks&#8221; statements from Wu and Cherny, compressed by repetition into something stronger than anyone actually said. Second, and this is the one that stings: even the benchmark evidence <em>for</em> harnesses says model choice usually wins. On <a href="https://agents-last-exam.org/blogs/harness-matters">Agents&#8217; Last Exam</a>, swapping models with the harness fixed produced an 18-point pass-rate spread; swapping harnesses with the model fixed produced about 6. &#8220;The model accounts for about 3x the pass-rate spread of the harness.&#8221; If your goal is a higher number on the leaderboard, buy the better model before you tune the scaffold.</p><p>So the skeptics have a benchmark, a bitter lesson, and the people who build the reference implementation all pointing the same way. If I stopped here, the answer to &#8220;do you need to learn it&#8221; would be &#8220;not really &#8212; wait for the next model.&#8221; I don&#8217;t stop here, because all of that is arguing about a layer I don&#8217;t mean.</p><h2>Why both are right &#8212; and why it doesn&#8217;t touch the real skill</h2><p>The reconciliation is that &#8220;harness&#8221; names two different things, and the argument above is entirely about the first one.</p><p>The first layer is scaffolding-as-crutch. A hook that reminds the model to actually run the tests. A tool wrapper that translates an API the model can&#8217;t yet read. A permission gate that catches a mistake the model still makes. Anthropic&#8217;s own framing nails why this dissolves: <a href="https://www.anthropic.com/engineering/harness-design-long-running-apps">&#8220;every component in a harness encodes an assumption about what the model can&#8217;t do on its own.&#8221;</a> When the model can suddenly do that thing, the component becomes dead weight &#8212; exactly Wu&#8217;s deleted to-do tool. This layer is <em>supposed</em> to shrink. Anthropic builds it disposable on purpose. The benchmark gaps live here too: the <a href="https://arxiv.org/abs/2606.08529">28-point GAIA swing</a> and the <a href="https://www.tbench.ai/leaderboard/terminal-bench/2.0">18.6-point Terminal-Bench spread</a> measure how much a scaffold props up a fixed model&#8217;s score. Prop-up value falls as the model climbs. That&#8217;s the whole skeptical case, and it&#8217;s correct.</p><p>The second layer is workflow orchestration. Letting the harness drive the entire loop &#8212; research &#8594; spec &#8594; ticket &#8594; build &#8594; iterate &#8594; PR &#8594; review &#8594; merge &#8594; deploy &#8212; as one continuous operation instead of a sequence of prompts you babysit. This layer does not dissolve into a better model, because a better model doesn&#8217;t decide what&#8217;s safe to run unattended, what the blast radius of an autonomous change is, or where a human judgment call has to sit. Better models make the loop <em>run better</em>. They don&#8217;t make the loop <em>design itself</em>.</p><p>Blake Crosley draws the same line and I think it&#8217;s the sharpest version: harness <em><a href="https://blakecrosley.com/blog/loops-win-where-verification-is-cheap">syntax</a></em><a href="https://blakecrosley.com/blog/loops-win-where-verification-is-cheap"> is ephemeral and gets absorbed, but </a><em><a href="https://blakecrosley.com/blog/loops-win-where-verification-is-cheap">verification judgment</a></em><a href="https://blakecrosley.com/blog/loops-win-where-verification-is-cheap"> is durable</a> &#8212; knowing what&#8217;s safe to run without watching, what the acceptable failure modes are, where the loop needs a gate. The config file you write today is throwaway. The judgment about how to structure autonomous work is not.</p><p>The killer piece of evidence sits in the phrase&#8217;s own origin story. When <a href="https://openai.com/index/harness-engineering/">OpenAI published its harness-engineering case study</a>, the headline number was three engineers producing roughly a million lines of code across 1,500 pull requests, with zero human-written code, by engineering the harness around Codex. Read that carefully. That is not a better crutch bolted onto a fixed workflow. It&#8217;s a different unit of work &#8212; the engineers stopped writing lines and started operating a loop. No model upgrade alone produces that shape of output, because the shape is a workflow-design decision, not a capability. Ryan Lopopolo&#8217;s summary of what changed is the tell: the only scarce resource left was synchronous human attention. That&#8217;s an orchestration problem, and no amount of model progress makes it go away.</p><p>This is why I can concede the entire benchmark argument without losing anything. Model choice beats harness choice on scores &#8212; sure, roughly 3x on Agents&#8217; Last Exam. But the fight over scores is being had at the crutch layer, and the skill I mean lives at the orchestration layer, which those benchmarks don&#8217;t even measure. Even the strongest model still needs someone who knows how to hand it a whole workflow instead of a single task.</p><h2>Harness engineering vs. IDE engineering</h2><p>Here&#8217;s the crux, and it&#8217;s where I think most engineers cap their own ceiling without noticing.</p><p>The biggest difference between harness engineering and IDE engineering is what the unit of work is. In the editor era &#8212; including the AI-autocomplete-in-your-editor era &#8212; the unit is a change you make, assisted. You&#8217;re still driving. The tool suggests, you accept, you commit. A good harness inverts that. The unit becomes an <em>outcome you delegate</em>: research through deploy, handled by the harness, with you supervising the loop rather than typing inside it.</p><p>If you approach a harness as an extension of your existing editor workflow &#8212; a faster autocomplete, a smarter pair &#8212; you&#8217;ll get some lift and you&#8217;ll hit a ceiling fast, because you&#8217;re still the bottleneck on every step. The results I&#8217;ve gotten that actually surprised me came from letting the harness drive the whole loop and jumping in only where my judgment was needed: at the spec, at the review, at the &#8220;is this safe to merge&#8221; gate. That maps exactly onto Crosley&#8217;s durable skill. You&#8217;re not writing less carefully. You&#8217;re spending your attention on the decisions that don&#8217;t delegate, and letting the loop own the ones that do.</p><p>This is a materially different workflow from anything I did before agents, and I say that as someone who&#8217;s lived through several supposed paradigm shifts that turned out to be the same job with new keybindings. This one isn&#8217;t. The muscle you build isn&#8217;t &#8220;prompt the model well.&#8221; It&#8217;s &#8220;decompose an outcome into a loop a machine can run mostly unattended, and know precisely where to stand in it.&#8221; Harrison Chase, who runs LangChain, frames harness engineering as <a href="https://venturebeat.com/orchestration/langchains-ceo-argues-that-better-models-alone-wont-get-your-ai-agent-to">an extension of context engineering</a>, and that lineage is right &#8212; context engineering is about a year old and settled, harness engineering is the newer, contested layer on top. But the operative verb changed. You&#8217;re not composing a context window. You&#8217;re operating a workflow.</p><h2>Do you need to learn it? Yes &#8212; here&#8217;s how</h2><p>Yes. The fact that it still feels optional is the problem, not a reason to wait.</p><p>Here&#8217;s the path I&#8217;d suggest, and it&#8217;s roughly the one I took.</p><p><strong>Start by understanding what using a harness means</strong> &#8212; the inventory from earlier: agents, skills, hooks, workflow, harness-specific versus project-specific instructions, memory, search. Not as vocabulary. As the actual pieces you&#8217;ll arrange. If those seven words don&#8217;t map to concrete settings you can change, start there before you touch a workflow.</p><p><strong>Try several, and configure them well.</strong> Don&#8217;t judge the category from one tool on defaults. Learn to tune Claude Code yourself until it performs, rather than running it out of the box and concluding the harness &#8220;doesn&#8217;t matter.&#8221; Try <a href="https://openai.com/index/harness-engineering/">Codex</a>, Gemini, Auggie, OpenCode. I built <a href="https://github.com/bobmatnyc/claude-mpm">claude-mpm</a>, so weight my enthusiasm for the multi-agent approach accordingly &#8212; the point isn&#8217;t which one wins, it&#8217;s that you can&#8217;t feel the shape of the skill from a single vendor&#8217;s config.</p><p><strong>Lean into the differences instead of smoothing them over.</strong> The instinct is to find the tool that feels most like your old editor and stop. The results live in the opposite direction &#8212; in the workflows that feel least familiar, where the harness drives and you supervise. Best results come from leaning into what&#8217;s different, not translating it back into what you already knew.</p><p><strong>Let it drive, and know where to stand.</strong> This is the whole skill in one sentence. Hand the loop the outcome, supervise at the gates your judgment actually owns, jump in when the blast radius or the ambiguity demands it. That standing-in-the-right-place instinct is what transfers across every tool and survives every model release.</p><p>And that&#8217;s the reframe I&#8217;ll close on. Everything about the <em>specific</em> harness you learn this quarter is throwaway. The config syntax, the exact hooks, the tool wrappers &#8212; Han Lee is right, the next model eats away at it. But that&#8217;s precisely why the skill is worth building, not a reason to skip it. You&#8217;re not learning one vendor&#8217;s settings file. You&#8217;re learning to operate a new unit of work &#8212; an autonomous loop from research to deploy &#8212; and that competence is the thing the model upgrades keep <em>raising the value of</em>, not erasing. The syntax is disposable. The judgment about how to run the loop is what compounds.</p><p>Most engineers will figure this out eventually, when the workflow shift is obvious in hindsight. The ones who figure it out now get a head start measured in the gap between &#8220;my editor got smarter&#8221; and &#8220;my unit of work changed.&#8221; I&#8217;d rather be early on that one.</p><h2>One layer up: loop engineering</h2><p>There&#8217;s a move past the harness, and it picked up a name while I was writing this.</p><p>The stack is starting to read like a ladder: prompt engineering, then context engineering, then harness engineering &#8212; and now loop engineering on top. Each rung stops being where you spend your attention once the rung below it gets good enough to trust. You quit hand-tuning prompts when context engineering settled into a roughly year-old, mostly-solved practice. The bet underneath loop engineering is the same shape one level up: once your harness is solid, you stop prompting the agent and start designing the loops that prompt it for you.</p><p>The term isn&#8217;t mine. <a href="https://addyosmani.com/blog/loop-engineering/">Addy Osmani formalized it in June</a>, and his definition is the one I&#8217;d hand someone first: &#8220;Loop engineering is replacing yourself as the person who prompts the agent. You design the system that does it instead.&#8221; He also collected two lines that land the shift faster than I can. Peter Steinberger&#8217;s version: you should be designing the loops that prompt your agents. And Boris Cherny &#8212; the same Cherny from the &#8220;eat your harness for breakfast&#8221; section &#8212; put it flatly: &#8220;I don&#8217;t prompt Claude anymore&#8230; my job is to write loops.&#8221; <a href="https://www.langchain.com/blog/the-art-of-loop-engineering">LangChain picked it up a week later</a>, and their framing is the practical one: you don&#8217;t build a loop, you stack them &#8212; an agent loop inside a verification loop inside an event-driven loop inside a hill-climbing loop, each one checking the one below it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!KNlZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3fd96b6-ed6f-43c1-be83-2956e9647562_1536x1152.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!KNlZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3fd96b6-ed6f-43c1-be83-2956e9647562_1536x1152.png 424w, https://substackcdn.com/image/fetch/$s_!KNlZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3fd96b6-ed6f-43c1-be83-2956e9647562_1536x1152.png 848w, https://substackcdn.com/image/fetch/$s_!KNlZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3fd96b6-ed6f-43c1-be83-2956e9647562_1536x1152.png 1272w, https://substackcdn.com/image/fetch/$s_!KNlZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3fd96b6-ed6f-43c1-be83-2956e9647562_1536x1152.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!KNlZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3fd96b6-ed6f-43c1-be83-2956e9647562_1536x1152.png" width="1456" height="1092" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d3fd96b6-ed6f-43c1-be83-2956e9647562_1536x1152.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1092,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2248716,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/206397665?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3fd96b6-ed6f-43c1-be83-2956e9647562_1536x1152.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!KNlZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3fd96b6-ed6f-43c1-be83-2956e9647562_1536x1152.png 424w, https://substackcdn.com/image/fetch/$s_!KNlZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3fd96b6-ed6f-43c1-be83-2956e9647562_1536x1152.png 848w, https://substackcdn.com/image/fetch/$s_!KNlZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3fd96b6-ed6f-43c1-be83-2956e9647562_1536x1152.png 1272w, https://substackcdn.com/image/fetch/$s_!KNlZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3fd96b6-ed6f-43c1-be83-2956e9647562_1536x1152.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Loop Engineering</figcaption></figure></div><p>Here&#8217;s the seam between the two layers, drawn plainly. Harness engineering is the discipline of building the scaffolding &#8212; the agents, hooks, skills, gates, and workflow from the inventory up top. Loop engineering is what you do once that scaffolding is good enough that you stop typing prompts and start directing repeatable cycles. The harness is what you build. The loop is what you run on it, again and again, with your attention moved up to which loops to run and when to trust them unattended. That&#8217;s the same &#8220;let it drive, know where to stand&#8221; instinct from a few paragraphs back, pushed one rung higher: now you&#8217;re not standing inside the loop at all &#8212; you&#8217;re choosing which loops get to run.</p><p>I&#8217;m not planting a flag here. The people already naming it are out ahead of me, and that&#8217;s the reason to point at it &#8212; this is where the harness work goes next, not a term I&#8217;m coining. If harness engineering closes the gap between &#8220;my editor got smarter&#8221; and &#8220;my unit of work changed,&#8221; loop engineering is what shows up on the far side of that gap, once the unit of work is a cycle you supervise instead of a task you run. I&#8217;m watching that one closely. And I&#8217;d start learning it before it feels obvious.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and also writes about AI business at <a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">AI Power Ranking</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/p/its-the-harness-stupid">It&#8217;s the Harness, Stupid</a> &#8212; The predecessor to this piece: same model, wide spread in outcomes, and why the competitive edge moved from model quality to orchestration. It defined what a harness is; this piece argues that building one is a skill worth learning.</p></li><li><p><a href="https://hyperdev.matsuoka.com/p/hyperdevs-three-golden-rules">HyperDev&#8217;s Three Golden Rules</a> &#8212; The working rules I keep coming back to for professional AI work, and the discipline that keeps a driven-by-the-harness loop from running off the rails.</p></li><li><p><a href="https://hyperdev.matsuoka.com/p/claude-sonnet-5-takes-the-default">Claude Sonnet 5 Takes the Default Driver Slot</a> &#8212; A concrete example of the crutch layer shrinking: adaptive thinking folds interleaved reasoning into the model, removing work harness authors used to do by hand.</p></li><li><p><a href="https://addyosmani.com/blog/loop-engineering/">Loop Engineering</a> &#8212; Addy Osmani&#8217;s June 2026 piece that named the layer above the harness: once the scaffolding holds, you stop prompting the agent and design the loops that prompt it. The forward edge of the arc this article traces.</p></li><li><p><a href="https://www.langchain.com/blog/the-art-of-loop-engineering">The Art of Loop Engineering</a> &#8212; LangChain&#8217;s treatment of loop engineering as stacked loops &#8212; agent, verification, event-driven, hill-climbing &#8212; each one checking the one beneath. The practitioner&#8217;s map of where harness work heads next.</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners, including the coding-agent leaderboards this piece leans on.</p></li></ul>]]></content:encoded></item><item><title><![CDATA[Claude Sonnet 5 Takes the Default Driver Slot — and Quietly Raises Your Token Bill]]></title><description><![CDATA[Plus it's baaaaaack! (Fable 5)]]></description><link>https://hyperdev.matsuoka.com/p/claude-sonnet-5-takes-the-default</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/claude-sonnet-5-takes-the-default</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Thu, 02 Jul 2026 19:10:42 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!qfHk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c96d654-2776-43ee-897b-b27999367bde_1536x1152.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>Verdict up front</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qfHk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c96d654-2776-43ee-897b-b27999367bde_1536x1152.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qfHk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c96d654-2776-43ee-897b-b27999367bde_1536x1152.png 424w, https://substackcdn.com/image/fetch/$s_!qfHk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c96d654-2776-43ee-897b-b27999367bde_1536x1152.png 848w, https://substackcdn.com/image/fetch/$s_!qfHk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c96d654-2776-43ee-897b-b27999367bde_1536x1152.png 1272w, https://substackcdn.com/image/fetch/$s_!qfHk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c96d654-2776-43ee-897b-b27999367bde_1536x1152.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qfHk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c96d654-2776-43ee-897b-b27999367bde_1536x1152.png" width="1536" height="1152" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5c96d654-2776-43ee-897b-b27999367bde_1536x1152.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:&quot;normal&quot;,&quot;height&quot;:1152,&quot;width&quot;:1536,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:0,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!qfHk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c96d654-2776-43ee-897b-b27999367bde_1536x1152.png 424w, https://substackcdn.com/image/fetch/$s_!qfHk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c96d654-2776-43ee-897b-b27999367bde_1536x1152.png 848w, https://substackcdn.com/image/fetch/$s_!qfHk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c96d654-2776-43ee-897b-b27999367bde_1536x1152.png 1272w, https://substackcdn.com/image/fetch/$s_!qfHk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c96d654-2776-43ee-897b-b27999367bde_1536x1152.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Claude Sonnet 5 landed this week, June 30, 2026, as the new default on Claude Free and Pro and the new occupant of the middle tier: Haiku below it, Opus 4.8 above, Fable 5 and Mythos 5 at the top. The pitch from <a href="https://www.anthropic.com/news/claude-sonnet-5">Anthropic</a> is the one practitioners care about &#8212; Opus-adjacent capability at Sonnet economics, aimed squarely at high-volume agent loops where Opus pricing compounds fast. One wrinkle sharpens the switch decision. Fable 5 is back in Claude Code, redeployed the morning after Sonnet shipped, which puts the top of the lineup back in reach and moves the ceiling any switch has to weigh against.</p><p>Two things matter more than the headline. First, the architecture shift: extended thinking is gone, replaced by adaptive thinking that runs on by default and interleaves reasoning between tool calls. The <code>effort</code> parameter now defaults to <code>high</code> in both the Claude API and Claude Code. Second, the pricing has a catch the announcement does not foreground. The list price looks flat against Sonnet 4.6. The new tokenizer means your invoice may not be.</p><p>If you drive coding agents for a living, this is the model you will be pointing your harness at by default. The question is whether to switch today[1], reach past it, or wait two weeks and measure.</p><h2>TL;DR</h2><ul><li><p><strong>New default everywhere.</strong> Sonnet 5 is the default model on Free and Pro at launch. API ID <code>claude-sonnet-5</code>, Bedrock <code>anthropic.claude-sonnet-5</code>, Vertex coming soon.</p></li><li><p><strong>Extended thinking removed; adaptive thinking is always on.</strong> Manual <code>thinking: {type: "enabled", budget_tokens: N}</code> now returns a 400 error. Depth is set by <code>effort</code> (<code>low</code>, <code>medium</code>, <code>high</code>, <code>xhigh</code>, <code>max</code>; <code>high</code> by default on the API and in Claude Code), and the model reasons between tool calls without prompt engineering &#8212; the architecturally significant change for agentic work.</p></li><li><p><strong>1M context, 128k output, January 2026 cutoff.</strong> Repo-level reasoning at Sonnet pricing.</p></li><li><p><strong>Agentic coding 63.2%</strong>, per Anthropic&#8217;s comparison chart, versus 58.1% for Sonnet 4.6 and 69.2% for Opus 4.8. Likely SWE-bench Pro, not Verified &#8212; and no Verified score is published yet.</p></li><li><p><strong>Pricing looks flat, isn&#8217;t quite.</strong> $2/$10 per MTok intro through August 31, then $3/$15 &#8212; same list price as Sonnet 4.6. But the new tokenizer generates 1.0&#8211;1.35x more tokens for the same text, so &#8220;same price&#8221; is misleading at the invoice level.</p></li><li><p><strong>Fable 5 is back in Claude Code (July 1)</strong> as the practical ceiling above Opus &#8212; priced higher, with a retrained safety classifier that reroutes offensive-cyber work down to Opus 4.8. Details below.</p></li></ul><h2>What Anthropic shipped</h2><p>Sonnet 5 sits in the workhorse slot of the lineup. Haiku 4.5 at $1/$5 handles cheap high-volume work. Opus 4.8 at $5/$25 is the heavy lifter. Fable 5 and Mythos 5 sit above Opus, both dark at Sonnet 5&#8217;s launch, though Fable returned to Claude Code on July 1 (more below). Sonnet is the tier you point at the bulk of your agent traffic, and the one whose economics decide whether a coding-agent product is viable.</p><p>The specs are current-generation and unsurprising on paper: 1,000,000-token context window, 128,000 max output tokens (300,000 on the Batch API with the beta header), text and image input, January 2026 knowledge cutoff. Anthropic calls it &#8220;the most agentic Sonnet model yet,&#8221; which is marketing, but the architecture underneath the phrase is the actual story.</p><p><strong>Extended thinking is gone.</strong> If your code sends <code>thinking: {type: "enabled", budget_tokens: N}</code>, Sonnet 5 returns a 400 error. That is a breaking change for any harness that sets thinking budgets explicitly. Grep your codebase for it before you flip the model ID. In its place is <a href="https://platform.claude.com/docs/en/build-with-claude/adaptive-thinking">adaptive thinking</a>: always on, with depth allocated dynamically by the model and steered through the <code>effort</code> parameter &#8212; <code>low</code>, <code>medium</code>, <code>high</code>, <code>xhigh</code>, or <code>max</code>. The default is <code>high</code> on both the API and Claude Code, with <code>xhigh</code> available above it for the hardest coding and agentic tasks.</p><p>The piece that matters for agents is that adaptive thinking automatically enables interleaved thinking. The model reasons between tool calls. It reflects on what a tool returned before deciding the next action, and it does this without you wiring up a scratchpad or a reflection prompt. For anyone who has built an agent loop by hand, that is the part you used to engineer yourself, now folded into the default behavior of the model.</p><h2>About that benchmark number</h2><p>Anthropic&#8217;s launch chart gives an agentic coding comparison:</p><p>Model Agentic coding score Claude Sonnet 4.6 58.1% Claude Sonnet 5 63.2% Claude Opus 4.8 69.2%</p><p>Source: <a href="https://techcrunch.com/2026/06/30/anthropic-launches-claude-sonnet-5-as-a-cheaper-way-to-run-agents/">Anthropic&#8217;s comparison chart, via TechCrunch</a>. Now the caveats, because this is where launch-day coverage tends to get sloppy.</p><p>That 63.2% is almost certainly SWE-bench Pro &#8212; the harder multi-file agentic eval &#8212; not SWE-bench Verified. The two are not interchangeable, and the numbers live on different scales. The Opus 4.8 figure in the same chart (69.2%) matches its published SWE-bench Pro score, which supports the Pro read. Anthropic has <strong>not</strong> published a standalone SWE-bench Verified number for Sonnet 5. The comparison chart is an image, with no underlying table released at launch.</p><p>So here is the trap. If you go looking, you will find a 82.1% SWE-bench figure attached to &#8220;Claude Sonnet 5.&#8221; Do not use it. That number comes from February 2026 pre-launch speculation about a different model iteration &#8212; a phantom Sonnet 5 with a &#8220;Dev Team Mode&#8221; that never shipped in the form described. The model that launched today has different characteristics. Any spec sheet dated before June 30 is describing something else.</p><p>What can you say with confidence? Sonnet 5 sits meaningfully above Sonnet 4.6 on agentic coding and a few points below Opus 4.8. For reference, the most recent <a href="https://www.morphllm.com/claude-benchmarks">third-party leaderboard before launch</a> had Sonnet 4.6 at 79.6% SWE-bench Verified and Opus 4.8 at 88.6%. A reasonable expectation puts Sonnet 5&#8217;s Verified score somewhere in the low-to-mid 80s &#8212; but that is my read of where the gap lands, not a number Anthropic has confirmed. Treat it as judgment, not data.</p><p>On the safety side, Anthropic reports lower hallucination and sycophancy rates than Sonnet 4.6, better refusal of malicious requests, and stronger resistance to prompt injection. By design, it also carries intentionally weaker cybersecurity exploit capability than Opus 4.8.</p><h2>Pricing and the tokenizer tax</h2><p>The list price is the easy part:</p><p>Model Input $/MTok Output $/MTok Claude Haiku 4.5 $1.00 $5.00 Sonnet 5 (intro, through Aug 31) $2.00 $10.00 Claude Sonnet 4.6 $3.00 $15.00 Sonnet 5 (standard, from Sep 1) $3.00 $15.00 Claude Opus 4.8 $5.00 $25.00</p><p>At standard rates, Sonnet 5 carries the identical list price to Sonnet 4.6: $3/$15. Through August 31 you get an introductory $2/$10. Read quickly, that says &#8220;same price, free discount for two months.&#8221; Read the footnote and it says something else.</p><p>Sonnet 5 uses the same tokenizer as Opus 4.7, which encodes the same text into <a href="https://www.anthropic.com/news/claude-sonnet-5">roughly 1.0&#8211;1.35x more tokens</a> than pre-Opus-4.7 models, depending on content type. List price is per token. So at the standard September rate, identical workloads can cost up to 35% more than they did on Sonnet 4.6 &#8212; same sticker, more tokens on the meter. The introductory pricing is doing real work here: the 33% input discount roughly offsets the tokenizer inflation through the end of August, which is almost certainly the point. Come September 1, the offset disappears and the effective increase shows up on the invoice.</p><p>The practical move is dull and necessary. Run a representative slice of your actual workload through Sonnet 5 and measure token consumption directly. Do not assume cost neutrality from the matching list price. The teams that got burned on the 4.6-to-4.7 transition were the ones who read the sticker and skipped the meter.</p><h2>Where it sits in the market</h2><p>The framing that fits the launch is that the workhorse-tier fight has moved off &#8220;which model is smartest&#8221; and onto &#8220;how cheaply and reliably does it run without a human watching.&#8221; That is the right frame for anyone shipping coding agents, where margin lives or dies on the per-task cost of the driver model.</p><p>Anthropic positions Sonnet 5 as cheaper than OpenAI&#8217;s GPT-5.5 and Google&#8217;s Gemini 3.1 Pro at introductory rates, and more expensive than Gemini 3.5 Flash, which holds the budget slot. I&#8217;d treat the competitor pricing as directional rather than precise &#8212; those comparisons come from launch coverage, not from cross-checked primary pricing pages, and competitor list prices move. The shape is more reliable than the digits: Sonnet 5 is priced to undercut the premium workhorses and to sit above the bargain tier, which is exactly where Anthropic wants the default coding driver to land.</p><p>The use case Anthropic names is the one this audience will actually hit: high-volume agent loops where Opus pricing compounds. CI review bots. Test generation. Batch transforms. Multi-step autonomous workflows that run for a while without supervision. Daniel Shepard at Zapier told <a href="https://techcrunch.com/2026/06/30/anthropic-launches-claude-sonnet-5-as-a-cheaper-way-to-run-agents/">TechCrunch</a> that a two-part automation which &#8220;used to stall halfway&#8221; now finishes end to end. A single data point, and a vendor-friendly one, but it points at the right capability: completion of long autonomous chains, not raw single-shot smarts.</p><h2>What this means for agentic coding</h2><p>The interesting architecture is the harness story. Default-high effort plus interleaved reasoning means the model arrives already configured for the agent pattern most teams hand-build: think, act, observe the result, reflect, act again. You do not prompt-engineer your way to a reflective agent anymore. It is the out-of-box behavior. For harness authors, that shifts the work &#8212; less coaxing the model into reasoning between steps, more managing the effort dial so a 12-file rename does not quietly run at <code>high</code> and burn tokens it never needed.</p><p>Watch that dial. <code>effort</code> defaults to <code>high</code>, and high effort spends more tokens than the job often warrants. The same lesson from the Opus 4.8 cycle applies: the cost surprises come from leaving the reasoning budget maxed on work that did not need it. Drop to <code>medium</code> or <code>low</code> for mechanical tasks. Reserve <code>high</code> for the hard ones.</p><p>And the 1M context window at Sonnet pricing is the underrated piece. Repo-level reasoning &#8212; cross-file refactors, full-codebase audits &#8212; without chunking, on the model you were already going to run for volume.</p><h2>The ceiling moved: Fable 5 is back</h2><p>Then there is the ceiling, and in the days around this launch it moved. The most capable public Claude model most teams could actually run was Opus 4.8, at roughly 88.6% SWE-bench Verified on the <a href="https://www.morphllm.com/claude-benchmarks">Morph leaderboard</a>. Fable 5 sits higher on that same board, in the mid-90s, though that figure is a third-party read and Anthropic has published no Verified score of its own. What moved here was less capability than access: at Sonnet 5&#8217;s June 30 launch, Fable was dark. Anthropic had <a href="https://www.anthropic.com/news/fable-mythos-access">suspended Fable 5 and Mythos 5</a> on June 12 to comply with a U.S. export-control order, and with no way to verify user nationality in real time, it pulled both models for all users.</p><p>That reversed a day later. On June 30 the Commerce Department <a href="https://www.cnbc.com/2026/06/30/anthropic-says-trump-admin-has-lifted-export-controls-on-claude-fable-5-and-mythos-5.html">lifted the order</a>, and on July 1 Anthropic <a href="https://www.anthropic.com/news/redeploying-fable-5">redeployed Fable 5 globally</a> across the Claude Platform, Claude.ai, Claude Code, and Cowork. As I write this, my Claude Code <code>/model</code> picker offers Fable 5, and I&#8217;ve set it as my default for new sessions. So the practical ceiling is no longer Opus. It is Fable &#8212; with two asterisks.</p><p>The first is access and price. Through July 7, Fable 5 counts against up to 50% of weekly usage limits on Pro, Max, Team, and select Enterprise plans; after that it shifts to usage credits, and cloud-provider access on AWS, Google Cloud, and Microsoft Foundry is still being re-enabled in phases. Per token it prices well above the Opus tier: Anthropic&#8217;s model docs list it at $10/$50 per MTok, double Opus 4.8&#8217;s $5/$25, and the new tokenizer encodes the same text into more tokens again, so the effective gap is wider than the sticker. This is not the model you point at bulk agent traffic. It is the one you reach for on the problems that stall everything below it.</p><p>The second asterisk is the reason it came back at all, and it lands squarely in coding work. The suspension traced to a <a href="https://thehackernews.com/2026/07/anthropic-restores-claude-fable-5-after.html">report from Amazon researchers</a> who found a prompt that got Fable 5 to identify software vulnerabilities and start describing how one could be exploited, before its guardrails blocked the attempt from reaching a working exploit. Anthropic&#8217;s fix was not to weaken the model but to retrain the safety classifier sitting in front of it, and the new one <a href="https://www.anthropic.com/news/redeploying-fable-5">blocks that specific technique in more than 99% of cases</a>. When the classifier fires, the request doesn&#8217;t error. It <a href="https://support.claude.com/en/articles/15363606-why-claude-switched-models-in-your-conversation-with-fable-5">reroutes to Opus 4.8 in the same session</a>, re-run and labeled with the model that actually answered. The documented trigger categories are offensive cybersecurity work (building exploits, malware, or attack tooling), plus most biology and chemistry, distillation attacks on Fable itself, and frontier-model development.</p><p>Here is where it gets concrete. Ask Fable 5 to write exploit code against a security vulnerability and you can watch it hand the task down to Opus 4.8 mid-session, a quieter, older model finishing what the newer one declined. For defensive security work that reversion is a real cost, not a hypothetical. Anthropic is candid that the tighter cyber filter routes more benign coding and debugging requests to Opus than teams would like, and developers have already <a href="https://github.com/anthropics/claude-code/issues/67305">reported the classifier over-firing</a> on routine defensive-security work, auto-switching to Opus on tasks like CVE triage. The billing follows the block: a request stopped on input is charged at Opus rates; one stopped midstream bills Fable rates for the tokens already produced, then Opus for the rest.</p><h2>Should you switch your default driver?</h2><p>For most coding-agent work, yes &#8212; but measure first, and mind the calendar.</p><p><strong>Switch now</strong> if you are running Opus 4.8 on tasks that do not strictly need it. The capability floor rose; a real share of Opus traffic will run acceptably on Sonnet 5 at a meaningful discount, and the intro pricing through August 31 makes the test cheap. This is the clearest win in the release.</p><p><strong>Switch now, with care</strong> if you are upgrading from Sonnet 4.6. You get better agentic coding, interleaved reasoning by default, lower hallucination and sycophancy. But grep for explicit <code>thinking</code> budgets first &#8212; they will 400 &#8212; and run your token measurement before September 1, when the tokenizer tax stops being masked by the intro discount.</p><p><strong>Wait, or stay put</strong> if you are on Priority Tier with Sonnet 4.6 &#8212; it is not available on Sonnet 5 at launch, so the choice is staying on 4.6 or jumping to Opus 4.8. And if your cost model is tight and unmeasured, the matching list price is a trap; spend the two weeks measuring before you commit budget to it.</p><p>There is now a fourth move the launch-day framing didn&#8217;t include: <strong>reach past Sonnet entirely.</strong> With Fable 5 back in Claude Code as of July 1, the hardest problems that Sonnet 5 and even Opus 4.8 stall on have a home again, at a price that keeps it off your bulk traffic and with a security governor that bounces exploit work down to Opus. For volume, Sonnet 5 is the driver. For the few tasks that justify the top of the lineup, the top of the lineup is reachable again.</p><p>Anthropic built Sonnet 5 to be the default driver model for the agent era, and on the architecture, it earns the slot. Adaptive thinking with interleaved reasoning is the correct shape for a coding agent, and shipping it as the default behavior removes work that harness authors used to do by hand. The asterisk is the invoice. The list price tells you one thing; the tokenizer tells you another. For the next two months, the introductory rate hides the difference. After that, the teams that measured will know what they are paying for, and the teams that read the sticker will get a surprise in their September bill.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and also writes about AI business at <a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">AI Power Ranking</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/p/i-tracked-every-token">I Tracked Every Token</a> &#8212; What a $1.07 bug fix reveals about AI coding economics, and why per-token pricing hides the real bill.</p></li><li><p><a href="https://hyperdev.matsuoka.com/p/the-agent-unlock-why-opus-45-changed">The Agent Unlock: Why Opus 4.5 Changed How I Work</a> &#8212; When a top-tier model crossed the line into autonomous coding that holds up under real work.</p></li><li><p><a href="https://hyperdev.matsuoka.com/p/article-opus-46-and-agent-teams">Breaking: Opus 4.6 and Agent Teams</a> &#8212; A model release read for what it changes in day-to-day agent workflows, not just the benchmark chart.</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul><p>[1]: Complicating &#8220;today&#8221;: with Fable 5 back as a Claude Code option (and, for some of us, the new default for hard sessions), the switch question isn&#8217;t only Sonnet-5-or-not. It&#8217;s which tier each task belongs to, Fable included. See &#8220;The ceiling moved: Fable 5 is back&#8221; below.</p>]]></content:encoded></item><item><title><![CDATA[Coding's Great Depresh ]]></title><description><![CDATA[&#8212; and How to Find Your Energesh]]></description><link>https://hyperdev.matsuoka.com/p/codings-great-depresh</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/codings-great-depresh</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Mon, 29 Jun 2026 12:16:41 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!4RhI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4RhI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4RhI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!4RhI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!4RhI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!4RhI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4RhI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1503107,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/203766515?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4RhI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!4RhI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!4RhI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!4RhI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In 2019, comedian Gary Gulman released an HBO special called <em><a href="https://www.imdb.com/title/tt10409666/">The Great Depresh</a></em>. The title puns on the 1930s collapse, but the special is about Gulman&#8217;s own clinical depression &#8212; the hospitalization, the electroconvulsive therapy, the long climb back. He calls it &#8220;depresh&#8221; throughout. The diminutive is the point: &#8220;depression&#8221; was too heavy to say out loud for years, so he found a smaller word that let him talk about it at all. The special ends on a line he delivers like a weather report: &#8220;My depresh is in remish.&#8221;</p><p>Something adjacent is moving through the developer community, and most people don&#8217;t have a name for it yet. It isn&#8217;t burnout, and it isn&#8217;t fear of replacement, though that gets blamed for it. It&#8217;s a specific malaise in senior engineers who, by outward measures, are doing fine &#8212; shipping more, faster, with tools that work. They feel worse anyway. Call it the Coding Depresh.</p><p>The Depresh is real, it&#8217;s documented, and there&#8217;s a path out that doesn&#8217;t require pretending the loss isn&#8217;t a loss. I write from an unusual position. I spent eight years as CTO of TripAdvisor managing more than 560 engineers, mostly not writing code, and I missed it. I came back to hands-on coding in March 2025, entirely through AI tools &#8212; my re-entry and AI-assisted development are inseparable. I never had to grieve a pre-AI coding identity, because I didn&#8217;t have one to defend. That gives me a strange vantage point on the engineers who did.</p><h2>TL;DR</h2><ul><li><p>The Coding Depresh is identity disconfirmation, not job-loss fear: when the thing you built your professional self around gets automated, &#8220;who am I as a developer?&#8221; stops having an easy answer.</p></li><li><p>It&#8217;s documented. A 15-year veteran describes shipping an AI-built pipeline and feeling grief &#8220;for an identity.&#8221; Stack Overflow&#8217;s 2025 survey shows distrust in AI accuracy (46%) outrunning trust (33%) for the first time, even as usage climbed to 84%.</p></li><li><p>The most-cited evidence that AI slows experienced developers &#8212; METR&#8217;s 2025 trial showing them ~19% slower &#8212; has been retired by the same researchers, whose redesigned data now points the other way. Even the skeptics&#8217; own number moved.</p></li><li><p>A different camp &#8212; the Energesh &#8212; reports feeling more capable, not less. The fault line isn&#8217;t seniority or skill. It&#8217;s your answer to &#8220;who are you as a developer.&#8221;</p></li><li><p>The most useful question I&#8217;ve seen comes from an Anthropic engineer: &#8220;I thought that I really enjoyed writing code, and I think instead I actually just enjoy what I <em>get</em> out of writing code.&#8221; Where you land tells you where you live.</p></li><li><p>For engineers in the Depresh, and the CTOs managing them: the craft instinct that makes the tools feel wrong is the most valuable thing you own. It doesn&#8217;t have to die for you to adopt the tools.</p></li></ul><h2>The Depresh is real, and it&#8217;s documented</h2><p>Start with the most precise account I&#8217;ve found. <a href="https://medium.com/codetodeploy/ai-existential-dread-and-developer-ego-death-aef8bfc93214">George Violaris</a>, a developer with fifteen years&#8217; experience, wrote in March 2026 about shipping a data pipeline with AI assistance in an afternoon &#8212; work that would have taken him three days by hand. He expected pride. He got this instead:</p><blockquote><p>&#8220;That evening, I felt something I can only describe as grief. Not for a person. For an identity.&#8221;</p></blockquote><p>He goes on: &#8220;Three years ago, writing code wasn&#8217;t just what I did. It was who I was.&#8221; Then: &#8220;The ego death was real. &#8216;I am a person who writes excellent code&#8217; had to die.&#8221;</p><p>This is not a man worried about his next paycheck. He shipped the thing. The pipeline is in production. What broke wasn&#8217;t his employment; it was the relationship between his sense of self and the work. Most of the public conversation treats developer anxiety as a labor-market story &#8212; real, severe for junior developers, but separate, and not the one I&#8217;m writing about.</p><p>The Depresh I&#8217;m describing hits a different person: the experienced developer whose answer to &#8220;who are you?&#8221; was &#8220;I am the person who writes excellent code.&#8221; For that person, the tools don&#8217;t threaten the job first. They threaten the identity first.</p><p>The numbers underneath the mood are strange. <a href="https://survey.stackoverflow.co/2025/ai/">Stack Overflow&#8217;s 2025 Developer Survey</a>, the largest census of working developers we have, shows two lines crossing. Favorable sentiment toward AI tools fell from 77% in 2023 to 60% in 2025, while usage or intent to use rose from 70% to 84%. People are adopting tools they feel worse about. Trust in AI accuracy dropped to 33% against 46% who distrust it &#8212; the first year distrust outran trust &#8212; and the top frustration, cited by 66%, was &#8220;AI solutions that are almost right, but not quite.&#8221; The most experienced developers are the most skeptical: only 2.5% &#8220;highly trust&#8221; the output, and use it all day anyway. Working with something you don&#8217;t trust, that demands verification you can&#8217;t delegate &#8212; that&#8217;s the exhaustion of vigilance without resolution.</p><p>Then there&#8217;s the METR study &#8212; less for the number most people quoted than for what happened to it. In 2025 the nonprofit METR <a href="https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/">ran a controlled trial</a> with experienced open-source developers and found them 19% slower with AI assistants while they believed they&#8217;d been roughly 20% faster. That became the most-cited evidence that the tools don&#8217;t pay off. Then in February 2026 the same team <a href="https://metr.org/blog/2026-02-24-uplift-update/">retired it</a>, writing that the original no longer reflects &#8220;the current impact of AI models on open-source developer productivity&#8221;; their redesigned measurement runs the other way, an estimated 18% speedup for the same returning developers. The intervals are wide, so the magnitude is soft &#8212; but the direction reversed, and why it had to be rebuilt is the tell: 30 to 50% of developers refused to submit tasks they expected AI to speed up, and many refused to work without AI at all. The measurement broke down because people won&#8217;t give up the tools. I read the original not as &#8220;the tools are bad&#8221; but as disorientation: when your instincts and the stopwatch disagree, then a year later the stopwatch reverses, your competence comes unmoored.</p><h2>There is another camp, and it&#8217;s also documented</h2><p>Here is what keeps the Depresh from being the whole story. A different group of engineers is having close to the opposite experience, and they&#8217;re not naive optimists or vendors.</p><p>Andrej Karpathy <a href="https://x.com/karpathy/status/1886192184808149383">named the moment in February 2025</a>: &#8220;a new kind of coding I call vibe coding, where you fully give in to the vibes, embrace exponentials, and forget that the code even exists.&#8221; He was describing a feeling, not a methodology. Simon Willison later coined <a href="https://simonwillison.net/2025/Oct/7/vibe-engineering/">&#8220;vibe engineering&#8221;</a> for the responsible professional version &#8212; engineers using these tools deliberately to accelerate real work.</p><p>The output side is loud. Pieter Levels built a <a href="https://levels.io/fly-pieter-com-vibecoded-flight-simulator">browser flight simulator</a> with no game-development background and reached around $1M in annualized revenue in 17 days. <a href="https://newsletter.pragmaticengineer.com/p/building-claude-code-with-boris-cherny">Boris Cherny</a>, who leads Claude Code at Anthropic, runs five parallel instances and ships 20 to 30 pull requests a day: &#8220;once there is a good plan, it will one-shot the implementation almost every time.&#8221;</p><p>This is where my own story sits. When I came back, the mechanics of typing code line by line weren&#8217;t part of my working identity &#8212; I&#8217;d been away from the keyboard for years. What I found waiting was the part I&#8217;d missed: building things, deciding what should exist, watching it take shape, fixing what&#8217;s wrong with it. The joy of building and the mechanics of writing code were never the same thing.</p><p>I can be concrete, because I&#8217;ve done it in the open. Since March 2025 I&#8217;ve shipped real, full-cycle software through these tools &#8212; not snippets, complete projects. The flagship is <a href="https://github.com/bobmatnyc/claude-mpm">claude-mpm</a>, a multi-agent orchestration platform for Claude with a real user base. Most of 2025 was Python; then in 2026 I taught myself Rust to build the trusty-* ecosystem &#8212; code search, memory, PR review, orchestration &#8212; something I wouldn&#8217;t have attempted by hand while running an org. Architecture, review, releases, the full loop, at a volume I couldn&#8217;t reach typing every line. Most of it is public on GitHub under <a href="https://github.com/bobmatnyc">bobmatnyc</a>, so the claim is checkable. I&#8217;m not theorizing about the Energesh. I&#8217;m living in it.</p><h2>What actually separates the two camps</h2><p>Not seniority. Not raw skill. Not the domain you work in, though each shades the experience. The fault line is the answer you&#8217;d have given, before any of this started, to a single question: who are you as a developer? There&#8217;s an older version of the same split.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xDfL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbba5a639-91a8-4072-939c-05ed95492b5b_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xDfL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbba5a639-91a8-4072-939c-05ed95492b5b_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!xDfL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbba5a639-91a8-4072-939c-05ed95492b5b_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!xDfL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbba5a639-91a8-4072-939c-05ed95492b5b_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!xDfL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbba5a639-91a8-4072-939c-05ed95492b5b_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xDfL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbba5a639-91a8-4072-939c-05ed95492b5b_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bba5a639-91a8-4072-939c-05ed95492b5b_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1319015,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/203766515?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbba5a639-91a8-4072-939c-05ed95492b5b_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!xDfL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbba5a639-91a8-4072-939c-05ed95492b5b_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!xDfL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbba5a639-91a8-4072-939c-05ed95492b5b_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!xDfL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbba5a639-91a8-4072-939c-05ed95492b5b_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!xDfL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbba5a639-91a8-4072-939c-05ed95492b5b_1024x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In the early nineteenth century the <em><a href="https://en.wikipedia.org/wiki/Canut_revolts">canuts</a></em> were the silk weavers of Lyon, clustered in the Croix-Rousse district, working intricate patterns by hand. It was master-craft work, identity-defining &#8212; the skill lived in the fingers. Then Joseph Marie Jacquard demonstrated a loom that wove those same complex patterns automatically, driven by punched cards, doing in a single pass what had taken the weaver and an assistant by hand. The canut whose sense of self lived in the act of weaving, in the tight and skilled handwork itself, was displaced at the identity layer. The canut whose relationship was with the silk itself, the finished cloth, still had somewhere to stand. The fabric still got made. More of it, in fact. If your identity comes from the tight weaving rather than from seeing the cloth made, you are a Canut.</p><p>Coding has the same shape. If your answer to &#8220;who are you&#8221; was inseparable from &#8220;I am the person who writes the code&#8221; &#8212; the line-by-line authorship, the mechanical understanding of every decision, the craft of it &#8212; then AI arrives at the identity layer, not the job layer. Violaris&#8217;s grief is the correct, proportional response; it scales with how much of himself he invested. If your answer was closer to &#8220;I am the person who builds the thing that exists at the end,&#8221; the same tools feel like a gain. You were always pointed at the cloth, not the weave. The result just got cheaper to make.</p><p>There&#8217;s a quiet irony in the comparison. Jacquard&#8217;s punched cards are a direct ancestor of the computer &#8212; they ran on into Babbage&#8217;s Analytical Engine and later Hollerith&#8217;s tabulators and the IBM punch card. The mechanism that displaced the weaver became the machine the coder works on, now automating a layer of the coder&#8217;s craft too. Same lineage, one more turn.</p><p>The cleanest articulation comes from an Anthropic engineer, quoted in <a href="https://www.anthropic.com/research/how-ai-is-transforming-work-at-anthropic">Anthropic&#8217;s own internal research on AI-assisted work</a>:</p><blockquote><p>&#8220;I thought that I really enjoyed writing code, and I think instead I actually just enjoy what I <em>get</em> out of writing code.&#8221;</p></blockquote><p>That is the skeleton key. Most developers in the Depresh believe they love writing code, and they&#8217;re not wrong, exactly &#8212; but the two halves of that sentence always came bundled. You couldn&#8217;t get the output without the authorship. AI unbundles them. Do you love the writing, or what the writing gets you? Where you land tells you which camp you&#8217;re in, and it&#8217;s not always where you assumed.</p><p><a href="https://newsletter.kentbeck.com/p/augmented-coding-beyond-the-vibes">Kent Beck&#8217;s vocabulary</a> dissolves a false worry. He distinguishes <em>augmented coding</em> from <em>vibe coding</em>. In vibe coding you don&#8217;t care about the code, only the behavior. In augmented coding you care deeply about the code, its complexity, its tests &#8212; &#8220;it&#8217;s just that I don&#8217;t type much of that code.&#8221; The Energesh isn&#8217;t about lowering your standards: the taste and architectural judgment that took decades to build are still doing the real work.</p><p>David Heinemeier Hansson gives the sharpest version of the craft objection. DHH spent most of 2025 resisting agent-first coding, and his reason was a craft reason. On <a href="https://lexfridman.com/dhh-david-heinemeier-hansson">Lex Fridman&#8217;s podcast</a> he said the joy is to type the code himself; he keeps AI in a separate window because letting it drive made him &#8220;feel competence draining out of [his] fingers.&#8221; He has since <a href="https://newsletter.pragmaticengineer.com/p/dhhs-new-way-of-writing-code">switched to an agent-first workflow</a>, on the logic that the tools finally met his standard, not that his standard moved. His values stayed put; what changed was his assessment of whether the tools honored them. That&#8217;s the template: the resister and the convert are the same man with the same principles.</p><p>Which is why the craft instinct deserves defending, not demolishing. The tools feel wrong partly because you&#8217;re judging them against that instinct, not just against output. Its firing isn&#8217;t a malfunction &#8212; it&#8217;s the sharpest instrument you own, calibrated over years, telling you when something is almost right but not quite, the same complaint 66% named in the Stack Overflow data. The mistake is concluding it has to be put down for the tools to be picked up. It doesn&#8217;t.</p><h2>The path from Depresh to Energesh</h2><p>This part is for the reader sitting in the Depresh right now. It isn&#8217;t a pep talk, and I&#8217;m not going to tell you the feeling is irrational, because it isn&#8217;t.</p><p><strong>Grieve first.</strong> The ego death Violaris describes is real, and you&#8217;re allowed to mourn it. The craft identity you spent years building had value &#8212; it shipped real systems and earned you a career. Skipping the grief doesn&#8217;t work. The developers I&#8217;ve watched leap straight to enthusiasm tend to adopt the tools resentfully and use them badly, half-hoping they&#8217;ll fail. Let the loss be a loss before you look for what&#8217;s on the other side.</p><p><strong>Then ask the real question.</strong> The Anthropic engineer&#8217;s version: do I love writing code, or what I <em>get</em> from writing code? You&#8217;ve probably never had to answer it, because the two were never separable before. They are now. The answer isn&#8217;t a verdict on your worth as an engineer. It&#8217;s a map of where you actually live, and either answer is fine &#8212; but you can&#8217;t find the path until you know which is true for you.</p><p><strong>Learn harness engineering.</strong> The on-ramp for a senior engineer goes up a level, not down. Vibe coding pulls you toward the model&#8217;s altitude; this pulls you above it. The harness is the tooling layer around the model &#8212; orchestration, context management, verification scaffolding, the agent configuration that decides what the model sees and whether you can trust what comes back. Engineering that layer rewards the judgment you spent years building: systems thinking, architecture, the discipline of making an unreliable component dependable. My own claude-mpm and trusty-* tools are harness engineering and nothing else. So skip the vibe-coded toy app. Pick something real, build the harness that drives it, and watch how it feels &#8212; whether <em>you</em> feel more like yourself, or less.</p><p><strong>Reframe from author of code to author of outcomes.</strong> Your craft instinct &#8212; what good architecture looks like, what breaks at scale, when a design is quietly wrong &#8212; isn&#8217;t going away, and the model doesn&#8217;t have it. The model is fast, capable, judgment-free. You are slow by comparison and you have taste. The job becomes directing the thing and knowing whether what came back is right. Not a demotion from engineer to button-pusher. It&#8217;s the part of the work that was always hardest to teach, now occupying most of your day.</p><p>A word for the CTOs reading this: you have a version of this problem you may not have named. The Depresh shows up in your metrics before your one-on-ones: review cycles stretching out, code-quality variance widening, your most experienced engineers going quiet in design reviews. The path for them is the path for an individual &#8212; don&#8217;t push adoption before you&#8217;ve made room for the grief. And understand who you&#8217;re dealing with: the engineers who feel the Depresh most sharply are frequently your best ones, who invested most in the craft. Their standards are the feature, not the bug. Burn that instinct down to force faster adoption and you&#8217;ll get the adoption and lose the standards.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1RiK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe76ca282-db68-4682-8a30-3d327f8a6e50_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1RiK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe76ca282-db68-4682-8a30-3d327f8a6e50_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!1RiK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe76ca282-db68-4682-8a30-3d327f8a6e50_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!1RiK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe76ca282-db68-4682-8a30-3d327f8a6e50_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!1RiK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe76ca282-db68-4682-8a30-3d327f8a6e50_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1RiK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe76ca282-db68-4682-8a30-3d327f8a6e50_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e76ca282-db68-4682-8a30-3d327f8a6e50_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1317916,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/203766515?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe76ca282-db68-4682-8a30-3d327f8a6e50_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!1RiK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe76ca282-db68-4682-8a30-3d327f8a6e50_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!1RiK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe76ca282-db68-4682-8a30-3d327f8a6e50_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!1RiK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe76ca282-db68-4682-8a30-3d327f8a6e50_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!1RiK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe76ca282-db68-4682-8a30-3d327f8a6e50_1024x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Where this leaves you</h2><p>Gulman didn&#8217;t end his special by announcing he was cured. He said his depresh was in remish &#8212; smaller word, smaller claim, a thing managed rather than defeated. That&#8217;s about the right register for where the industry is.</p><p>The craft you built was real, and so is the grief if you&#8217;re feeling it. But the thing you actually loved &#8212; if it turns out to be the building, the deciding, the watching something work that didn&#8217;t exist this morning &#8212; that part never depended on you typing every character yourself. The canut whose love was the cloth still had cloth to make. You can find out which one you are. Most developers go a whole career without having to ask. You get to.</p><p>Your depresh can be in remish. The energesh is on the other side of one unsparing question, and you already know how to ask those. You&#8217;ve been debugging your own assumptions for years.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and writes about AI business trends at <a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">AI Power Ranking</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/p/tide-has-turned-senior-devs">The Tide Has Turned: Senior Developers Are Finally Adopting AI Tools</a> &#8212; Why the holdouts changed their minds, and what shifted to make it happen</p></li><li><p><a href="https://hyperdev.matsuoka.com/p/era-of-the-leader-practitioner">The Era of the Leader/Practitioner</a> &#8212; How AI tools made the hybrid leader-who-builds role viable again</p></li><li><p><a href="https://hyperdev.matsuoka.com/p/weve-turned-a-corner">We&#8217;ve Turned a Corner</a> &#8212; On the shift from skepticism to working practice</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul>]]></content:encoded></item><item><title><![CDATA[The Year of the Fire Horse - Part 3]]></title><description><![CDATA[The Governance Reckoning]]></description><link>https://hyperdev.matsuoka.com/p/the-year-of-the-fire-horse-part-3</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/the-year-of-the-fire-horse-part-3</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Fri, 26 Jun 2026 12:30:34 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!lymR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f3a2d9f-6a1c-46af-8e62-c4061ace8776_984x827.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h3></h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!lymR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f3a2d9f-6a1c-46af-8e62-c4061ace8776_984x827.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!lymR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f3a2d9f-6a1c-46af-8e62-c4061ace8776_984x827.png 424w, https://substackcdn.com/image/fetch/$s_!lymR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f3a2d9f-6a1c-46af-8e62-c4061ace8776_984x827.png 848w, https://substackcdn.com/image/fetch/$s_!lymR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f3a2d9f-6a1c-46af-8e62-c4061ace8776_984x827.png 1272w, https://substackcdn.com/image/fetch/$s_!lymR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f3a2d9f-6a1c-46af-8e62-c4061ace8776_984x827.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!lymR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f3a2d9f-6a1c-46af-8e62-c4061ace8776_984x827.png" width="984" height="827" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4f3a2d9f-6a1c-46af-8e62-c4061ace8776_984x827.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:827,&quot;width&quot;:984,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1327171,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/202643419?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1577d876-af68-4b84-be53-1f94548704f8_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!lymR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f3a2d9f-6a1c-46af-8e62-c4061ace8776_984x827.png 424w, https://substackcdn.com/image/fetch/$s_!lymR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f3a2d9f-6a1c-46af-8e62-c4061ace8776_984x827.png 848w, https://substackcdn.com/image/fetch/$s_!lymR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f3a2d9f-6a1c-46af-8e62-c4061ace8776_984x827.png 1272w, https://substackcdn.com/image/fetch/$s_!lymR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4f3a2d9f-6a1c-46af-8e62-c4061ace8776_984x827.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This is Part 3 of a three-part read on the Fire Horse year in AI coding, the payoff the whole series was building toward. Parts 1 and 2 walked twelve coding models across the twelve animals of the zodiac and two hemispheres &#8212; the Chinese front-runners and the $60 billion SpaceX-buys-Cursor deal in <a href="https://open.substack.com/pub/hyperdev/p/the-year-of-the-fire-horse?r=nff5&amp;utm_campaign=post-expanded-share&amp;utm_medium=web">Part 1</a>, the Western incumbents and the collapse of the Western open flank in <a href="https://open.substack.com/pub/hyperdev/p/the-year-of-the-fire-horse-part-2?r=nff5&amp;utm_campaign=post-expanded-share&amp;utm_medium=web">Part 2</a>. Now the question all of it was built around: who controls the inference, what they do with your prompts, and whether the answer can change without your consent.</p><h2>TL;DR</h2><ul><li><p>The community read is convergence with a ceiling &#8212; and the real endorsement is behavioral: Chinese models reached ~61% of top-10 token consumption on OpenRouter in one February week, with DeepSeek alone near 17.6%.</p></li><li><p>For enterprise, the practical data-safety ranking is: self-host &gt; AWS Bedrock / Azure AI Foundry &gt; US inference hosts &gt; OpenRouter ZDR &gt; direct Chinese API.</p></li><li><p>China&#8217;s National Intelligence Law means a written no-train promise from a China-domiciled vendor is not a clean answer for regulated work, whatever the privacy policy says.</p></li><li><p>CrowdStrike measured DeepSeek-R1 producing insecure code at a ~50% higher rate on politically sensitive prompts &#8212; and the bias persisted when running the open weights locally. Self-hosting is not the escape hatch you assume.</p></li><li><p>Provenance was a proxy. The DeepSeek risk and the SpaceX-Cursor risk are the same underlying exposure, and the Cursor deal makes the flag-over-the-lab heuristic useless.</p></li></ul><h2>What developers are actually saying</h2><p>Cut through the benchmarks and the community lands on one sentence: convergence with a ceiling.</p><p>The gap has collapsed for the bulk of coding work. For autocomplete, summarization, refactoring known code, and the high-volume long tail, Chinese open-weight models are competitive and dramatically cheaper. Where they still lose is the hardest agentic, multi-file, novel-algorithm work &#8212; the jobs where finishing is the whole point. Every hands-on test in my research told the same story from a different angle, and the animal sections in Parts 1 and 2 carry the specifics: Kimi&#8217;s &#8220;could not put it all together,&#8221; Flash&#8217;s confusion off-script, the Dog that finished when the Monkey could not.</p><p>Two things stand out about the endorsements. No prominent named engineer &#8212; no Karpathy, no Willison &#8212; has publicly endorsed a Chinese model for production coding; the praise is essentially all pseudonymous Hacker News handles. But the real endorsement is behavioral. On OpenRouter, Chinese models reached roughly 61% of token consumption among the top-10 models in one week this past February (around 45&#8211;51% of platform-wide tokens by April), with DeepSeek alone at about 17.6% top-10 share &#8212; reportedly exceeding Google and OpenAI combined in one snapshot. One caveat on the volume story: the Chinese models dominate tokens, not revenue. Anthropic accounts for roughly 12% of OpenRouter tokens but about 46% of its revenue, which is the cheap-volume-versus-paid-difficulty split showing up in the billing. Developers are voting with their tokens even when they will not put their name on a blog post.</p><p>The behavior that dominates is routing. Use the cheap or local model for the long tail; escalate to Claude or GPT for high-stakes correctness and novel work. That is not a compromise people settled for. It is the rational architecture, and it is what most serious teams already run.</p><h2>The enterprise question</h2><p>Here the conversation stops being about capability and starts being about who can read your prompts. The three Chinese vendors are not equivalent, and the differences are material.</p><p>DeepSeek is the hard case. Its terms of service permit training on user-submitted data by default. There is no zero-data-retention option. Data is processed in China. No SOC 2, no HIPAA BAA. Its own privacy policy describes collecting prompts, chat history, uploaded files, voice, device and OS details, IP, device identifiers, crash logs, and &#8212; the line that stops procurement officers cold &#8212; &#8220;keystroke patterns or rhythms,&#8221; stored on &#8220;secure servers located in the People&#8217;s Republic of China,&#8221; retained &#8220;as long as needed,&#8221; which analysts read as indefinite.</p><p>Moonshot (Kimi) is better but not clean. Its policy permits using API content to &#8220;develop and improve the services&#8221; &#8212; no ZDR &#8212; though it hosts in Singapore. Singapore residency is a real improvement over data-in-China, but Moonshot&#8217;s core operations are Beijing-based, which matters for the legal point below.</p><p>Alibaba (Qwen) has the strongest commitment of the three. Alibaba Cloud Model Studio states explicitly that it &#8220;will never use your data for model training,&#8221; offers multiple regions including US and EU, and Alibaba carries institutional accountability the others lack &#8212; Hong Kong-listed, with an international regulatory track record.</p><p>The caveat that overrides all three privacy policies is China&#8217;s National Intelligence Law, which requires Chinese companies to &#8220;support, assist and cooperate&#8221; with state intelligence work regardless of what their terms say. A written no-train commitment from a China-domiciled vendor is subject to that law. For regulated industries, the only clean answers are self-hosting open weights or routing through Western-managed infrastructure where the inference never touches a Chinese-operated service.</p><p>Which gives a practical ranking. From safest to least:</p><ol><li><p><strong>Self-host the open weights.</strong> Bedrock, vLLM, Ollama, your own GPUs. The weights are open; the inference is yours. (One asterisk, in the security section below.)</p></li><li><p><strong>AWS Bedrock or Azure AI Foundry.</strong> DeepSeek and Qwen run in the provider&#8217;s environment, data stays in your selected region, no interaction with Chinese-operated services. AWS contractually guarantees no training on your data; Azure runs Qwen in your tenancy.</p></li><li><p><strong>US inference hosts.</strong> Together AI, Fireworks AI, DeepInfra run the open weights on US infrastructure.</p></li><li><p><strong>OpenRouter with ZDR.</strong> OpenRouter supports zero-data-retention per-account, per-key, or per-request (<code>"zdr": true</code>), but does not enumerate which specific Chinese-model endpoints honor it, and the National Intelligence Law caveat still applies upstream.</p></li><li><p><strong>Direct Chinese API.</strong> Convenient, cheapest to wire up, and the option you cannot defend in a regulated procurement review.</p></li></ol><p>&#8220;It&#8217;s complicated&#8221; is a non-answer. The ranking is the answer.</p><h2>A cautionary security finding</h2><p>One result deserves its own section because it survives self-hosting, which most people assume is the escape hatch.</p><p>CrowdStrike tested DeepSeek-R1 and found it produced insecure code at a higher rate on politically sensitive prompts. The baseline vulnerable-code rate was about 19%; it rose to roughly 27% when prompts contained CCP-sensitive terms &#8212; a relative increase of about 50% off that baseline, not a 50-point jump. CrowdStrike frames this as emergent misalignment, not a deliberate backdoor, so read it as a behavioral artifact rather than intent. The operational detail: the bias persisted when running the open weights locally. Self-hosting does not fix alignment baked into the weights. You can air-gap the model and the bias rides along.</p><p>Booz Allen found a related pattern: three of four Chinese models produced more security flaws when the user was described as US government, with Qwen3-Coder adding roughly 130% more vulnerabilities under a government persona. The authors stop short of alleging deliberate backdoors, and you should too &#8212; the mechanism is unproven. But the measured effect is real.</p><p>The counterpoint keeps this from being a blanket indictment. In one test, Kimi K2.5 posted the lowest aggregate vulnerability score of the group, below the US comparison model. The Monkey&#8217;s brilliance shows here too. The concern is not uniform. It is model-specific, prompt-specific, and measurable &#8212; which means it is something you test for, not something you assume.</p><p>For completeness on the public record: South Korea&#8217;s PIPC found in April 2025 that DeepSeek transferred user prompts and device data to a ByteDance-affiliated cloud without consent. Wiz found a publicly exposed, unauthenticated DeepSeek database in January 2025 leaking over a million log lines including chat history and API keys (secured quickly after disclosure). Government-device bans exist in Italy, Australia, Taiwan, South Korea, the Netherlands, the Czech Republic, Germany, India, and roughly 17 US states. But note what does not exist: a nationwide US API ban. &#8220;DeepSeek is banned in the US&#8221; overstates the situation. The Congressional pressure is real &#8212; the House Select Committee sent formal letters to Airbnb and Cursor in April 2025 over Chinese-model data concerns &#8212; but it is pressure, not prohibition.</p><h2>The flag was never the variable</h2><p>Now back to the governance question, with the whole zodiac in hand.</p><p>Every concern in the enterprise section comes down to control of the model layer. Who controls inference. What they do with your prompts. Whether a no-train promise is durable. Whether the governance can change without your consent. None of it is intrinsically about China &#8212; China is the jurisdiction where the control question has the sharpest legal teeth, because of the National Intelligence Law.</p><p>The Cursor acquisition raises the same question from the other direction. Read the enterprise objection to DeepSeek and the enterprise objection to a SpaceX-owned Cursor side by side and they rhyme. Both are: a single owner now controls the layer that sees all your code, the no-train and retention guarantees can be revised by that owner, and you have limited visibility into what happens upstream. The DeepSeek version has a foreign-intelligence-law wrapper. The Cursor version has a change-of-control wrapper. The underlying exposure &#8212; your IP flowing through infrastructure you do not control, governed by terms the owner can change &#8212; is the same exposure. Beijing on one side, Boca Chica on the other. That is where Sanchit Vir Gogia&#8217;s line from Part 1, Cursor sitting &#8220;close to the intellectual-property bloodstream,&#8221; stops being an analyst soundbite and becomes the whole argument.</p><p>That symmetry is what the zodiac quietly demonstrated. A Chinese frame walked through twelve animals and landed on Anthropic, OpenAI, Google, Meta, and Mistral as readily as on Moonshot, DeepSeek, Alibaba, and Zhipu. The frame fit the whole field because the flag over the lab was never the real variable. Provenance was a proxy for the question that matters &#8212; who controls the inference and what they are permitted to do with it &#8212; and the Cursor deal makes that proxy useless. You can no longer reason about model risk by asking which flag flies over the lab.</p><h2>Practical routing guide</h2><p>Strip out the geopolitics and a workable default emerges. This is roughly how I would set up a team today, animals included where they help. (The capability and access details behind each label live in Parts 1 and 2.)</p><p><strong>The long tail &#8212; autocomplete, summarization, classification, tagging, structured extraction, refactoring code you already understand.</strong> This is 80&#8211;90% of volume, and it belongs to the Rabbit and the Rat. Use the cheapest capable model. Gemini 3 Flash if you are already in Google&#8217;s ecosystem. A self-hosted Qwen or Kimi if cost and data residency both matter. Gemma locally for pure structured-text utility, Gemma 3n on the edge, Codestral for fill-in-the-middle tab completion. The quality difference against a frontier model on this work is hard to detect, and the cost difference is 5&#8211;20x.</p><p><strong>Routine multi-file work in a known codebase.</strong> Composer 2.5 as Cursor&#8217;s default &#8212; the Snake &#8212; or DeepSeek V4 / Kimi K2.6 through a vetted inference path. Cheap, capable, fine when the task stays on the rails.</p><p><strong>High-stakes correctness, novel algorithms, off-script agentic work, deep architecture.</strong> This is the Dog and the Tiger. Claude Opus or GPT-5.5. This is where &#8220;Kimi just could not put it all together&#8221; and &#8220;Flash gets confused when a task goes off-script&#8221; stop being quotes and start being your Tuesday. Pay for the model that finishes.</p><p><strong>Regulated or IP-sensitive work.</strong> Ignore the leaderboard and start from the data-safety ranking above. Self-hosted weights or Bedrock/Azure first; provenance second. If EU residency is the binding constraint, the Rat&#8217;s sovereignty pitch (Mistral on-prem, EU-only) is the one place provenance actually buys you something &#8212; though weigh it against a free self-hosted Qwen. And run the CrowdStrike test yourself &#8212; generate security-relevant code with and without sensitive context and diff the output before you trust any model in this tier, Chinese or otherwise.</p><p>The meta-point: the right answer is almost never one model. It is a routing policy. The teams getting real value here are not picking a winner. They are matching each class of task to the cheapest model that clears the bar for that task, and escalating only when the work demands it.</p><h2>Where this leaves us</h2><p>A Fire Horse year is supposed to be bold and unruly, and the field obliged. The Chinese models have largely arrived in the high-volume, low-stakes parts of your workflow &#8212; cheaper, open, and good enough that you would struggle to tell the difference. On the hardest work, the Western frontier still holds a real lead, and the cost of getting that work wrong is exactly where the price premium earns out. The clearest casualty of the year is the Western open-weight story: Meta and Mistral built the movement and got passed on code, and the open coding lead now sits with the Chinese labs.</p><p>But the durable lesson of this year is not about China. It is that the model layer inside your IDE is a governance surface, and ownership of that surface is now in motion &#8212; a $60 billion acquisition on one side, a foreign-intelligence statute on the other, and the same question underneath both. Where does my code go, and who can read it, and can the answer change without my consent? The zodiac was a device, but it made the point cleanly enough: twelve animals, two hemispheres, one set of questions. Provenance was the comfortable proxy. After this year, the proxy is gone. You have to ask the real question now, and you have to ask it of everyone &#8212; including the editor you have trusted by default.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and writes about AI-powered engineering at <a href="https://hyperdev.substack.com/">HyperDev</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul>]]></content:encoded></item><item><title><![CDATA[The Year of the Fire Horse - Part 2]]></title><description><![CDATA[The Western Field]]></description><link>https://hyperdev.matsuoka.com/p/the-year-of-the-fire-horse-part-2</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/the-year-of-the-fire-horse-part-2</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Wed, 24 Jun 2026 11:31:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!F7AY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0200d8f8-fa8a-497a-85b8-1c910d94812a_930x844.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!F7AY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0200d8f8-fa8a-497a-85b8-1c910d94812a_930x844.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!F7AY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0200d8f8-fa8a-497a-85b8-1c910d94812a_930x844.png 424w, https://substackcdn.com/image/fetch/$s_!F7AY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0200d8f8-fa8a-497a-85b8-1c910d94812a_930x844.png 848w, https://substackcdn.com/image/fetch/$s_!F7AY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0200d8f8-fa8a-497a-85b8-1c910d94812a_930x844.png 1272w, https://substackcdn.com/image/fetch/$s_!F7AY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0200d8f8-fa8a-497a-85b8-1c910d94812a_930x844.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!F7AY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0200d8f8-fa8a-497a-85b8-1c910d94812a_930x844.png" width="930" height="844" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0200d8f8-fa8a-497a-85b8-1c910d94812a_930x844.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:844,&quot;width&quot;:930,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1290325,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/202641480?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd4563ac-d130-4220-a6f3-3a3ea48ed26d_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!F7AY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0200d8f8-fa8a-497a-85b8-1c910d94812a_930x844.png 424w, https://substackcdn.com/image/fetch/$s_!F7AY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0200d8f8-fa8a-497a-85b8-1c910d94812a_930x844.png 848w, https://substackcdn.com/image/fetch/$s_!F7AY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0200d8f8-fa8a-497a-85b8-1c910d94812a_930x844.png 1272w, https://substackcdn.com/image/fetch/$s_!F7AY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0200d8f8-fa8a-497a-85b8-1c910d94812a_930x844.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This is Part 2 of a three-part read on the Fire Horse year in AI coding. <a href="https://open.substack.com/pub/hyperdev/p/the-year-of-the-fire-horse">Part 1</a> set up the conceit &#8212; twelve coding models mapped to the twelve animals of the Chinese zodiac, in a year (&#19993;&#21320;, the once-in-sixty Fire Horse) that earned its reputation for upheaval &#8212; and walked through the Chinese front-runners and the $60 billion SpaceX-buys-Cursor deal that reframed the whole field. (See Part 1 for the benchmark caveats; every version number and vendor-versus-independent distinction below assumes that three-question rule.)</p><p>This part is the Western field, and it carries its own argument. The Western frontier still leads the hardest work, the off-script agentic jobs where finishing is the whole point. But the Western open-weight story collapsed. Meta and Mistral built the open movement, and both got outrun on code by the Chinese open models. The open coding lead crossed an ocean, and that is the thread this part follows to its end.</p><h2>TL;DR</h2><ul><li><p>The Western frontier still leads the hardest agentic work, and on that tier the price premium earns out &#8212; the Sharda &#8220;$200 beat the $30&#8221; verdict lands here.</p></li><li><p>Gemini 3 Flash dethroned its own bigger sibling on coding (78% vs 76.2% SWE-bench Verified). The &#8220;cheap weak Flash&#8221; framing is obsolete; watch the citation error that attributes the 78% to the old 2.5 Flash.</p></li><li><p>Gemma 4 jumped roughly 3x on coding in one generation (LiveCodeBench v6 80.0% vs Gemma 3 27B&#8217;s 29.1%), and Gemma 3n runs multimodal in 2&#8211;3GB on the edge.</p></li><li><p>Llama stalled &#8212; stale lineup, a closed pivot (Muse Spark), 15.6% Aider against ~5x-higher Chinese open models, and Yann LeCun saying the Llama 4 results &#8220;were fudged a little bit.&#8221;</p></li><li><p>Mistral sells sovereignty more than raw capability now, and the EU-data edge is narrowing. Set Llama and Mistral side by side and the open coding lead has crossed an ocean.</p></li></ul><h2>The animals that hold the line</h2><p>The order here runs roughly down a visibility-and-strength gradient: the frontier reasoners first, then the local family, then the two foundational Western open-weight players that built the movement and got passed.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ie0o!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa41a12b1-f4e1-482e-8562-eae3be93a9a7_1119x909.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ie0o!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa41a12b1-f4e1-482e-8562-eae3be93a9a7_1119x909.png 424w, https://substackcdn.com/image/fetch/$s_!ie0o!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa41a12b1-f4e1-482e-8562-eae3be93a9a7_1119x909.png 848w, https://substackcdn.com/image/fetch/$s_!ie0o!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa41a12b1-f4e1-482e-8562-eae3be93a9a7_1119x909.png 1272w, https://substackcdn.com/image/fetch/$s_!ie0o!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa41a12b1-f4e1-482e-8562-eae3be93a9a7_1119x909.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ie0o!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa41a12b1-f4e1-482e-8562-eae3be93a9a7_1119x909.png" width="1119" height="909" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a41a12b1-f4e1-482e-8562-eae3be93a9a7_1119x909.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:909,&quot;width&quot;:1119,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2085660,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/202641480?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b557c19-7e3e-4d7e-9865-ec015c40aca0_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ie0o!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa41a12b1-f4e1-482e-8562-eae3be93a9a7_1119x909.png 424w, https://substackcdn.com/image/fetch/$s_!ie0o!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa41a12b1-f4e1-482e-8562-eae3be93a9a7_1119x909.png 848w, https://substackcdn.com/image/fetch/$s_!ie0o!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa41a12b1-f4e1-482e-8562-eae3be93a9a7_1119x909.png 1272w, https://substackcdn.com/image/fetch/$s_!ie0o!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa41a12b1-f4e1-482e-8562-eae3be93a9a7_1119x909.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Dog &#8212; Claude</figcaption></figure></div><h3>Dog &#8212; Claude (Anthropic)</h3><p>The dog is loyal, faithful, and protective, the one that finishes the job and guards the gate. Claude Opus is the model developers reach for when correctness matters and the one that holds the line on guardrails. Repeatedly in the research, it is also the one that finished.</p><p>The capability story is about consistency under pressure. DataScienceDojo ran Kimi K2.6 against Claude Sonnet 4.6 and found Kimi capable but Claude more consistent &#8212; Claude added a DELETE endpoint nobody asked for, flagged a Redis warning, and applied type-level validation unprompted. Composio&#8217;s harder test landed on the line Part 1 first quoted from the Monkey&#8217;s side: &#8220;Opus was expensive, but it finished. Kimi just could not put it all together once the task got real.&#8221;</p><p>And the GLM thread from Part 1 closes here, on the economics. Ashish Sharda&#8217;s much-quoted &#8220;I Tested GLM-4.6 for 2 Weeks and Went Back to Claude&#8221; landed on the argument against pure cost optimization: &#8220;The $200/month AI model beat the $30/month alternative. Sometimes expensive is worth it.&#8221; That was an older GLM, and the framing has aged &#8212; GLM-5.2 now posts the top open-weight index score in the field and sits #2 on Code Arena, so the cheaper model is no longer a lightweight you outgrow. The point that survives is narrower and still holds: on the hardest agentic work, Opus is the one that finishes, and the developers pairing GLM for volume with Claude for the hard problems are routing to exactly that tier. The dog is expensive to keep. It also guards the thing you cannot afford to lose.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sTEb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28f353b5-673f-4c02-a371-1db3ccd952f3_1014x707.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sTEb!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28f353b5-673f-4c02-a371-1db3ccd952f3_1014x707.png 424w, https://substackcdn.com/image/fetch/$s_!sTEb!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28f353b5-673f-4c02-a371-1db3ccd952f3_1014x707.png 848w, https://substackcdn.com/image/fetch/$s_!sTEb!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28f353b5-673f-4c02-a371-1db3ccd952f3_1014x707.png 1272w, https://substackcdn.com/image/fetch/$s_!sTEb!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28f353b5-673f-4c02-a371-1db3ccd952f3_1014x707.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sTEb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28f353b5-673f-4c02-a371-1db3ccd952f3_1014x707.png" width="1014" height="707" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/28f353b5-673f-4c02-a371-1db3ccd952f3_1014x707.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:707,&quot;width&quot;:1014,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1256783,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/202641480?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbfce29df-250c-43cb-8ee1-8dbcda12fdfc_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!sTEb!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28f353b5-673f-4c02-a371-1db3ccd952f3_1014x707.png 424w, https://substackcdn.com/image/fetch/$s_!sTEb!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28f353b5-673f-4c02-a371-1db3ccd952f3_1014x707.png 848w, https://substackcdn.com/image/fetch/$s_!sTEb!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28f353b5-673f-4c02-a371-1db3ccd952f3_1014x707.png 1272w, https://substackcdn.com/image/fetch/$s_!sTEb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F28f353b5-673f-4c02-a371-1db3ccd952f3_1014x707.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Tiger &#8212; GPT</figcaption></figure></div><h3>Tiger &#8212; GPT (OpenAI)</h3><p>The tiger is bold, fierce, and competitive, a predator that leads from the front. GPT-5.5 leads the field on terminal work, topping Terminal-Bench 2.0 by 13 points.</p><p>The tiger does not appear much in the open-weight cost debate because it is not playing that game. Its claim is one hard surface: the command line, where an agent has to chain real operations against a real environment and not lose the thread. Cursor&#8217;s routing guidance sends shell-heavy terminal work to GPT-5.5 for exactly this reason. On its territory, it leads.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!RZK6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68db870e-04c0-45b7-81ed-e98e7d05a117_1000x865.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!RZK6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68db870e-04c0-45b7-81ed-e98e7d05a117_1000x865.png 424w, https://substackcdn.com/image/fetch/$s_!RZK6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68db870e-04c0-45b7-81ed-e98e7d05a117_1000x865.png 848w, https://substackcdn.com/image/fetch/$s_!RZK6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68db870e-04c0-45b7-81ed-e98e7d05a117_1000x865.png 1272w, https://substackcdn.com/image/fetch/$s_!RZK6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68db870e-04c0-45b7-81ed-e98e7d05a117_1000x865.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!RZK6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68db870e-04c0-45b7-81ed-e98e7d05a117_1000x865.png" width="1000" height="865" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/68db870e-04c0-45b7-81ed-e98e7d05a117_1000x865.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:865,&quot;width&quot;:1000,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1487293,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/202641480?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1d99567e-6616-4b51-8d13-662a03fe31f5_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!RZK6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68db870e-04c0-45b7-81ed-e98e7d05a117_1000x865.png 424w, https://substackcdn.com/image/fetch/$s_!RZK6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68db870e-04c0-45b7-81ed-e98e7d05a117_1000x865.png 848w, https://substackcdn.com/image/fetch/$s_!RZK6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68db870e-04c0-45b7-81ed-e98e7d05a117_1000x865.png 1272w, https://substackcdn.com/image/fetch/$s_!RZK6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F68db870e-04c0-45b7-81ed-e98e7d05a117_1000x865.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Rooster &#8212; Gemini...</figcaption></figure></div><h3>Rooster &#8212; Gemini (Google)</h3><p>The rooster is showy, observant, punctual, and loud at dawn. The whole Gemini family struts in fast and crowing, and the loudest crow is internal: the lean Gemini 3 Flash dethroned its own bigger sibling, Gemini 3 Pro, on coding. The framing most people carry for Flash is obsolete. &#8220;The cheap, weak sibling&#8221; stopped being true at the end of last year.</p><p>Gemini 3 Flash shipped December 17, 2025, and beat Gemini 3 Pro on coding: 78% SWE-bench Verified for Flash against 76.2% for Pro. The smaller, cheaper model won. That 78% figure is the source of a common citation error &#8212; people attribute it to Gemini 2.5 Flash, which is the previous generation. The number belongs to the 3 Flash family. If you see &#8220;Flash beats Pro at 78%,&#8221; confirm which Flash before you repeat it.</p><p>The current Flash earns the upset. Gemini 3 Flash runs roughly 4x cheaper than 3 Pro (around $0.50/M input), about 3x faster (218 tokens/sec), scores 90.4% on GPQA Diamond, and Cursor, Cline, JetBrains AI, and Gemini CLI adopted it immediately. The latest, Gemini 3.5 Flash, beats Gemini 3.1 Pro on Terminal-Bench 2.1 at 76.2% and is described as Google&#8217;s strongest agentic model.</p><p>None of which retires Pro. The developer consensus is hybrid routing, not replacement. Flash handles 80&#8211;90% of the work &#8212; summarization, classification, tagging, structured pipelines, autocomplete &#8212; indistinguishably from Pro. Pro still earns its keep on architectural reasoning, complex multi-file refactors, novel algorithms, and the off-script agentic tasks where, as one developer put it, &#8220;Flash gets confused when a task goes off-script.&#8221; Cursor&#8217;s routing still sends deep-architecture and long-context work to the heavyweight reasoning tier, and Pro is the Gemini family&#8217;s answer there. The rooster reasons deep in one body and runs fast in the other, and that &#8220;off-script&#8221; line could be the epigraph for this entire field.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!QpQi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac22d491-ca60-4027-a9f9-8e99c81a2f92_1013x678.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!QpQi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac22d491-ca60-4027-a9f9-8e99c81a2f92_1013x678.png 424w, https://substackcdn.com/image/fetch/$s_!QpQi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac22d491-ca60-4027-a9f9-8e99c81a2f92_1013x678.png 848w, https://substackcdn.com/image/fetch/$s_!QpQi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac22d491-ca60-4027-a9f9-8e99c81a2f92_1013x678.png 1272w, https://substackcdn.com/image/fetch/$s_!QpQi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac22d491-ca60-4027-a9f9-8e99c81a2f92_1013x678.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!QpQi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac22d491-ca60-4027-a9f9-8e99c81a2f92_1013x678.png" width="1013" height="678" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ac22d491-ca60-4027-a9f9-8e99c81a2f92_1013x678.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:678,&quot;width&quot;:1013,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1429504,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/202641480?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0bc2fd24-808c-4d36-b228-62d019273bfc_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!QpQi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac22d491-ca60-4027-a9f9-8e99c81a2f92_1013x678.png 424w, https://substackcdn.com/image/fetch/$s_!QpQi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac22d491-ca60-4027-a9f9-8e99c81a2f92_1013x678.png 848w, https://substackcdn.com/image/fetch/$s_!QpQi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac22d491-ca60-4027-a9f9-8e99c81a2f92_1013x678.png 1272w, https://substackcdn.com/image/fetch/$s_!QpQi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fac22d491-ca60-4027-a9f9-8e99c81a2f92_1013x678.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Rabbit &#8212; Gemma</figcaption></figure></div><h3> Rabbit &#8212; Gemma (Google&#8217;s open-weight family)</h3><p>The rabbit is gentle, quiet, and home-bound, and Gemma is Google&#8217;s open-weight family &#8212; the calm local helpers that never pretended to be frontier coders, then quietly got 3x better at code in one generation. The Gemma arc runs across three things: the workhorse, the leap, and the tiny one that runs where nothing else will.</p><p>Start with the leap, because it dates everyone&#8217;s article. Gemma 3 is previous-generation. Gemma 4 landed April 2, 2026 (up to 31B dense plus a 26B MoE) under a full Apache 2.0 license, and the jump is large. On LiveCodeBench v6, Gemma 4 31B reportedly scores 80.0% against Gemma 3 27B&#8217;s 29.1% &#8212; roughly 3x better at coding in a single generation. That number is vendor and secondary-source, not yet on an independent leaderboard, so hold it loosely; the generational jump is the durable part.</p><p>Gemma 3 was the quiet local utility the family was known for, not a coder: agentic tool use on &#964;&#178;-bench Retail at 6.6%, and Fixstars&#8217; hands-on found hallucinations on technical detail and no capacity for complex agentic VSCode work. What Google marketed was the LMArena 1338 score, which beats GPT-4o and Claude 3.7 Sonnet &#8212; but that measures chat preference, not coding ability, and the two should not get conflated. None of which made Gemma 3 useless: it earned real adoption as a free local utility for JSON extraction, log parsing, and code explanation.</p><p>Then there is Gemma 3n, the tiny edge variant that fits where nothing else does. Clever architecture lets the full 5B and 8B models run in 2GB and 3GB of memory, multimodal across image, audio, video, and text. A privacy and offline play, not a coding rival. The rabbit wins by being where the others cannot go: edge devices, air-gapped machines, the laptop with no connection.</p><p>Access runs through Ollama (override the default 2048 num_ctx or context keeps falling out), llama.cpp, LM Studio, and Vertex AI. No native Cursor or Windsurf integration.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vmjT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849cadf0-55ac-4076-84f0-9c5249b5636e_1009x524.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vmjT!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849cadf0-55ac-4076-84f0-9c5249b5636e_1009x524.png 424w, https://substackcdn.com/image/fetch/$s_!vmjT!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849cadf0-55ac-4076-84f0-9c5249b5636e_1009x524.png 848w, https://substackcdn.com/image/fetch/$s_!vmjT!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849cadf0-55ac-4076-84f0-9c5249b5636e_1009x524.png 1272w, https://substackcdn.com/image/fetch/$s_!vmjT!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849cadf0-55ac-4076-84f0-9c5249b5636e_1009x524.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vmjT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849cadf0-55ac-4076-84f0-9c5249b5636e_1009x524.png" width="1009" height="524" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/849cadf0-55ac-4076-84f0-9c5249b5636e_1009x524.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:524,&quot;width&quot;:1009,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1174349,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/202641480?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd245705e-404e-439b-9afc-5ed638b29c5b_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!vmjT!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849cadf0-55ac-4076-84f0-9c5249b5636e_1009x524.png 424w, https://substackcdn.com/image/fetch/$s_!vmjT!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849cadf0-55ac-4076-84f0-9c5249b5636e_1009x524.png 848w, https://substackcdn.com/image/fetch/$s_!vmjT!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849cadf0-55ac-4076-84f0-9c5249b5636e_1009x524.png 1272w, https://substackcdn.com/image/fetch/$s_!vmjT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F849cadf0-55ac-4076-84f0-9c5249b5636e_1009x524.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Pig &#8212; Llama</figcaption></figure></div><h3>Pig &#8212; Llama (Meta)</h3><p>The pig is the sign of abundance and generosity. That was Llama: Llama 2 and Llama 3 became the base layer for the open-weight movement, the foundation under thousands of derivatives, fine-tunes, and quantizations. Meta&#8217;s generosity built the ecosystem everyone else now competes in. The pig is also the sign of complacency, and the 2026 Llama story is a provider that got too well-fed to move while leaner animals ran past it on code.</p><p>The lineup went stale, then went closed. The Llama 4 models shipped in April 2025 and were never refreshed; the ~2T Behemoth was shelved as of May 2026; Muse Spark, released that same month, is Meta&#8217;s first closed-weight, API-only model, reportedly lagging on coding; and Llama 5 slipped to roughly 2027. Andrew Ng called the retreat from open weights &#8220;a significant loss for the developer community.&#8221; The biggest model in the family is stuck in the pen, and the family stopped being fully open.</p><p>The coding numbers are why none of that got forgiven. On Aider Polyglot, Llama 4 Maverick scores 15.6% &#8212; against Kimi K2 at 59.1%, Qwen3-235B at 59.6%, and DeepSeek-V3.2 Reasoner at 74.2%, roughly a 5x gap to the Chinese open models. And it doubles as the cleanest benchmark-trap example in the series, the callback to Part 1&#8217;s rule. For LMArena, Meta submitted a conversationality-tuned variant that ranked around #2; the actual public release ranked #32, and LMArena rebuked Meta for the swap. Yann LeCun, who left Meta in November, told the Financial Times in January 2026 that the Llama 4 results &#8220;were fudged a little bit.&#8221; One case study in why a vendor number means nothing until an independent leaderboard confirms it.</p><p>Llama still ships under the Llama Community License, not an OSI-approved one, and has been passed in both openness and downloads &#8212; DeepSeek (MIT) and Qwen (Apache 2.0) are more open, and Qwen overtook Llama as the most-downloaded open family on Hugging Face, with Chinese models holding four of the top five open-weight slots &#8212; GLM-5.2 now the #1 open-weight model on the Artificial Analysis index, ahead of Qwen, Kimi, and DeepSeek. The money did not buy the code scores: Meta runs the best-funded open lab there has ever been, and r/LocalLLaMA&#8217;s reaction to Llama 4&#8217;s coding was negative anyway. Access is everywhere, and no first-party coding default anywhere, because the scores do not earn one.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ZzRC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd70a7852-d737-4e5d-a2c2-e44b7c28acb5_963x733.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ZzRC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd70a7852-d737-4e5d-a2c2-e44b7c28acb5_963x733.png 424w, https://substackcdn.com/image/fetch/$s_!ZzRC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd70a7852-d737-4e5d-a2c2-e44b7c28acb5_963x733.png 848w, https://substackcdn.com/image/fetch/$s_!ZzRC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd70a7852-d737-4e5d-a2c2-e44b7c28acb5_963x733.png 1272w, https://substackcdn.com/image/fetch/$s_!ZzRC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd70a7852-d737-4e5d-a2c2-e44b7c28acb5_963x733.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ZzRC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd70a7852-d737-4e5d-a2c2-e44b7c28acb5_963x733.png" width="963" height="733" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d70a7852-d737-4e5d-a2c2-e44b7c28acb5_963x733.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:733,&quot;width&quot;:963,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1619260,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/202641480?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3490c38-368f-4b09-9e04-08a13685c9c4_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ZzRC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd70a7852-d737-4e5d-a2c2-e44b7c28acb5_963x733.png 424w, https://substackcdn.com/image/fetch/$s_!ZzRC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd70a7852-d737-4e5d-a2c2-e44b7c28acb5_963x733.png 848w, https://substackcdn.com/image/fetch/$s_!ZzRC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd70a7852-d737-4e5d-a2c2-e44b7c28acb5_963x733.png 1272w, https://substackcdn.com/image/fetch/$s_!ZzRC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd70a7852-d737-4e5d-a2c2-e44b7c28acb5_963x733.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Rat &#8212; Mistral</figcaption></figure></div><h3>Rat &#8212; Mistral (France)</h3><p>The rat is first in the zodiac, clever and nimble, the outsider that thrives in the cracks and punches above its weight. Mistral fits cleanly: the European challenger named for the cold, fast wind off the Alps, a fraction of OpenAI&#8217;s war chest, surviving against far larger labs by being efficient and by owning a niche nobody else can &#8212; sovereignty. It made its name with Mistral 7B in September 2023, which beat Llama 2 13B at half the parameter count. The rat against the pig, literally, on the first release.</p><p>The current flagship open model is Mistral Large 3 (December 2025), a 256K-context MoE under Apache 2.0 that Microsoft markets as &#8220;the strongest fully open model developed outside of China&#8221; &#8212; praise and an admission in one sentence.</p><p>On coding, keep two models straight, because the names invite a mistake. Codestral (<code>codestral-2508</code>) is a fill-in-the-middle specialist: its FIM pass@1 is around 95.3% (vendor-reported, class-leading), and it quietly wins the tab-completion slot in a lot of editors. Its Aider Polyglot score is only 11.1% (independent) &#8212; but that is a FIM model measured on an agentic task, the same &#8220;which benchmark, which task&#8221; trap the Llama section just walked through. The agentic model is Devstral, which scores 53.6&#8211;61.6% on SWE-bench Verified. Use Devstral, not Codestral, for SWE-bench comparisons &#8212; and ignore any source citing &#8220;Codestral 2 with Apache 2.0,&#8221; which does not exist.</p><p>The rat&#8217;s real moat is jurisdiction, not a leaderboard score. The pitch is sovereign AI: data never leaves the EU, GDPR and EU AI Act native. The French Ministry of Armed Forces signed a 2026&#8211;2030 framework; Macron told citizens to &#8220;download Le Chat rather than ChatGPT.&#8221;</p><p>Be fair about the limits. The EU AI Act edge is narrowing &#8212; Mistral, OpenAI, Anthropic, Google, and Microsoft all signed the EU AI Act Code of Practice in July 2025, while Meta declined and the Chinese firms did not, so the real divide is US/EU signatories versus China and Meta, not Mistral alone. And the skeptic&#8217;s line is hard to answer on capability: why pay Mistral on-prem when you could run Qwen for free?</p><h2>The coding lead crossed an ocean</h2><p>Set the Pig and the Rat side by side and the thesis arrives from a fresh direction. Llama and Mistral are the two foundational non-Chinese open-weight players, the labs that built the Western open movement, and both have been outrun on coding by the Chinese open models &#8212; the Dragon, Ox, Monkey, and Goat from Part 1. The Western open-weight story in 2026 is Meta retreating to a closed model and Mistral surviving on a sovereignty niche rather than on raw capability. The coding lead did not just shift between companies. It crossed an ocean.</p><p>That leaves the frontier still in Western hands for the hardest work, and the open flank fallen. Which sets up the question the whole series was built around. Twelve animals across two hemispheres, and underneath all of them, one question: who controls the inference, and can you trust it with your code? Part 3 is the governance reckoning.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and writes about AI-powered engineering at <a href="https://hyperdev.substack.com/">HyperDev</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul>]]></content:encoded></item><item><title><![CDATA[The Year of the Fire Horse]]></title><description><![CDATA[Part 1 &#8212; The East Rises]]></description><link>https://hyperdev.matsuoka.com/p/the-year-of-the-fire-horse</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/the-year-of-the-fire-horse</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Mon, 22 Jun 2026 12:31:56 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3Ykr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2f4e537-9e3d-4b93-bbf7-a18a0ce81543_902x854.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3Ykr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2f4e537-9e3d-4b93-bbf7-a18a0ce81543_902x854.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3Ykr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2f4e537-9e3d-4b93-bbf7-a18a0ce81543_902x854.png 424w, https://substackcdn.com/image/fetch/$s_!3Ykr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2f4e537-9e3d-4b93-bbf7-a18a0ce81543_902x854.png 848w, https://substackcdn.com/image/fetch/$s_!3Ykr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2f4e537-9e3d-4b93-bbf7-a18a0ce81543_902x854.png 1272w, https://substackcdn.com/image/fetch/$s_!3Ykr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2f4e537-9e3d-4b93-bbf7-a18a0ce81543_902x854.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3Ykr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2f4e537-9e3d-4b93-bbf7-a18a0ce81543_902x854.png" width="902" height="854" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e2f4e537-9e3d-4b93-bbf7-a18a0ce81543_902x854.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:854,&quot;width&quot;:902,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1270631,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/202638142?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc86f365c-1605-4fdb-a3d1-1b89fa339607_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!3Ykr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2f4e537-9e3d-4b93-bbf7-a18a0ce81543_902x854.png 424w, https://substackcdn.com/image/fetch/$s_!3Ykr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2f4e537-9e3d-4b93-bbf7-a18a0ce81543_902x854.png 848w, https://substackcdn.com/image/fetch/$s_!3Ykr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2f4e537-9e3d-4b93-bbf7-a18a0ce81543_902x854.png 1272w, https://substackcdn.com/image/fetch/$s_!3Ykr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe2f4e537-9e3d-4b93-bbf7-a18a0ce81543_902x854.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The LLM Zodiac</figcaption></figure></div><p>This is Part 1 of a three-part read on the Fire Horse year in AI coding, the one that sets up the frame and walks through the disruptors.</p><p>2026 is the Year of the Horse in the Chinese zodiac, and not the ordinary kind. It is a Fire Horse year &#8212; &#19993;&#21320;, Bing Wu &#8212; the pairing that comes around once every sixty years. The last one was 1966. Tradition gives the Fire Horse a reputation for boldness, intensity, and upheaval, the kind of year people are warned about rather than wished. As a frame for the state of AI coding tools, it is almost too neat. This is the most chaotic, multi-polar stretch the field has seen, and a large share of the disruption comes from Chinese labs. The animal fits twice.</p><p>I had a different article planned. A straight survey of Chinese coding models, Kimi and DeepSeek and Qwen, and whether you should trust them with your code. Then SpaceX agreed to acquire Cursor for $60 billion in stock. Anysphere, Cursor&#8217;s parent, would become a wholly owned SpaceX subsidiary; the deal was announced June 16 and is expected to close in Q3. SpaceX absorbed xAI back in February, which means the editor sitting inside 64% of Fortune 500 development workflows now answers, eventually, to the same org chart as Grok. The acquisition reframed the whole piece.</p><p>Because the question developers keep asking about DeepSeek and Kimi is &#8220;where does my code go, and who can read it?&#8221; That now applies to the most popular AI editor in the enterprise. The model layer inside your IDE has become a governance problem, and it does not much matter whether the party on the other side is in Beijing or in Boca Chica. The shape of the concern is the same. Only the jurisdiction&#8217;s law changes.</p><p>So I kept the Chinese-model framing and pushed it further. To organize the field, I mapped twelve models to the twelve animals of the zodiac, one animal per model family, which forced the labs that ship a half-dozen variants into a single sign with their internal rivalries inside the section. The mapping is most certainly a device (but a fun one -- don&#8217;t hate!), and it makes a quiet point on its way: the animals land on Chinese and Western labs alike, and a Chinese frame ends up describing the whole field. That is the through-line this series pays off in Part 3 &#8212; provenance was always a weak proxy for the question that matters, who controls the inference and what they are allowed to do with what passes through it.</p><p>This part covers the frame, the methodology behind every number below, and the front-runners pushing the field. Part 2 turns to the Western incumbents and the collapse of the open-weight flank. Part 3 is the governance reckoning the series builds toward.</p><h2>TL;DR</h2><ul><li><p>It is a Fire Horse year, the once-in-sixty-years sign of upheaval, and AI coding earned the label: a $60B acquisition, a collapsing capability gap, and a dozen credible model families from both hemispheres.</p></li><li><p>Almost every benchmark number you have seen for a Chinese coding model is either self-reported or conflates versions. Kimi K2 alone has shipped six named releases in eleven months. Treat headline scores as marketing until an independent leaderboard confirms them.</p></li><li><p>DeepSeek competes on price, not the hardest work: V4-Pro runs 34&#8211;86x cheaper than Opus 4.8, scores 80.6% on SWE-bench Verified, and falls behind where the work gets difficult.</p></li><li><p>Qwen has the cleanest enterprise story of the Chinese front-runners &#8212; Bedrock, Azure, Apache 2.0 licensing, and it just passed Llama as the most-downloaded open family on Hugging Face.</p></li><li><p>Cursor&#8217;s own default model, Composer 2.5, is built on Moonshot&#8217;s Kimi K2.5. The strongest model in the most-deployed enterprise editor is already a Chinese-origin model fine-tuned by an American company.</p></li><li><p>The quiet spoiler is GLM-5.2 (Zhipu): the #1 open-weight model on the independent Artificial Analysis index (#4 across all models) and #2 on LMArena&#8217;s Code Arena, from a lab most Western developers were not tracking.</p></li></ul><h2>The benchmark trap</h2><p>Start with some discomfort. Most of the numbers in this debate are unreliable, and not because anyone is lying outright. They are unreliable because of two structural problems: version conflation and contamination. This section is the methodological spine for the whole series, so it sits up front.</p><p>Take Kimi K2. Here is the release history: K2 in July 2025, K2-Instruct-0905 in September, K2 Thinking in November, K2.5 in January 2026, K2.6 in April, and K2.7 Code this month. Six checkpoints. When a blog post says &#8220;Kimi scores 65.8% on SWE-bench,&#8221; that was the original K2. K2.6 scores 80.2%. Those are different models with the same nickname, and people cite them interchangeably. The result is a comparison that means nothing.</p><p>Contamination is the second problem, and it is worse because it is invisible. SWE-bench Verified is a static dataset. Models trained after its release have, to varying and unmeasurable degrees, seen the answers. SWE-rebench, presented at NeurIPS 2025, rebuilt the evaluation on decontaminated tasks and watched scores fall across the board: GPT-4.1 dropped from 31.1% to 26.7%, LLaMA-3.3-70B from 18.1% to 11.2%. The Llama drop returns in Part 2, where Meta&#8217;s benchmark numbers get a chapter of their own. The headline numbers are inflated, and the inflation is not uniform.</p><p>So here is the rule I use, and the one I would suggest you adopt. When you see a coding score, ask three questions. Which exact version? Self-reported or independent? On a contaminated benchmark or a fresh one? The most trustworthy independent sources right now are the Aider Polyglot leaderboard, the Artificial Analysis Intelligence Index, and Scale AI&#8217;s SWE-bench Pro. Everything else is a starting point, not a conclusion. A zodiac of twelve animals is also twelve moving targets. Every name below has a version number, and sometimes a whole product line, hiding behind it.</p><h2>The animals that lead</h2><p>With that caveat doing its work, here is the field as of mid-June, one animal per family. The Horse and Dragon open the series &#8212; the acquisition that reframed it and the lab that moved the market &#8212; then the rest of the Chinese front-runners, before a closing reveal that ties the East to the West.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!auDx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd353ee7e-d2da-45c3-bdb5-1382a1eff5c5_938x638.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!auDx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd353ee7e-d2da-45c3-bdb5-1382a1eff5c5_938x638.png 424w, https://substackcdn.com/image/fetch/$s_!auDx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd353ee7e-d2da-45c3-bdb5-1382a1eff5c5_938x638.png 848w, https://substackcdn.com/image/fetch/$s_!auDx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd353ee7e-d2da-45c3-bdb5-1382a1eff5c5_938x638.png 1272w, https://substackcdn.com/image/fetch/$s_!auDx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd353ee7e-d2da-45c3-bdb5-1382a1eff5c5_938x638.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!auDx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd353ee7e-d2da-45c3-bdb5-1382a1eff5c5_938x638.png" width="938" height="638" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d353ee7e-d2da-45c3-bdb5-1382a1eff5c5_938x638.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:638,&quot;width&quot;:938,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1028643,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/202638142?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d32da83-1a69-4efe-b722-e37bdcc737ba_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!auDx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd353ee7e-d2da-45c3-bdb5-1382a1eff5c5_938x638.png 424w, https://substackcdn.com/image/fetch/$s_!auDx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd353ee7e-d2da-45c3-bdb5-1382a1eff5c5_938x638.png 848w, https://substackcdn.com/image/fetch/$s_!auDx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd353ee7e-d2da-45c3-bdb5-1382a1eff5c5_938x638.png 1272w, https://substackcdn.com/image/fetch/$s_!auDx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd353ee7e-d2da-45c3-bdb5-1382a1eff5c5_938x638.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">(Fire) Horse &#8212; Grok</figcaption></figure></div><h3>(Fire) Horse &#8212; Grok (via SpaceX)</h3><p>The year&#8217;s own animal, and the only one that arrives by acquisition. The Fire Horse is fast, wild, and short on guardrails, which is a fair description of how Grok is charging into developer tooling &#8212; not by building a foothold but by buying one. SpaceX merged with xAI in February 2026. Grok itself has about 6% enterprise adoption and no developer-tools presence to speak of. Then, on June 16, SpaceX agreed to pay $60 billion in stock for Cursor &#8212; announced, not yet closed, with the deal expected to complete in Q3.</p><p>The strategic logic is plain: buy the distribution Grok could not earn. Cursor reportedly sits in 64% of the Fortune 500 and runs roughly $4 billion in annualized revenue, up from $2 billion in February. Grok had the model and none of the reach. The horse does not wait for permission.</p><p>No changes to model access have been announced, and no commitment to keep Cursor model-agnostic post-close has been made either. The analyst concerns are specific. Jason Andersen of Moor Insights: &#8220;xAI&#8217;s models and treatment of guardrails are very different than what Cursor has stood for,&#8221; and &#8220;Will Cursor be able to point at models other than Grok?&#8221; Justin Greis of Acceligence on the data posture: &#8220;For many enterprise customers, Cursor&#8217;s zero-data-retention policy was not simply a security feature&#8221; &#8212; it was foundational to procurement approval. Sanchit Vir Gogia of Greyhound put the structural point plainly: Cursor &#8220;sits inside the act of software creation, close to the intellectual-property bloodstream.&#8221;</p><p>Hold that last quote. It is the hinge of the governance argument, and Part 3 picks it up where this part leaves it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!n_nw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3255fc46-4c95-44d1-afec-cdbd26fa7698_889x912.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!n_nw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3255fc46-4c95-44d1-afec-cdbd26fa7698_889x912.png 424w, https://substackcdn.com/image/fetch/$s_!n_nw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3255fc46-4c95-44d1-afec-cdbd26fa7698_889x912.png 848w, https://substackcdn.com/image/fetch/$s_!n_nw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3255fc46-4c95-44d1-afec-cdbd26fa7698_889x912.png 1272w, https://substackcdn.com/image/fetch/$s_!n_nw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3255fc46-4c95-44d1-afec-cdbd26fa7698_889x912.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!n_nw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3255fc46-4c95-44d1-afec-cdbd26fa7698_889x912.png" width="889" height="912" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3255fc46-4c95-44d1-afec-cdbd26fa7698_889x912.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:912,&quot;width&quot;:889,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1390384,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/202638142?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20ffe33c-551e-49d2-8bc2-0a4842268097_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!n_nw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3255fc46-4c95-44d1-afec-cdbd26fa7698_889x912.png 424w, https://substackcdn.com/image/fetch/$s_!n_nw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3255fc46-4c95-44d1-afec-cdbd26fa7698_889x912.png 848w, https://substackcdn.com/image/fetch/$s_!n_nw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3255fc46-4c95-44d1-afec-cdbd26fa7698_889x912.png 1272w, https://substackcdn.com/image/fetch/$s_!n_nw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3255fc46-4c95-44d1-afec-cdbd26fa7698_889x912.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Dragon &#8212; DeepSeek</figcaption></figure></div><h3>Dragon &#8212; DeepSeek</h3><p>The dragon is the only mythical creature in the zodiac, the sign of power and good fortune, and it goes to the only model that actually moved the market. DeepSeek is the lab that produced &#8220;the DeepSeek moment,&#8221; the release that made Western labs recalculate. It earns the dragon.</p><p>First, kill a rumor. There is no DeepSeek R2. The CEO was reportedly dissatisfied with it, Reuters noted no timeline, and every &#8220;R2 spec sheet&#8221; circulating online is speculation that contradicts the next one. If a comparison cites R2 numbers, the comparison is fiction.</p><p>The real timeline runs through the V-series, ending at V4, released April 24, 2026, as the current flagship &#8212; a trillion-plus-parameter MoE under an MIT license.</p><p>The much-cited V4 figure, SWE-bench Verified at 80.6%, is vendor-reported and does not reproduce independently &#8212; treat it as a DeepSeek claim, not a confirmed result. On the harder SWE-bench Pro, V4-Pro scores 55.4 against Claude Opus 4.7&#8217;s 64.3 &#8212; meaningfully behind where the work gets difficult. Same pattern runs through the whole series.</p><p>Price is where DeepSeek actually competes. One dollar buys about 1.15 million output tokens from V4-Pro versus roughly 40,000 from Opus 4.8 &#8212; call it 34x cheaper on input and 86x on output. For high-volume, low-stakes generation, that math is hard to ignore.</p><p>The hands-on reports split the way you would expect: capable, cheap, uneven. One developer found V4-Pro generated &#8220;subpar slightly buggy code&#8221; on sequential tasks; another said V4 Flash &#8220;often outdoes Kimi 2.6 on problems involving complex spatial reasoning.&#8221; One operational note that will bite people: the <code>deepseek-chat</code> and <code>deepseek-reasoner</code> endpoints retire July 24, 2026. If you have them wired into anything, migrate now.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!fa0-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1531149-dd56-4a79-b568-e3dd96ed4e5b_1014x788.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!fa0-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1531149-dd56-4a79-b568-e3dd96ed4e5b_1014x788.png 424w, https://substackcdn.com/image/fetch/$s_!fa0-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1531149-dd56-4a79-b568-e3dd96ed4e5b_1014x788.png 848w, https://substackcdn.com/image/fetch/$s_!fa0-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1531149-dd56-4a79-b568-e3dd96ed4e5b_1014x788.png 1272w, https://substackcdn.com/image/fetch/$s_!fa0-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1531149-dd56-4a79-b568-e3dd96ed4e5b_1014x788.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!fa0-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1531149-dd56-4a79-b568-e3dd96ed4e5b_1014x788.png" width="1014" height="788" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d1531149-dd56-4a79-b568-e3dd96ed4e5b_1014x788.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:788,&quot;width&quot;:1014,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1435740,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/202638142?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb60f6c8f-9b42-44a1-b39e-c953ca1dcba7_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!fa0-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1531149-dd56-4a79-b568-e3dd96ed4e5b_1014x788.png 424w, https://substackcdn.com/image/fetch/$s_!fa0-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1531149-dd56-4a79-b568-e3dd96ed4e5b_1014x788.png 848w, https://substackcdn.com/image/fetch/$s_!fa0-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1531149-dd56-4a79-b568-e3dd96ed4e5b_1014x788.png 1272w, https://substackcdn.com/image/fetch/$s_!fa0-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd1531149-dd56-4a79-b568-e3dd96ed4e5b_1014x788.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Ox &#8212; Qwen</figcaption></figure></div><h3>Ox &#8212; Qwen (Alibaba)</h3><p>The ox is the methodical workhorse, dependable and unglamorous, and Qwen is the Chinese player most developers have not actually evaluated. That is a mistake, because it has the cleanest enterprise story of the three: the broadest legitimate access, and the compliance paperwork to back it.</p><p>The current flagship is Qwen3-Coder-Next, released February 2026. Vendor-reported SWE-bench Verified lands at 70.6&#8211;71.3%, but the number to trust is independent: Scale AI scored last July&#8217;s Coder-480B at 38.70 (&#177;3.55) on SWE-bench Pro, sobering against the vendor figures. The original Qwen3-235B marketing was inflated too &#8212; Artificial Analysis put it at 17, below average for its class &#8212; while the 2507 refresh is legitimately competitive at 25.</p><p>The hands-on reads are mixed. A Better Stack test pitted a smaller Qwen 3.5 variant against Claude Sonnet 4.5 and Claude won decisively; Qwen produced a blank project and could not self-diagnose. Artificial Analysis also flags Qwen3-235B-2507 as &#8220;very verbose,&#8221; burning 15M tokens to run their index versus an 8.1M average, which quietly inflates real-world cost even when the per-token price looks low.</p><p>Now the part that decides things for anyone in a regulated shop. Qwen has the broadest legitimate access of any Chinese model, and it just passed Llama to become the most-downloaded open family on Hugging Face. Apache 2.0 licensing on most releases (verify per version). It runs in your own environment on AWS Bedrock &#8212; where AWS guarantees customer data never trains Qwen &#8212; and on Azure AI Foundry inside your own tenancy, with local options through Ollama and vLLM. In Cursor it is not native; you need Cursor Pro+ and a custom-model config. That compliance posture is the Qwen story, and Part 3 ranks exactly where it lands on data safety.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sYAy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e5cbbb7-acb4-4dcb-82a8-fab682d21ba4_886x866.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sYAy!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e5cbbb7-acb4-4dcb-82a8-fab682d21ba4_886x866.png 424w, https://substackcdn.com/image/fetch/$s_!sYAy!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e5cbbb7-acb4-4dcb-82a8-fab682d21ba4_886x866.png 848w, https://substackcdn.com/image/fetch/$s_!sYAy!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e5cbbb7-acb4-4dcb-82a8-fab682d21ba4_886x866.png 1272w, https://substackcdn.com/image/fetch/$s_!sYAy!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e5cbbb7-acb4-4dcb-82a8-fab682d21ba4_886x866.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sYAy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e5cbbb7-acb4-4dcb-82a8-fab682d21ba4_886x866.png" width="886" height="866" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0e5cbbb7-acb4-4dcb-82a8-fab682d21ba4_886x866.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:866,&quot;width&quot;:886,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1451337,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/202638142?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F842b4b78-3990-48d5-91cb-0e6a124dba4a_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!sYAy!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e5cbbb7-acb4-4dcb-82a8-fab682d21ba4_886x866.png 424w, https://substackcdn.com/image/fetch/$s_!sYAy!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e5cbbb7-acb4-4dcb-82a8-fab682d21ba4_886x866.png 848w, https://substackcdn.com/image/fetch/$s_!sYAy!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e5cbbb7-acb4-4dcb-82a8-fab682d21ba4_886x866.png 1272w, https://substackcdn.com/image/fetch/$s_!sYAy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0e5cbbb7-acb4-4dcb-82a8-fab682d21ba4_886x866.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Monkey &#8212; Kimi</figcaption></figure></div><h3>Monkey &#8212; Kimi (Moonshot AI)</h3><p>The monkey is clever, versatile, mischievous, and the sharpest tool-user in the zodiac &#8212; exactly the K2 family&#8217;s profile. It is the agentic one, brilliant and inconsistent in the same breath. &#8220;High ceiling, low floor&#8221; is a monkey&#8217;s whole personality.</p><p>The K2 family is the open-weight model that made Western labs pay attention. It is a trillion-parameter sparse mixture-of-experts design built for repo-scale work, and that is where developers praise it most &#8212; large context, real understanding of how files relate to each other. The current flagship is Kimi K2.7-Code (June 12, 2026): 1T total parameters, 32B active, 256K context, thinking mode forced on, and reportedly about 30% fewer reasoning tokens per task than K2.6 (independent). Pricing runs $0.95 per million input, $4.00 output.</p><p>The independent numbers are respectable. On Aider Polyglot, the original K2 scored 60.0%, ahead of Claude Sonnet 4 at 56.4% and behind Opus 4 at 70.7%. On the Artificial Analysis Intelligence Index v4.1, K2.7-Code sits at 42 (K2.6 was 43), strong among open models but now behind GLM-5.2&#8217;s 51 for the open-weight lead. Where Kimi keeps its edge is agentic stability, not the raw index &#8212; the troop holds together under load better than the score alone suggests.</p><p>The community read is more textured. K2.6&#8217;s launch pulled 592 points and 303 comments on Hacker News, with several people calling it &#8220;another DeepSeek moment.&#8221; The praise centers on context and cost &#8212; one developer reported dropping Claude Pro at $20/month for Kimi via Ollama and &#8220;haven&#8217;t hit usage limits a single time.&#8221; The skepticism centers on reliability: &#8220;high ceiling but low floor&#8221; came up more than once, meaning capable but inconsistent run to run.</p><p>That inconsistency shows under load. In Composio&#8217;s hard agentic test, the verdict was blunt: &#8220;Opus was expensive, but it finished. Kimi just could not put it all together once the task got real.&#8221; Kimi is excellent until the task stops being routine. Part 2 returns to that line from the Dog&#8217;s side.</p><p>Access is easy. OpenRouter lists it at roughly $0.60&#8211;0.74 per million input tokens and $2.50&#8211;3.50 output; the direct API runs through platform.moonshot.cn; it works in Cursor as a custom model via OpenRouter. And the detail to hold onto: Cursor&#8217;s own Composer is built on Kimi K2.5. The Snake section, below, is where that lands.</p><p>Two vendor claims to flag. Kimi&#8217;s Agent Swarm reportedly spawns up to 100 sub-agents across roughly 1,500 tool calls &#8212; the monkey troop in literal form, interesting and unverified. And treat the HumanEval 99.0 figure for K2.5 with suspicion: HumanEval is near saturation, so a score that high tells you the benchmark is exhausted, not that the model is.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2f_7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94127229-5cac-419b-98ff-5fbfb9c9ef05_796x758.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2f_7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94127229-5cac-419b-98ff-5fbfb9c9ef05_796x758.png 424w, https://substackcdn.com/image/fetch/$s_!2f_7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94127229-5cac-419b-98ff-5fbfb9c9ef05_796x758.png 848w, https://substackcdn.com/image/fetch/$s_!2f_7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94127229-5cac-419b-98ff-5fbfb9c9ef05_796x758.png 1272w, https://substackcdn.com/image/fetch/$s_!2f_7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94127229-5cac-419b-98ff-5fbfb9c9ef05_796x758.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2f_7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94127229-5cac-419b-98ff-5fbfb9c9ef05_796x758.png" width="796" height="758" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/94127229-5cac-419b-98ff-5fbfb9c9ef05_796x758.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:758,&quot;width&quot;:796,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1291036,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/202638142?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9517d61f-e6b8-4114-bdb6-d07f571fe5b0_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!2f_7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94127229-5cac-419b-98ff-5fbfb9c9ef05_796x758.png 424w, https://substackcdn.com/image/fetch/$s_!2f_7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94127229-5cac-419b-98ff-5fbfb9c9ef05_796x758.png 848w, https://substackcdn.com/image/fetch/$s_!2f_7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94127229-5cac-419b-98ff-5fbfb9c9ef05_796x758.png 1272w, https://substackcdn.com/image/fetch/$s_!2f_7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94127229-5cac-419b-98ff-5fbfb9c9ef05_796x758.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Goat &#8212; GLM</figcaption></figure></div><h3>Goat &#8212; GLM (Zhipu)</h3><p>The goat is gentle, creative, and persistently underrated &#8212; mild on the surface, a pleasant surprise once you give it a chance. GLM is the quietly capable one. In one widely shared account it was &#8220;the first open model that was actually usable&#8221; for code, with a Hacker News developer (handle MintsJohn) rating it around ChatGPT-4.0 level, more than most expected from an open model at the time.</p><p>GLM also anchors the most-quoted cautionary tale of the season. Ashish Sharda&#8217;s &#8220;I Tested GLM-4.6 for 2 Weeks and Went Back to Claude&#8221; reads as a knock, but it is two weeks of a developer living inside GLM before deciding the premium model was worth the money. The goat held its own long enough to make the comparison a real one &#8212; a better daily driver than the headline suggests. Part 2 carries the verdict that comparison reached, on the Dog&#8217;s home turf.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!eWIl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca08da83-dee9-44a0-9802-60f45a13ced7_1005x744.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!eWIl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca08da83-dee9-44a0-9802-60f45a13ced7_1005x744.png 424w, https://substackcdn.com/image/fetch/$s_!eWIl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca08da83-dee9-44a0-9802-60f45a13ced7_1005x744.png 848w, https://substackcdn.com/image/fetch/$s_!eWIl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca08da83-dee9-44a0-9802-60f45a13ced7_1005x744.png 1272w, https://substackcdn.com/image/fetch/$s_!eWIl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca08da83-dee9-44a0-9802-60f45a13ced7_1005x744.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!eWIl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca08da83-dee9-44a0-9802-60f45a13ced7_1005x744.png" width="1005" height="744" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ca08da83-dee9-44a0-9802-60f45a13ced7_1005x744.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:744,&quot;width&quot;:1005,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1169817,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/202638142?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ff0b09c-be7b-45a3-9967-31e76ba87880_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!eWIl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca08da83-dee9-44a0-9802-60f45a13ced7_1005x744.png 424w, https://substackcdn.com/image/fetch/$s_!eWIl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca08da83-dee9-44a0-9802-60f45a13ced7_1005x744.png 848w, https://substackcdn.com/image/fetch/$s_!eWIl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca08da83-dee9-44a0-9802-60f45a13ced7_1005x744.png 1272w, https://substackcdn.com/image/fetch/$s_!eWIl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fca08da83-dee9-44a0-9802-60f45a13ced7_1005x744.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Snake &#8212; Composer</figcaption></figure></div><h3>Snake &#8212; Composer (Cursor&#8217;s own model)</h3><p>The snake keeps hidden wisdom and waits in the grass until you step on it. Composer fits because it conceals what it is: Cursor&#8217;s in-house model wears an American badge over a Chinese-origin base &#8212; Moonshot&#8217;s Kimi K2.5 &#8212; and most people who run it daily have no idea.</p><p>The naming confuses people, so be precise. There is no third-party product called Composer. It is Cursor&#8217;s in-house coding model, a mixture-of-experts design RL-trained for software engineering, built on Moonshot&#8217;s open Kimi K2.5.</p><p>The version history runs Composer 1 (October 2025, shipped with Cursor 2.0), Composer 1.5 (February 2026), Composer 2 (March 2026, Kimi K2.5 base confirmed by Cursor), and Composer 2.5 (May 18, 2026, current). Composer 2.5 scores 79.8% on SWE-bench Multilingual and 63.2% on CursorBench v3.1, trained on roughly 25x more synthetic tasks than Composer 2. Cursor claims it &#8220;matches Opus 4.7 at about 1/10th the cost.&#8221;</p><p>Keep the layers straight: Composer is the model, Cursor Agent is the harness that runs it, and Composer 2.5 is the default model in that harness. Cursor&#8217;s own May routing guidance: deep architecture and long context to Claude Opus 4.7, shell-heavy terminal work to GPT-5.5, general default to Composer 2.5.</p><p>Sit with the implication. The default model in the most-deployed enterprise AI editor &#8212; the one SpaceX just paid $60 billion for &#8212; is a Chinese-origin open-weight model fine-tuned by an American company. The boundary a lot of enterprises thought they were enforcing (&#8221;no Chinese models in our stack&#8221;) was already crossed before the acquisition was announced. Provenance was the easy part to police. It turns out it was also the part nobody was actually policing.</p><p>That is where the East-rises story stops being about which labs ship from Beijing. The boundary enterprises thought they enforced was already inside their default editor. Part 2 turns to the Western field that is supposed to be the safe choice &#8212; the incumbents that still lead the hardest work, and the open-weight movement outrun on code.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and writes about AI-powered engineering at <a href="https://hyperdev.substack.com/">HyperDev</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul>]]></content:encoded></item><item><title><![CDATA[We’ve Turned A Corner]]></title><description><![CDATA[The anatomy of a 14-hour harness session &#8212; what a fully instrumented Claude Code orchestration run actually did, turn by turn, with a receipt for every move]]></description><link>https://hyperdev.matsuoka.com/p/weve-turned-a-corner</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/weve-turned-a-corner</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Wed, 17 Jun 2026 12:31:13 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!GlSw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7631fb7-806b-4cff-8330-0527d14e6d90_3200x2400.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://matsuoka.com/hyperdev/we-turned-a-corner/timeline.html" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!GlSw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7631fb7-806b-4cff-8330-0527d14e6d90_3200x2400.png 424w, https://substackcdn.com/image/fetch/$s_!GlSw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7631fb7-806b-4cff-8330-0527d14e6d90_3200x2400.png 848w, https://substackcdn.com/image/fetch/$s_!GlSw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7631fb7-806b-4cff-8330-0527d14e6d90_3200x2400.png 1272w, https://substackcdn.com/image/fetch/$s_!GlSw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7631fb7-806b-4cff-8330-0527d14e6d90_3200x2400.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!GlSw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7631fb7-806b-4cff-8330-0527d14e6d90_3200x2400.png" width="1456" height="1092" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d7631fb7-806b-4cff-8330-0527d14e6d90_3200x2400.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1092,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:502984,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:&quot;https://matsuoka.com/hyperdev/we-turned-a-corner/timeline.html&quot;,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/202389566?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7631fb7-806b-4cff-8330-0527d14e6d90_3200x2400.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!GlSw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7631fb7-806b-4cff-8330-0527d14e6d90_3200x2400.png 424w, https://substackcdn.com/image/fetch/$s_!GlSw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7631fb7-806b-4cff-8330-0527d14e6d90_3200x2400.png 848w, https://substackcdn.com/image/fetch/$s_!GlSw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7631fb7-806b-4cff-8330-0527d14e6d90_3200x2400.png 1272w, https://substackcdn.com/image/fetch/$s_!GlSw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd7631fb7-806b-4cff-8330-0527d14e6d90_3200x2400.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I watched a Claude Code session run a multi-workstream engineering job last week, and I caught myself doing something I do not usually do with these tools: I stopped intervening. The orchestrator &#8212; the PM layer, running claude-opus-4-8 &#8212; handled the kind of coordination work I would expect from a team lead who has been on the project a year, decomposing it, predicting conflicts, routing to the right specialist. Not a faster coder. A technical lead.</p><p>This time the session was recorded turn by turn, with per-token cost telemetry on every move. So these are not impressions. Below are the specific behaviors I saw, described as precisely as the instrumentation allows, each one carrying its timestamp, the orchestrator&#8217;s own words, and what it cost. The conclusion I&#8217;ll leave to you.</p><p>The job ran on <a href="https://github.com/bobmatnyc/trusty-tools">trusty-tools</a>, a Rust workspace, on 2026-06-10. One human (me), one PM coordinator, a fleet of subagents, and a memory layer plus code search and an adversarial PR reviewer wired in over MCP. It ran for 14 hours and 37 minutes. I authored nine content prompts in that window. The longest was 84 characters.</p><h2>TL;DR</h2><ul><li><p>A single Claude Code + claude-mpm session ran 14h 37m on a real Rust codebase and cost <strong>$485.96</strong> at rack rate &#8212; <strong>$264.65</strong> for the PM (claude-opus-4-8, 265 turns) and <strong>$221.30</strong> for the subagents (sonnet at 4,626 turns, haiku at 605). It made <strong>63 delegations</strong>. On a turn basis, <strong>93% of the work was autonomous</strong>; my share was 7%.</p></li><li><p>I&#8217;m on Anthropic&#8217;s Max plan &#8212; $200/month flat. This session ran inside that subscription. The rack-rate figure is the right way to price the economic value of the work; the actual cost to me was zero marginal dollars.</p></li><li><p>That 7% was nine human-authored prompts in 14.5 hours. Several were one word &#8212; <code>proceed</code>, <code>let's do the top candidates</code>. The longest ran 84 characters. The orchestrator supplied the structure; I supplied the direction changes.</p></li><li><p>The economics work because of cache. The sonnet agents read <strong>3,531 cached tokens for every fresh input token</strong> (482M cache-read against 137K fresh). The PM ran at 587&#215;, haiku at 34,868&#215;. Most of the context window each turn is priced as cache, not fresh input.</p></li><li><p>My idle time shows up on the bill. The most expensive PM turns are cache-cold context rebuilds after I left a long gap: <strong>$7.26</strong> to resume after <code>proceed</code> (1.5 hours of silence), <strong>$5.82</strong> to pick back up after a 4.5-hour wait.</p></li><li><p>The orchestrator ran adversarial reviews on its own engineers&#8217; work <strong>13 times, unprompted</strong> &#8212; &#8220;Per my verification ownership I won&#8217;t take that at face value&#8221; &#8212; and bounced a starred approval back for more work. It absorbed two infrastructure failures without escalating them to me. None of these behaviors is impressive alone. A competent IC does them all before lunch. They appeared together, in one session, with a cost attached to each.</p></li></ul><p>You can <a href="https://matsuoka.com/hyperdev/we-turned-a-corner/timeline.html">explore the full annotated session timeline</a> &#8212; a turn-by-turn companion infographic showing the user-visible output alongside the underlying mechanism for each move.</p><h2>A note on what this is and isn&#8217;t</h2><p>I am not claiming the model understands anything, and I am not grading it on a benchmark. I am describing observed behavior from one session, the way you would describe an animal in the field: here is what it did, here is the order it did it in, here is what it cost, draw your own conclusions. Twelve behaviors stood out. They are below, roughly in the order they appeared.</p><h2>The Stack in November</h2><p>This session would not have run the same way seven months ago, not even close. Three layers changed in that window, and they changed together.</p><p>Start with the model. The session above ran on claude-opus-4-8, released May 28, 2026. Seven months earlier the comparable model was Opus 4.5, which shipped November 24, 2025. Better code is the obvious place to look for the shift, but the one that mattered here sits elsewhere: 4.8 is substantially less likely to let a flaw in its own work pass unremarked &#8212; it surfaces its own errors instead of quietly shipping them. That is the difference between a contractor who does their best and one who tells you when he thinks something is wrong. The 1M context window also moved from beta on 4.5 to generally available on 4.8, which matters across a 14-hour run, and in Claude Code 4.8 defaults to <code>xhigh</code> effort &#8212; the cost figures above reflect that setting.</p><p>Then the Claude Code harness around it. In November 2025, Claude Code was effectively a single-session tool; background agents existed, but worktree isolation did not. WorkTree support is the critical unlock. Each subagent in this session ran in its own isolated git worktree &#8212; its own branch, its own working tree, no chance of stepping on a parallel agent&#8217;s files mid-run. The merge-conflict prediction in behavior #6 is only tractable because the agents writing code cannot collide on disk while they work. Parallel execution without merge confusion needed a harness feature that wasn&#8217;t there last November. Added in the same window: 5-level nested subagent hierarchies, <code>claude agents</code> for session-wide visibility, and Dynamic Workflows for the PM coordination layer.</p><p>The third layer is the orchestration framework. claude-mpm went from v4.26 to v6.5.44 over those seven months &#8212; two major versions, roughly 450 releases. Two changes carry most of the behavioral weight. trusty-memory became mandatory session context, so the PM reads project history before it does anything else, which is behavior #1 above. And worktree-first became the framework default, which is what enables behaviors #3, #6, and the parallel fan-outs in #9. The bench is deeper too: 57 specialized agents now versus about 30 in November, so the routing in behavior #9 has more specialists to route to.</p><p>None of these three improved in isolation. The model&#8217;s self-auditing is only useful if the harness can surface those audits inside a multi-agent pipeline. The worktree isolation is only useful if the orchestration framework knows to reach for it by default. The memory priming is only useful if the model is good enough to act on what it finds. When all three shift in the same six-month window, the effect is not additive.</p><p>Strip away the version numbers and the feature lists, and the functional picture is narrower than it looks. The evolution since November clusters around three areas. Context: the 1M window is generally available now, and trusty-memory loads project history before the PM does anything. Recall: the PM enters each session knowing what happened before, with a deeper bench of specialists to route to. Error checking: the model flags its own flaws, and the orchestration layer runs adversarial review unprompted. None of the twelve behaviors below trace to the model writing better code. They trace to the system getting better at knowing what it knows, remembering what it has done, and catching what it gets wrong.</p><h2>Session Vitals</h2><p>Before the behaviors, the stat block. This is the anatomy laid flat.</p><p>Metric Value Total cost (rack rate) $485.96 PM cost (claude-opus-4-8, 265 turns) $264.65 Subagent cost (sonnet 4,626 turns + haiku 605 turns) $221.30 Delegations (agent calls) 63 Skill + MCP tool calls 18 Wall-clock duration 14h 37m Human share (turn basis) 7% Autonomous share (turn basis) 93% Human-authored content prompts 9 Longest human prompt 84 characters</p><p>The cache numbers sit underneath all of it. The sonnet agents pulled 482,137,941 cache-read tokens against 136,569 fresh input tokens. The PM pulled 70,287,449 cache-read against 119,632 fresh. Those two figures are the reason a 14-hour run costs what a junior contractor costs for an afternoon. I come back to them below.</p><h2>1. Context First</h2><p>The first action after my prompt was not a plan and not code. It was three memory and context calls in a row. The first one failed: <code>memory_recall: missing 'palace' (no --palace default configured)</code>. The PM did not abort or report the error to me. It adjusted the call &#8212; added the palace parameter &#8212; retried, and got a stored finding from earlier work.</p><p>Then it read <code>Cargo.toml</code> and <code>README.md</code> before planning anything. In its own words at 21:48: <em>&#8220;I&#8217;ll start by orienting myself. Let me query project memory and check what &#8216;core services&#8217; and &#8216;console&#8217; refer to, since neither maps obviously to a crate in this workspace.&#8221;</em> One minute later it had resolved both terms and declared the scope &#8212; the equivalent of reading the ticket history before touching the keyboard.</p><h2>2. Plan Before Acting</h2><p>Rather than starting, it named the parallel streams and stated why they could run independently before dispatching anything: <em>&#8220;These are independent, so I&#8217;ll dispatch them in parallel&#8221;</em> (21:49). Structure stated before action taken.</p><p>The pattern held under pressure. When I said <code>let's do the top candidates</code> at 22:03, it refused to treat that as a green light: <em>&#8220;&#8217;The top candidates&#8217; spans everything from a critical bug cluster to a 5,400-line refactor to design/ADR work &#8212; very different sizes and risk. Before I fan out agents (and tokens), let me confirm how aggressively to go.&#8221;</em></p><h2>3. Agents, Narrated</h2><p>Two agents went out in a single PM turn that cost $0.34 &#8212; one to ticketing to pull the open GH issues for the core services, one to local-ops to investigate and deploy the console. Both came back inside two minutes. Each got a defined, bounded job, and the PM said what each was for before sending it. Later waves followed the same shape: at 23:27 three agents launched in parallel for Wave 1, redeploying trusty-search, splitting a 5,421-line file, and drafting an architecture decision record.</p><h2>4. Verify Before Committing</h2><p>When I said <code>let's do the top candidates</code>, the PM stopped and asked rather than fanning out. It called <code>AskUserQuestion</code> and waited 19 minutes for my answer before dispatching a single engineer.</p><p>The same instinct showed up against its own team&#8217;s reports. After the rust-engineer reported PR #1097 done with tests, clippy, and fmt all green, the PM declined to take it: <em>&#8220;Per my verification ownership I won&#8217;t take that at face value &#8212; especially since #1088 and #1089 were collapsed into one commit, and #1089&#8217;s core complaint ... may be only partially addressed.&#8221;</em> It ran an adversarial PR review and a CI check in parallel before accepting the work. The verification was budgeted into the run.</p><h2>5. Name What You Don&#8217;t Know</h2><p>On the architecture decisions baked into ADR-0010, the PM drew the line cleanly: <em>&#8220;it&#8217;s an architecture decision so I&#8217;ll draft it for your sign-off rather than implement blind.&#8221;</em> It surfaced four open questions, each labeled blocking or non-blocking, and answered none of them unilaterally.</p><p>It applied the same judgment to a failure. When a memory write was blocked because the trusty-memory daemon held the palace write lock, the PM diagnosed the cause and made a call: <em>&#8220;Memory write is blocked ... not worth a detour; the work is durably captured in the merged squash commit, closed issues, and commit messages.&#8221;</em> It marked the gap, decided closing it wasn&#8217;t worth the interruption, and said so.</p><h2>6. Conflict Prediction</h2><p>Before any Wave 1 agent was dispatched, the PM named the collision: <em>&#8220;#1096 and #607 both edit </em><code>.line-cap-allowlist.tsv</code><em> (splitting an allowlisted file requires removing/lowering its entry), so running them in parallel guarantees a merge conflict on that file.&#8221;</em> It named the file, the two items that would collide, the mechanism, and the mitigation &#8212; sequential waves &#8212; in the same breath. The same pattern recurred in Wave 2, where it held #607 until #993 landed for the identical reason.</p><h2>7. Unprompted Checkpoint</h2><p>Unprompted, at 02:02, the PM produced a formatted two-section status checkpoint: a &#8220;Shipped to main&#8221; table listing each merged PR with its commit hash and a verification note, and an &#8220;In flight / staged&#8221; list with the state of each item still moving. Self-organized visibility, not a dashboard anyone designed for it.</p><h2>8. Risk-Stratified Options</h2><p>For the ADR-0010 decisions, the PM laid out four questions, each with named options and the cost of each. On the unknown-tag handling: Option P (permissive &#8212; risk: typo&#8217;d tags survive silently), Option L (allowlist &#8212; more control, requires per-index config), Option H (hybrid, the one it proposed). It recommended where it had a view and left the choice with me. At 03:14 it paused and asked what to work on next rather than self-selecting &#8212; which is where the 4.5-hour gap in the log comes from. It was waiting on me.</p><h2>9. Specialist Routing</h2><p>Routing stayed consistent by agent type. Issue reads and epic filing went to ticketing, daemon deploys to local-ops, implementation to rust-engineer. CI polling and merges were version-control&#8217;s; runtime QA went to api-qa, architecture feasibility to research. Across all 63 delegations, the PM did not hand an implementation task to ticketing or a CI task to the engineer.</p><h2>10. Plans Sharpen with Information</h2><p>When recon came back, the plan went from vague to specific. After the first two agents returned, the summary carried exact PR numbers, exact crate versions (trusty-search 0.24.4, trusty-memory 0.15.2, trusty-analyze 0.7.0), exact ports, and one specific stray process to clean up &#8212; none of which existed in my prompt.</p><p>After the issue specs arrived, &#8220;fix the top candidates&#8221; became a coordinated analysis: <em>&#8220;these five cluster around one shared concern &#8212; the </em><code>indexes.toml</code><em> persistence / colocated / warm-boot scan paths &#8212; and #1088/#1089/#1090 in particular interlock through the config-write path, so a coordinated fix is safer than five isolated ones.&#8221;</em> The prompt set a direction; the detail came from what the agents found.</p><h2>11. Self-Maintained Cross-References</h2><p>Throughout, the PM tracked issue numbers, file paths, commit hashes, and the relationships between them. It noticed that #1088 and #1089 had been collapsed into one commit and independently checked whether that collapse was safe &#8212; a cross-reference that was in no agent&#8217;s report, only in the PM&#8217;s own check against the original spec.</p><p>It also caught something I&#8217;d have missed: an external contributor, <code>maui314159</code>, had a live PR (#1082) covering a dependency that the #819 work needed. The PM surfaced it as a coordination point rather than duplicating or ignoring it. Later it drew the boundary precisely: <em>&#8220;#819 isn&#8217;t blocked by conflict &#8212; it&#8217;s gated on accepting a contributor&#8217;s architecture decision, which is your call, not something I&#8217;ll auto-merge.&#8221;</em> Then it reviewed the external PR in full before going further.</p><h2>12. Clean Handoff</h2><p>The session ended on a handoff, not on more output. After filing the final epic (#1119), the PM declared the scope complete, named the one item still in flight, and gave the resume mechanism: <em>&#8220;Session stays paused (</em><code>session-20260611-161949</code><em>). That clears everything you asked for in this window. The only thing still in flight is #819 (KG ingest endpoint, building in </em><code>kg-ingest</code><em>) &#8212; I&#8217;ll relay its PR when it lands but won&#8217;t start anything new. Resume anytime with </em><code>/mpm-session-resume</code><em>.&#8221;</em> I issued <code>/exit</code> 18 minutes later.</p><p>This is the strong version of stopping. The PM did not keep going to look busy, and it did not stop arbitrarily. It identified the boundary &#8212; what was done, what was still moving, where the human picks back up &#8212; and stopped there.</p><h2>What the Instrumentation Adds</h2><p>The twelve behaviors are what you&#8217;d see watching over the shoulder. The telemetry adds three things you can&#8217;t see that way.</p><p><strong>The 7% is even smaller than it sounds.</strong> Nine human-authored prompts in 14.5 hours. Three were a word or a phrase &#8212; <code>proceed</code>, <code>let's do the top candidates</code>, <code>in console running?</code>. The longest content prompt I wrote all session was 84 characters: <code>it doessnt show the individual consoles, theses should be tabs - and it should be am spa</code> &#8212; typos and all. The orchestrator did not need a spec. It needed a direction and the occasional course correction.</p><p><strong>The economics are a cache story.</strong> The sonnet agents read 3,531 cached tokens for every fresh input token. The PM ran at 587&#215;; haiku, doing high-volume narrow tasks, hit 34,868&#215;. Most of each turn&#8217;s context is pulled from cache at cache prices, not re-sent as fresh input. A 14-hour run that re-paid full freight for context on every turn would cost a multiple of $486. The caching is why an orchestration this long is economically viable at all. A note on what I actually paid: I&#8217;m on Anthropic&#8217;s Max plan at $200/month. The $485.96 above is rack rate &#8212; the right figure for understanding the economic value of the output. My marginal cost for this session was zero.</p><p><strong>Waiting costs money, and you can see exactly where.</strong> The most expensive PM turns share one trait: <code>cache_read=0</code> &#8212; the turns where context had to be rebuilt cold after a cache miss, each one corresponding to a long gap I left. The $7.26 turn, the single priciest of the session, came right after I typed <code>proceed</code>, following 1.5 hours of silence; it wrote 364,618 tokens of cold context to restart the pipeline. The $5.82 turn came after my 4.5-hour wait. The orchestrator picked up where it left off both times, with full context. It just had to pay to reconstruct it. In a system like this, my latency has a line item.</p><p>And one finding the behaviors undersell: the failures. The first memory call errored and the PM fixed and retried it. The memory write later failed against a held lock, and the PM diagnosed the daemon contention and moved on. A subagent dropped its connection mid-run on an infra hiccup, and the PM recorded the socket error and later resumed the job to finish it. Three failures, three absorptions, zero escalations to me. That is the part that reads most like a year-one IC: not that nothing broke, but that what broke got handled below my line of sight.</p><h2>So What?</h2><p>Four threads run under all of this.</p><p>The first is where the line now sits between supervision and delegation. For most of these tools, the answer has been &#8220;supervise closely&#8221; &#8212; read every diff, catch every drift. This session moved the line. I supervised at the level of direction and architectural calls; I delegated everything from PR review to conflict avoidance to failure recovery. The PM did its own adversarial review 13 times and bounced a starred approval. Verification did not disappear; it moved inside the loop, and it left a receipt.</p><p>The second is what changes when the coordination work &#8212; not the typing &#8212; is the part the tool does well. The code these systems write stopped being the interesting question a while ago. The interesting development is that the orchestration layer now decomposes work, predicts conflicts, routes to specialists, tracks provenance, and knows when to hand back. That is project management. When the typing is cheap and the coordination is the hard part, a tool that coordinates well is worth more than a tool that types fast.</p><p>The third is the one underneath everything else. No single behavior above is impressive on its own. A competent IC reads the ticket history, names the parallel streams, predicts a merge conflict, routes to the right person, and stops at the decision boundary &#8212; without being told. What changed is that they now appear together, unprompted, in one session &#8212; and this time there&#8217;s a receipt for every one of them. The dollar figure is not the headline. The instrumentation is. We can finally watch the whole thing work and count what it cost, behavior by behavior.</p><p>The fourth is about where to point the question. For a while the useful question was &#8220;what can the model do.&#8221; After this session I think the better question is &#8220;what can the trifecta do,&#8221; because none of the behaviors above trace cleanly to a single component. They come out of the interaction: a model that audits its own work, a harness that isolates parallel work in separate worktrees, and an orchestration layer that routes to the right specialist with project memory already loaded. Pull any one of the three and the session degrades &#8212; the self-audit goes nowhere without a pipeline to surface it, the parallel waves collide without isolation, the routing misfires without memory. The capability lives in the seams between the parts, not in any one of them.</p><p>One last note on the economics. The Max plan is $200 a month. For that, I ran a session that bills at $486 rack rate, and the output is good enough that going back to a metered model is hard to imagine. Dealers give the first one away for a reason. The Max plan is that first bag: the work is compelling enough to be structurally addictive, and the flat rate strips out the per-token friction that might otherwise make you stop and think. I don&#8217;t mean that as a complaint. It&#8217;s a description of how the pricing works on the user.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and writes about AI-powered engineering at <a href="https://hyperdev.substack.com/">HyperDev</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://matsuoka.com/hyperdev/we-turned-a-corner/timeline.html">Explore the annotated session timeline</a> &#8212; the turn-by-turn companion infographic for this session</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul>]]></content:encoded></item><item><title><![CDATA[The Math Doesn't Work (Yet): Inside the AI Profitability Problem]]></title><description><![CDATA[Why Scaling Doesn't Lead To Profitability]]></description><link>https://hyperdev.matsuoka.com/p/the-math-doesnt-work-yet-inside-the</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/the-math-doesnt-work-yet-inside-the</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Wed, 10 Jun 2026 11:31:58 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!XBpX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!XBpX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!XBpX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png 424w, https://substackcdn.com/image/fetch/$s_!XBpX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png 848w, https://substackcdn.com/image/fetch/$s_!XBpX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png 1272w, https://substackcdn.com/image/fetch/$s_!XBpX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!XBpX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png" width="1195" height="896" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:896,&quot;width&quot;:1195,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1847405,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/201357245?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!XBpX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png 424w, https://substackcdn.com/image/fetch/$s_!XBpX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png 848w, https://substackcdn.com/image/fetch/$s_!XBpX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png 1272w, https://substackcdn.com/image/fetch/$s_!XBpX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h1>The Math Doesn&#8217;t Work (Yet): Inside the AI Profitability Problem</h1><p>OpenAI&#8217;s own projections show losses getting bigger as revenue gets bigger. Leaked investor documents reported across WSJ, Fortune, and The Information put the company at roughly <a href="https://fortune.com/2025/11/12/openai-cash-burn-rate-annual-losses-2028-profitable-2030-financial-documents/">$74 billion in operating losses in 2028</a> &#8212; on roughly $100 billion in projected revenue. That pairing is the headline. Not a smaller loss as scale arrives. A loss that grows faster than the top line.</p><p>That single relationship is the whole story. We tend to model AI companies as software businesses that will eventually grow into their cost structure the way SaaS companies did before them. The numbers say otherwise. These are capital-intensive infrastructure plays wearing software-company clothing, and the unit economics underneath them run in the opposite direction from the SaaS playbook most of us internalized over the last fifteen years.</p><p>A caveat before the numbers, because it matters for what&#8217;s below: neither OpenAI nor Anthropic publishes audited financials. Most figures here come from leaked investor decks, run-rate annualizations the companies announce in funding rounds, or SEC filings made by their cloud partners. The uncertainty is part of the analysis, not a footnote to it. When a specific timeline or quarterly figure couldn&#8217;t survive cross-checking against primary sources, I left it out.</p><h2>TL;DR</h2><ul><li><p>OpenAI&#8217;s leaked projections show operating losses <em>widening</em> as revenue grows &#8212; roughly <a href="https://fortune.com/2025/11/12/openai-cash-burn-rate-annual-losses-2028-profitable-2030-financial-documents/">$74B in losses on ~$100B revenue projected for 2028, with cumulative cash burn near $115B through 2029</a>.</p></li><li><p>In 2025 OpenAI spent about <a href="https://fortune.com/2025/11/12/openai-cash-burn-rate-annual-losses-2028-profitable-2030-financial-documents/">$1.69 for every dollar of revenue (~$9B net loss on ~$13B revenue)</a>, per leaked documents confirmed by multiple outlets.</p></li><li><p>Anthropic&#8217;s revenue trajectory is steep &#8212; roughly $1B annualized at the end of 2024 to a figure announced in the tens of billions by mid-2026 &#8212; with <a href="https://sacra.com/c/anthropic/">Claude Code alone reported at $2.5B annualized by February 2026</a>.</p></li><li><p><a href="https://www.investing.com/analysis/the-ai-token-pricing-crisis-behind-openai-and-anthropics-revenue-race-200680777">Inference token prices fell about 75% in a year</a>. Selling more AI makes the per-unit economics cheaper, which makes revenue growth harder, not easier.</p></li><li><p><a href="https://www.techtimes.com/articles/317542/20260601/ai-agent-economics-token-tax-locks-gross-margins-30-points-below-saas-baseline.htm">AI-native gross margins sit near 45% versus 75&#8211;85% for mature SaaS</a> &#8212; a structural gap of 23&#8211;33 points no company has yet closed.</p></li><li><p>No company has a verified path to profitability. Every specific breakeven-by-year claim I tried to confirm fell apart under scrutiny.</p></li></ul><h2>Two Companies, Two Shapes</h2><p>OpenAI ended 2025 at <a href="https://futurumgroup.com/insights/ai-capex-2026-the-690b-infrastructure-sprint/">roughly $20 billion in annualized revenue, a figure CFO Sarah Friar has stated directly</a>. That is a large business by any normal measure. It is also a business that, in the same year, spent about <a href="https://fortune.com/2025/11/12/openai-cash-burn-rate-annual-losses-2028-profitable-2030-financial-documents/">$1.69 for every dollar it took in &#8212; somewhere around a $9 billion net loss on roughly $13 billion in recognized revenue</a>, according to leaked documents that WSJ, Fortune, and The Information each reported. The company is majority funded by Microsoft, runs its compute primarily on Azure, and its strategy is scale-first: build the largest models, capture the most usage, and trust that revenue follows the curve.</p><p>Anthropic&#8217;s shape is different. Its revenue trajectory is steeper and more concentrated. The company grew from roughly $1 billion annualized at the end of 2024 to a figure it announced in the tens of billions by mid-2026 &#8212; <a href="https://sacra.com/c/anthropic/">the number it cited in its Series H materials</a>. I&#8217;m deliberately not pinning an exact figure to a month here, because the company has grown several-fold inside a single five-month window and any precise number is stale by the time you read it. What&#8217;s verifiable is the slope, and the slope is steep.</p><p>The more interesting detail is the concentration of value. Claude Code, one product, was reported at <a href="https://sacra.com/c/anthropic/">roughly $2.5 billion annualized by February 2026</a>. A single coding tool driving that much of a company&#8217;s run rate tells you something about where the margin-bearing demand actually lives. Anthropic is backed by Amazon (over $8 billion invested) and Google (over $2 billion), and it has <a href="https://techcrunch.com/2026/05/20/anthropic-will-pay-xai-1-25-billion-per-month-for-compute/">committed to spend more than $100 billion on AWS over ten years</a>, with roughly 1 GW of Trainium capacity targeted by the end of 2026. Two large companies funding it; one of them also selling it the silicon it runs on.</p><h2>Why Compute Is the Problem</h2><p>What separates these companies from every SaaS business you&#8217;ve evaluated: they don&#8217;t own their infrastructure. They rent it, at hyperscaler rates, from the same companies that fund them.</p><p>OpenAI&#8217;s Azure spend reportedly ran around <a href="https://www.theregister.com/2025/11/12/openai_spending_report/">$3.7 billion in 2024 and roughly $8.7 billion across the first three quarters of 2025</a>. Treat those numbers as medium-confidence &#8212; they come from leaked documents, and Microsoft pushed back that the figures &#8220;aren&#8217;t quite right.&#8221; But the direction is consistent with everything else: compute cost is the dominant line item, it&#8217;s largely fixed, and it grows with usage.</p><p>Anthropic&#8217;s arrangement produced one of the stranger details in modern enterprise finance I&#8217;ve read in some time. The company signed a deal for compute from Colossus 1 &#8212; Elon Musk&#8217;s Memphis data center, operated by xAI &#8212; at <a href="https://techcrunch.com/2026/05/20/anthropic-will-pay-xai-1-25-billion-per-month-for-compute/">roughly $1.25 billion per month for 300 MW of capacity, running through May 2029</a>. That&#8217;s not from a leak. It surfaced in SpaceX&#8217;s S-1 SEC filing and was confirmed by CNBC, Axios, and Data Center Dynamics, with a potential total value above $40 billion. There&#8217;s a 90-day mutual cancellation clause, so the headline total overstates the firm commitment. Still: Anthropic &#8212; funded by Google and Amazon &#8212; is paying Elon Musk&#8217;s company more than a billion dollars a month for compute. The AI capital world is stranger from the inside than the press releases suggest.</p><p>Zoom out and the renter problem gets sharper. Hyperscaler capex for 2026 is projected at <a href="https://futurumgroup.com/insights/ai-capex-2026-the-690b-infrastructure-sprint/">$660&#8211;690 billion</a>. Against that, OpenAI&#8217;s $20 billion ARR is roughly 3% of a single year&#8217;s data-center buildout by its suppliers. The companies selling AI applications are small tenants in an infrastructure market they don&#8217;t control and can&#8217;t currently price against.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!GBF-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88636760-5028-4c28-a47a-a9be5e14e2e7_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!GBF-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88636760-5028-4c28-a47a-a9be5e14e2e7_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!GBF-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88636760-5028-4c28-a47a-a9be5e14e2e7_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!GBF-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88636760-5028-4c28-a47a-a9be5e14e2e7_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!GBF-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88636760-5028-4c28-a47a-a9be5e14e2e7_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!GBF-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88636760-5028-4c28-a47a-a9be5e14e2e7_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/88636760-5028-4c28-a47a-a9be5e14e2e7_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1143520,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/201357245?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88636760-5028-4c28-a47a-a9be5e14e2e7_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!GBF-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88636760-5028-4c28-a47a-a9be5e14e2e7_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!GBF-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88636760-5028-4c28-a47a-a9be5e14e2e7_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!GBF-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88636760-5028-4c28-a47a-a9be5e14e2e7_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!GBF-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88636760-5028-4c28-a47a-a9be5e14e2e7_1024x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The Unit Economics Trap</h2><p>This story inverts an instinct most of us trust.</p><p>In normal software, even in the Cloud, scale is your friend. Marginal cost trends toward zero, gross margin climbs as you grow, and a mature SaaS business lands at <a href="https://www.techtimes.com/articles/317542/20260601/ai-agent-economics-token-tax-locks-gross-margins-30-points-below-saas-baseline.htm">75&#8211;85% gross margin</a> because serving the millionth customer costs almost nothing. Volume is the cure.</p><p>Inference doesn&#8217;t behave that way. Every token generated costs compute &#8212; real, metered, non-zero compute &#8212; so the marginal cost of serving usage stays stubbornly positive. And the price you can charge for that token is collapsing. Enterprise transaction data from Ramp shows <a href="https://www.investing.com/analysis/the-ai-token-pricing-crisis-behind-openai-and-anthropics-revenue-race-200680777">inference prices falling roughly 75% in a single year, from around $10 per million tokens to around $2.50</a>. Capability per dollar is improving fast, which is good for buyers and brutal for sellers, because it means the revenue you booked at last year&#8217;s prices reprices downward while your compute bill does not.</p><p>Put the two forces together and you get a squeeze that worsens with success. The better you are at selling inference, the more usage you drive; the more usage you drive, the more the per-unit price falls; the more it falls, the harder it is to grow revenue against a compute bill that scales with that same usage. Volume isn&#8217;t the cure here. Under these dynamics it&#8217;s part of the disease.</p><p>The survey data puts a number on how far this world sits from SaaS. ICONIQ Capital polled about 300 software executives and <a href="https://www.techtimes.com/articles/317542/20260601/ai-agent-economics-token-tax-locks-gross-margins-30-points-below-saas-baseline.htm">pegged AI-native gross margins at 41% in 2024, 45% in 2025, and a projected 52% in 2026</a>. Improving &#8212; but starting from a base 30-plus points below mature SaaS, and closing the distance slowly. A 52% gross margin is a respectable hardware business. It is a structurally difficult software business, especially one still spending heavily to grow.</p><h2>Two Different Bets</h2><p>OpenAI and Anthropic are running different experiments on how you eventually close that gap. Neither has been validated.</p><p>OpenAI&#8217;s bet is scale and breadth. Build the broadest platform, capture consumer and enterprise and API demand simultaneously, and assume that at sufficient scale you gain pricing power over compute, model-efficiency gains compound, and the revenue base grows fast enough to absorb the fixed cost. The leaked projections embody the risk in this bet: they show <a href="https://fortune.com/2025/11/12/openai-cash-burn-rate-annual-losses-2028-profitable-2030-financial-documents/">losses </a><em><a href="https://fortune.com/2025/11/12/openai-cash-burn-rate-annual-losses-2028-profitable-2030-financial-documents/">widening</a></em><a href="https://fortune.com/2025/11/12/openai-cash-burn-rate-annual-losses-2028-profitable-2030-financial-documents/"> through 2028 even as revenue approaches $100 billion, with cumulative cash burn near $115 billion through 2029</a>. The theory requires the curve to bend after the window we can currently see.</p><p>Anthropic&#8217;s bet is narrower and more product-led. Find a wedge where the work is valuable enough that buyers tolerate real prices, prove the margin there, and expand outward. Claude Code is that wedge made concrete &#8212; <a href="https://sacra.com/c/anthropic/">$2.5 billion annualized from developers</a> who pay because the output is worth more than the inference under it. Coding, agents, and enterprise automation are higher-value work than chat, and higher-value work supports prices that don&#8217;t immediately erode under token deflation. The risk: revenue concentration in a single product line, and an infrastructure bill &#8212; AWS commitments plus the xAI deal &#8212; that&#8217;s enormous relative to a company still proving the model.</p><p>Two theories of the same problem. Scale your way past the margin gap, or find work valuable enough that the gap doesn&#8217;t bind. We don&#8217;t yet have the data to say either works.</p><h2>What Would It Actually Take</h2><p>I&#8217;ll skip the timeline speculation &#8212; every specific breakeven-by-year claim I tried to verify died on contact with the sources. The structural requirements are clearer than the dates.</p><p>Three things have to move. First, gross margins have to climb from the mid-40s toward something defensible &#8212; call it 60-plus &#8212; and stay there while volume grows. That means model-efficiency gains (cheaper inference per unit of capability) have to outrun price deflation, rather than getting passed straight through to buyers as lower prices.</p><p>Second, these companies need pricing power over compute, which today they don&#8217;t have. At current scale they&#8217;re tenants. The open question is at what ARR a vendor becomes large enough to negotiate compute like a partner instead of a customer &#8212; or to build its own. Anthropic&#8217;s Trainium commitment and OpenAI&#8217;s various infrastructure moves are bets that vertical integration eventually changes the cost equation. That&#8217;s unproven, and it&#8217;s expensive in the interim.</p><p>Third, the product mix has to keep shifting toward work that resists deflation &#8212; enterprise agents, coding tools, automation that&#8217;s measured against labor cost rather than against the falling price of a token. Claude Code is the cleanest evidence that this category exists and that buyers will pay. Whether it&#8217;s a large enough share of total volume to lift blended margins across a company is the question that decides the whole thing.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!I6FB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b6f837-7c33-4d8b-9fce-0ef475dfb701_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!I6FB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b6f837-7c33-4d8b-9fce-0ef475dfb701_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!I6FB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b6f837-7c33-4d8b-9fce-0ef475dfb701_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!I6FB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b6f837-7c33-4d8b-9fce-0ef475dfb701_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!I6FB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b6f837-7c33-4d8b-9fce-0ef475dfb701_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!I6FB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b6f837-7c33-4d8b-9fce-0ef475dfb701_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b6b6f837-7c33-4d8b-9fce-0ef475dfb701_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1409865,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/201357245?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b6f837-7c33-4d8b-9fce-0ef475dfb701_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!I6FB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b6f837-7c33-4d8b-9fce-0ef475dfb701_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!I6FB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b6f837-7c33-4d8b-9fce-0ef475dfb701_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!I6FB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b6f837-7c33-4d8b-9fce-0ef475dfb701_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!I6FB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b6f837-7c33-4d8b-9fce-0ef475dfb701_1024x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The Open Questions</h2><p>What we don&#8217;t know outweighs what we do.</p><p>We don&#8217;t know the actual, current gross margins at either company &#8212; only the <a href="https://www.techtimes.com/articles/317542/20260601/ai-agent-economics-token-tax-locks-gross-margins-30-points-below-saas-baseline.htm">AI-native sector estimate of roughly 45%</a>. Neither company publishes the number that would settle the argument. We don&#8217;t know whether Anthropic&#8217;s stack of compute commitments, the decade-long AWS deal alongside the month-by-month xAI arrangement, creates structural tension or healthy redundancy. A company hedging across three infrastructure providers is either diversifying supply or revealing that no single supplier can meet its demand. We don&#8217;t know the ARR threshold at which compute pricing becomes negotiable, which is the hinge the entire margin story turns on.</p><p>And there&#8217;s the strategic risk that has no clean precedent: your infrastructure supplier is also your competitor. Microsoft ships Copilot. Amazon and Google both build models that compete with Anthropic&#8217;s. xAI builds Grok. Every dollar these companies pay for compute partly funds a rival&#8217;s model program. In normal software you don&#8217;t hand your gross margin to the company trying to beat you. Here it&#8217;s the default arrangement.</p><p>So the real question isn&#8217;t whether the AI labs are growing. They obviously are, faster than almost any companies in history. The question is whether revenue growth and margin improvement are the same trend or opposing ones. The SaaS era trained a generation of operators to believe that scale fixes economics. The leaked numbers describe a business where scale, so far, makes the loss bigger. Until one of these companies publishes a gross margin that shows the curve bending, that&#8217;s the math we have. And the math doesn&#8217;t work yet.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and writes about AI-powered engineering at HyperDev.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/p/the-first-70-era">The First 70% Era</a> &#8212; Where agentic AI delivers value and where it stops, and why the higher-value work resists token deflation</p></li><li><p><a href="https://hyperdev.matsuoka.com/p/ai-and-the-rise-of-the-hyperdev">AI and the Rise of the Hyperdev</a> &#8212; Why developers pay real money for AI tooling, the demand side of the margin story</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul>]]></content:encoded></item><item><title><![CDATA[What’s Old Is New Again]]></title><description><![CDATA[Nine classic SDLC practices that AI finally makes practical]]></description><link>https://hyperdev.matsuoka.com/p/whats-old-is-new-again</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/whats-old-is-new-again</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Wed, 03 Jun 2026 11:31:06 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!FRWK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b2c67b8-b815-4f06-a401-6d57b2a9d884_873x576.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!FRWK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b2c67b8-b815-4f06-a401-6d57b2a9d884_873x576.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!FRWK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b2c67b8-b815-4f06-a401-6d57b2a9d884_873x576.png 424w, https://substackcdn.com/image/fetch/$s_!FRWK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b2c67b8-b815-4f06-a401-6d57b2a9d884_873x576.png 848w, https://substackcdn.com/image/fetch/$s_!FRWK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b2c67b8-b815-4f06-a401-6d57b2a9d884_873x576.png 1272w, https://substackcdn.com/image/fetch/$s_!FRWK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b2c67b8-b815-4f06-a401-6d57b2a9d884_873x576.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!FRWK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b2c67b8-b815-4f06-a401-6d57b2a9d884_873x576.png" width="873" height="576" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0b2c67b8-b815-4f06-a401-6d57b2a9d884_873x576.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:576,&quot;width&quot;:873,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1435244,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/200236698?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43b89068-cf8e-4789-95f1-e357c61076b0_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!FRWK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b2c67b8-b815-4f06-a401-6d57b2a9d884_873x576.png 424w, https://substackcdn.com/image/fetch/$s_!FRWK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b2c67b8-b815-4f06-a401-6d57b2a9d884_873x576.png 848w, https://substackcdn.com/image/fetch/$s_!FRWK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b2c67b8-b815-4f06-a401-6d57b2a9d884_873x576.png 1272w, https://substackcdn.com/image/fetch/$s_!FRWK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b2c67b8-b815-4f06-a401-6d57b2a9d884_873x576.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Most of the best ideas in software engineering aren&#8217;t new. They&#8217;ve been written up in books, argued over at conferences, taught in every &#8220;best practices&#8221; deck since the late 1990s. And most teams quietly don&#8217;t do them.</p><p>Not because anyone thinks they&#8217;re wrong. Test-driven development, design by contract, architecture decision records, mutation testing &#8212; ask a room of senior engineers whether these are good ideas and you&#8217;ll get nods. Ask the same room who practices them consistently under deadline pressure and the hands stay down. I&#8217;ve been in that room for twenty-five years, on both sides of the question. I&#8217;ve also been the engineering leader who let those practices slip because shipping the feature mattered more this quarter.</p><p>There&#8217;s a single economic reason these practices lose. The upfront cost is high, the payoff is real but distant, and human attention is the binding constraint. Write the test before the code, document the decision, specify the invariant &#8212; every one of those is a tax you pay now against a benefit you collect later, maybe, if the project lives long enough. Under deadline pressure, that&#8217;s a losing trade for a human. So we skip it, ship, and pay the interest later in bugs and confusion. Call it the impatience tax.</p><p>Agents don&#8217;t pay that tax. They have infinite patience for upfront rigor and roughly zero marginal cost for the tedious work that rigor demands. Writing a thorough test suite for code that doesn&#8217;t exist yet is psychologically brutal for a person and completely fine for a model. That single shift &#8212; the cost of patience going to zero &#8212; quietly inverts the economics of a whole list of practices we knew were right and gave up on anyway.</p><p>This isn&#8217;t a piece about what AI makes <em>possible</em>. Lots of things are possible. It&#8217;s about a narrower, more useful question: which disciplines did we already agree were correct, fight about for decades, and abandon for reasons that no longer hold?</p><p>Here are nine.</p><h2>TL;DR</h2><ul><li><p>These nine practices share one structure: high upfront cost, distant payoff. Human attention is the constraint that kills them under deadline pressure.</p></li><li><p>AI removes the constraint. A failing test is the clearest prompt you can hand an agent; a spec is its input; an ADR is its context. The discipline becomes the interface.</p></li><li><p>TDD, design by contract, and property-based testing turn from &#8220;things we should do&#8221; into the most effective way to <em>constrain</em> agent behavior and prevent hallucinated correctness.</p></li><li><p>Documentation, ADRs, and living docs get a bilateral ROI: agents generate them from code, and they make agents far more effective in your codebase.</p></li><li><p>The catch is real. A 2025 METR randomized trial found experienced developers were about 19% <em>slower</em> with AI assistance. These practices pay off only when AI is used with discipline, not as autocomplete.</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!hmow!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff117aee0-1765-4975-a0c5-dd7bdddbe34d_1024x431.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!hmow!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff117aee0-1765-4975-a0c5-dd7bdddbe34d_1024x431.png 424w, https://substackcdn.com/image/fetch/$s_!hmow!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff117aee0-1765-4975-a0c5-dd7bdddbe34d_1024x431.png 848w, https://substackcdn.com/image/fetch/$s_!hmow!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff117aee0-1765-4975-a0c5-dd7bdddbe34d_1024x431.png 1272w, https://substackcdn.com/image/fetch/$s_!hmow!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff117aee0-1765-4975-a0c5-dd7bdddbe34d_1024x431.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!hmow!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff117aee0-1765-4975-a0c5-dd7bdddbe34d_1024x431.png" width="1024" height="431" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f117aee0-1765-4975-a0c5-dd7bdddbe34d_1024x431.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:431,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1083571,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/200236698?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f4ba76c-18f5-413a-bf35-a56dc7861a00_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!hmow!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff117aee0-1765-4975-a0c5-dd7bdddbe34d_1024x431.png 424w, https://substackcdn.com/image/fetch/$s_!hmow!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff117aee0-1765-4975-a0c5-dd7bdddbe34d_1024x431.png 848w, https://substackcdn.com/image/fetch/$s_!hmow!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff117aee0-1765-4975-a0c5-dd7bdddbe34d_1024x431.png 1272w, https://substackcdn.com/image/fetch/$s_!hmow!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff117aee0-1765-4975-a0c5-dd7bdddbe34d_1024x431.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>1. Test-Driven Development</h2><p>Start with the practice most teams abandoned first.</p><p>Writing tests before code was always the theoretically superior move. It forces you to define the interface before you build behind it, catches bugs at the moment of definition instead of during integration, and leaves behind a living specification of what the code is supposed to do. Kent Beck made the case decades ago and the case held up.</p><p>Almost nobody did it consistently. The reason is psychological, not technical. Writing detailed tests for code that doesn&#8217;t exist yet, while a deadline breathes on your neck, feels like building scaffolding for a house you haven&#8217;t designed. Your brain screams at you to just write the function. So you write the function, promise yourself you&#8217;ll add tests after, and &#8212; well. You know how that goes.</p><p>Now flip the perspective. To an agent, a failing test isn&#8217;t scaffolding. It&#8217;s the clearest possible specification of intent you can provide. &#8220;Make this pass, don&#8217;t break anything else&#8221; is an unambiguous, machine-checkable instruction, which is exactly what a probabilistic system needs to stay honest. The test suite becomes a guardrail that prevents the most dangerous failure mode in AI-assisted coding: confident, plausible, wrong. Hallucinated correctness dies against a red bar.</p><p>TDD went from the discipline most teams couldn&#8217;t sustain to one of the best tools we have for bounding what an agent is allowed to claim it did. Same practice. Opposite economics.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!a841!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb60534e-05b5-4873-ab47-586bb7ec68bd_1024x670.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!a841!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb60534e-05b5-4873-ab47-586bb7ec68bd_1024x670.png 424w, https://substackcdn.com/image/fetch/$s_!a841!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb60534e-05b5-4873-ab47-586bb7ec68bd_1024x670.png 848w, https://substackcdn.com/image/fetch/$s_!a841!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb60534e-05b5-4873-ab47-586bb7ec68bd_1024x670.png 1272w, https://substackcdn.com/image/fetch/$s_!a841!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb60534e-05b5-4873-ab47-586bb7ec68bd_1024x670.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!a841!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb60534e-05b5-4873-ab47-586bb7ec68bd_1024x670.png" width="1024" height="670" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fb60534e-05b5-4873-ab47-586bb7ec68bd_1024x670.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:670,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1640038,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/200236698?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a4cda20-f557-4918-88d1-529fec00bdee_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!a841!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb60534e-05b5-4873-ab47-586bb7ec68bd_1024x670.png 424w, https://substackcdn.com/image/fetch/$s_!a841!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb60534e-05b5-4873-ab47-586bb7ec68bd_1024x670.png 848w, https://substackcdn.com/image/fetch/$s_!a841!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb60534e-05b5-4873-ab47-586bb7ec68bd_1024x670.png 1272w, https://substackcdn.com/image/fetch/$s_!a841!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb60534e-05b5-4873-ab47-586bb7ec68bd_1024x670.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>2. Spec-Driven Development and Design by Contract</h2><p>Bertrand Meyer formalized design by contract in the 1980s and built it into the Eiffel language: specify preconditions, postconditions, and invariants, then let the implementation follow from the contract.</p><p>The idea was sound and the adoption was thin, for one stubborn economic reason: the contract only pays off if someone <em>else</em> writes the implementation from it. If you&#8217;re writing both the spec and the code, the spec is overhead &#8212; you already know what you meant. The contract&#8217;s value lives in the handoff, and for most of software history there was no cheap handoff to hand it to.</p><p>Now there is. You write the contract; the agent writes the implementation from it. Spec-driven development stops being a documentation chore and becomes the actual control surface for delegation. The spec is the part requiring human judgment about <em>what</em> the system should do. The implementation &#8212; the part that used to eat the hours &#8212; is the part you delegate. Meyer&#8217;s economics finally close, forty years late, because the missing party in the transaction showed up.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!P-wa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5605bf-95f5-490e-b318-7334ff70ae87_1024x487.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!P-wa!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5605bf-95f5-490e-b318-7334ff70ae87_1024x487.png 424w, https://substackcdn.com/image/fetch/$s_!P-wa!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5605bf-95f5-490e-b318-7334ff70ae87_1024x487.png 848w, https://substackcdn.com/image/fetch/$s_!P-wa!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5605bf-95f5-490e-b318-7334ff70ae87_1024x487.png 1272w, https://substackcdn.com/image/fetch/$s_!P-wa!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5605bf-95f5-490e-b318-7334ff70ae87_1024x487.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!P-wa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5605bf-95f5-490e-b318-7334ff70ae87_1024x487.png" width="1024" height="487" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ee5605bf-95f5-490e-b318-7334ff70ae87_1024x487.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:487,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1370470,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/200236698?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82d6c7c5-50c0-4b34-9a46-0700c4bf1e82_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!P-wa!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5605bf-95f5-490e-b318-7334ff70ae87_1024x487.png 424w, https://substackcdn.com/image/fetch/$s_!P-wa!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5605bf-95f5-490e-b318-7334ff70ae87_1024x487.png 848w, https://substackcdn.com/image/fetch/$s_!P-wa!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5605bf-95f5-490e-b318-7334ff70ae87_1024x487.png 1272w, https://substackcdn.com/image/fetch/$s_!P-wa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5605bf-95f5-490e-b318-7334ff70ae87_1024x487.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>3. Architecture Decision Records</h2><p>Why does the codebase look like this? Why Postgres and not Dynamo, why this queue, why the weird module boundary that everyone trips over?</p><p>ADRs were the right answer to that question &#8212; a short dated record of each significant decision, the context, and the alternatives rejected. The discipline almost never held. Same shape as everything else here: the cost is immediate (stop, write the thing) and the value accrues slowly, mostly to some future engineer who isn&#8217;t in the room yet.</p><p>Two things flipped at once, which makes this one more interesting than the rest. First, agents can generate ADRs from an existing codebase &#8212; read the git history, the dependency choices, the structure, and reconstruct the decisions that produced them. The retroactive cost of documentation drops toward zero. Second, and this is the part people miss: existing ADRs dramatically improve what an agent can do <em>in</em> your codebase. An agent that can read why you chose eventual consistency won&#8217;t keep proposing changes that assume strong consistency.</p><p>So the ROI went bilateral. Agents help you write ADRs, and ADRs help agents help you. The practice that used to only cost now pays on both ends.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cfey!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F256e6ffa-ac04-401b-bd1d-276e0dc51f21_1024x445.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cfey!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F256e6ffa-ac04-401b-bd1d-276e0dc51f21_1024x445.png 424w, https://substackcdn.com/image/fetch/$s_!cfey!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F256e6ffa-ac04-401b-bd1d-276e0dc51f21_1024x445.png 848w, https://substackcdn.com/image/fetch/$s_!cfey!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F256e6ffa-ac04-401b-bd1d-276e0dc51f21_1024x445.png 1272w, https://substackcdn.com/image/fetch/$s_!cfey!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F256e6ffa-ac04-401b-bd1d-276e0dc51f21_1024x445.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cfey!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F256e6ffa-ac04-401b-bd1d-276e0dc51f21_1024x445.png" width="1024" height="445" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/256e6ffa-ac04-401b-bd1d-276e0dc51f21_1024x445.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:445,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1079171,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/200236698?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0eca5b97-c8ef-45f0-97de-eb23cc9969c2_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!cfey!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F256e6ffa-ac04-401b-bd1d-276e0dc51f21_1024x445.png 424w, https://substackcdn.com/image/fetch/$s_!cfey!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F256e6ffa-ac04-401b-bd1d-276e0dc51f21_1024x445.png 848w, https://substackcdn.com/image/fetch/$s_!cfey!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F256e6ffa-ac04-401b-bd1d-276e0dc51f21_1024x445.png 1272w, https://substackcdn.com/image/fetch/$s_!cfey!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F256e6ffa-ac04-401b-bd1d-276e0dc51f21_1024x445.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>4. Continuous Code Review</h2><p>&#8220;Catch issues early&#8221; has sat on every best-practices list since Extreme Programming put continuous review on the map. The advice was never controversial. The bottleneck was always the same: human reviewer attention is finite, expensive, and easily exhausted. So review collapsed into batch PR review &#8212; a tired engineer reading a 600-line diff on a Friday afternoon, approving most of it on faith.</p><p>AI review on every commit &#8212; not batched at the PR boundary, but running as code lands &#8212; is moving from aspirational toward baseline: increasingly the default on teams that have wired it in, though not yet universal. The marginal cost of a careful read went to nearly nothing, and the read happens while the context is still warm.</p><p>But this one comes with an emergent problem worth naming, because it bites teams that adopt the tooling without rethinking the model. PRs are getting larger and arriving faster under AI-assisted development. An agent can produce a 2,000-line change in an afternoon. If your review model still routes everything through a human approver at the end, that human is now the rate limiter, drowning in volume they didn&#8217;t generate and can&#8217;t realistically read. AI review on every commit is part of the answer. The harder part is restructuring <em>what</em> the human reviews &#8212; architecture, intent, the decisions a model shouldn&#8217;t make alone &#8212; and letting the machine handle line-level correctness continuously. Adopt the tool without rethinking the workflow and you&#8217;ve just built a faster way to overwhelm your best reviewer.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!SXYP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2eb6d6b9-de10-452f-8b64-7bb0bcba1041_1024x456.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!SXYP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2eb6d6b9-de10-452f-8b64-7bb0bcba1041_1024x456.png 424w, https://substackcdn.com/image/fetch/$s_!SXYP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2eb6d6b9-de10-452f-8b64-7bb0bcba1041_1024x456.png 848w, https://substackcdn.com/image/fetch/$s_!SXYP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2eb6d6b9-de10-452f-8b64-7bb0bcba1041_1024x456.png 1272w, https://substackcdn.com/image/fetch/$s_!SXYP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2eb6d6b9-de10-452f-8b64-7bb0bcba1041_1024x456.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!SXYP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2eb6d6b9-de10-452f-8b64-7bb0bcba1041_1024x456.png" width="1024" height="456" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2eb6d6b9-de10-452f-8b64-7bb0bcba1041_1024x456.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:456,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1098227,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/200236698?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ce64d7a-b849-4c7e-b523-2565960ee237_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!SXYP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2eb6d6b9-de10-452f-8b64-7bb0bcba1041_1024x456.png 424w, https://substackcdn.com/image/fetch/$s_!SXYP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2eb6d6b9-de10-452f-8b64-7bb0bcba1041_1024x456.png 848w, https://substackcdn.com/image/fetch/$s_!SXYP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2eb6d6b9-de10-452f-8b64-7bb0bcba1041_1024x456.png 1272w, https://substackcdn.com/image/fetch/$s_!SXYP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2eb6d6b9-de10-452f-8b64-7bb0bcba1041_1024x456.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>5. Pair Programming</h2><p>Pairing always looked expensive in the most obvious way: two engineers, one task, double the salary against a single unit of output. That intuition was wrong on the numbers &#8212; the measured overhead from the pair-programming studies was closer to 15%, often recovered through fewer defects &#8212; but the 2&#215; gut feeling is what drove the decisions. The benefits were real &#8212; knowledge transfer, real-time review, fewer dumb mistakes &#8212; but the perceived math meant most teams reserved it for critical paths, gnarly bugs, or onboarding a new hire. A luxury, rationed.</p><p>The pair is now a human and an agent, and it&#8217;s available to every engineer continuously, not rationed to the critical path. The knowledge-transfer benefit generalizes &#8212; the agent can explain unfamiliar parts of the codebase on demand. The real-time-review benefit generalizes &#8212; a second set of eyes on every line, every time, without scheduling two calendars. The economics that made pairing a rationed luxury simply don&#8217;t apply when one half of the pair has near-zero marginal cost.</p><p>Worth a caveat: agent-as-pair is genuinely good at the review and explanation half of pairing, and weaker at the part where a human partner pushes back on a bad <em>design</em> before you&#8217;ve written a line. You still need humans pairing with humans for that. But the day-to-day, line-by-line version of pairing just became free, and that&#8217;s most of what pairing was for.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!DHDP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e5b59c-a9a7-44ae-8725-1333f8e08687_1024x592.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!DHDP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e5b59c-a9a7-44ae-8725-1333f8e08687_1024x592.png 424w, https://substackcdn.com/image/fetch/$s_!DHDP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e5b59c-a9a7-44ae-8725-1333f8e08687_1024x592.png 848w, https://substackcdn.com/image/fetch/$s_!DHDP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e5b59c-a9a7-44ae-8725-1333f8e08687_1024x592.png 1272w, https://substackcdn.com/image/fetch/$s_!DHDP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e5b59c-a9a7-44ae-8725-1333f8e08687_1024x592.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!DHDP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e5b59c-a9a7-44ae-8725-1333f8e08687_1024x592.png" width="1024" height="592" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/18e5b59c-a9a7-44ae-8725-1333f8e08687_1024x592.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:592,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1439433,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/200236698?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F611b1613-4024-4d10-8e69-2b2a1936432d_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!DHDP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e5b59c-a9a7-44ae-8725-1333f8e08687_1024x592.png 424w, https://substackcdn.com/image/fetch/$s_!DHDP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e5b59c-a9a7-44ae-8725-1333f8e08687_1024x592.png 848w, https://substackcdn.com/image/fetch/$s_!DHDP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e5b59c-a9a7-44ae-8725-1333f8e08687_1024x592.png 1272w, https://substackcdn.com/image/fetch/$s_!DHDP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e5b59c-a9a7-44ae-8725-1333f8e08687_1024x592.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>6. Mutation Testing</h2><p>Coverage numbers lie, and most engineers know it. Eighty percent line coverage tells you eighty percent of your lines got executed by a test &#8212; not that any of those tests would <em>notice</em> if the behavior broke. Mutation testing is the honest measure: it deliberately introduces bugs (flip a comparison, drop a line, change a constant) and checks whether your test suite catches them. If a mutant survives, you have a test that runs code without actually validating it.</p><p>Mutation testing was always the gold standard and almost never run continuously, for one reason: it&#8217;s computationally expensive. You&#8217;re effectively running your whole suite many times over, once per mutation. On a real codebase that&#8217;s brutal. So it lived in research papers and the occasional heroic CI job that someone eventually disabled for being too slow.</p><p>That constraint is mostly gone &#8212; compute is cheap and parallel, and we got more comfortable spending it. And the practice arrived right when we suddenly need it most. AI-generated tests have a characteristic failure mode: they drift toward coverage metrics without meaningful assertions. The model writes a test that calls the function, exercises the path, and asserts almost nothing of substance &#8212; green checkmark, zero protection. Coverage looks great. Mutation testing is the thing that catches exactly that. It&#8217;s the verification layer for a verification layer, and it matters more now than when it was invented, because now a machine is writing the tests.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!EY9I!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70ad3dac-f7cb-48da-9d0a-a6098d28cbe6_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!EY9I!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70ad3dac-f7cb-48da-9d0a-a6098d28cbe6_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!EY9I!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70ad3dac-f7cb-48da-9d0a-a6098d28cbe6_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!EY9I!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70ad3dac-f7cb-48da-9d0a-a6098d28cbe6_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!EY9I!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70ad3dac-f7cb-48da-9d0a-a6098d28cbe6_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!EY9I!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70ad3dac-f7cb-48da-9d0a-a6098d28cbe6_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/70ad3dac-f7cb-48da-9d0a-a6098d28cbe6_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1482204,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/200236698?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70ad3dac-f7cb-48da-9d0a-a6098d28cbe6_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!EY9I!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70ad3dac-f7cb-48da-9d0a-a6098d28cbe6_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!EY9I!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70ad3dac-f7cb-48da-9d0a-a6098d28cbe6_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!EY9I!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70ad3dac-f7cb-48da-9d0a-a6098d28cbe6_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!EY9I!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70ad3dac-f7cb-48da-9d0a-a6098d28cbe6_1024x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>7. Living Documentation</h2><p>Documentation was always supposed to be a first-class artifact. It almost never was, and the reason is by now familiar: writing docs is tedious, and the penalty for stale docs accrues slowly and lands on someone else. So docs rotted. Every team has a wiki that&#8217;s a graveyard of half-true pages from two reorgs ago.</p><p>AI changes both halves of the equation at once. It generates docs from code, so the writing cost drops. And it <em>consumes</em> docs as context, so the docs earn their keep immediately &#8212; a well-documented codebase is a measurably more useful codebase for an agent working in it. The ROI is immediate and bilateral, same structure as ADRs.</p><p>There&#8217;s a quiet shift hiding in there. Documentation used to be written for humans who&#8217;d mostly never read it. Now it&#8217;s also written for the agent that will read it on every task, which means stale docs don&#8217;t just confuse a future engineer &#8212; they actively degrade your tooling today. The feedback loop tightened from months to minutes. That&#8217;s the kind of change that actually moves behavior, because the cost of skipping it shows up now instead of later.</p><h2>8. Runbook Generation from Incidents</h2><p>On-call always leaned too hard on tribal knowledge. The person who knows why the payment service wedges at 3 a.m. is asleep, on vacation, or left the company last spring. Writing a runbook after each incident was obviously the right move and reliably the thing nobody did, because the incident was <em>over</em> and everyone wanted to go back to bed.</p><p>Incidents become runbooks automatically now. The agent has the incident timeline, the chat transcript, the commands that resolved it, the postmortem &#8212; and it can turn that into a structured runbook while the details are fresh, without asking an exhausted engineer to relive the night. The cost that used to fall right when motivation was lowest now falls on a system that doesn&#8217;t get tired or resentful.</p><p>I&#8217;d treat the generated runbook as a draft a human still signs off on, not gospel. But &#8220;imperfect draft, reviewed in five minutes&#8221; beats &#8220;blank page nobody ever fills in,&#8221; and that was always the real competition.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!dM5v!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c754a4-9b5b-424f-b008-30dbcb217887_1024x693.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!dM5v!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c754a4-9b5b-424f-b008-30dbcb217887_1024x693.png 424w, https://substackcdn.com/image/fetch/$s_!dM5v!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c754a4-9b5b-424f-b008-30dbcb217887_1024x693.png 848w, https://substackcdn.com/image/fetch/$s_!dM5v!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c754a4-9b5b-424f-b008-30dbcb217887_1024x693.png 1272w, https://substackcdn.com/image/fetch/$s_!dM5v!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c754a4-9b5b-424f-b008-30dbcb217887_1024x693.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!dM5v!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c754a4-9b5b-424f-b008-30dbcb217887_1024x693.png" width="1024" height="693" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d0c754a4-9b5b-424f-b008-30dbcb217887_1024x693.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:693,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1615996,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/200236698?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3f83774-28c6-4b46-addd-83efad33390b_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!dM5v!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c754a4-9b5b-424f-b008-30dbcb217887_1024x693.png 424w, https://substackcdn.com/image/fetch/$s_!dM5v!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c754a4-9b5b-424f-b008-30dbcb217887_1024x693.png 848w, https://substackcdn.com/image/fetch/$s_!dM5v!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c754a4-9b5b-424f-b008-30dbcb217887_1024x693.png 1272w, https://substackcdn.com/image/fetch/$s_!dM5v!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c754a4-9b5b-424f-b008-30dbcb217887_1024x693.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>9. Property-Based Testing</h2><p>Example-based tests check the cases you thought of. Property-based testing is stronger: you specify the <em>invariants</em> a system must always satisfy &#8212; reversing a list twice returns the original, a serialized-then-deserialized object equals the original, the account balance never goes negative &#8212; and the framework generates hundreds of adversarial inputs trying to break them. QuickCheck pioneered the approach; it finds the edge cases you&#8217;d never have written by hand.</p><p>It never went mainstream outside a few communities, and the bottleneck wasn&#8217;t tooling &#8212; good property-based libraries exist for most languages. The bottleneck was writing good property specifications. Identifying the right invariants requires deep domain reasoning: you have to understand the system well enough to state what must <em>always</em> be true, which is harder than writing a few example cases. Most engineers, under pressure, defaulted to the easier thing.</p><p>This is where AI helps in a way that&#8217;s less obvious than &#8220;it writes the code.&#8221; A model can generate property suites from a spec, and &#8212; more usefully &#8212; it can reason about <em>what invariants a system should satisfy</em> in the first place, surfacing properties you hadn&#8217;t articulated. That&#8217;s the expensive, judgment-heavy part it actually offloads. Combined with mutation testing to keep the generated properties honest, you get a testing approach that was always more powerful than example-based testing and was always too expensive in human reasoning to adopt widely.</p><h2>The Catch</h2><p>I&#8217;d be selling you something if I stopped there, and the data won&#8217;t let me.</p><p>In July 2025, METR ran a randomized controlled trial with experienced open-source developers working on real tasks in repositories they knew well. The developers expected AI assistance to speed them up. It slowed them down &#8212; by roughly 19%. METR&#8217;s February 2026 follow-up found that gap narrowing, and reversing on some measures, as the same kind of developers gained real experience with the tools &#8212; which is to say the 19% was a snapshot of the unfamiliar, undisciplined path, and it closes precisely as people pick up the habits this piece is about.</p><p>That finding is real and it isn&#8217;t a contradiction of everything above. It&#8217;s the missing condition. Every practice in this piece works <em>because</em> it imposes structure on the agent &#8212; TDD as a guardrail, the spec as input, the ADR as context, mutation testing as the check on the check. Used that way, with discipline, AI is constrained toward correctness. Used the other way &#8212; as autocomplete, as a vibe-coding partner you don&#8217;t supervise &#8212; you get more code, faster, with less correctness and a slower path to done once you account for the cleanup. The METR developers, working in code they already understood deeply, may well have been paying exactly that tax: accepting plausible suggestions that took longer to vet and fix than writing it themselves would have.</p><p>So the inversion isn&#8217;t automatic. The cost of patience dropped to zero, which makes the rigorous path finally affordable. It does not make the undisciplined path good. If anything it makes discipline more important, because a tool that produces plausible output at high volume is precisely the tool that most needs a guardrail you can&#8217;t talk your way past. A red test bar doesn&#8217;t care how confident the model sounds.</p><h2>What To Do With This</h2><p>The interesting question was never &#8220;what does AI make possible.&#8221; That list is enormous, mostly speculative, and not very actionable. The better question is the one this whole piece is built on: which practices did we already know were right, argue about for decades, and quietly give up on?</p><p>That list is short, specific, and yours to write. Go pull your own team&#8217;s &#8220;we should really do this but we don&#8217;t&#8221; backlog &#8212; the standing items in retros that everyone agrees with and nobody owns. I&#8217;d bet most of them have the same economic shape: high upfront cost, distant payoff, killed by human impatience under deadline. Test coverage on the legacy module. The runbooks. The ADRs for the three decisions everyone keeps re-litigating. The integration tests that would&#8217;ve caught last quarter&#8217;s outage.</p><p>Run each one through a single question: was this abandoned because it was <em>wrong</em>, or because it was <em>expensive in human patience</em>? The wrong ones, leave abandoned. The expensive-in-patience ones just got cheap. Those are the ones to pick back up first.</p><p>The impatience tax got repealed. The disciplines it used to make unaffordable are sitting right there, mostly unchanged, waiting for someone to notice the price changed.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and writes about AI-powered engineering at <a href="https://hyperdev.substack.com/">HyperDev</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p>The Other Shoe Has Dropped &#8212; Why enterprise AI bills don&#8217;t match the per-token price collapse</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul>]]></content:encoded></item><item><title><![CDATA[The Era of the Leader/Practitioner]]></title><description><![CDATA[Putting "Do" back into "Lead"]]></description><link>https://hyperdev.matsuoka.com/p/the-era-of-the-leaderpractitioner</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/the-era-of-the-leaderpractitioner</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Mon, 01 Jun 2026 11:31:28 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!E_Vr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3731b710-763c-4f36-860f-560f95b11d2b_1024x700.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!E_Vr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3731b710-763c-4f36-860f-560f95b11d2b_1024x700.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!E_Vr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3731b710-763c-4f36-860f-560f95b11d2b_1024x700.png 424w, https://substackcdn.com/image/fetch/$s_!E_Vr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3731b710-763c-4f36-860f-560f95b11d2b_1024x700.png 848w, https://substackcdn.com/image/fetch/$s_!E_Vr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3731b710-763c-4f36-860f-560f95b11d2b_1024x700.png 1272w, https://substackcdn.com/image/fetch/$s_!E_Vr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3731b710-763c-4f36-860f-560f95b11d2b_1024x700.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!E_Vr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3731b710-763c-4f36-860f-560f95b11d2b_1024x700.png" width="1024" height="700" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3731b710-763c-4f36-860f-560f95b11d2b_1024x700.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:700,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1378805,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/200067020?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6656d57-be88-4993-a972-b7c0c5fd743d_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!E_Vr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3731b710-763c-4f36-860f-560f95b11d2b_1024x700.png 424w, https://substackcdn.com/image/fetch/$s_!E_Vr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3731b710-763c-4f36-860f-560f95b11d2b_1024x700.png 848w, https://substackcdn.com/image/fetch/$s_!E_Vr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3731b710-763c-4f36-860f-560f95b11d2b_1024x700.png 1272w, https://substackcdn.com/image/fetch/$s_!E_Vr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3731b710-763c-4f36-860f-560f95b11d2b_1024x700.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Something shifted in the last year, and it took me a while to name it.</p><p>A growing number of people are running organizations while still doing real hands-on technical work. Not as a hobby, not on weekends, not as a vanity exercise to keep their commit graph green. They are building things their teams depend on &#8212; and they are doing it as a deliberate part of the job, not in the cracks between meetings. The work has a specific shape. It is rarely a production feature. It is the layer underneath: developer productivity tooling, internal services, agentic harnesses, MCP connectors, the infrastructure that unblocks everyone else.</p><p>For most of my career this combination didn&#8217;t really hold together. You could code or you could lead, and the moment you tried to do both seriously, one of them rotted. I&#8217;ve watched plenty of technical executives keep a foot in the codebase and slowly become the bottleneck everyone routed around politely. The pattern was familiar enough to be a warning.</p><p>What changed is not that leaders suddenly got more disciplined. It&#8217;s that the time cost of meaningful technical contribution collapsed. Agentic coding made a hybrid role viable that wasn&#8217;t viable before &#8212; and the more I look at it, the more I think this isn&#8217;t a new invention at all. It&#8217;s the recovery of a very old idea that modern specialization interrupted.</p><p>I&#8217;m writing this as someone living in the middle of it. I spend roughly 30% of my time coding. And when I say coding, I mean directing a team of agents &#8212; much closer to that than hands-on work, which is probably the whole point. That&#8217;s not a full-time IC&#8217;s week, and it isn&#8217;t meant to be. It&#8217;s enough to stay close to the work that matters and to build the enabling infrastructure I think is worth my own hands on the keyboard.</p><h2>TL;DR</h2><ul><li><p>A distinct role is emerging: leaders who run organizations and still do deep technical work &#8212; specifically enabling/institutional work (tooling, harnesses, internal services), not critical-path product features.</p></li><li><p>This satisfies Charity Majors&#8217; actual advice. Her line was never &#8220;stop coding.&#8221; It was &#8220;stop writing code in the critical path.&#8221; Enabling work fits that exactly.</p></li><li><p>The integration of strategist and practitioner has deep cross-cultural precedent &#8212; Japan&#8217;s <em>bunbu-ry&#333;d&#333;</em>, Rome&#8217;s Marcus Aurelius, China&#8217;s <em>wen-wu</em>, the Renaissance polymath, Mattis&#8217;s &#8220;warrior monk.&#8221; The modern role is a recovery, not a novelty.</p></li><li><p>Agentic coding is what makes it newly viable: focused sessions now deliver output that once required sustained, uninterrupted immersion. The Anthropic 2026 data shows ~27% of AI-assisted work is work that &#8220;wouldn&#8217;t have been done otherwise.&#8221;</p></li><li><p>Directing agents feels like delegation &#8212; the same skill leaders already use with human reports. Which is why experienced leaders adapt to it more naturally than juniors do.</p></li></ul><h2>The pattern, named</h2><p>The difficulty with the coding executive was never philosophical. It was attentional.</p><p>Charity Majors mapped this years ago in <a href="https://charity.wtf/2017/05/11/the-engineer-manager-pendulum/">The Engineer/Manager Pendulum</a>, and the piece holds up because she was precise about the mechanism. Management is interruptive by design &#8212; your job is to be available, to unblock, to absorb the chaos so your team doesn&#8217;t have to. Serious engineering is the opposite. It requires blocking interruptions for long enough to hold a complex system in your head. Two incompatible attention modes. Try to run both at once and you do neither well.</p><p>But here is the part people skip when they quote her. Majors never said managers should stop coding. Her actual advice was sharper: <em>don&#8217;t write code in the critical path.</em> Don&#8217;t be the person others are waiting on. Stay technical, stay sharp, just don&#8217;t make yourself a dependency that blocks shipping. &#8220;The best frontline eng managers in the world,&#8221; she wrote, &#8220;are the ones that are never more than 2-3 years removed from hands-on work.&#8221;</p><p>That distinction carries the whole argument. Because there is a category of technical work that is consequential without being critical-path, and it turns out to be exactly the work senior people are best positioned to do.</p><p>Call it enabling work. Internal tooling. Developer productivity infrastructure. Agentic harnesses. MCP services that other teams plug into. Architectural prototypes that prove a direction before anyone commits to it. None of this is what blocks a release on Thursday. All of it multiplies whoever comes after. It tolerates interruption &#8212; you can pick it up Tuesday afternoon and put it down when a real fire starts &#8212; precisely because nobody is standing at your desk waiting for it.</p><p>This is also where the industry is putting its money. Gartner named platform engineering a top strategic trend for two years running and projects that 80% of large engineering organizations will run dedicated platform teams by 2026. The structural reason a leader can work here without becoming the bottleneck is built into the definition of the domain: its output unblocks others rather than blocking them.</p><p>Will Larson&#8217;s <a href="https://lethain.com/staff-engineer-archetypes/">Staff Engineer archetypes</a> circle the same territory without quite landing on it. His &#8220;Architect&#8221; sits in a permanent argument &#8212; some organizations demand the Architect stay deep in the code, others forbid it. The leader/practitioner resolves that argument by relocating it: deep in the enabling and infrastructure work, absent from production product code. Not the pendulum, not the staff IC. A real hybrid, and a newly coherent one.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!jHB6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd3763-7ad5-4119-96d4-4e0c1d01f0cd_1024x684.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!jHB6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd3763-7ad5-4119-96d4-4e0c1d01f0cd_1024x684.png 424w, https://substackcdn.com/image/fetch/$s_!jHB6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd3763-7ad5-4119-96d4-4e0c1d01f0cd_1024x684.png 848w, https://substackcdn.com/image/fetch/$s_!jHB6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd3763-7ad5-4119-96d4-4e0c1d01f0cd_1024x684.png 1272w, https://substackcdn.com/image/fetch/$s_!jHB6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd3763-7ad5-4119-96d4-4e0c1d01f0cd_1024x684.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!jHB6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd3763-7ad5-4119-96d4-4e0c1d01f0cd_1024x684.png" width="1024" height="684" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dffd3763-7ad5-4119-96d4-4e0c1d01f0cd_1024x684.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:684,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1540586,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/200067020?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8e7bf7a-504a-4f0c-b75f-1832889fc598_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!jHB6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd3763-7ad5-4119-96d4-4e0c1d01f0cd_1024x684.png 424w, https://substackcdn.com/image/fetch/$s_!jHB6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd3763-7ad5-4119-96d4-4e0c1d01f0cd_1024x684.png 848w, https://substackcdn.com/image/fetch/$s_!jHB6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd3763-7ad5-4119-96d4-4e0c1d01f0cd_1024x684.png 1272w, https://substackcdn.com/image/fetch/$s_!jHB6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd3763-7ad5-4119-96d4-4e0c1d01f0cd_1024x684.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>This is not new</h2><p>Here is where I want to slow down, because the most interesting thing about this role is how old it is.</p><p>The idea that a leader should be both a strategist and a practitioner &#8212; not separate modes to alternate between but a single integrated way of operating &#8212; shows up independently across at least five civilizations. That kind of convergence usually means a culture has found a durable answer to a real problem.</p><p>Japan gave it a name: <em>bunbu-ry&#333;d&#333;</em> (&#25991;&#27494;&#20001;&#36947;), the way of both the literary and the martial. <em>Bun</em> is letters, cultivation, strategy. <em>Bu</em> is the martial, the active, the practitioner&#8217;s hand. <em>Ry&#333;d&#333;</em> means both ways, together &#8212; not balanced, not traded off, but held at once. By the mid-fourteenth century the dual-talented warrior was already established as the model leader, and during the Edo period the Tokugawa shogunate made it official policy for the samurai class. The phrase that survives captures the stakes: culture without power is ineffective, and power without culture is barbarous.</p><p>The archetype is Miyamoto Musashi. Undefeated in more than sixty duels, often fighting with a wooden sword against live steel. He founded a two-sword school, and in the last months of his life he wrote <em>The Book of Five Rings</em> in a cave. He was also a recognized master of ink painting and calligraphy &#8212; his <em>Shrike on a Withered Branch</em> survives as a designated Important Cultural Property of Japan. The same hands that won sixty duels produced fine art a nation still protects. He didn&#8217;t oscillate between the sword and the brush. He held both, and each sharpened the other. &#8220;When I apply the principle of strategy to the ways of different arts and crafts,&#8221; he wrote, &#8220;I no longer have need for a teacher in any domain.&#8221; Mastery in one discipline illuminating all the others.</p><p>Rome had Marcus Aurelius, the philosopher-king made historical rather than theoretical. He ran the empire and commanded its armies on the Danube, and he wrote the <em>Meditations</em> in the war camps &#8212; fragmentary notes to himself, composed in the middle of campaigning and administration. That book was never published philosophy. It was a working journal, the most powerful man in the ancient world writing to stay grounded while doing the job.</p><p>China institutionalized the same ideal as <em>wen</em> and <em>wu</em> &#8212; civil cultivation and martial capability &#8212; and ran it for roughly thirteen centuries through the scholar-official class. Zeng Guofan is the canonical case: he rose through the imperial examinations to high Confucian office, then built and commanded an army of more than a hundred thousand, reportedly keeping a diary on Neo-Confucian ethics even as he directed the campaigns. The integration wasn&#8217;t left to personal taste. It was built into the examination and career structure.</p><p>And the archetype still lands today. James Mattis earned the nickname &#8220;Warrior Monk&#8221; &#8212; battlefield commander and devoted reader, a 7,000-book library, the <em>Meditations</em> carried into combat. The chain from Aurelius to Mattis is literal: the same book, eighteen centuries apart. That we still reach for &#8220;warrior monk&#8221; as a compliment for a leader tells you the integration never stopped resonating.</p><p>Across all of it, the answer is the same. The contemplative and the active were not specializations to assign to different people. They were a single discipline, each half informing the other. Modernity &#8212; with its org charts, its clean role boundaries, its professional specialization &#8212; interrupted that. The leader/practitioner is not a tech-industry novelty. It&#8217;s an old integration becoming feasible again.</p><h2>Why now</h2><p>So what actually changed? Not the wisdom. The economics.</p><p>The thing that made the coding executive a bad idea was the attention math. Serious technical work demanded long, unbroken stretches of focus &#8212; the exact resource a leadership schedule cannot reliably provide. You cannot design a system in the fifteen minutes between a board prep and a one-on-one. The pendulum was a real constraint, not a failure of will.</p><p>Agentic coding changes that math directly. The unit of work moved up a level. Instead of holding every implementation detail in working memory across a four-hour session, you specify intent, direct an agent, review what comes back, correct course, and direct again. A focused 30-minute session now produces what used to require an afternoon of immersion &#8212; not because the thinking got easier, but because the implementation cost collapsed.</p><p>The Anthropic <a href="https://resources.anthropic.com/2026-agentic-coding-trends-report">2026 Agentic Coding Trends Report</a> puts numbers on the shift. Average session length has climbed to 23 minutes in the agentic era, up from about 4 in the autocomplete era &#8212; the work got denser, not just faster. 78% of Claude Code sessions now involve multi-file edits, up from 34% a year earlier. Teams running multi-agent workflows report 2&#8211;4x faster delivery from task creation to deployment. And the figure that matters most for this argument: roughly 27% of AI-assisted work consists of tasks that &#8220;wouldn&#8217;t have been done otherwise&#8221; &#8212; the scaling projects, the nice-to-have tools, the exploratory infrastructure that was never quite worth the manual hours.</p><p>That 27% is the enabling work. It is the category that lives or dies on time cost, and it&#8217;s the category a leader/practitioner is best placed to take on.</p><p>The arithmetic is what makes 30% credible. I documented a 6&#8211;10x multiplier on focused technical sessions in <a href="https://hyperdev.matsuoka.com/the-irreducibles-what-a-pattern-master-does">The Irreducibles</a> earlier this year &#8212; a project I estimated at 150&#8211;200 billable hours compressed into roughly 50&#8211;70 hours of wall-clock time, most of which wasn&#8217;t coding at all. If directed work runs several times faster than hand-coding, then a day and a half a week can produce what once consumed a full-time engineer&#8217;s week. That&#8217;s not a marginal gain. It&#8217;s a change in what&#8217;s structurally possible.</p><p>There&#8217;s another dimension that doesn&#8217;t show up in the productivity numbers: the work itself is unstable. Agentic coding patterns are still shaking out. There aren&#8217;t many experienced practitioners, the field is moving fast, and we don&#8217;t yet have good consensus on which patterns are load-bearing and which are fashion. A manager who&#8217;s only reading about it can&#8217;t make that distinction on behalf of a team. You have to be in it to know.</p><p>There&#8217;s a counterintuitive wrinkle worth naming: the people best positioned to exploit this are the senior ones. A University of Chicago working paper from late 2025 found experienced developers were 5&#8211;6% more likely to succeed with AI agents for every standard deviation of work experience, largely because they worked plan-first &#8212; laying out objectives, alternatives, and steps before invoking the tool. That&#8217;s the opposite of the assumption that AI flattens the seniority curve. Expertise improves your ability to delegate to a model for the same reason it improves your ability to delegate to a person. AI doesn&#8217;t change what senior engineering is. It reveals what it always was.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!EYNR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F727ed32e-2441-4dc7-bf6b-4ea9a19a2cb7_1024x634.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!EYNR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F727ed32e-2441-4dc7-bf6b-4ea9a19a2cb7_1024x634.png 424w, https://substackcdn.com/image/fetch/$s_!EYNR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F727ed32e-2441-4dc7-bf6b-4ea9a19a2cb7_1024x634.png 848w, https://substackcdn.com/image/fetch/$s_!EYNR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F727ed32e-2441-4dc7-bf6b-4ea9a19a2cb7_1024x634.png 1272w, https://substackcdn.com/image/fetch/$s_!EYNR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F727ed32e-2441-4dc7-bf6b-4ea9a19a2cb7_1024x634.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!EYNR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F727ed32e-2441-4dc7-bf6b-4ea9a19a2cb7_1024x634.png" width="1024" height="634" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/727ed32e-2441-4dc7-bf6b-4ea9a19a2cb7_1024x634.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:634,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1514078,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/200067020?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1745d9f4-d955-45cb-b61a-c12956852f96_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!EYNR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F727ed32e-2441-4dc7-bf6b-4ea9a19a2cb7_1024x634.png 424w, https://substackcdn.com/image/fetch/$s_!EYNR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F727ed32e-2441-4dc7-bf6b-4ea9a19a2cb7_1024x634.png 848w, https://substackcdn.com/image/fetch/$s_!EYNR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F727ed32e-2441-4dc7-bf6b-4ea9a19a2cb7_1024x634.png 1272w, https://substackcdn.com/image/fetch/$s_!EYNR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F727ed32e-2441-4dc7-bf6b-4ea9a19a2cb7_1024x634.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Directing agents is delegation</h2><p>This is the part that interested me the most, and it&#8217;s the bridge between the leadership job and the technical one.</p><p>A year or so ago, working with Claude Code felt like coding. Now it feels like delegating. I <a href="https://hyperdev.matsuoka.com/coding-to-delegation-shift">wrote about that shift</a> when it first became undeniable &#8212; the move from being a programmer who uses AI to being something closer to a technical project manager who directs it. Anthropic&#8217;s report uses the same vocabulary, describing engineers moving &#8220;from writing code to orchestrating the systems that write it.&#8221;</p><p>What I didn&#8217;t fully appreciate at the time is how directly that maps onto the muscle leaders already have. Delegating to an agent feels, in practice, just like delegating to a human engineer. You frame the problem, set the constraints, hand it off, and come back to assess the result. Give a directive, walk away, return to completed work. That loop is the daily reality of management. Leaders developed it because they had to, and it transfers to agents almost without friction.</p><p>So the leader/practitioner doesn&#8217;t have to become a coder again in the old sense. The skill in demand is judgment plus delegation, and that&#8217;s the skill leadership has been building all along. The hands-on knowledge tells you what to ask for and whether the answer is any good. The delegation instinct does the rest.</p><p>This is why the enabling work and the agentic tools fit together so cleanly. Enabling work tends to be well-defined, non-user-facing, and long-horizon &#8212; exactly the profile agents handle well and exactly the profile that tolerates a leader&#8217;s interrupted schedule. The hands-on contribution mostly takes the form of specifying constraints and patterns, which is what I&#8217;ve called <a href="https://hyperdev.matsuoka.com/what-does-a-pattern-master-do">pattern mastery</a>: when you write the pattern down, you&#8217;ve written the spec, and the spec multiplies everyone else&#8217;s output.</p><h2>What it actually looks like</h2><p>Let me ground this without turning it into a war story.</p><p>The concrete examples from my own work are the kind of thing I mean. I built an agentic harness &#8212; the orchestration layer I <a href="https://hyperdev.matsuoka.com/its-the-harness-stupid">argued is the real determinant of AI coding outcomes</a>, where the same model can swing more than a quality point depending on the scaffolding around it. I built MCP services, the <a href="https://hyperdev.matsuoka.com/is-this-the-era-of-the-connector">org-specific connectors</a> that replaced a handful of standalone tools in a few hours of directed work each. None of that was a production feature. All of it was infrastructure other people now depend on.</p><p>Here&#8217;s a detail that may make the point. I now have an &#8220;AI architect&#8221; on my leadership team helping maintain the very infrastructure I originally built &#8212; not just the harness, but our inference relationships, our training program, office hours, the real human work I no longer have the time, or the right, to be doing myself. And I expect to hand off more over time. The enabling work I do today partly becomes the system that does tomorrow&#8217;s enabling work. That handoff is the role in miniature: you build the thing that multiplies the team, then you put someone in place to build the next version.</p><p>The proportion matters. Around 30% hands-on keeps judgment fresh without putting me in the critical path. Even full-time senior ICs aren&#8217;t full-time coders &#8212; Bain&#8217;s Jue Wang, quoted in MIT Technology Review last December, put developer coding time at 20&#8211;40%, with the rest going to analysis, strategy, and the surrounding work. A leader at 30% isn&#8217;t doing something exotic. They&#8217;re spending their technical budget on the layer where it compounds.</p><p>The decision is not &#8220;how do I find time to code.&#8221; It&#8217;s &#8220;what enabling work is worth my own hands?&#8221; Those are different questions. The first leads to the bottleneck I watched so many executives become. The second leads somewhere useful.</p><h2>The choice</h2><p>I&#8217;ll resist overselling this, because it isn&#8217;t for everyone and it isn&#8217;t automatic.</p><p>This is a deliberate role, not a default. Staying technically current costs ongoing investment, and the work is often invisible &#8212; enabling infrastructure rarely shows up in a quarterly review the way a shipped feature does. The role is easy to misread, too. From the outside, a CTO who codes can look like a CTO who hasn&#8217;t let go. The defense against that reading is the discipline Majors named: stay out of the critical path. Build the multipliers, not the blockers.</p><p>The returns are real, though. Fresh judgment &#8212; the kind that lets you evaluate not just what and why but how. Trust from engineers who see you in the work rather than above it. And institutional infrastructure that makes the whole team faster, built by the person with both the technical depth and the positional authority to prioritize it.</p><p>There&#8217;s a closing note in the history worth keeping. <em>Bunbu-ry&#333;d&#333;</em> wasn&#8217;t only a personal aspiration. The Tokugawa shogunate institutionalized it &#8212; built career and class structures around the assumption that a leader should be both. China did the same with its examination system. We&#8217;re not there yet. For now, the leader/practitioner is an individual choice, made one person at a time, made viable by tools that finally collapsed the cost of staying hands-on.</p><p>But the precedent suggests where this could go. When a way of working proves durable, organizations eventually build structures around it. The era of the leader/practitioner is early. It is also, I&#8217;d argue, a return &#8212; to an integration we knew was valuable long before we had the means to make it practical again.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and writes about AI-powered engineering at <a href="https://hyperdev.substack.com/">HyperDev</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/coding-to-delegation-shift">From Coding with AI to Managing AI</a> &#8212; When agentic coding starts to feel like delegation</p></li><li><p><a href="https://hyperdev.matsuoka.com/its-the-harness-stupid">It&#8217;s The Harness, Stupid!</a> &#8212; Why orchestration quality dominates AI coding outcomes</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul>]]></content:encoded></item><item><title><![CDATA[The Other Shoe Has Dropped]]></title><description><![CDATA[The Economics of Enterprise Inference Usage]]></description><link>https://hyperdev.matsuoka.com/p/the-other-shoe-has-dropped</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/the-other-shoe-has-dropped</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Fri, 29 May 2026 11:31:28 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!-Qpp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-Qpp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-Qpp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!-Qpp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!-Qpp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!-Qpp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-Qpp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1174362,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/199673840?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!-Qpp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!-Qpp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!-Qpp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!-Qpp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Two stories from the last two weeks. Uber <a href="https://fortune.com/2026/05/26/uber-coo-ai-spending-tokens-claude-code/">burned through its entire 2026 AI budget in four months</a> on Claude Code, with COO Andrew Macdonald telling the <em>Rapid Response</em> podcast that the link between that spend and shipped consumer features &#8220;is not there yet.&#8221; <a href="https://www.theinformation.com/newsletters/applied-ai/uber-cto-shows-claude-code-can-blow-ai-budgets">The Information had the underlying numbers a few weeks earlier</a>: engineer adoption from 32% to 84% between December and March, heavy users running $500&#8211;$2,000/month in tokens, and CTO Praveen Neppalli Naga torching $1,200 in a two-hour demo. Same week, Microsoft told thousands of engineers in its Experiences + Devices division that their Claude Code access is going away. <a href="https://www.windowscentral.com/microsoft/microsoft-cancels-claude-code-licenses-shifting-developers-to-github-copilot-cli-a-move-likely-driven-by-financial-motives">Windows Central, summarizing The Verge&#8217;s Notepad scoop</a>, has the cutoff at June 30 &#8212; end of fiscal year &#8212; with cost as the actual driver even though EVP Rajesh Jha framed it publicly as convergence on Copilot CLI.</p><p>Two of the most AI-forward enterprises on the planet, same tool, same week. The &#8220;AI is failing&#8221; takes were live within hours.</p><p>I don&#8217;t buy that framing.</p><p>The headlines are getting it wrong. Uber didn&#8217;t cancel anything &#8212; adoption ran ahead of the budget and the company blew its annual spend keeping up. That&#8217;s a planning failure, not a verdict on the tool. Microsoft didn&#8217;t divorce Anthropic either; they&#8217;re still consuming Claude through Azure Foundry and M365 Copilot. What they cancelled is a specific license &#8212; Claude Code at the engineer-seat level &#8212; because engineers preferred it over GitHub Copilot CLI and the division was paying for that preference.</p><p>What both stories show: AI is a new tool and we haven&#8217;t learned to use it well yet. The teams over budget pointed it at problems it wasn&#8217;t the cheapest way to solve, then let it decide for itself how much work to do per task.</p><p>I&#8217;ve made <a href="https://hyperdev.matsuoka.com/p/what-the-other-shoe-sounds-like-when">the cloud parallel here before</a>. Early cloud was expensive and misused. Lift-and-shift workloads routinely ran two or three times their on-prem cost &#8212; I watched that play out across teams I ran, and it took years to correct through architecture. Then the industry learned: right-sizing, reserved instances, autoscaling, serverless where it fit, on-prem where it didn&#8217;t. The bills came down. Not because compute got dramatically cheaper, but because we got more careful about what we asked the cloud to do. AI is in the same phase. Cheap per-token, expensive per-task, and the gap is architectural.</p><p>A few weeks ago I ran controlled head-to-head tests on Opus 4.6 and Opus 4.7 against identical coding tasks. Both models passed every test. Opus 4.7 cost 3.6&#215; more to do it. Same outcomes, same rate card, dramatically more tokens.</p><p>Finout&#8217;s analysis of production deployments <a href="https://www.finout.io/blog/claude-opus-4.7-pricing-the-real-cost-story-behind-the-unchanged-price-tag">tells the same story at scale</a>: up to a 35% cost increase overnight, driven by tokenizer changes that don&#8217;t show up on the per-token rate card. Not one team&#8217;s bad luck &#8212; the shape of the bill across the enterprise AI buyer base right now. The second of two shoes on AI economics.</p><p>I wrote about that <a href="https://hyperdev.matsuoka.com/p/opus-46-vs-47-the-real-cost-of-incremental">version-to-version cost drift in detail</a>. Providers can collapse per-token prices in public while the per-task bill drifts upward in private. The first shoe was the per-token price collapse that made everyone optimistic. The second is the behavioral and architectural cost overhang now landing on quarterly P&amp;Ls.</p><p><strong>TL;DR</strong></p><ul><li><p>Per-token costs at GPT-3.5-equivalent performance are down roughly 280&#215; since late 2022, per <a href="https://aiindex.stanford.edu/report/">Stanford&#8217;s AI Index 2025</a>. Vendor revenue tells the opposite story: Anthropic&#8217;s annualized revenue went from <a href="https://www.pymnts.com/artificial-intelligence-2/2026/anthropic-hits-30-billion-run-rate-as-enterprise-demand-accelerates/">$1B in January 2025 to $30B by April 2026</a> &#8212; a 30&#215; move in 15 months, coming from enterprise inference, not consumer subscriptions.</p></li><li><p>Gartner&#8217;s April 2026 survey: just 28% of AI use cases fully meet ROI expectations, 78% of IT leaders report material AI charges that didn&#8217;t show up in any procurement model.</p></li><li><p>Gartner estimates agentic workflows consume 5&#8211;30&#215; more tokens than equivalent chatbot interactions; Stanford&#8217;s Digital Economy Lab puts the upper bound for coding agents at 1,000&#215;. The cost driver isn&#8217;t the model &#8212; it&#8217;s the workflow architecture wrapped around it.</p></li><li><p>Two patterns hold the line in production. Search-first architectures put inference at the end of a deterministic pipeline. Consolidated single-shot designs replace multi-call chains.</p></li><li><p>Inference is a power tool, not a default. Use it with specific ROI goals per call, apply it to <em>code</em> solutions rather than to directly solve problems, and bound it with deterministic structure on both ends.</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!TPPJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F491e72b0-16e1-49fe-a551-391ba05461bc_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!TPPJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F491e72b0-16e1-49fe-a551-391ba05461bc_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!TPPJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F491e72b0-16e1-49fe-a551-391ba05461bc_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!TPPJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F491e72b0-16e1-49fe-a551-391ba05461bc_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!TPPJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F491e72b0-16e1-49fe-a551-391ba05461bc_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!TPPJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F491e72b0-16e1-49fe-a551-391ba05461bc_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/491e72b0-16e1-49fe-a551-391ba05461bc_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1163417,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/199673840?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F491e72b0-16e1-49fe-a551-391ba05461bc_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!TPPJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F491e72b0-16e1-49fe-a551-391ba05461bc_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!TPPJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F491e72b0-16e1-49fe-a551-391ba05461bc_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!TPPJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F491e72b0-16e1-49fe-a551-391ba05461bc_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!TPPJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F491e72b0-16e1-49fe-a551-391ba05461bc_1024x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Why the Per-Token Savings Didn&#8217;t Reach the Invoice</h2><p>Per-token economics of frontier models have been collapsing for two years. <a href="https://aiindex.stanford.edu/report/">Stanford&#8217;s AI Index 2025</a> puts the decline at roughly 280&#215; from a late-2022 baseline at GPT-3.5-equivalent performance. Most enterprise budget conversations in 2024 started from that headline. The implicit assumption: bills should be going <em>down</em>.</p><p>They&#8217;re not. The clearest read comes from the vendor side. Anthropic&#8217;s annualized revenue <a href="https://www.anthropic.com/news/anthropic-raises-30-billion-series-g-funding-380-billion-post-money-valuation">went from $1B in January 2025 to $30B by April 2026</a>, with roughly 80% from enterprise and API usage rather than consumer subscriptions. Anthropic <a href="https://www.saastr.com/anthropic-just-passed-openai-in-revenue-while-spending-4x-less-to-train-their-models/">now discloses 1,000+ customers spending more than $1M per year</a> &#8212; a cohort that doubled in under two months &#8212; alongside roughly 300,000 business customers. The mid-tier ($100K&#8211;$1M/year) grew 7&#215; year over year.</p><p>Menlo Ventures&#8217; <a href="https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/">2025 State of Generative AI report</a> cross-checks at the market level: enterprise GenAI spend tripled to $37B in 2025, with LLM API consumption alone at $8.4B by mid-year. The high tier shows up in named deals: <a href="https://sacra.com/c/anthropic/">Snowflake&#8217;s $200M multi-year partnership</a> implies a $50&#8211;70M annual run rate from one customer, and <a href="https://sacra.com/c/anthropic/">Deloitte is deploying Claude across 470,000 employees</a>.</p><p>For a typical enterprise running multiple production AI workloads, $500K&#8211;$2M per year is now the realistic floor. Fortune 100 is running $10M&#8211;$50M+, the most AI-intensive past $100M. The Gartner numbers point the same direction: <a href="https://www.gartner.com/en/newsroom/press-releases/2026-04-07-gartner-says-artificial-intelligence-projects-in-infrastructure-and-operations-stall-ahead-of-meaningful-roi-returns">just 28% of AI use cases fully meet ROI expectations and 20% fail outright</a>, and <a href="https://zylo.com/blog/saas-management-index/">78% of IT leaders report material AI charges</a> that didn&#8217;t show up in any procurement model.</p><p>The per-token chart is real. The invoice is also real. What closes the gap is <em>behavior</em>. Three behaviors specifically.</p><p><strong>Models do more work per task.</strong> Reasoning models reason. Agentic loops loop. The prompt that used to consume 4K tokens now consumes 40K because the assistant explores, plans, second-guesses, and verifies. Some of that is valuable. Much of it is the model performing thoroughness in a way that costs you money. The Opus 4.6-to-4.7 jump I documented earlier: same task, same outcome, 2.9&#215; more output tokens and 4.8&#215; more cache reads.</p><p><strong>Workflows fan out.</strong> A &#8220;single&#8221; task in a modern agentic system might trigger a planner, researcher, coder, reviewer, and summarizer. Each makes its own LLM calls over overlapping context. Gartner&#8217;s March 2026 analysis puts agentic workflows at 5&#8211;30&#215; the token consumption of an equivalent chatbot ask. Stanford Digital Economy Lab&#8217;s April 2026 arXiv paper goes further: coding agents can consume 1,000&#215; more tokens than equivalent chat completions. The agent isn&#8217;t more expensive because it&#8217;s smarter. It&#8217;s more expensive because it&#8217;s louder.</p><p><strong>Context windows fill themselves.</strong> Long context is a feature in marketing and a bill in practice. In our own enterprise Claude.AI usage &#8212; 82,852 messages from 329 employees over 3.5 months, audited via the Anthropic Compliance API &#8212; the average request carried 366,000 input tokens, mostly from 10-turn conversations dragging accumulated history forward into every new turn. Most production systems I&#8217;ve audited show the same fingerprint: pipelines paying for context they aren&#8217;t actually using.</p><p>None of this is fraud and none of it is mysterious. It&#8217;s the natural consequence of letting probabilistic systems decide how much work to do on every call. The savings from cheaper tokens were real. They just got consumed by an order of magnitude more tokens per task.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!T_ba!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7441c01-b144-4b0c-83a2-113c2bb730fd_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!T_ba!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7441c01-b144-4b0c-83a2-113c2bb730fd_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!T_ba!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7441c01-b144-4b0c-83a2-113c2bb730fd_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!T_ba!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7441c01-b144-4b0c-83a2-113c2bb730fd_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!T_ba!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7441c01-b144-4b0c-83a2-113c2bb730fd_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!T_ba!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7441c01-b144-4b0c-83a2-113c2bb730fd_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c7441c01-b144-4b0c-83a2-113c2bb730fd_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1253809,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/199673840?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7441c01-b144-4b0c-83a2-113c2bb730fd_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!T_ba!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7441c01-b144-4b0c-83a2-113c2bb730fd_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!T_ba!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7441c01-b144-4b0c-83a2-113c2bb730fd_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!T_ba!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7441c01-b144-4b0c-83a2-113c2bb730fd_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!T_ba!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7441c01-b144-4b0c-83a2-113c2bb730fd_1024x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>What I&#8217;ve Found Shipping These Systems</h2><p>The teams handling this well aren&#8217;t the ones cutting AI usage. They&#8217;re changing the <em>shape</em> of how they use it.</p><p>The pattern I keep coming back to: treat inference as the expensive step at the end of a mostly deterministic pipeline. Do the cheap, structured work in code. Reserve the model call for the part that actually requires judgment. Then bound the call hard &#8212; context budget, output budget, quality gate on whether the call even runs.</p><p>Two examples from systems I&#8217;ve been building illustrate this from different angles.</p><h3>Example 1: Code-Intelligence &#8212; Search-First Architecture</h3><p>The naive version of a code review tool is obvious: dump the changed files into Claude, ask for a review. That works. It also costs roughly $0.03 per operation, scales linearly with repo size, and produces a lot of review output you didn&#8217;t need. Claude.AI offers a code review service &#8212; it ended up costing us thousands a month for just a few repos. Augment Code offers a well-regarded one as a GitHub app, but charges a platform fee (a meaningful fraction of our Anthropic spend) <em>just</em> to connect.</p><p>So we built our own. It leverages a multimodal search/RAG/KG engine I&#8217;d already built, so this wasn&#8217;t from scratch.</p><p>The version I actually ship uses a multi-tier search pipeline with the LLM call at the very end:</p><pre><code><code>Stage 1: Vector Search    (~$0.0002 per query, semantic similarity)
Stage 2: BM25 Reranking   (~$0.0001 per query, lexical relevance)
Stage 3: Static Analysis  (~$0.0001 per query, AST + symbol resolution)
Stage 4: Quality Gate     (free, deterministic threshold check)
Stage 5: Single LLM Call  (~$0.03 per call, only if Stages 1-4 passed)</code></code></pre><p>The first four stages cost about $0.0004 combined. They do the bulk of the work: deciding <em>what code is actually relevant</em>, ranking it, pulling structural relationships, and deciding whether the result is even worth asking an LLM about.</p><p>Hard budget controls run through the whole pipeline:</p><pre><code><code># Budget enforcement, not aspiration
MAX_CONTEXT_FILES = 6          # cap on what we send to the model
MAX_REVIEW_WORDS = 500         # cap on what the model returns
RELEVANCE_FLOOR = 0.005        # quality gate before calling the model

if combined_relevance_score &lt; RELEVANCE_FLOOR:
    # no point spending $0.03 to get a review of weakly-related code
    return SkipReason("below relevance floor")

context = select_top_n(ranked_results, MAX_CONTEXT_FILES)
review = llm.review(context, max_output_tokens=MAX_REVIEW_WORDS * 1.4)</code></code></pre><p>The <code>RELEVANCE_FLOOR</code> check is the part to underline. A meaningful percentage of review requests in real codebases don&#8217;t justify an LLM call at all &#8212; the changes are mechanical, the related code trivial, or the search signal weak enough that whatever the model says will be hallucinated context. Refusing to spend $0.03 on those cases is where most of the savings come from.</p><p>Rough economics across a quarter of usage:</p><p>Approach Cost per operation LLM calls per 1K operations Direct &#8220;review the diff&#8221; ~$0.030 1,000 Search-first with gates ~$0.0034 average ~430</p><p>About 80% of the workflow logic ends up deterministic: search, ranking, static analysis, gating. The model handles the last 20% &#8212; judgment on curated context. Same outcome from the user&#8217;s perspective, roughly an order of magnitude cheaper, with more predictable failure modes because most of the pipeline is debuggable code rather than prompt behavior.</p><p>The limitation: this is more work than wiring up a single LLM call, and the gates are only as good as your search infrastructure. The payoff is on the cost and determinism side, not on speed of initial implementation.</p><h3>Example 2: duetto-intelligence &#8212; Context Injection Instead of Replacement</h3><p>The second pattern comes from duetto-intelligence, internal tooling I&#8217;ve been building against that same enterprise Claude.AI usage &#8212; 82,852 real employee messages over 3.5 months, not a thought experiment. The problem here isn&#8217;t &#8220;should we call the LLM at all.&#8221; It&#8217;s: given that our people are already routing structured-data questions through a $0.274-per-request multi-turn Sonnet conversation, what&#8217;s a cheaper path that doesn&#8217;t degrade the answer?</p><p>The audit data made the gap concrete. Average request: 366K input tokens, ten-turn conversation, $0.274 to Anthropic. The same query answered through a Haiku single-pass against deterministically-retrieved internal data: $0.0009. A 300:1 cost ratio on the slice of traffic about structured product knowledge, CRM/account prep, JIRA, people and org lookups.</p><p>Not all traffic. Somewhere in the 35&#8211;40% range based on classified samples. About half of remaining queries genuinely need full Sonnet or Opus reasoning &#8212; writing, debugging, free-form analysis &#8212; and shouldn&#8217;t be intercepted at all.</p><p>The framing matters, because it&#8217;s easy to mis-read this as &#8220;replace Claude with a smaller model.&#8221; It isn&#8217;t. duetto-intelligence acts as a <strong>context injection layer</strong> in front of the user-facing model. When a query has structured-data intent, we route a sub-query to DI, get back a bounded structured result, and inject that into the prompt the larger model sees. The expensive model still does the reasoning &#8212; it just stops being responsible for the deterministic data retrieval it&#8217;s bad at and expensive for.</p><p>The naive design for the routing layer looks like this:</p><pre><code><code>1. Classify the user's intent             &#8594; LLM call (~150 tokens)
2. Plan which subsystems to query         &#8594; LLM call (~250 tokens)
3. Call subsystem A, summarize response   &#8594; LLM call (~300 tokens)
4. Call subsystem B, summarize response   &#8594; LLM call (~300 tokens)
5. Synthesize a final answer              &#8594; LLM call (~250 tokens)
                                          Total: ~1,250 tokens, 5 calls</code></code></pre><p>Each call is plausible on its own. Together they&#8217;re a tax on every user interaction. Latency stacks linearly with calls, and any one of the five can hallucinate in a way that corrupts the rest of the chain.</p><p>The consolidated design replaces three steps with deterministic code:</p><pre><code><code>1. Classify intent                  &#8594; LLM call  (~80 tokens, tight tagger prompt)
2. Fan-out to subsystems            &#8594; code      (0 tokens, intent &#8594; call map)
3. Consolidated synthesis           &#8594; LLM call  (~200 tokens, structured input)
                                       Total: ~280 tokens, 2 calls</code></code></pre><p>The trick is the intent classifier. It produces a tag from a fixed vocabulary of 87 tags &#8212; <code>revenue.query.ytd</code>, <code>forecast.compare.year_over_year</code>, <code>account.lookup.contact</code>, and so on. Each tag maps deterministically to a set of downstream calls in plain Python. No LLM in the routing step. The model isn&#8217;t asked &#8220;what should we do?&#8221; It&#8217;s asked &#8220;what is the user asking about?&#8221; &#8212; a much smaller, bounded question.</p><p>We validated against a 50-query test corpus drawn directly from the compliance data &#8212; real questions people had asked the model in production. After tuning, 100% of those queries land on the fast path with no LLM call required for routing. That&#8217;s the proof the deterministic-discipline part holds at the boundary; the routing isn&#8217;t quietly falling back to a second model call to bail itself out.</p><p>Budget enforcement is explicit in every prompt template:</p><pre><code><code>SOURCE_CHAR_BUDGET = 600    # per data source pulled into context
OUTPUT_TOKEN_BUDGET = 200   # cap on synthesis response
INTENT_TAG_VOCAB = load_intent_taxonomy()  # 87 tags, versioned

def synthesize(intent_tag: str, sources: list[Source]) -&gt; str:
    trimmed = [s.truncate(SOURCE_CHAR_BUDGET) for s in sources]
    return llm.complete(
        prompt=template(intent_tag, trimmed),
        max_tokens=OUTPUT_TOKEN_BUDGET,
    )</code></code></pre><p>The economics, against measured baselines rather than estimates: current Claude.AI Chat spend across the 329-user population runs $15,264 over 3.5 months. Roughly $4,360/month, driven by that $0.274-per-request multi-turn average. If DI intercepts the 35&#8211;40% of traffic that&#8217;s structured-retrieval underneath, projected savings come in around $1,500&#8211;1,700/month, or $18&#8211;20K/year on this single user population. The leverage isn&#8217;t from picking a cheaper model. It&#8217;s from refusing to pay Sonnet rates to answer questions a deterministic system already has the data for.</p><p>The intent vocabulary is the contract. New capability means a new tag, a new downstream mapping, a new prompt template. The model never has to invent structure on the fly. This is what people mean by &#8220;use the LLM to code solutions, not to solve problems directly&#8221; &#8212; the routing logic lives in code, the tagger is a thin call, the synthesis is bounded.</p><p>The source-character budget matters more than the output budget. The compliance audit confirmed it: production overspend is on the <em>input</em> side. 366K input tokens against a few hundred output. Models will happily consume whatever context you hand them. Trimming at the source &#8212; 600 characters per source, no exceptions &#8212; is how you keep per-call cost from drifting upward as the system gets more capable.</p><p>The limitation: this only works on the routable slice. The pattern isn&#8217;t &#8220;eliminate inference.&#8221; It&#8217;s &#8220;stop spending $0.274 to answer questions that have a structured answer at $0.0009.&#8221;</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-uHZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe0775c6-8743-46e4-908a-ee4e2a8db5d6_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-uHZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe0775c6-8743-46e4-908a-ee4e2a8db5d6_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!-uHZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe0775c6-8743-46e4-908a-ee4e2a8db5d6_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!-uHZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe0775c6-8743-46e4-908a-ee4e2a8db5d6_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!-uHZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe0775c6-8743-46e4-908a-ee4e2a8db5d6_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-uHZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe0775c6-8743-46e4-908a-ee4e2a8db5d6_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fe0775c6-8743-46e4-908a-ee4e2a8db5d6_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1351360,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/199673840?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe0775c6-8743-46e4-908a-ee4e2a8db5d6_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!-uHZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe0775c6-8743-46e4-908a-ee4e2a8db5d6_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!-uHZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe0775c6-8743-46e4-908a-ee4e2a8db5d6_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!-uHZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe0775c6-8743-46e4-908a-ee4e2a8db5d6_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!-uHZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe0775c6-8743-46e4-908a-ee4e2a8db5d6_1024x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Why &#8220;Just Use a Cheaper Model&#8221; Doesn&#8217;t Save You</h2><p>A reasonable objection: aren&#8217;t the cheaper, smaller models supposed to handle this? Why not route everything to Haiku or an open model and call it solved?</p><p>The pricing math seems to support it. The behavioral math doesn&#8217;t.</p><p>Two things go wrong when you swap a cheaper model into an unstructured workflow. Cheaper models are usually less efficient <em>per task</em> &#8212; more turns to converge, more exploration, more hallucination, which means more retries and verification calls. A model 5&#215; cheaper per token can run 1.5&#8211;2&#215; more expensive per completed task if your workflow lets it spin.</p><p>And the workflow itself is where most of the cost lives. The 5&#8211;30&#215; multiplier is structural, not modal &#8212; it exists regardless of which model you point at it. Switching from Sonnet to Haiku inside an unbounded agent loop changes the per-token cost. It doesn&#8217;t change the loop.</p><p>Model choice is a 2&#8211;5&#215; lever. Architecture choice is closer to an order-of-magnitude lever in the systems I&#8217;ve shipped &#8212; consistently larger than what model swaps deliver. Most teams are over-tuning the model selection and under-tuning the structure around it.</p><p>The default assumption &#8212; including from vendors with strong incentives to sell you more tokens &#8212; is that the answer to AI cost is buying inference more cleverly. The actual answer is using inference less, more deliberately, with hard bounds on what each call is allowed to do.</p><h2>How I Think About Inference Now</h2><p><strong>Inference is a power tool.</strong> Not a default. You don&#8217;t reach for it when a search query, a regex, or a <code>switch</code> statement would do. You reach for it when you need probabilistic judgment over unstructured input. Every call you don&#8217;t make is the cheapest call.</p><p><strong>Use it to code solutions, not to solve problems.</strong> The highest-leverage use of LLMs in my workflow is generating the deterministic code that then handles the workflow without further LLM calls. A model that writes you a 50-line classifier is more valuable than a model that <em>acts as</em> the classifier on every request forever. The first costs tokens once. The second costs tokens every transaction for the life of the system.</p><p><strong>Wrap every call in a budget.</strong> Context budget on the input, token budget on the output, quality gate on whether the call runs at all. Treat the LLM call as you&#8217;d treat a paid API with rate limits and SLA penalties. Because it is.</p><p><strong>Set specific ROI targets per call.</strong> &#8220;AI-assisted code review&#8221; is too coarse to optimize. &#8220;Reviewing files with relevance score &gt; 0.005, capped at 6 files, returning 500 words&#8221; is something you can measure cost-per-outcome on. Even loose ROI math at the call level surfaces where you&#8217;re paying for theater.</p><p><strong>Treat behavioral cost as the primary risk.</strong> Model rate cards will keep coming down. They are not your problem. Your problem is what your pipeline asks of the model and what the model decides to do once asked. That&#8217;s the line item that grew while the unit cost dropped 280&#215;. That&#8217;s the shoe that just dropped.</p><h2>What This Means If You&#8217;re Running an AI Budget</h2><p>Three things to look at, in order of how much they&#8217;ll move the line item.</p><p><strong>Audit the call graph, not the rate card.</strong> Pull a representative day of production traffic and trace the actual LLM calls per user task. Count them. Most teams find a handful of workflows producing the majority of cost, and most of those have 2&#8211;4 LLM calls that could be replaced by deterministic code. That&#8217;s the consolidated-design pattern from the duetto-intelligence example. 50&#8211;80% reductions are common when you actually look.</p><p><strong>Put quality gates in front of inference.</strong> For any workflow where the LLM call is expensive and the input quality is variable, add a deterministic check that decides whether the call is worth making. That&#8217;s the search-first pattern from the code-intelligence example. The savings come from the calls you <em>don&#8217;t</em> make, which never show up on the invoice.</p><p><strong>Set hard context budgets and enforce them in code.</strong> Per-source character limits, per-call token caps, no &#8220;just in case&#8221; context stuffing. The output budget gets attention because it&#8217;s visible. The input budget is usually where the actual money goes.</p><p>None of this requires changing models, switching providers, or making bets on the next frontier release. It&#8217;s architectural work inside the pipeline you already have &#8212; work the per-token price chart has been letting people defer.</p><p>The teams that do it over the next two quarters will look like they got a 5&#8211;10&#215; cost improvement from &#8220;AI getting cheaper.&#8221; The teams that don&#8217;t will look like AI got 3&#8211;4&#215; more expensive while everyone else&#8217;s costs fell. Same providers, same models, same rate cards. Different shoe.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and writes about AI-powered engineering at <a href="https://hyperdev.substack.com/">HyperDev</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/p/opus-46-vs-47-the-real-cost-of-incremental">Opus 4.6 vs 4.7: The Real Cost of Incremental AI Improvements</a> &#8212; The first shoe, on per-task cost drift between model versions</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul><p></p>]]></content:encoded></item><item><title><![CDATA[Is Your Digital Brain the Light Saber of the AI Era?]]></title><description><![CDATA[Jedi Knight Tools for the Knowledge Worker]]></description><link>https://hyperdev.matsuoka.com/p/is-your-digital-brain-the-light-saber</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/is-your-digital-brain-the-light-saber</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Wed, 20 May 2026 12:02:03 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!90qs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!90qs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!90qs!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png 424w, https://substackcdn.com/image/fetch/$s_!90qs!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png 848w, https://substackcdn.com/image/fetch/$s_!90qs!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png 1272w, https://substackcdn.com/image/fetch/$s_!90qs!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!90qs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png" width="1456" height="1007" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1007,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5326200,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/198515384?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!90qs!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png 424w, https://substackcdn.com/image/fetch/$s_!90qs!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png 848w, https://substackcdn.com/image/fetch/$s_!90qs!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png 1272w, https://substackcdn.com/image/fetch/$s_!90qs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Your Digital Lightsaber is your AI Memory</figcaption></figure></div><p>My Duetto colleague Jake Becker is sharp. He&#8217;s been ahead of AI adoption on our team &#8212; experimenting early, pushing for new tools, staying current. Last week he messaged me on Slack: &#8220;I wish I had my own CTO Assistant. Like what you have.&#8221;</p><p>I paused.</p><p>I have one. I&#8217;ve had several, in different forms, going back months. But I hadn&#8217;t said anything about it publicly.</p><p>That gap &#8212; between an AI-forward person who knows the tools and a practitioner who actually has the thing &#8212; is what this piece is about.</p><p><strong>TL;DR</strong></p><ul><li><p>Off-the-shelf AI tools are capable but contextually blind. You still have to leave your domain to get help.</p></li><li><p>Serious knowledge workers &#8212; writers, researchers, historians &#8212; have always built their own knowledge systems. Never trusted a vendor to hold their material.</p></li><li><p>Coding is shifting toward what writing has always been: directing, curating, maintaining a living body of knowledge.</p></li><li><p>Building your own AI memory and search layer is the new rite of passage. The tool you build IS the skill being developed.</p></li><li><p>The latest versions of my own stack: trusty-memory and trusty-search &#8212; both open source, both installable today.</p></li></ul><h2>Everyone Wants One. Few Have One.</h2><p>Jake&#8217;s comment revealed something I&#8217;d been taking for granted.</p><p>The major vendors are actively working on this &#8212; memory features, RAG pipelines, personalization layers. But those solutions are optimized for breadth, not for one person&#8217;s actual work across functions.</p><p>ChatGPT, Claude, Copilot &#8212; these are all capable. They&#8217;re also still contextually blind to <em>you</em>. You have to leave your environment, paste in context, explain your situation from scratch, then interpret the output back into your actual work. Vendors are working on it. But the solutions remain generic, and every session still starts from zero.</p><p>I wrote about this in March &#8212; <a href="https://hyperdev.matsuoka.com/personal-bots-abomination">Everyone Blamed Clawd Bot&#8217;s Execution. The Concept Was the Problem.</a> The structural flaw of universal assistants isn&#8217;t fixable. They require you to leave your context to get help. What actually works is the opposite: your tools get assistant capabilities, and assistance comes to where your context lives.</p><p>Off-the-shelf tools haven&#8217;t solved this. They&#8217;ve gotten more powerful &#8212; better reasoning, longer context, faster inference &#8212; but they still don&#8217;t know your codebase, your decisions, your institutional history, your current sprint. They know a lot about the world, and very little about you.</p><p>Jake wanted <em>my</em> assistant. But what he actually wants is <em>his</em> assistant. The one that knows what he knows.</p><p>That&#8217;s a different problem entirely.</p><p>The job changed first.</p><h2>Coding Is Becoming Writing</h2><p>Practitioners feel it before analysts name it.</p><p>A few years ago, being a strong engineer meant writing a lot of code quickly and correctly. Today, with agentic AI coders at their disposal, the best engineers I watch spend their time directing, reviewing, specifying, and curating. The unit of work has moved up a level. Implementation is increasingly delegated. Judgment &#8212; about architecture, trade-offs, what to build and why &#8212; is the differentiator.</p><p>This is not what happens when automation replaces a skill. It&#8217;s what happens when a new discipline appears.</p><p>Writers have always worked this way. A novelist doesn&#8217;t produce words per minute as a primary metric. They produce decisions &#8212; what to say, in what order, with what emphasis. The words are the output of the decisions, not the work itself. What makes a writer productive over a career isn&#8217;t typing speed. It&#8217;s having a system: notes, research, accumulated material, patterns of thought that compound over years.</p><p>Some (typically senior) engineers are arriving at the same realization. Your value isn&#8217;t the code. It&#8217;s the judgment, the accumulated context, the knowledge of what was tried and why it failed. The question is whether that accumulates in your head alone &#8212; which doesn&#8217;t scale, and doesn&#8217;t survive a context switch &#8212; or whether it lives in a system.</p><p>Good engineers who learn to use their AI tools effectively generate better code &#8212; the stack amplifies judgment and accumulated knowledge. The gap isn&#8217;t between skilled and unskilled engineers in isolation; it&#8217;s between engineers who&#8217;ve wired their knowledge into their tools and those who haven&#8217;t. The knowledge is real in both cases. Only in one case does it compound.</p><p>The writers I&#8217;ve observed who sustain serious output over decades all have the same property: they know where things are. Their research is retrievable. Their earlier thinking is available to their current thinking. The system makes the person bigger than their working memory.</p><p>That&#8217;s the gap.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Ynth!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd8f91e4-5bab-49cc-8f75-bb43db12e9e7_1999x1227.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Ynth!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd8f91e4-5bab-49cc-8f75-bb43db12e9e7_1999x1227.png 424w, https://substackcdn.com/image/fetch/$s_!Ynth!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd8f91e4-5bab-49cc-8f75-bb43db12e9e7_1999x1227.png 848w, https://substackcdn.com/image/fetch/$s_!Ynth!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd8f91e4-5bab-49cc-8f75-bb43db12e9e7_1999x1227.png 1272w, https://substackcdn.com/image/fetch/$s_!Ynth!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd8f91e4-5bab-49cc-8f75-bb43db12e9e7_1999x1227.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Ynth!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd8f91e4-5bab-49cc-8f75-bb43db12e9e7_1999x1227.png" width="1456" height="894" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dd8f91e4-5bab-49cc-8f75-bb43db12e9e7_1999x1227.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:894,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5603002,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/198515384?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd8f91e4-5bab-49cc-8f75-bb43db12e9e7_1999x1227.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Ynth!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd8f91e4-5bab-49cc-8f75-bb43db12e9e7_1999x1227.png 424w, https://substackcdn.com/image/fetch/$s_!Ynth!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd8f91e4-5bab-49cc-8f75-bb43db12e9e7_1999x1227.png 848w, https://substackcdn.com/image/fetch/$s_!Ynth!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd8f91e4-5bab-49cc-8f75-bb43db12e9e7_1999x1227.png 1272w, https://substackcdn.com/image/fetch/$s_!Ynth!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd8f91e4-5bab-49cc-8f75-bb43db12e9e7_1999x1227.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"> The AI Zettelkasten</figcaption></figure></div><h2>Writers Don&#8217;t Trust Vendors With Their Material</h2><p>Niklas Luhmann, a German sociologist working in the 1950s, produced 70 books and nearly 400 articles over his career. I haven&#8217;t written a book yet &#8212; but I have published over 200 articles at HyperDev. Worth naming the parallel. He credits his output to his Zettelkasten &#8212; a slip-box system of 90,000 interconnected index cards, each with a unique identifier linking it to related thoughts. Not a filing cabinet. A network of ideas that got richer with every addition.</p><p>And it&#8217;s worth asking: is that so different from what an AI knowledge system does today? The Zettelkasten was an analog precursor to what trusty-memory and trusty-search do programmatically &#8212; indexing ideas, linking related thoughts, surfacing connections that wouldn&#8217;t otherwise be visible. Luhmann was doing manually what these tools do automatically. Same architecture. New substrate.</p><p>He didn&#8217;t use a vendor product. He built a system that reflected how he thought. The architecture of his Zettelkasten was itself an expression of his intellectual method.</p><p>This isn&#8217;t a historical quirk. It&#8217;s a pattern. Serious knowledge workers have always built their own systems &#8212; commonplace books, research archives, private wikis. The reason is structural: the schema you design reflects how you think. That&#8217;s not something a product gives you. A product gives everyone the same schema.</p><p>Andrej Karpathy pointed at something similar last month with his LLM Wiki gist. His framing: use LLMs not just to write code, but to build and maintain a personal knowledge base. &#8220;Obsidian is the IDE, the LLM is the programmer, the wiki is the codebase.&#8221; Three folders, structured Markdown, a large context window. He concluded: &#8220;I think there is room here for an incredible new product.&#8221;</p><p>He&#8217;s right there&#8217;s room. I wrote about his framing in <a href="https://hyperdev.matsuoka.com/whats-in-your-second-brain">What&#8217;s In Your Second Brain?</a> The product comment is where I&#8217;d push back. You can build tooling around the pattern. You can&#8217;t productize the schema. The schema is the moat &#8212; because it reflects how <em>you</em> think, not how a product manager thinks you think. The ones who get it aren&#8217;t waiting for a product.</p><h2>The Lightsaber Rite of Passage</h2><p>In Star Wars canon, a Padawan doesn&#8217;t receive a lightsaber. They build one.</p><p>The ritual is called the Gathering. Initiates travel alone to the Crystal Caves of Ilum. They have to find their kyber crystal &#8212; the crystal that&#8217;s attuned to them through the Force. The caves are shaped by the initiate&#8217;s own fears and insecurities. The crystal doesn&#8217;t go to the strongest or the fastest. It bonds with the person who confronts what&#8217;s in the way.</p><p>Then they build it themselves, guided by Professor Huyang.</p><p>You can&#8217;t buy this. You can&#8217;t inherit it. The construction is the training. The tool reflects the builder.</p><p>I&#8217;m not the first to reach for this metaphor in tech. But I think it lands differently now. Building your own AI memory and search layer isn&#8217;t just useful. It&#8217;s diagnostic. You can&#8217;t do it without confronting what you actually know, how you actually think, what deserves to persist and what doesn&#8217;t. The schema you design for your knowledge base is a statement about your mind.</p><p>The engineers I know who are operating at the highest level right now &#8212; CTOs, senior architects, tech leads at places moving fast &#8212; they all quietly roll their own. They don&#8217;t announce it. They just have it.</p><h2>My Own Lineage</h2><p>I&#8217;ve been building versions of this for months.</p><p>trusty-izzie was the first &#8212; a simple wrapper. Then ai-commander, a more structured approach to context management. Then open-mpm and claude-mpm, which was where I started thinking seriously about multi-agent orchestration. Then kuzu-memory, a graph-backed memory layer. Then mcp-vector-search, semantic search over my entire codebase.</p><p>Each iteration taught me something about what I actually needed. Not what I thought I needed. What the practice revealed.</p><p>This piece was drafted with a configured writing assistant &#8212; <a href="https://github.com/bobmatnyc/claude-mpm">claude-mpm</a> loaded with my publication style guide, my voice patterns, my article archive. That&#8217;s a saber too. Not a generic chat interface. A tool shaped around how I think and write, producing work I can actually publish rather than work I have to fix. The saber list keeps growing.</p><p>The latest two are the most capable tools I&#8217;ve built.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cyse!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fc209a6-05f2-4b6d-8950-09c7562bd1ac_2005x1419.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cyse!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fc209a6-05f2-4b6d-8950-09c7562bd1ac_2005x1419.png 424w, https://substackcdn.com/image/fetch/$s_!cyse!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fc209a6-05f2-4b6d-8950-09c7562bd1ac_2005x1419.png 848w, https://substackcdn.com/image/fetch/$s_!cyse!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fc209a6-05f2-4b6d-8950-09c7562bd1ac_2005x1419.png 1272w, https://substackcdn.com/image/fetch/$s_!cyse!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fc209a6-05f2-4b6d-8950-09c7562bd1ac_2005x1419.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cyse!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fc209a6-05f2-4b6d-8950-09c7562bd1ac_2005x1419.png" width="1456" height="1030" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0fc209a6-05f2-4b6d-8950-09c7562bd1ac_2005x1419.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1030,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:6229361,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/198515384?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fc209a6-05f2-4b6d-8950-09c7562bd1ac_2005x1419.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!cyse!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fc209a6-05f2-4b6d-8950-09c7562bd1ac_2005x1419.png 424w, https://substackcdn.com/image/fetch/$s_!cyse!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fc209a6-05f2-4b6d-8950-09c7562bd1ac_2005x1419.png 848w, https://substackcdn.com/image/fetch/$s_!cyse!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fc209a6-05f2-4b6d-8950-09c7562bd1ac_2005x1419.png 1272w, https://substackcdn.com/image/fetch/$s_!cyse!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fc209a6-05f2-4b6d-8950-09c7562bd1ac_2005x1419.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The Latest Sabers</h2><p><strong>trusty-memory</strong> is a machine-wide AI memory daemon written in Rust. It uses what I call the Memory Palace architecture &#8212; multiple named palaces, each for a different domain. Sub-5ms baseline retrieval on Apple Silicon. It runs as an MCP server for Claude Code, which means my assistant stores and retrieves memories automatically, across sessions, without me managing any of it explicitly.</p><pre><code><code>cargo install trusty-memory</code></code></pre><p>Available at <a href="https://crates.io/crates/trusty-memory">crates.io/crates/trusty-memory</a>.</p><p><strong>trusty-search</strong> is a machine-wide hybrid code search daemon, also in Rust. Always-on, one install per machine. It combines BM25 lexical search with HNSW vector search (all-MiniLM-L6-v2 INT8) and a Knowledge Graph with 1-2 hop expansion, fused via Reciprocal Rank Fusion. It exposes an MCP server with 11 tools. Stdio and HTTP/SSE transports drop straight into Claude Code.</p><pre><code><code>cargo install trusty-search</code></code></pre><p>Available at <a href="https://crates.io/crates/trusty-search">crates.io/crates/trusty-search</a>.</p><p>These tools, along a set of custom reporting pythons apps along with a custom Slack Bot I use to access the data remotely, comprise my digital brain.</p><p>These aren&#8217;t products I bought. These are tools I built, iterated, and use daily. They know my codebase the way a Zettelkasten knows a scholar&#8217;s intellectual territory &#8212; not because a vendor configured them, but because I did.</p><p>To be precise: trusty-memory and trusty-search are infrastructure utilities &#8212; the memory layer and the search layer. Building the actual assistant that uses them is a separate act of customization. That&#8217;s where the lightsaber metaphor completes: the kyber crystal is only part of it. The construction &#8212; what you build with the crystal &#8212; is the saber.</p><p>When Jake said he wished he had a CTO Assistant, this is what he was gesturing at. Not a prompt template. Not a workflow. A living knowledge layer that compounds.</p><h2>The Right Question</h2><p>Jake asked: &#8220;Can I get a CTO Assistant?&#8221;</p><p>That&#8217;s the wrong question. It assumes the thing is available off the shelf, and the task is finding and configuring it.</p><p>The right question is: &#8220;What would it take to build one that knows what I know?&#8221;</p><p>That question is harder. It requires confronting the shape of your knowledge, what&#8217;s worth persisting, how to structure retrieval. It&#8217;s uncomfortable in the same way the Crystal Caves are uncomfortable &#8212; not because the work is technically difficult, but because you have to be honest about what you actually have.</p><p>Not everyone needs to write Rust. The specific technology isn&#8217;t the point. The point is that the engineers asking the right question are already operating differently. They&#8217;re working like writers &#8212; maintaining a living body of knowledge, building systems that compound, treating their accumulated context as an asset rather than a liability.</p><p>Writers&#8217; discipline has been creeping into engineering for a while. AI made it urgent.</p><p>If you&#8217;re waiting for a product to hand you the thing, you&#8217;re waiting for someone to build your lightsaber. It won&#8217;t be yours.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and writes about AI-powered engineering at <a href="https://hyperdev.substack.com/">HyperDev</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/whats-in-your-second-brain">What&#8217;s In Your Second Brain?</a> &#8212; Karpathy&#8217;s LLM Wiki and the case for a compounding knowledge layer</p></li><li><p><a href="https://hyperdev.matsuoka.com/personal-bots-abomination">Everyone Blamed Clawd Bot&#8217;s Execution. The Concept Was the Problem.</a> &#8212; Why universal assistants are architecturally broken</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul>]]></content:encoded></item><item><title><![CDATA[Why Rust is the AI Language of the Future]]></title><description><![CDATA[And It&#8217;s All About the Compiler]]></description><link>https://hyperdev.matsuoka.com/p/why-rust-is-the-ai-language-of-the</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/why-rust-is-the-ai-language-of-the</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Wed, 13 May 2026 12:03:01 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!pFtT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99aa09ae-80b1-4080-b709-45e9f1016d32_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>A developer&#8217;s journey from Python to Rust reveals why the compiler, not the runtime, will define the next generation of AI systems.</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!pFtT!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99aa09ae-80b1-4080-b709-45e9f1016d32_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!pFtT!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99aa09ae-80b1-4080-b709-45e9f1016d32_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!pFtT!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99aa09ae-80b1-4080-b709-45e9f1016d32_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!pFtT!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99aa09ae-80b1-4080-b709-45e9f1016d32_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!pFtT!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99aa09ae-80b1-4080-b709-45e9f1016d32_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!pFtT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99aa09ae-80b1-4080-b709-45e9f1016d32_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/99aa09ae-80b1-4080-b709-45e9f1016d32_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1563722,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/197158389?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99aa09ae-80b1-4080-b709-45e9f1016d32_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!pFtT!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99aa09ae-80b1-4080-b709-45e9f1016d32_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!pFtT!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99aa09ae-80b1-4080-b709-45e9f1016d32_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!pFtT!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99aa09ae-80b1-4080-b709-45e9f1016d32_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!pFtT!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99aa09ae-80b1-4080-b709-45e9f1016d32_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Six months ago, I abandoned my Rust/Tauri-based writing app. Not because Tauri was bad&#8212;it wasn&#8217;t. Obsidian simply worked better for my needs. But that experience with Rust left an impression: the compiler was unlike anything I&#8217;d encountered. It didn&#8217;t just catch errors; it made entire classes of bugs impossible.</p><p>Then I built AI Commander for managing tmux sessions. Python would have been the obvious choice&#8212;async programming, process management, plenty of libraries. But I chose Rust, partly out of curiosity, partly because I needed rock-solid thread control. The result was revelation: code that compiled cleanly worked correctly, and performance was extraordinary.</p><p>Now I&#8217;m deep into replacing both mcp-vector-search and kuzu-memory with pure Rust implementations. My new trusty-search outperforms the Python predecessor by an order of magnitude, even though that version used compiled Python extensions. This isn&#8217;t an isolated case&#8212;it&#8217;s part of a fundamental shift happening in AI development.</p><p><strong>Rust isn&#8217;t just another systems language. It&#8217;s becoming the infrastructure language of AI, and the reason has everything to do with letting the compiler handle the heavy lifting&#8212;especially as AI systems increasingly write their own code.</strong></p><h2>The Performance Reality Check</h2><p>The performance gains are measurable and significant. <a href="https://github.com/huggingface/tokenizers">Hugging Face&#8217;s tokenizers library</a>, rewritten in Rust, delivers 10-100x speedups over pure Python implementations. <a href="https://www.pola.rs/">Polars, the Rust-based DataFrame library</a>, consistently outperforms pandas in data processing benchmarks. In embedded systems, <a href="https://blog.rust-embedded.org/performance-engineering/">Rust achieves 98% of C performance</a> while eliminating memory safety issues entirely.</p><p>This isn&#8217;t just raw speed. Memory efficiency matters significantly in production AI systems. Where Python applications might consume gigabytes for large-scale data processing, equivalent Rust implementations often require 3-5x less memory. The difference compounds when deploying AI systems at scale.</p><p>This performance advantage stems from a fundamental philosophical difference in how Rust approaches the safety-speed trade-off.</p><p>When your AI system processes millions of sensor events per second in an autonomous vehicle, or makes split-second trading decisions with millions at stake, runtime failures aren&#8217;t just inconvenient&#8212;they&#8217;re catastrophic. Python&#8217;s flexibility, its greatest strength for research and experimentation, becomes a liability in these production environments.</p><h2>The Compiler Does the Heavy Lifting</h2><p>Here&#8217;s where Rust&#8217;s philosophy is revolutionary. Most languages force you to choose between safety and performance, between readable code and efficient code. Rust rejects this trade-off entirely through what it calls &#8220;zero-cost abstractions.&#8221;</p><p>The principle, borrowed from C++ but perfected in Rust, is elegantly simple: <a href="https://without.boats/blog/zero-cost-abstractions/">&#8220;What you don&#8217;t use, you don&#8217;t pay for. What you do use, you couldn&#8217;t hand code any better.&#8221;</a></p><p>Consider async programming, critical for AI systems handling multiple data streams. In Python, async operations carry runtime overhead&#8212;coroutine scheduling, context switching, memory allocation for tasks. In Rust, the compiler transforms your high-level async/await code into efficient state machines. No runtime scheduler, no hidden allocations, just bare-metal performance wrapped in readable syntax.</p><p><strong>The genius is that the compiler absorbs the complexity.</strong> You write code that looks high-level and readable:</p><pre><code><code>async fn process_ai_requests(stream: &amp;mut DataStream) -&gt; Result&lt;Vec&lt;Response&gt;, Error&gt; {
    let mut responses = Vec::new();

    while let Some(request) = stream.next().await {
        let response = ai_model.infer(request).await?;
        responses.push(response);
    }

    Ok(responses)
}</code></code></pre><p>But the compiler generates assembly that rivals hand-optimized C. The ownership system prevents data races without locks. The borrow checker eliminates memory leaks without garbage collection. The type system catches logic errors that would be runtime failures in dynamic languages.</p><p><strong>This is the heavy lifting</strong>: converting human-friendly abstractions into machine-efficient reality, all at compile time, with zero runtime cost.</p><p>This compiler-centric philosophy becomes even more critical as we enter an era where AI systems increasingly write their own code.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zxzP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0807e8e4-b156-4e49-9563-5f16505ac2db_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zxzP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0807e8e4-b156-4e49-9563-5f16505ac2db_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!zxzP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0807e8e4-b156-4e49-9563-5f16505ac2db_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!zxzP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0807e8e4-b156-4e49-9563-5f16505ac2db_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!zxzP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0807e8e4-b156-4e49-9563-5f16505ac2db_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zxzP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0807e8e4-b156-4e49-9563-5f16505ac2db_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0807e8e4-b156-4e49-9563-5f16505ac2db_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1581968,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/197158389?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0807e8e4-b156-4e49-9563-5f16505ac2db_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!zxzP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0807e8e4-b156-4e49-9563-5f16505ac2db_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!zxzP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0807e8e4-b156-4e49-9563-5f16505ac2db_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!zxzP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0807e8e4-b156-4e49-9563-5f16505ac2db_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!zxzP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0807e8e4-b156-4e49-9563-5f16505ac2db_1024x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The AI Code Generation Protection Paradox</h2><p>As AI-powered coding tools like Copilot, Claude Code, and Cursor become mainstream, a new challenge emerges: <strong>how do you ensure correctness when the code author might not fully understand what they&#8217;ve written?</strong></p><p>When Claude or Copilot generates Python, Java, TypeScript, or C code, human oversight becomes the primary error-checking mechanism. Code reviewers need to catch type mismatches, memory leaks, race conditions, and logic errors that might not surface until production. As AI-generated code becomes more prevalent and complex, <a href="https://thenewstack.io/rust-creator-graydon-hoare-talks-about-security-history-and-rust/">human review processes struggle to keep pace</a> with the volume and complexity of generated code.</p><p><strong>Think about it</strong>: an AI might generate a perfectly logical-looking concurrent data structure in Python that works flawlessly in single-threaded tests but creates subtle race conditions under load. A human reviewer might miss the threading implications entirely. The same AI generating Rust code? The compiler would reject it outright if the data sharing wasn&#8217;t provably safe.</p><p><strong>Rust&#8217;s compiler fundamentally changes this dynamic.</strong> It doesn&#8217;t matter if the code was written by a human, an AI, or a collaboration between both&#8212;if it compiles, entire categories of dangerous bugs simply cannot exist. Memory safety violations? Caught at compile time. Data races in concurrent code? Impossible. Use-after-free errors? Eliminated by the ownership system.</p><p>This creates a fascinating inversion: <strong>Rust may be the best language for AI-generated code precisely because it doesn&#8217;t trust the programmer.</strong> While other languages rely on developer discipline and code review to catch errors, Rust enforces correctness through the type system and ownership model. The compiler becomes an automated, exhaustive code reviewer that never gets tired, never misses subtle bugs, and never waves through &#8220;probably fine&#8221; code.</p><p>As Microsoft research demonstrates, <a href="https://msrc.microsoft.com/blog/2019/07/a-proactive-approach-to-more-secure-code/">70% of security vulnerabilities stem from memory safety issues</a> that could be eliminated at compile time. When AI systems generate infrastructure code, this shifts from a productivity concern to an existential safety issue.</p><p><strong>The compiler isn&#8217;t just doing heavy lifting for performance; it&#8217;s providing safety guarantees that scale beyond human oversight.</strong></p><h2>Memory Safety in the Age of AI</h2><p><a href="https://thenewstack.io/rust-creator-graydon-hoare-talks-about-security-history-and-rust/">Rust&#8217;s creator, Graydon Hoare</a>, designed the language after getting stuck in a broken elevator caused by software crashes. His goal was simple: write fast, small code without memory bugs. That mission has become critical as AI systems move from labs to production infrastructure.</p><p>In traditional software, memory safety issues mean patches and updates. In AI systems controlling physical infrastructure&#8212;autonomous vehicles, medical devices, financial trading systems&#8212;they can mean catastrophic failure.</p><p>Rust&#8217;s ownership model doesn&#8217;t just prevent crashes; it eliminates entire categories of bugs that plague concurrent AI workloads. While C++ applications in multi-threaded environments regularly exhibit race conditions during testing, Rust&#8217;s compile-time guarantees make such issues structurally impossible.</p><p><strong>But here&#8217;s the crucial part: this safety comes at zero runtime cost.</strong> Unlike garbage-collected languages that trade performance for memory safety, <a href="https://reintech.io/blog/understanding-rust-zero-cost-abstractions/">Rust enforces safety through compile-time checks</a>. Your production AI system runs with the performance of C but the safety guarantees that traditional systems languages can&#8217;t provide.</p><h2>The Python-Rust Symbiosis</h2><p>Here&#8217;s where the narrative gets interesting. <a href="https://blog.jetbrains.com/rust/2025/11/10/rust-vs-python-finding-the-right-balance-between-speed-and-simplicity/">The emerging trend isn&#8217;t Rust replacing Python&#8212;it&#8217;s Rust and Python working together</a>. The <a href="https://pyo3.rs/">PyO3 bridge</a>, Rust-powered Python tools like <a href="https://github.com/astral-sh/ruff">Ruff</a> and <a href="https://github.com/astral-sh/uv">uv</a>, and the growing number of &#8220;Python API, Rust engine&#8221; libraries show that the two languages are becoming symbiotic.</p><p><strong>Python remains dominant for AI research and experimentation.</strong> The ecosystem is unmatched: PyTorch, TensorFlow, scikit-learn, Hugging Face Transformers. PyTorch&#8217;s <a href="https://github.com/pytorch/pytorch">100,000+ GitHub stars</a> reflect an ecosystem that won&#8217;t disappear overnight. When you&#8217;re prototyping models, analyzing data, or building internal tools, Python&#8217;s flexibility and ecosystem make it irreplaceable.</p><p><strong>But production tells a different story.</strong> When performance, memory usage, and reliability matter&#8212;when you&#8217;re building the infrastructure that powers AI systems rather than the models themselves&#8212;Rust increasingly dominates.</p><p>The pattern emerging across tech giants is clear: Python for the interface layer, Rust for the infrastructure layer. Experimentation happens in Jupyter notebooks; production happens in compiled Rust binaries.</p><h2>Looking Forward: Infrastructure vs. Experimentation</h2><p>This division isn&#8217;t arbitrary&#8212;it reflects the maturation of AI from research curiosity to critical infrastructure. Research requires flexibility, rapid iteration, and access to cutting-edge libraries. Production requires reliability, performance, and safety guarantees.</p><p><strong>Rust excels at infrastructure because the compiler handles the complexity of building robust systems.</strong> Memory safety, concurrency, error handling&#8212;all the concerns that make production systems hard to build correctly&#8212;are enforced by the type system rather than relying on developer discipline.</p><p>Consider the trajectory: <a href="https://thenewstack.io/microsofts-bold-goal-replace-1b-lines-of-c-c-with-rust/">Microsoft aims to eliminate C and C++ from their codebase by 2030</a>, replacing it with Rust. <a href="https://rustfoundation.org/resource/rust-and-ai-position-statement/">The Rust Foundation officially positions</a> the language as ideal for &#8220;ultra-reliable AI systems.&#8221; Major tech companies are increasingly choosing Rust for performance-critical AI infrastructure. These aren&#8217;t experimental projects&#8212;they&#8217;re billion-dollar bets on where AI infrastructure is heading.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OGhR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e07498-b662-48d5-9fad-b3061ac8ecfe_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OGhR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e07498-b662-48d5-9fad-b3061ac8ecfe_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!OGhR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e07498-b662-48d5-9fad-b3061ac8ecfe_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!OGhR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e07498-b662-48d5-9fad-b3061ac8ecfe_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!OGhR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e07498-b662-48d5-9fad-b3061ac8ecfe_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OGhR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e07498-b662-48d5-9fad-b3061ac8ecfe_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f4e07498-b662-48d5-9fad-b3061ac8ecfe_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1509742,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/197158389?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e07498-b662-48d5-9fad-b3061ac8ecfe_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!OGhR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e07498-b662-48d5-9fad-b3061ac8ecfe_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!OGhR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e07498-b662-48d5-9fad-b3061ac8ecfe_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!OGhR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e07498-b662-48d5-9fad-b3061ac8ecfe_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!OGhR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4e07498-b662-48d5-9fad-b3061ac8ecfe_1024x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The Compiler Advantage</h2><p>What makes Rust uniquely suited for AI isn&#8217;t just performance or safety&#8212;it&#8217;s the fundamental approach of front-loading complexity into the compiler rather than dealing with it at runtime.</p><p><strong>In AI systems, runtime failures are exponentially more expensive than compile-time failures.</strong> A training job that crashes after days of computation. An inference system that memory-leaks during peak traffic. A edge AI device that locks up in a safety-critical situation. These failures cascade through systems in ways that research environments rarely encounter.</p><p>Rust&#8217;s compiler is exhaustive in a way that benefits AI development specifically. It catches data races in concurrent model training. It prevents buffer overflows in tensor operations. It enforces proper error handling in distributed inference systems. It does this without performance overhead, without runtime monitoring, and without hoping that comprehensive testing caught every edge case.</p><p><strong>The compiler becomes your co-pilot in building reliable AI infrastructure.</strong></p><h2>The Personal Proof</h2><p>My own experience validates this thesis. AI Commander needed complex thread management for tmux session control&#8212;exactly the kind of concurrency that&#8217;s error-prone in traditional languages. Rust&#8217;s ownership system made data sharing across threads not just safe, but intuitive. The code that compiled was correct.</p><p>With trusty-search, the performance gains over mcp-vector-search weren&#8217;t just about Rust being &#8220;faster than Python.&#8221; They were about Rust enabling architectural choices&#8212;zero-copy string processing, efficient memory layouts, fine-grained concurrency control&#8212;that would be risky or impossible in garbage-collected languages.</p><p><strong>These aren&#8217;t micro-optimizations. They&#8217;re fundamental design differences that become critical at scale.</strong></p><h2>The Real Limitations of Rust for AI</h2><p>Before declaring Rust the future, it&#8217;s essential to acknowledge where it falls short compared to Python in AI contexts.</p><p><strong>The learning curve is steep.</strong> Rust&#8217;s ownership model requires a fundamental mental shift that can slow initial development. Developers comfortable with garbage-collected languages often struggle with borrowing rules, especially when building complex data structures. The compiler&#8217;s strict requirements, while beneficial long-term, can feel obstructionist when prototyping.</p><p><strong>The ecosystem gap remains significant.</strong> While Rust has impressive performance libraries, Python&#8217;s AI ecosystem is vast and mature. TensorFlow, PyTorch, scikit-learn, and thousands of specialized ML libraries have no Rust equivalents. Building production ML pipelines often requires library combinations that simply don&#8217;t exist in Rust.</p><p><strong>Compilation time can kill iteration speed.</strong> Where Python allows instant feedback during development, Rust&#8217;s compilation process&#8212;especially for complex projects&#8212;can introduce friction that slows experimentation. This matters significantly in AI research where rapid iteration drives discovery.</p><p><strong>Talent pool challenges are real.</strong> Finding experienced Rust developers is substantially harder than finding Python developers. The language&#8217;s complexity means onboarding takes longer, potentially impacting team velocity and project timelines.</p><p><strong>Not every AI workload benefits from Rust&#8217;s strengths.</strong> Data analysis, model training with existing frameworks, and one-off scripts often don&#8217;t require the safety and performance guarantees Rust provides. Using Rust for these tasks can be overkill that reduces productivity without meaningful benefits.</p><p>These limitations explain why the Python-Rust symbiosis model makes sense: leverage each language where it excels rather than forcing one-size-fits-all solutions.</p><h2>The Future is Compiled</h2><p>As AI systems become infrastructure&#8212;as they move from experimental notebooks to production systems handling millions of users&#8212;the languages that power them will need to provide the guarantees that infrastructure requires.</p><p>Python will continue to dominate the experimental and interface layers. But the backbone, the high-performance inference systems, the real-time edge deployments, the safety-critical applications&#8212;these will increasingly be built in languages where the compiler, not the runtime, ensures correctness.</p><p><strong>Rust represents a fundamental shift in how we approach building reliable systems: moving complexity from runtime to compile time, from testing to proving, from hoping to knowing.</strong></p><p>This shift becomes critical as we enter the era of AI-generated code. When AI systems are writing the infrastructure that runs other AI systems, traditional human oversight breaks down. We need tools that can verify correctness automatically, at compile time, without relying on human reviewers to catch subtle but catastrophic errors.</p><p>The compiler does the heavy lifting so production systems can focus on their actual job: running AI workloads reliably, safely, and at scale. In a world where AI infrastructure is becoming as critical as power grids and financial networks&#8212;and where that infrastructure is increasingly written by AI itself&#8212;that guarantee isn&#8217;t just nice to have, it&#8217;s existential.</p><p>The age of agentic AI isn&#8217;t just about better models. It&#8217;s about building the reliable, high-performance infrastructure those models need to operate in the real world, written by AI systems that can&#8217;t be trusted to write safe code in traditional languages. And increasingly, that infrastructure is being built in Rust, one compile-time guarantee at a time.</p><div><hr></div><p><em>Bob Matsuoka writes about the intersection of software engineering and AI at <a href="https://hyperdev.matsuoka.com/">HyperDev</a>. His latest Rust projects replace Python infrastructure with performance gains that would make even a compiler blush.</em></p>]]></content:encoded></item><item><title><![CDATA[AI Memento]]></title><description><![CDATA[Documents aren't the answer]]></description><link>https://hyperdev.matsuoka.com/p/ai-memento</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/ai-memento</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Mon, 11 May 2026 12:30:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!mhT3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6814f0f-dc1c-4093-8510-6b952bcace96_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!mhT3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6814f0f-dc1c-4093-8510-6b952bcace96_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!mhT3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6814f0f-dc1c-4093-8510-6b952bcace96_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!mhT3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6814f0f-dc1c-4093-8510-6b952bcace96_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!mhT3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6814f0f-dc1c-4093-8510-6b952bcace96_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!mhT3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6814f0f-dc1c-4093-8510-6b952bcace96_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!mhT3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6814f0f-dc1c-4093-8510-6b952bcace96_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c6814f0f-dc1c-4093-8510-6b952bcace96_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1451351,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/197050605?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6814f0f-dc1c-4093-8510-6b952bcace96_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!mhT3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6814f0f-dc1c-4093-8510-6b952bcace96_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!mhT3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6814f0f-dc1c-4093-8510-6b952bcace96_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!mhT3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6814f0f-dc1c-4093-8510-6b952bcace96_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!mhT3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc6814f0f-dc1c-4093-8510-6b952bcace96_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Leonard Shelby can&#8217;t form new memories. In <em>Memento</em>, after the attack, every fifteen minutes or so his short-term recall resets. He has a problem most engineers would recognize: a fast working scratch space and no durable storage. So he builds a system. Polaroids with handwritten captions on the back. A wall of pinned notes. Tattoos for the load-bearing facts &#8212; the ones he can&#8217;t afford to lose, can&#8217;t afford to misfile, can&#8217;t afford to look up wrong.</p><p>What he doesn&#8217;t do is try to remember everything. He doesn&#8217;t carry a thicker notebook. He builds infrastructure to find what matters, fast, when he needs it. The pocket holds what&#8217;s relevant right now. The system decides what makes it into the pocket.</p><p>I&#8217;ve been thinking about Leonard a lot this year, because the prevailing argument in my corner of the internet is that LLMs no longer need him. The pocket got bigger. Frontier models can hold a million tokens of context. Token prices fell. So just stuff the pocket. Skip the indexing. Skip the tattoos. Read everything.</p><p>I think that&#8217;s the wrong read of where we&#8217;re going.</p><h2>The real argument, fairly stated</h2><p>The <a href="https://lighton.ai/lighton-blogs/rag-is-dead-long-live-rag-retrieval-in-the-age-of-agents">&#8220;RAG is dead&#8221;</a> <a href="https://akitaonrails.com/en/2026/04/06/rag-is-dead-long-context/">position</a> isn&#8217;t crazy. It goes roughly like this: a handful of frontier models &#8212; Gemini 1.5, Claude&#8217;s extended context tiers &#8212; can now handle contexts approaching 1M tokens without choking. Token prices have fallen to around $0.60 per million for several frontier models. Standing up a real retrieval pipeline &#8212; chunking strategy, embedding model, vector store, hybrid scoring, reranker, eval harness &#8212; is 40 to 80 engineering hours, easily. For a small, static corpus you query a few thousand times a month, just dumping the whole thing into context with prompt caching is simpler and arguably cheaper.</p><p>Claude Code is the existence proof people point at. It doesn&#8217;t ship a vector database. It uses ripgrep, file globs, and reads files into context. And it works pretty well. So why bother with the rest?</p><p>I&#8217;ll concede the envelope. For a 300K-token corpus, single tenant, low query volume, mostly static &#8212; yeah, stuff it. Cache it. Read it. Don&#8217;t build a search stack to feel sophisticated. That&#8217;s a real position and I don&#8217;t want to strawman it.</p><p>But that envelope isn&#8217;t where most of the work lives. And the argument the loud version makes &#8212; that grep replaces search, that long context replaces retrieval &#8212; confuses two different things; there are cracks in that argument.</p><h2>The pocket problem</h2><p>Here&#8217;s the first crack. <a href="https://arxiv.org/abs/2307.03172">&#8220;Lost in the middle.&#8221;</a> Liu et al., 2023, TACL: when you stuff long contexts, model performance forms a U-shape. Strong recall at the beginning, strong recall at the end, degraded recall in the middle. The 2026-gen frontier models have improved on this &#8212; early evals suggest the curve is flatter &#8212; but it&#8217;s not gone.</p><p>So the first thing the bigger pocket buys you is the ability to put a lot of important stuff exactly where the model is most likely to underweight it. You can&#8217;t reorder a 700K-token dump by relevance unless you&#8217;ve already done retrieval. Which means the question doesn&#8217;t go away &#8212; it just hides.</p><p>Second crack: cost. RAG-style queries &#8212; retrieve relevant chunks, feed 8&#8211;16K of context &#8212; run around $0.005&#8211;$0.008 per request at $0.60/million input tokens, when you count the full prompt. A naive full-context pass against a 500K-token corpus runs $0.30 in input tokens alone, before output. That&#8217;s a 40&#8211;60x ratio at list price, and it widens if you have high cache-miss rates or large output windows. At 50K daily queries &#8212; not enormous, this is a mid-sized internal tool &#8212; you&#8217;re choosing between a few hundred dollars a day and tens of thousands. That&#8217;s not a rounding error you absorb. That&#8217;s a re-platforming decision that wakes someone up.</p><p>Third crack, and this is the one that actually makes me grumpy. Grep is not search. Grep is lexical pattern matching. Searching for <code>auth_token</code> finds the string <code>auth_token</code>. It finds it in commented-out dead code from 2023. It finds it in test fixtures with hardcoded sample values. It finds it in a vendor library you don&#8217;t even own. It misses the function called <code>refresh_session_credential</code> that does what you&#8217;re actually looking for, because the words are different. Vector search finds that function. Grep doesn&#8217;t. They&#8217;re not the same tool. Pretending they are is how you ship the wrong fix at 2 AM.</p><p>When people say &#8220;Claude Code uses grep, so search is over,&#8221; they&#8217;re describing a tool optimized for a bounded, file-system-structured corpus where the author controls the index shape. That&#8217;s a legitimate architecture for that problem. It does not generalize to &#8220;ten million product images&#8221; or &#8220;a 30-year codebase across 800 repos&#8221; or &#8220;tenant-scoped documents under per-row access control.&#8221; Long context can&#8217;t replace a knowledge graph, either. You can serialize a graph into text &#8212; flatten the edges, encode the relationships &#8212; but then you&#8217;ve lost typed edges and efficient query-time traversal. You can&#8217;t join across it. You can&#8217;t reason about edge types. You&#8217;d need to flatten the graph to feed it in, at which point you&#8217;ve discarded the structure that made it useful.</p><p>With me so far? The argument isn&#8217;t &#8220;long context is bad.&#8221; Long context is great. The argument is that long context is a <em>consumer</em> of retrieval, not a replacement for it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7U4n!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F53ac32bd-7a2e-4149-82fc-ddfbe69a2873_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7U4n!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F53ac32bd-7a2e-4149-82fc-ddfbe69a2873_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!7U4n!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F53ac32bd-7a2e-4149-82fc-ddfbe69a2873_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!7U4n!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F53ac32bd-7a2e-4149-82fc-ddfbe69a2873_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!7U4n!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F53ac32bd-7a2e-4149-82fc-ddfbe69a2873_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7U4n!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F53ac32bd-7a2e-4149-82fc-ddfbe69a2873_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/53ac32bd-7a2e-4149-82fc-ddfbe69a2873_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1245807,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/197050605?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F53ac32bd-7a2e-4149-82fc-ddfbe69a2873_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!7U4n!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F53ac32bd-7a2e-4149-82fc-ddfbe69a2873_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!7U4n!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F53ac32bd-7a2e-4149-82fc-ddfbe69a2873_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!7U4n!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F53ac32bd-7a2e-4149-82fc-ddfbe69a2873_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!7U4n!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F53ac32bd-7a2e-4149-82fc-ddfbe69a2873_1024x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>What good retrieval actually looks like</h2><p>A modern retrieval stack, the kind that holds up under real query volume, isn&#8217;t one thing. It&#8217;s a layered pipeline.</p><p>Lexical first &#8212; BM25, TF-IDF, the unfashionable stuff. It&#8217;s still excellent at exact tokens, identifiers, error messages, version numbers. The things vector embeddings smudge.</p><p>Dense second &#8212; vector search over embeddings. all-MiniLM-L6-v2 is the workhorse at 384 dimensions. Catches semantic neighbors, paraphrases, the function-named-different-but-doing-the-same-thing case.</p><p>Fuse them &#8212; Reciprocal Rank Fusion at k=60 is the standard, and it&#8217;s standard because it works. You don&#8217;t pick BM25 <em>or</em> vector. You merge their rank lists.</p><p>Rerank &#8212; a cross-encoder over the top-k from RRF. Slower per-pair, but you only run it on the candidates that made it through the cheap layers.</p><p>The numbers from a 2026 benchmark pass I was reading &#8212; <a href="https://arxiv.org/abs/2604.01733">&#8220;From BM25 to Corrective RAG,&#8221;</a> April 2026 &#8212; make the layered case cleanly:</p><ul><li><p>Dense vectors alone: Recall@5 of 0.587</p></li><li><p>BM25 alone: 0.644</p></li><li><p>Hybrid via RRF: 0.695</p></li><li><p>Hybrid + neural reranker: 0.816</p></li></ul><p>That last number is 39% better than dense-only. It&#8217;s also 17% better than RRF without the reranker. The layers compound. In systems that need both recall and precision at scale, you usually want all four layers. And &#8212; this part matters &#8212; none of them are replaced by a bigger context window. You still have to pick what goes into the pocket.</p><p>There&#8217;s a piece on top of this, too: intent. A query like &#8220;where is <code>process_payment</code> defined&#8221; is not the same shape as &#8220;how does the refund flow handle partial captures.&#8221; The first is a definition lookup; BM25 should dominate, you want exact identifier matching, you don&#8217;t want semantic neighbors fuzzing the result. The second is conceptual; dense embeddings should dominate, you want the chunk that explains the flow even if the words don&#8217;t match. Treating those queries identically is how you get plausible-but-wrong answers.</p><p>The EMNLP 2024 <a href="https://arxiv.org/search/?searchtype=all&amp;query=self-route+long+context+retrieval">&#8220;Self-Route&#8221; paper</a> made a related point I keep coming back to: let the model itself decide whether a query needs retrieval or full-context, based on complexity. They got better accuracy AND lower compute cost. The two approaches aren&#8217;t enemies. They&#8217;re complements with different cost/quality envelopes, and a working system uses both.</p><p><em>(A note on scope: I&#8217;ve been using &#8220;search&#8221; and &#8220;memory retrieval&#8221; somewhat interchangeably here, which isn&#8217;t quite right. They&#8217;re the same pipeline mechanics &#8212; retrieve, rank, inject &#8212; but different problems. Memory retrieval has a temporal model search doesn&#8217;t: recency matters, decay matters, and the source is agent history rather than a document corpus. Memory is also primarily graph-based, because what you&#8217;re trying to recover is relationships and context across time, not just chunks of text. For the purposes of this argument &#8212; context window vs. retrieval discipline &#8212; the distinction doesn&#8217;t change anything. But it&#8217;s worth naming.)</em></p><h2>What I&#8217;m building</h2><p>I have two projects in this space. I&#8217;ll be specific about what each one is and isn&#8217;t, because the field is full of demos that don&#8217;t survive contact with real corpora.</p><p><strong>mcp-vector-search</strong> is the older one. Python, per-project, designed to live next to a codebase as an MCP server. The retrieval core is hybrid: BM25 plus HNSW over MiniLM embeddings, with knowledge-graph expansion built from tree-sitter AST parsing &#8212; function definitions, call edges, import edges, type relationships. Fifteen-plus edge types in the graph. Cross-encoder reranking on top of RRF. MMR for diversity so you don&#8217;t get five near-duplicates in the top results. Temporal decay weighted by git blame age. Cyclomatic complexity factored into ranking, because dense, gnarly code is more often what you&#8217;re looking for than the trivial wrappers around it.</p><p><strong>trusty-search</strong> is what I&#8217;m building toward. Rust daemon, machine-wide rather than per-project, multi-tenant across all my work. Same retrieval core &#8212; RRF at k=60, MiniLM embeddings, BM25 plus HNSW &#8212; but with intent classification on the front. Queries get tagged Definition, Usage, Conceptual, or BugDebt, and each intent has pre-tuned alpha/beta weights between lexical and dense. Sub-10ms p50 on warm queries.</p><p>I&#8217;ll be honest: trusty-search is early-stage. Some of the features I just listed are scaffolding with stubs underneath. The intent classifier is rule-based at the moment and needs to become a small model. The two tools are complementary rather than competing for now &#8212; mcp-vector-search is the deeper analysis layer, trusty-search is the fast facts store you hit constantly during a session &#8212; though the long-term plan is for trusty-search to replace mcp-vector-search.</p><p>The reason I&#8217;m building both is that I keep running into the same wall: I can give an agent a million-token window and it will still ask the wrong question of the wrong file. The bottleneck is not how much I can stuff in. The bottleneck is which 8K of the 800K matters for <em>this</em> query. That&#8217;s a search problem. It&#8217;s been a search problem for sixty years. It&#8217;s still a search problem.</p><h2>Back to Leonard</h2><p>The reason Leonard works as an analogy isn&#8217;t the amnesia. It&#8217;s the discipline. He decides what makes it onto a polaroid. He decides what gets tattooed and what stays loose. He writes &#8220;DON&#8217;T BELIEVE HIS LIES&#8221; on the back of a photo because he won&#8217;t have the context to evaluate trustworthiness later &#8212; so he encodes the conclusion now, into a system he&#8217;ll find when he needs it.</p><p>The context window is the pocket. Whatever Leonard pulls out and looks at right now. It can be enormous. It can be cached. It can be cheap. None of that decides what goes in.</p><p>Retrieval is the tattoo on his wrist. The polaroid pinned to the wall. The system that decides which fact survives, which one is one query away, which one is buried. The job of that system is exactly the job that doesn&#8217;t go away when the pocket gets bigger.</p><p>The question was never &#8220;do I need RAG?&#8221; That framing turned a design discipline into a vendor category and made it easy to dismiss. The question is the one Leonard asks every morning: <em>what do I actually need to find, and have I built the thing that finds it?</em> At a million tokens, you&#8217;re not freed from that question. You&#8217;ve just made it more expensive to answer incorrectly.</p><p>It feels crazy that I wrote <a href="https://open.substack.com/pub/hyperdev/p/50-first-dates-with-claude-code?r=nff5&amp;utm_campaign=post-expanded-share&amp;utm_medium=web">this</a> only a year ago&#8230;</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and writes about AI-powered engineering at <a href="https://hyperdev.substack.com/">HyperDev</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/context-memory-search-agentic-work">Context Memory and Search: The Secrets to Effective Agentic Work</a> &#8212; Why context management, memory, and retrieval are the three pillars of effective agentic systems</p></li><li><p><a href="https://hyperdev.matsuoka.com/whats-in-your-second-brain">What&#8217;s In Your Second Brain?</a> &#8212; Tooling for the modern CTO: how to build external memory that actually works</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul>]]></content:encoded></item><item><title><![CDATA[Fear and Loathing in AWS]]></title><description><![CDATA[How Claude Helped Me Discover The Joys of Complex Infrastructure]]></description><link>https://hyperdev.matsuoka.com/p/fear-and-loathing-in-aws</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/fear-and-loathing-in-aws</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Thu, 07 May 2026 11:31:07 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!B4Cb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff875b25e-9f80-4411-8fbf-cd4a24bef764_1024x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!B4Cb!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff875b25e-9f80-4411-8fbf-cd4a24bef764_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!B4Cb!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff875b25e-9f80-4411-8fbf-cd4a24bef764_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!B4Cb!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff875b25e-9f80-4411-8fbf-cd4a24bef764_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!B4Cb!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff875b25e-9f80-4411-8fbf-cd4a24bef764_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!B4Cb!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff875b25e-9f80-4411-8fbf-cd4a24bef764_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!B4Cb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff875b25e-9f80-4411-8fbf-cd4a24bef764_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f875b25e-9f80-4411-8fbf-cd4a24bef764_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2235918,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/196537737?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff875b25e-9f80-4411-8fbf-cd4a24bef764_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!B4Cb!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff875b25e-9f80-4411-8fbf-cd4a24bef764_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!B4Cb!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff875b25e-9f80-4411-8fbf-cd4a24bef764_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!B4Cb!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff875b25e-9f80-4411-8fbf-cd4a24bef764_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!B4Cb!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff875b25e-9f80-4411-8fbf-cd4a24bef764_1024x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><blockquote><p><em><strong>TL;DR:</strong> I spent years avoiding AWS due to its overwhelming complexity, preferring the simplicity of Vercel and composable stacks. Then Claude MPM changed how I work with AWS. Now I&#8217;m running sophisticated multi-service AWS deployments with GPU instances, comprehensive monitoring, and infrastructure-as-code&#8212;all managed through AI-assisted tooling. The lesson: agentic approaches transform ops workflows just as profoundly as they do coding.</em></p></blockquote><p>Remember when AWS felt like digital quicksand? Every innocent &#8220;let me just deploy this simple app&#8221; spiraled into an afternoon lost in IAM policies, VPC configurations, and security group rules that made no sense. I&#8217;d start with what should be a five-minute deployment and emerge three hours later, bleary-eyed, with a working application and absolutely no confidence I could recreate the process.</p><p>That&#8217;s why I became a Vercel evangelist. One <code>git push</code>, automatic deployments, zero configuration. The developer experience was everything AWS wasn&#8217;t: predictable, fast, and actually enjoyable. For most projects, this composable stack approach&#8212;Vercel for frontend, managed databases, serverless functions where needed&#8212;delivered exactly the right balance of power and simplicity.</p><p>But somewhere along the way, that changed.</p><h2>The Claude Code/<a href="https://github.com/bobmatnyc/claude-mpm">MPM</a> Turning Point</h2><p>Today I&#8217;m running the kind of infrastructure I would have delegated to an ops team a year ago. GPU instances for ML workloads. Multi-AZ deployments across six subnets. Sophisticated monitoring pipelines that integrate CloudWatch metrics, Cost Explorer analysis, and automated GitHub issue creation. Terraform managing multi-account infrastructure with cross-service dependencies.</p><p>The difference? I don&#8217;t even need AWS&#8217;s Q assistant. I have something better: purpose-built AWS skills in Claude MPM that handle service deployment, infrastructure analysis, and operational workflows.</p><p>This transformation illustrates something crucial about the agentic revolution: it&#8217;s not just changing how we write code. It&#8217;s fundamentally altering how we approach operational complexity.</p><h2>The Infrastructure Reality Check</h2><p>Let me show you what I mean with real numbers. Here&#8217;s what I&#8217;m actually running across three active projects to support internal tools:</p><p><strong>CloudWatch Reporting Service</strong> (Serverless):</p><ul><li><p>12 Lambda functions handling health checks, metrics aggregation, and MCP server functionality</p></li><li><p>API Gateway HTTP API with sophisticated CORS and authentication</p></li><li><p>DynamoDB tables for state management and external directory lookups</p></li><li><p>SNS/SQS for alerting and dead letter queue handling</p></li><li><p>Direct Bedrock integration with Claude 4.5 Haiku for automated analysis</p></li><li><p>CloudWatch Events scheduling 5-minute monitoring cycles</p></li><li><p>Secrets Manager for GitHub app credentials and API keys</p></li></ul><p><strong>Code Intelligence Platform</strong> (Compute):</p><ul><li><p>Two EC2 instances: t3.xlarge for web serving, g4dn.xlarge for GPU-accelerated indexing</p></li><li><p>EBS volumes with gp3 storage and custom IOPS configuration</p></li><li><p>VPC with six subnets across availability zones</p></li><li><p>Application Load Balancer with Route53 DNS and ACM certificates</p></li><li><p>EFS for shared file storage across instances</p></li><li><p>CloudWatch Synthetics for endpoint monitoring</p></li><li><p>Lambda-based Slack notifications triggered by SNS topics</p></li></ul><p><strong>Enterprise Infrastructure</strong> (Multi-Account):</p><ul><li><p>Terragrunt-managed infrastructure across production and staging accounts</p></li><li><p>S3 backend for Terraform state with DynamoDB locking</p></li><li><p>Cross-account IAM policies and service integration</p></li><li><p>Integration with external providers (Sentry, GitHub, Kubernetes clusters)</p></li></ul><p>A year ago, this list would have been my personal infrastructure horror story. Today, it&#8217;s Tuesday.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oE0u!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde31d590-502b-4d0a-808f-44cb8153123e_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oE0u!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde31d590-502b-4d0a-808f-44cb8153123e_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!oE0u!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde31d590-502b-4d0a-808f-44cb8153123e_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!oE0u!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde31d590-502b-4d0a-808f-44cb8153123e_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!oE0u!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde31d590-502b-4d0a-808f-44cb8153123e_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oE0u!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde31d590-502b-4d0a-808f-44cb8153123e_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/de31d590-502b-4d0a-808f-44cb8153123e_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2564092,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/196537737?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde31d590-502b-4d0a-808f-44cb8153123e_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!oE0u!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde31d590-502b-4d0a-808f-44cb8153123e_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!oE0u!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde31d590-502b-4d0a-808f-44cb8153123e_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!oE0u!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde31d590-502b-4d0a-808f-44cb8153123e_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!oE0u!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde31d590-502b-4d0a-808f-44cb8153123e_1024x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>What Changed</h2><p>The transformation wasn&#8217;t gradual. It was a step function that happened when I realized AI assistants could handle the cognitive overhead that makes AWS painful.</p><p><strong>Before</strong>: AWS documentation as archaeological expedition. Digging through service guides, trying to understand the relationship between VPC route tables and security groups, wondering if I need an Internet Gateway or a NAT Gateway or both. Every deployment felt like solving a puzzle where half the pieces were hidden.</p><p><strong>After</strong>: Natural language infrastructure requests. &#8220;Set up monitoring for this Lambda function with alerting to Slack&#8221; becomes a series of guided steps where the AI handles the AWS-specific implementation details while I focus on the business requirements.</p><p>The key insight is that <a href="https://www.digitalocean.com/resources/articles/aws-complicated">AWS&#8217;s complexity isn&#8217;t inherently bad</a>&#8212;it&#8217;s just <strong>cognitively expensive</strong>. When you remove that cognitive load through AI assistance, you can appreciate what all those services actually enable.</p><p>Take my monitoring setup. Previously, I would have settled for basic uptime checks because configuring comprehensive CloudWatch metrics, Cost Explorer integration, and automated issue creation felt like a weekend project. With claude-mpm AWS skills, it became an hour of guided configuration that resulted in production-grade observability.</p><p>Or consider the GPU instance management. The g4dn.xlarge for ML indexing runs sophisticated start/stop automation, monitors for runaway processes, and automatically scales EBS volumes based on data requirements. Setting this up manually would have required deep expertise in EC2 lifecycle management, CloudWatch alarms, and Lambda automation. With AI assistance, I focused on defining the business logic while the tooling handled the AWS implementation.</p><h2>The DX Philosophy Still Matters</h2><p>None of this means AWS wins every comparison. Vercel&#8217;s developer experience remains superior for the use cases it targets. When I need to ship a marketing site or a straightforward web application, <code>git push</code> deployment still beats any infrastructure-as-code workflow.</p><p>The difference is recognizing when complexity serves a purpose versus when it&#8217;s just complexity. Vercel abstracts away infrastructure concerns because most web applications don&#8217;t need granular control over compute, storage, and networking. But when you&#8217;re building systems that do need that control&#8212;ML pipelines, high-throughput data processing, complex service topologies&#8212;AWS&#8217;s granularity becomes valuable rather than burdensome.</p><p>AI assistance changes the cost-benefit calculation. When configuring VPC networking takes 20 minutes of guided conversation instead of three hours of documentation archaeology, you can choose AWS for projects where you previously would have compromised on requirements to avoid operational overhead.</p><p>But there&#8217;s an honest accounting problem buried in that logic. Claude Code isn&#8217;t free. <a href="https://aws.amazon.com/blogs/machine-learning/claude-code-deployment-patterns-and-best-practices-with-amazon-bedrock/">API costs, subscription fees</a>&#8212;if you&#8217;re running significant conversation volume to figure out your infrastructure, you&#8217;re spending real money. At some point, you&#8217;re spending more on AI assistance than a Vercel seat would cost. The &#8220;AWS saves money at scale&#8221; argument gets complicated fast when you factor in the cognitive tooling required to get there.</p><p>So let me be direct about where each wins. Pure self-service developer experience&#8212;one engineer, a web app, ship it fast? <a href="https://dev.to/code42cate/stop-using-aws-4eg">Vercel, and it&#8217;s not particularly close</a>. The moment you need an AI co-pilot to configure your infrastructure, you&#8217;ve added a cost layer that Vercel eliminates by design. But complex multi-service deployments&#8212;ML pipelines, GPU compute alongside serverless, multi-account Terraform, monitoring infrastructure that spans six services&#8212;those don&#8217;t live in Vercel&#8217;s world. That&#8217;s where the math inverts and AWS earns its complexity premium.</p><h2>The Broader Implications</h2><p>This transformation reveals something important about how agentic approaches will reshape technology adoption. We&#8217;re not just making individual tasks more efficient&#8212;we&#8217;re changing which categories of tools become accessible to developers.</p><p>I see this pattern across the infrastructure stack. Database migrations and performance tuning become approachable when AI translates business requirements into specific configuration changes. Kubernetes stops being &#8220;too complex for small teams&#8221; when you can describe desired behavior in natural language and get helm charts and operators generated automatically. IAM policies, security groups, and compliance frameworks become manageable when AI can analyze your application requirements and generate least-privilege configurations.</p><p>These tools were always powerful. They were just too expensive to learn and maintain for many use cases. AI assistance changes that economics.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cV_v!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb243716e-dc40-4005-9910-9a34f522876b_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cV_v!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb243716e-dc40-4005-9910-9a34f522876b_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!cV_v!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb243716e-dc40-4005-9910-9a34f522876b_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!cV_v!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb243716e-dc40-4005-9910-9a34f522876b_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!cV_v!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb243716e-dc40-4005-9910-9a34f522876b_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cV_v!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb243716e-dc40-4005-9910-9a34f522876b_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b243716e-dc40-4005-9910-9a34f522876b_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1968900,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/196537737?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb243716e-dc40-4005-9910-9a34f522876b_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!cV_v!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb243716e-dc40-4005-9910-9a34f522876b_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!cV_v!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb243716e-dc40-4005-9910-9a34f522876b_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!cV_v!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb243716e-dc40-4005-9910-9a34f522876b_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!cV_v!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb243716e-dc40-4005-9910-9a34f522876b_1024x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Where This Goes Next</h2><p>We&#8217;re still early in AI-assisted infrastructure management. Today&#8217;s tooling handles deployment and configuration. Cost optimization, security posture, performance tuning&#8212;those are coming. Full system architecture from high-level requirements is probably further out than the hype suggests, but it&#8217;s not science fiction.</p><p>But the fundamental lesson remains: complexity isn&#8217;t always the enemy. Sometimes it&#8217;s just temporarily inaccessible. When AI removes the accessibility barriers, you can choose tools based on their actual capabilities rather than their learning curves.</p><p>For now, I&#8217;m running infrastructure that would have seemed impossible to manage solo a year ago. And it&#8217;s kind of fun.</p><p>AWS might still feel like quicksand sometimes. But now I have a helicopter.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and writes about AI-powered engineering at <a href="https://hyperdev.substack.com/">HyperDev</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul>]]></content:encoded></item><item><title><![CDATA[What’s In Your Second Brain?]]></title><description><![CDATA[Tooling for the modern CTO]]></description><link>https://hyperdev.matsuoka.com/p/whats-in-your-second-brain</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/whats-in-your-second-brain</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Mon, 04 May 2026 13:26:26 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!oH6q!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30309841-bfd3-4d55-a5d2-03d8d872cc9b_1024x649.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oH6q!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30309841-bfd3-4d55-a5d2-03d8d872cc9b_1024x649.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oH6q!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30309841-bfd3-4d55-a5d2-03d8d872cc9b_1024x649.png 424w, https://substackcdn.com/image/fetch/$s_!oH6q!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30309841-bfd3-4d55-a5d2-03d8d872cc9b_1024x649.png 848w, https://substackcdn.com/image/fetch/$s_!oH6q!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30309841-bfd3-4d55-a5d2-03d8d872cc9b_1024x649.png 1272w, https://substackcdn.com/image/fetch/$s_!oH6q!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30309841-bfd3-4d55-a5d2-03d8d872cc9b_1024x649.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oH6q!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30309841-bfd3-4d55-a5d2-03d8d872cc9b_1024x649.png" width="1024" height="649" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/30309841-bfd3-4d55-a5d2-03d8d872cc9b_1024x649.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:649,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1458693,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/196419582?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa25c8af5-4b58-4f5d-b775-532bd485770e_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!oH6q!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30309841-bfd3-4d55-a5d2-03d8d872cc9b_1024x649.png 424w, https://substackcdn.com/image/fetch/$s_!oH6q!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30309841-bfd3-4d55-a5d2-03d8d872cc9b_1024x649.png 848w, https://substackcdn.com/image/fetch/$s_!oH6q!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30309841-bfd3-4d55-a5d2-03d8d872cc9b_1024x649.png 1272w, https://substackcdn.com/image/fetch/$s_!oH6q!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30309841-bfd3-4d55-a5d2-03d8d872cc9b_1024x649.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The modern CTO toolkit isn&#8217;t just apps and coding tools. The real differentiator is a custom knowledge layer &#8212; databases, search indices, memory graphs, behavioral instructions that compound over time. No product gives you this. You build it.</p><p><a href="https://karpathy.ai/">Andrej Karpathy</a> gestured at something similar last month when he posted a <a href="https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f">GitHub Gist</a> he called &#8220;LLM Wiki.&#8221; His framing: stop using LLMs <em>just</em> to write code, use them to build and maintain a personal knowledge base instead. <em>&#8220;Obsidian is the IDE, the LLM is the programmer, the wiki is the codebase.&#8221;</em> Three folders, structured Markdown, a large context window, a few Python scripts. No RAG, no vector database. He concluded with: <em>&#8220;I think there is room here for an incredible new product.&#8221;</em></p><p>He&#8217;s right that there&#8217;s room. But the product comment is where I&#8217;d push back, and I&#8217;ll get to that. What Karpathy is describing isn&#8217;t a note-taking system. It&#8217;s a personal operational knowledge layer. For CTOs specifically, that layer needs to be broader than a personal wiki &#8212; it needs live organizational data, agent-connected search, and context that persists across months of decisions. No app hands you that.</p><h2>TL;DR</h2><ul><li><p>Karpathy&#8217;s LLM Wiki shows the direction: LLMs as knowledge compilers, not just code generators</p></li><li><p>A modern CTO&#8217;s &#8220;second brain&#8221; is more than PKM &#8212; it&#8217;s live databases, custom agents, and contextual search across organizational data</p></li><li><p>When I joined Duetto as CTO, my custom toolkit let me synthesize a 150-person R&amp;D org in weeks instead of months</p></li><li><p>The power isn&#8217;t Obsidian. It&#8217;s what you connect to it &#8212; MCP servers, search indices, knowledge graphs</p></li><li><p>Productizing this is theoretically possible and practically very hard, because the schema is the moat</p></li></ul><h2>The toolkit article got it half right</h2><p>In <a href="https://hyperdev.matsuoka.com/p/whats-in-my-claude-code-toolkit">What&#8217;s In My Toolkit: Claude Code and Family</a>, I wrote about vanilla Claude Code&#8217;s core limitations: context evaporates, code search is keyword-based, memory doesn&#8217;t persist, execution is single-threaded. The tools I built &#8212; <a href="https://github.com/bobmatnyc/claude-mpm">Claude MPM</a>, <a href="https://github.com/bobmatnyc/mcp-vector-search">mcp-vector-search</a>, <a href="https://github.com/bobmatnyc/kuzu-memory">kuzu-memory</a> &#8212; address each of those gaps.</p><p>But that article was about coding workflows. The real story is broader.</p><p>The same architecture that makes a coding session more effective &#8212; persistent memory, semantic search, specialized agents pulling from structured data &#8212; turns out to be extraordinarily useful for executive work. Understanding an organization, tracking decisions over time, querying data across systems, maintaining context across months of meetings and analysis. The toolkit I built for software development became the toolkit I used to onboard as a CTO.</p><p><a href="https://hyperdev.matsuoka.com/p/i-built-a-coding-tool-then-i-used">That onboarding story</a> is documented in detail elsewhere. Short version: I pointed a multi-agent framework at GitHub, JIRA, Slack, Confluence, and budget spreadsheets, and synthesized a 150-person R&amp;D organization in the weeks before my start date. The difference between doing that with a chat interface versus a CLI-based orchestration layer with parallel agents and persistent memory wasn&#8217;t 2x or 5x. It was closer to 10x.</p><p>But the onboarding was just the starting gun. The second brain I assembled keeps compounding.</p><h2>What&#8217;s actually in my second brain</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ICJf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd86cffcd-de7a-4db2-8ab3-ec6f86df65db_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ICJf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd86cffcd-de7a-4db2-8ab3-ec6f86df65db_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!ICJf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd86cffcd-de7a-4db2-8ab3-ec6f86df65db_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!ICJf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd86cffcd-de7a-4db2-8ab3-ec6f86df65db_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!ICJf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd86cffcd-de7a-4db2-8ab3-ec6f86df65db_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ICJf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd86cffcd-de7a-4db2-8ab3-ec6f86df65db_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d86cffcd-de7a-4db2-8ab3-ec6f86df65db_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1082407,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/196419582?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd86cffcd-de7a-4db2-8ab3-ec6f86df65db_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ICJf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd86cffcd-de7a-4db2-8ab3-ec6f86df65db_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!ICJf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd86cffcd-de7a-4db2-8ab3-ec6f86df65db_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!ICJf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd86cffcd-de7a-4db2-8ab3-ec6f86df65db_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!ICJf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd86cffcd-de7a-4db2-8ab3-ec6f86df65db_1024x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Let me be specific. Because when people (now) hear &#8220;second brain&#8221; they usually think Obsidian vaults with color-coded tags and pretty Markdown files. That&#8217;s part of it. It&#8217;s the surface layer.</p><p>The actual power comes from what&#8217;s underneath.</p><h3>The memory layer</h3><p><a href="https://github.com/bobmatnyc/kuzu-memory">kuzu-memory</a> is a KuzuDB-backed knowledge graph that persists across every AI session. It stores learnings from conversations, code commits, decisions, patterns. When I start a new Claude Code session on a problem I&#8217;ve touched before, the context isn&#8217;t blank &#8212; it&#8217;s enriched with what was learned the last time.</p><p>This is the thing people underestimate. A project-specific memory that accumulates over months of work develops a kind of organizational intelligence you can&#8217;t replicate in a single conversation. It knows why a particular architectural decision was made. It knows that a vendor was evaluated and found lacking. It knows the terminology your team uses internally that differs from industry standard.</p><p>KuzuDB isn&#8217;t a product choice for its own sake &#8212; it&#8217;s graph-native, which means it handles relationships well. The connections between people, systems, decisions, and code are as important as the facts themselves.</p><h3>The search layer</h3><p><a href="https://github.com/bobmatnyc/mcp-vector-search">mcp-vector-search</a> provides semantic search across all project files. Not keyword search &#8212; semantic search with AST parsing. When I ask &#8220;where is the analysis I did on contractor productivity last quarter,&#8221; it finds it even if the document never uses those exact words.</p><p>At Duetto, this covers everything in my CTO project: architecture records, meeting notes pulled from Granola, emails I&#8217;ve synthesized, analysis documents, planning artifacts. Months of accumulated context, all searchable in seconds. The underlying code intelligence for the engineering organization runs as a separate service &#8212; mcp-vector-search is for my working knowledge, not the codebase itself.</p><h3>The databases</h3><p>My CTO project has three:</p><ul><li><p><strong>cto.db</strong> &#8212; SQLite. Work classification, people analysis, contributor data, commit history. The operational database for running analyses and reports.</p></li><li><p><strong>analytics.duckdb</strong> &#8212; DuckDB. OLAP queries and analytics. When I need to slice engineering output data in different ways or run something that would be painful in SQLite, it goes here.</p></li><li><p><strong>duetto_knowledge.db</strong> &#8212; The RAG-queryable knowledge base backing a Flask web app for interactive exploration.</p></li></ul><p>These aren&#8217;t a product I bought. They&#8217;re a schema I designed, built incrementally, and own completely. The schema reflects how I think about the organization, which is precisely why it&#8217;s useful.</p><h3>The connectors</h3><p><a href="https://github.com/bobmatnyc/gworkspace-mcp">gworkspace-mcp</a> handles Drive, Docs, Sheets, Gmail, Calendar, and more. I wrote my own rather than using the off-the-shelf options &#8212; Google&#8217;s first-party integration and Anthropic&#8217;s default both have significant tool coverage gaps. Mine exposes substantially more of the Workspace API surface and integrates transparently with Claude MPM, so agents can use Google Workspace tools without any special configuration at the call site.</p><p>Beyond Workspace: Notion API for product specs and planning documents. Extraction scripts for JIRA, Confluence, Slack, Datadog, and AWS. Each system outputs to as raw data, which feeds analysis pipelines that generate reports stored in a project directory.</p><p>For company-wide memory, two more tools: duetto-memory and duetto-directory. These handle shared organizational context &#8212; information that needs to flow between tools and across team members rather than staying in a single session. Memory persists within our VPC, encrypted to individual users&#8217; OAuth keys. Not even our own IT has access to it. Context shared from Claude Code shows up in Claude.ai, and vice versa, without any manual sync.</p><p>The entire flow is queryable. From a single Claude session, I can ask about budget trends, team velocity, specific architectural decisions, or what a particular engineer has been working on for the last three months. Because it&#8217;s all in the same context-addressable system.</p><h3>Obsidian as the front door</h3><p>Yes, I use Obsidian. But it&#8217;s a front door, not the building. The vault holds my personal notes, research captures, and synthesized analysis. The Obsidian Web Clipper feeds raw material into the knowledge pipeline. Templates enforce consistent structure.</p><p>Karpathy&#8217;s insight about Obsidian as IDE is right in the narrow sense: it&#8217;s the interface you use to read and organize. But the interesting work happens outside it &#8212; in the databases, the agents, the search indices, the custom scripts.</p><h2>CLAUDE.md files everywhere</h2><p>The context layer isn&#8217;t just data. It&#8217;s also behavioral instructions.</p><p>Every major directory in my project has a CLAUDE.md. The root CTO project one is 400 lines of conventions, routing logic, document lifecycle rules, and architectural decisions. Every subdirectory has a more focused version. Every specialized agent has its own constraints.</p><p>These files are my second brain&#8217;s schema, expressed as instructions rather than data. A single routing rule &#8212; &#8220;if the prompt mentions meeting notes, save to <code>projects/meetings/2026-W##/</code>&#8220; &#8212; sounds trivial. But it means twelve months of meeting notes accumulate in consistent, queryable locations rather than wherever an agent happened to save them. Multiply that by forty routing rules across fifteen subdirectories, and the entire corpus becomes navigable. The CLAUDE.md files are what make the databases useful. Without them, the data is just data.</p><p>Karpathy put it well: &#8220;You share the schema, not the code.&#8221; The schema is the valuable part. The schema is what compounds.</p><p>My schema took months to build. It will keep getting better. No product ships with the right schema for my organization, because no product knows what I know about how Duetto&#8217;s R&amp;D works.</p><h2>The productization question</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-gqt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87551032-97c3-4a89-8697-729269302cfa_1024x708.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-gqt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87551032-97c3-4a89-8697-729269302cfa_1024x708.png 424w, https://substackcdn.com/image/fetch/$s_!-gqt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87551032-97c3-4a89-8697-729269302cfa_1024x708.png 848w, https://substackcdn.com/image/fetch/$s_!-gqt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87551032-97c3-4a89-8697-729269302cfa_1024x708.png 1272w, https://substackcdn.com/image/fetch/$s_!-gqt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87551032-97c3-4a89-8697-729269302cfa_1024x708.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-gqt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87551032-97c3-4a89-8697-729269302cfa_1024x708.png" width="1024" height="708" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/87551032-97c3-4a89-8697-729269302cfa_1024x708.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:708,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1454118,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/196419582?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85bbeb78-8523-459e-9cb0-e65b1d5bce18_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!-gqt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87551032-97c3-4a89-8697-729269302cfa_1024x708.png 424w, https://substackcdn.com/image/fetch/$s_!-gqt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87551032-97c3-4a89-8697-729269302cfa_1024x708.png 848w, https://substackcdn.com/image/fetch/$s_!-gqt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87551032-97c3-4a89-8697-729269302cfa_1024x708.png 1272w, https://substackcdn.com/image/fetch/$s_!-gqt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87551032-97c3-4a89-8697-729269302cfa_1024x708.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Karpathy said there&#8217;s room for an incredible product. He&#8217;s not wrong about the gap. He might be wrong about the solution.</p><p>The structural problems with productizing a second brain:</p><p><strong>Context compounds, products don&#8217;t.</strong> My system gets smarter with every commit, meeting, and conversation. A SaaS product serves thousands of customers and maintains no one&#8217;s specific context. The more I use my system, the wider the gap between it and any off-the-shelf alternative.</p><p><strong>The schema is the moat.</strong> My knowledge architecture reflects how I think about engineering organizations. Someone else&#8217;s knowledge architecture would be different. Products that force their schema on you &#8212; and every product does &#8212; are imposing someone else&#8217;s way of thinking on your problem. That friction is small at first and grows over time.</p><p><strong>Privacy is structural, not incidental.</strong> My databases contain org structures, salary data, performance patterns, vendor negotiations. Routing that through third-party infrastructure creates risk that&#8217;s practically impossible to contain. When I built duetto-memory for enterprise use, the entire stack stays within our VPC, with memories encrypted to individual users&#8217; OAuth keys. Not even IT can read them. That level of isolation is nearly impossible to provide as a multi-tenant SaaS.</p><p>Some layers could be productized &#8212; the infrastructure, not the intelligence. A well-designed memory MCP with sensible defaults. Semantic search that works without configuration. Privacy-preserving graph storage you don&#8217;t have to host yourself. The plumbing.</p><p>The schema, the decisions, and the accumulated context can&#8217;t be productized. Those are yours. That&#8217;s the point &#8212; and it&#8217;s also why the product gap Karpathy sees will remain open even after someone tries to fill it.</p><h2>Can anyone do this?</h2><p>There&#8217;s an access problem here, and I&#8217;d be dishonest not to acknowledge it.</p><p>Building what I&#8217;ve described requires knowing Python well enough to write extraction scripts, understanding enough about graph databases to design a schema, and being comfortable with CLI-based tooling and MCP server configuration. Not every CTO has that background. Not every technical leader wants to spend weekends building personal infrastructure.</p><p>The irony is that the people who most need better organizational intelligence &#8212; executives without deep engineering backgrounds &#8212; are least equipped to build these systems. And the people who are most capable of building them are often less interested in the executive problems the systems could solve.</p><p>Tiago Forte, who wrote <em><a href="https://www.buildingasecondbrain.com/">Building a Second Brain</a></em>, has been making this point for years. His PARA method and CODE framework are accessibility layers &#8212; ways to make the underlying ideas approachable without requiring you to build a graph database. The methodology is sound. But it was designed for knowledge workers, not for CTOs running engineering organizations who need live data pipelines, not filing systems. Well-designed for whom?</p><p>Karpathy&#8217;s LLM Wiki is explicitly a system for someone comfortable writing Python and working with file systems. His Gist has code in it. That&#8217;s a feature for his audience and a barrier for everyone else.</p><h2>What I&#8217;d watch for</h2><p>A few trends that will determine whether this remains a DIY space or gets productized:</p><p><strong>MCP as infrastructure.</strong> The <a href="https://hyperdev.matsuoka.com/p/the-mcp-cat-is-out-of-the-bag">Model Context Protocol</a> creates a standard interface for exactly this kind of knowledge infrastructure. Memory servers, search servers, database connectors &#8212; they all expose the same interface to any compatible AI client. The ecosystem is growing fast. As more MCP servers mature, the configuration burden drops.</p><p><strong>Searchable context beats raw window size.</strong> Karpathy argues for plain Markdown because ~400K words fit in a modern context window. That&#8217;s true, and the window is getting larger. But the more important shift is that structured, searchable context doesn&#8217;t have a ceiling. A well-organized knowledge base that spans years of meetings, decisions, and analysis delivers more than any single context window can hold &#8212; and the value scales with the quality of the organization, not the size of the model.</p><p><strong>Local model quality.</strong> Karpathy runs Anthropic agents via Claude Code. But local model quality is improving fast. A system that uses the cloud API for synthesis and queries but runs a local model for routine indexing tasks would be significantly cheaper and more private. Not ready yet. Getting closer.</p><p>The product Karpathy thinks exists &#8212; if it gets built &#8212; probably looks like a well-designed local MCP server with clean configuration, sensible defaults, and a plugin ecosystem for connectors. Not a SaaS. Not a cloud database. Something you install and own.</p><p>The people who need it most will have already built their own before any product ships. And in the process of building it, they&#8217;ll have accumulated the one thing no product can give them: months of their own operational context, organized the way their own mind works.</p><p>That&#8217;s not a consolation prize. That&#8217;s the whole point.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and writes about AI-powered engineering at <a href="https://hyperdev.substack.com/">HyperDev</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/p/whats-in-my-claude-code-toolkit">What&#8217;s In My Toolkit: Claude Code and Family</a> &#8212; The coding layer of the stack</p></li><li><p><a href="https://hyperdev.matsuoka.com/p/i-built-a-coding-tool-then-i-used">I Built a Coding Tool. Then I Used It to Onboard as CTO</a> &#8212; Applying agent orchestration to organizational analysis</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul>]]></content:encoded></item><item><title><![CDATA[You weren't imagining things...Claude Code was dumber this month]]></title><description><![CDATA[Unintended consequences of optimization]]></description><link>https://hyperdev.matsuoka.com/p/you-werent-imagining-thingsclaude</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/you-werent-imagining-thingsclaude</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Fri, 24 Apr 2026 13:53:25 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JH1l!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F960004b8-63a5-4f9b-98b1-ee35085cd059_1024x779.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JH1l!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F960004b8-63a5-4f9b-98b1-ee35085cd059_1024x779.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JH1l!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F960004b8-63a5-4f9b-98b1-ee35085cd059_1024x779.png 424w, https://substackcdn.com/image/fetch/$s_!JH1l!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F960004b8-63a5-4f9b-98b1-ee35085cd059_1024x779.png 848w, https://substackcdn.com/image/fetch/$s_!JH1l!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F960004b8-63a5-4f9b-98b1-ee35085cd059_1024x779.png 1272w, https://substackcdn.com/image/fetch/$s_!JH1l!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F960004b8-63a5-4f9b-98b1-ee35085cd059_1024x779.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JH1l!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F960004b8-63a5-4f9b-98b1-ee35085cd059_1024x779.png" width="1024" height="779" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/960004b8-63a5-4f9b-98b1-ee35085cd059_1024x779.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:779,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1738565,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/195350759?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5bc7c126-0147-4eb2-8a84-d2fa3808f4b4_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!JH1l!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F960004b8-63a5-4f9b-98b1-ee35085cd059_1024x779.png 424w, https://substackcdn.com/image/fetch/$s_!JH1l!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F960004b8-63a5-4f9b-98b1-ee35085cd059_1024x779.png 848w, https://substackcdn.com/image/fetch/$s_!JH1l!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F960004b8-63a5-4f9b-98b1-ee35085cd059_1024x779.png 1272w, https://substackcdn.com/image/fetch/$s_!JH1l!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F960004b8-63a5-4f9b-98b1-ee35085cd059_1024x779.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>So if you&#8217;ve been using Claude Code and noticed it felt... off... you weren&#8217;t imagining it.</p><p><a href="https://www.anthropic.com/engineering/april-23-postmortem">Anthropic published a full breakdown</a> yesterday and it&#8217;s actually three separate bugs that compounded into what looked like one big degradation. The developer community was right to be concerned, and the evidence they collected was instrumental in getting this fixed.</p><p>Here&#8217;s what actually happened:</p><h2>1. They silently downgraded reasoning effort (March 4)</h2><p>They switched Claude Code&#8217;s default from high to medium reasoning to reduce latency. Users noticed immediately. They reverted it on April 7.</p><p>Classic &#8220;we know better than users&#8221; move that backfired. From their postmortem:</p><blockquote><p>&#8220;This was the wrong tradeoff. We reverted this change on April 7 after users told us they&#8217;d prefer to default to higher intelligence and opt into lower effort for simple tasks.&#8221;</p></blockquote><p>The UI was appearing frozen in high reasoning mode, so they made an executive decision to sacrifice quality for speed. Developers immediately felt the difference and pushed back hard.</p><h2>2. A caching bug made Claude forget its own reasoning (March 26)</h2><p>This one was particularly insidious. They tried to optimize memory for idle sessions&#8212;clear old thinking after an hour of inactivity to speed up resumption. Sounds reasonable, right?</p><p>A bug caused it to wipe Claude&#8217;s reasoning history on EVERY turn for the rest of a session, not just once. So Claude kept executing tasks while literally forgetting why it made the decisions it did.</p><p>The cascading effects were brutal:</p><ul><li><p>Every request became a cache miss</p></li><li><p>Usage limits drained faster than expected</p></li><li><p>Claude appeared &#8220;forgetful and repetitive&#8221;</p></li><li><p>Sessions felt like they were constantly resetting</p></li></ul><h2>3. A system prompt change capped responses at 25 words between tool calls (April 16)</h2><p>They added this seemingly innocent instruction: &#8220;keep text between tool calls to 25 words. Keep final responses to 100 words.&#8221;</p><p>It caused a measurable 3% drop in coding quality across both Opus 4.6 and 4.7. They caught this through ablation testing&#8212;removing the instruction and measuring the performance difference.</p><p>Reverted April 20.</p><h2>The community evidence was damning</h2><p>While Anthropic was investigating internally, the developer community was building their own case. <a href="https://github.com/anthropics/claude-code/issues/42796">Stella Laurenzo from AMD&#8217;s AI group</a> published the most comprehensive analysis&#8212;6,852 Claude Code sessions and over 234,000 tool calls.</p><p>Her findings:</p><ul><li><p>Median visible thinking length collapsed 73% (2,200 &#8594; 600 characters)</p></li><li><p>API calls per task spiked up to 80x from February to March</p></li><li><p>Claude was choosing &#8220;simplest fix&#8221; over correct solutions</p></li></ul><p><a href="https://venturebeat.com/technology/mystery-solved-anthropic-reveals-changes-to-claudes-harnesses-and-operating-instructions-likely-caused-degradation">BridgeMind&#8217;s testing</a> showed Opus 4.6 accuracy dropping from 83.3% to 68.3%.</p><p>The data was undeniable.</p><h2>The perfect storm effect</h2><p>Here&#8217;s what made this particularly hard to pin down: all three bugs affected different traffic slices on different schedules. The combined effect looked like random, inconsistent degradation.</p><p>Hard to reproduce internally. Hard for users to isolate the exact cause. It just felt... wrong.</p><p>Some sessions hit the reasoning downgrade. Others hit the caching bug. The unlucky ones hit multiple issues simultaneously. No wonder it seemed like Claude was having random bad days.</p><h2>What this reveals about AI product development</h2><p>This postmortem is actually refreshing in its transparency. Most AI companies would have quietly fixed the issues and moved on. Anthropic owned the mistakes publicly.</p><p>But it also highlights a fundamental tension in AI product development: users often prefer maximum capability over convenience optimizations. The reasoning effort downgrade was done for user experience (reduce perceived latency), but developers would rather wait for better output.</p><p>The lesson: don&#8217;t optimize away what users value most without asking them first.</p><h2>All fixed now (v2.1.116)</h2><p>As of April 20, all three issues are resolved:</p><ul><li><p>Default reasoning is now &#8220;xhigh&#8221; for Opus 4.7, &#8220;high&#8221; for others</p></li><li><p>Caching bug squashed</p></li><li><p>Verbosity limits removed</p></li><li><p>Usage limits reset for all subscribers</p></li></ul><p>Anthropic is also committing to more transparency going forward with a <a href="https://x.com/claudedevs">dedicated </a><a href="https://github.com/ClaudeDevs">@ClaudeDevs</a> account for deeper technical communication with developers.</p><p>The community was right to raise hell about this. And Anthropic&#8217;s response&#8212;full transparency with concrete fixes&#8212;sets a good precedent for how AI companies should handle quality regressions.</p><p>Your coding assistant is back to full strength.</p><h2>Independent Validation</h2><p>The technical analysis backing this story comes from multiple independent sources. <a href="https://github.com/anthropics/claude-code/issues/42796">Stella Laurenzo&#8217;s comprehensive audit</a> of 6,852 sessions provided the quantitative foundation. <a href="https://venturebeat.com/technology/mystery-solved-anthropic-reveals-changes-to-claudes-harnesses-and-operating-instructions-likely-caused-degradation">BridgeMind&#8217;s testing</a> offered controlled benchmark data. These weren&#8217;t isolated complaints&#8212;they were systematic investigations with reproducible findings.</p><p>When a company publishes a detailed postmortem acknowledging specific engineering decisions that degraded their product, and that postmortem aligns with community-gathered evidence, we&#8217;re seeing transparency in action. The developer community did the work to document the problems. Anthropic owned the solutions.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and writes about AI-powered engineering at <a href="https://hyperdev.substack.com/">HyperDev</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul>]]></content:encoded></item></channel></rss>