<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Hyperdev: Articles]]></title><description><![CDATA[Feature articles on all things agentic coding—
From prompt-driven development to autonomous dev workflows, this section dives into the tools, techniques, and thinking behind AI-first programming. Real-world use cases, architecture breakdowns, code experiments, and practical insight into building with agents instead of just writing code line by line.]]></description><link>https://hyperdev.matsuoka.com/s/hyperdev</link><image><url>https://substackcdn.com/image/fetch/$s_!j9a7!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab665959-5546-4469-9e93-9e1518976e2b_1024x1024.png</url><title>Hyperdev: Articles</title><link>https://hyperdev.matsuoka.com/s/hyperdev</link></image><generator>Substack</generator><lastBuildDate>Sat, 05 Sep 2026 17:26:58 GMT</lastBuildDate><atom:link href="https://hyperdev.matsuoka.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Robert Matsuoka]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[hyperdev@matsuoka.com]]></webMaster><itunes:owner><itunes:email><![CDATA[hyperdev@matsuoka.com]]></itunes:email><itunes:name><![CDATA[Robert Matsuoka]]></itunes:name></itunes:owner><itunes:author><![CDATA[Robert Matsuoka]]></itunes:author><googleplay:owner><![CDATA[hyperdev@matsuoka.com]]></googleplay:owner><googleplay:email><![CDATA[hyperdev@matsuoka.com]]></googleplay:email><googleplay:author><![CDATA[Robert Matsuoka]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[We Need to Teach Delegation as an Engineering Skill]]></title><description><![CDATA[idea -> communication -> outcome]]></description><link>https://hyperdev.matsuoka.com/p/we-need-to-teach-delegation-as-an</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/we-need-to-teach-delegation-as-an</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Fri, 04 Sep 2026 11:31:26 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!JOIZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50bb55de-6fed-4fde-87ba-f5f1917ac61f_1024x426.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JOIZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50bb55de-6fed-4fde-87ba-f5f1917ac61f_1024x426.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JOIZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50bb55de-6fed-4fde-87ba-f5f1917ac61f_1024x426.png 424w, https://substackcdn.com/image/fetch/$s_!JOIZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50bb55de-6fed-4fde-87ba-f5f1917ac61f_1024x426.png 848w, https://substackcdn.com/image/fetch/$s_!JOIZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50bb55de-6fed-4fde-87ba-f5f1917ac61f_1024x426.png 1272w, https://substackcdn.com/image/fetch/$s_!JOIZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50bb55de-6fed-4fde-87ba-f5f1917ac61f_1024x426.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JOIZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50bb55de-6fed-4fde-87ba-f5f1917ac61f_1024x426.png" width="1024" height="426" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/50bb55de-6fed-4fde-87ba-f5f1917ac61f_1024x426.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:426,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1066990,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/214072948?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F97d87c64-64eb-4974-9a4b-be5d2e7881b0_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!JOIZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50bb55de-6fed-4fde-87ba-f5f1917ac61f_1024x426.png 424w, https://substackcdn.com/image/fetch/$s_!JOIZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50bb55de-6fed-4fde-87ba-f5f1917ac61f_1024x426.png 848w, https://substackcdn.com/image/fetch/$s_!JOIZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50bb55de-6fed-4fde-87ba-f5f1917ac61f_1024x426.png 1272w, https://substackcdn.com/image/fetch/$s_!JOIZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F50bb55de-6fed-4fde-87ba-f5f1917ac61f_1024x426.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Engineering as orchestration</figcaption></figure></div><p>Some of the strongest engineers I know won&#8217;t work with coding agents, at least not at the level I&#8217;d expect. A bit of Cursor yes, some Claude -- but it&#8217;s an add-on, not fundamental harness loop engineering. Just as likely, they are uncomfortable with it because they have never had to practice the skill it takes.</p><p>They see the problem quickly. They know the part of the codebase that needs to change, the shortcut that will fail in production, and the test that will catch it. Then they give an agent two sentences, watch it head in the wrong direction, and take the keyboard back. Ten minutes later the change is done.</p><p>Their conclusion is reasonable: I can do this better myself.</p><p>For that one task, they may be right. But they have skipped a different piece of engineering. They moved from idea to code. Agentic development inserts a middle layer:</p><p><strong>idea -&gt; communication -&gt; outcome</strong></p><p>That middle layer is <strong>delegation</strong>. It means deciding which context another actor needs, which decisions belong to them, what cannot change, how the result will be evaluated, and when the work needs to come back. A manager has to do this with people. Coding agents have made it an individual-contributor skill too.</p><p>We teach engineers how to solve problems. We teach design, decomposition, testing, and code review. We spend far less time teaching them how to make their understanding usable by somebody else.</p><p>I think that explains a meaningful share of the gap between engineers who get substantial output from agents and engineers who find them irritating. The evidence does not prove delegation skill is the sole cause. It does show that expertise can make instruction more abstract, that experienced developers do not get an automatic productivity gain from AI, and that successful agent use depends heavily on problem framing and evaluation.</p><h2>TL;DR</h2><ul><li><p>Solving a problem yourself and specifying it for another capable actor are different forms of work. Agentic development makes engineers practice the second one.</p></li><li><p>Experts can be poor instructors because their knowledge has become compressed. A Stanford study found experts gave more abstract, less concrete instructions than beginners.</p></li><li><p>In roughly 400,000 Claude Code sessions, people made about 70% of planning decisions while Claude made about 80% of execution decisions. The human job did not disappear. It moved toward intent and evaluation.</p></li><li><p>Useful delegation defines the outcome, relevant context, decision authority, evidence of success, and a return path. It leaves implementation room inside those boundaries.</p></li><li><p>Engineering organizations should teach and assess delegation alongside system design, testing, and code review.</p></li></ul><h2>Strong Engineers Short-Circuit the Middle</h2><p>Expertise is compression.</p><p>A senior engineer does not consciously replay every step required to diagnose a stale cache, unwind a dependency cycle, or recognize that an innocent schema change will break an old consumer. Years of experience have turned those steps into pattern recognition. That is part of what makes the engineer fast.</p><p>It can also make the engineer a poor source of instructions.</p><p>In a 2001 <a href="https://pubmed.ncbi.nlm.nih.gov/11768064/">Journal of Applied Psychology study</a>, Pamela Hinds, Michael Patterson, and Jeffrey Pfeffer asked experts and beginners to give novices instructions for an electronic circuit-wiring task. The experts used more abstract and advanced statements, with fewer concrete statements. Novices taught by beginners performed better on the same task and reported fewer problems with the instructions.</p><p>The expert instructions had a different advantage. Novices taught by experts transferred what they learned more successfully to another task in the same domain. Abstraction was useful, but it was pitched at the wrong level for immediate execution.</p><p>That is a close description of many failed agent sessions. The engineer supplies the concept because the missing steps no longer feel like steps. The agent supplies its own version of those missing steps. Then the engineer calls the output stupid.</p><p>Sometimes it is. Coding agents make bad decisions, miss local conventions, and produce code that passes a narrow test while violating the system around it. But &#8220;the agent should know that&#8221; often means &#8220;I knew that and failed to communicate it.&#8221;</p><p>The distinction gets lost because explaining the work feels like overhead to somebody who can already do it. Delegation is slower than execution until the delegated system starts producing more than one person&#8217;s hands can.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!u7AK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb12d8c33-92e4-4e78-98ef-442c50ce5a26_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!u7AK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb12d8c33-92e4-4e78-98ef-442c50ce5a26_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!u7AK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb12d8c33-92e4-4e78-98ef-442c50ce5a26_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!u7AK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb12d8c33-92e4-4e78-98ef-442c50ce5a26_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!u7AK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb12d8c33-92e4-4e78-98ef-442c50ce5a26_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!u7AK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb12d8c33-92e4-4e78-98ef-442c50ce5a26_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b12d8c33-92e4-4e78-98ef-442c50ce5a26_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1218071,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/214072948?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb12d8c33-92e4-4e78-98ef-442c50ce5a26_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!u7AK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb12d8c33-92e4-4e78-98ef-442c50ce5a26_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!u7AK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb12d8c33-92e4-4e78-98ef-442c50ce5a26_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!u7AK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb12d8c33-92e4-4e78-98ef-442c50ce5a26_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!u7AK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb12d8c33-92e4-4e78-98ef-442c50ce5a26_1024x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Seniority Does Not Produce an Automatic AI Gain</h2><p>A strong counterweight to AI productivity anecdotes remains METR&#8217;s randomized study of early-2025 coding tools. <a href="https://metr.org/Early_2025_AI_Experienced_OS_Devs_Study-paper.pdf">Sixteen experienced open-source developers completed 246 tasks</a> in repositories where they averaged five years of prior experience. When AI tools were allowed, completion time increased by 19%.</p><p>The developers expected a 24% speedup before the work. Afterward, they still estimated that AI had made them 20% faster.</p><p>This is a historical result, not a verdict on current models. The study mostly used Cursor with Claude 3.5 and 3.7 Sonnet. METR later <a href="https://metr.org/blog/2026-02-24-uplift-update/">changed its follow-up design</a> after more developers declined to participate because they did not want to work without AI, which introduced selection bias. The researchers also warn against treating their original participants as representative of software development as a whole.</p><p>Still, the result closes off one comforting assumption: being good at a codebase does not make somebody good at directing an agent through it. The participants knew their systems. They had substantial AI experience. They could choose whether to use the tools. And in that setting, the tools cost them time while feeling faster.</p><p>METR did not test my delegation thesis. It measured the gap I am trying to explain.</p><h2>Agentic Development Splits Planning From Execution</h2><p>Anthropic&#8217;s June 2026 study of <a href="https://www.anthropic.com/research/claude-code-expertise">roughly 400,000 Claude Code sessions</a> offers a clearer picture of the work split. Its classifiers attributed about 70% of planning decisions to the person and about 80% of execution decisions to Claude.</p><p>Planning included what to do, which approach to take, and what counts as done. Execution included which files to change, what code to write, and which commands to run.</p><p>This is delegation in software form. The person retains the problem and the acceptance decision. The agent gets a bounded decision space inside it.</p><p>Anthropic also found that novice-rated prompts led to around five agent actions and 600 words of output, while expert-rated prompts led to about 12 actions and 3,200 words. Sessions rated intermediate or above reached Anthropic&#8217;s strict verified-success measure 28% to 33% of the time, compared with 15% for novice-rated sessions. When sessions ran into trouble, expert-rated users recovered more often.</p><p>There is a circularity in those numbers. Anthropic&#8217;s expertise classifier partly looked at how precisely a person framed directions, what they asked Claude to verify, and whether they corrected the model. Those are delegation behaviors. The study shows that the behaviors travel with longer agent runs and higher measured success, but it cannot tell us how much each behavior caused the result.</p><p>The occupational result is still suggestive. In code-producing sessions, management occupations finished slightly above software occupations on the strict verified-success measure, though Anthropic says managers may be more likely to state explicitly that they got what they wanted. A skill learned by directing people may transfer to directing software agents.</p><p>Small qualitative studies point the same way. In a 2026 mixed-methods study of senior and junior engineers, Dana Feng, Bhada Yun, and April Yi Wang found that <a href="https://arxiv.org/abs/2602.00496">senior engineers maintained control through detailed delegation</a>. They scoped changes, supplied context, delegated smaller units, asked for minimal reviewable diffs, and refined the work through feedback. The junior engineers moved between over-reliance and cautious avoidance.</p><p>Ten juniors and ten seniors are not a population. But the described behavior will be familiar to anybody who has watched one person steer an agent while another person alternates between accepting everything and taking the work back.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nXzJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80bb502b-db63-4fd1-bc1d-7cdf0ab21947_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nXzJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80bb502b-db63-4fd1-bc1d-7cdf0ab21947_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!nXzJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80bb502b-db63-4fd1-bc1d-7cdf0ab21947_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!nXzJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80bb502b-db63-4fd1-bc1d-7cdf0ab21947_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!nXzJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80bb502b-db63-4fd1-bc1d-7cdf0ab21947_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nXzJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80bb502b-db63-4fd1-bc1d-7cdf0ab21947_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/80bb502b-db63-4fd1-bc1d-7cdf0ab21947_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1081125,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/214072948?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80bb502b-db63-4fd1-bc1d-7cdf0ab21947_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!nXzJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80bb502b-db63-4fd1-bc1d-7cdf0ab21947_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!nXzJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80bb502b-db63-4fd1-bc1d-7cdf0ab21947_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!nXzJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80bb502b-db63-4fd1-bc1d-7cdf0ab21947_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!nXzJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F80bb502b-db63-4fd1-bc1d-7cdf0ab21947_1024x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>From Hub to System</h2><p>An engineering leader recently described his problem to me as being &#8220;too much in the hub.&#8221; Teams waited for him to make decisions. Work crossed his desk because he knew the history, the risks, and how the parts fit together. His competence had become part of the system architecture.</p><p>So he started moving ownership outward. One leader took a product area. Another took day-to-day operational work. The change was not a pile of redistributed tickets. Each person received a decision space and a return path.</p><p>Senior individual contributors create the same topology when every difficult change routes through them. Doing the work faster reinforces it. Soon the engineer is both the source of expertise and the queue in front of it.</p><p>Coding agents make that topology visible because they are always available to receive work. If nothing can move without the engineer touching the implementation, availability was not the constraint. The system had no interface for the engineer&#8217;s judgment.</p><p>Delegation builds that interface. A useful handoff answers five questions:</p><p>Decision Question Outcome What observable state should exist when the work is done? Context Which system facts and prior decisions affect the work? Authority What may the delegate decide or change without asking? Evidence Which tests, traces, screenshots, or artifacts establish success? Return path When should the delegate stop, ask, or escalate?</p><p>This applies to a person, an agent, or a group of agents. The amount of detail changes. The information contract does not.</p><h2>Delegate Outcomes, Not Motions</h2><p>&#8220;Implement this endpoint&#8221; delegates motion. The instruction names an activity and leaves the purpose, production bar, ownership boundary, and failure conditions unstated.</p><p>An outcome-level delegation describes the state the system must reach. It supplies the constraints that protect surrounding systems and says how to determine whether the result is acceptable. The delegate still chooses the implementation inside that space.</p><p>OpenAI&#8217;s <a href="https://openai.com/index/harness-engineering/">Codex harness-engineering case study</a> includes two good examples: service startup must finish in under 800 milliseconds, and no span in four named user journeys may exceed two seconds. Those are outcomes an agent can test. The repository exposes the application, logs, metrics, architecture rules, and remediation instructions required to pursue them.</p><p>The team did not achieve that by writing longer chat messages. It made the environment carry repeatable context. Plans became versioned artifacts. Architecture constraints became linters and structural tests. The agent could inspect the same evidence used to judge its work.</p><p>That team reported around one million lines across code, infrastructure, tooling, and documentation, with roughly 1,500 pull requests opened and merged over five months by three engineers directing Codex. OpenAI estimates the product took one-tenth the hand-written time. The more durable lesson is in their account of the slow start: progress was poor while the environment was underspecified.</p><p>I have seen the same pattern in my own work. <a href="https://github.com/bobmatnyc/trusty-tools">Trusty Tools</a>, the Rust-based multi-agent system I have been building since May 19, reached 3,200 commits and 3,426 closed GitHub issues by September 3. The current checkout contains roughly 1.46 million lines across code, documentation, and configuration formats.</p><p>Those are activity metrics. They do not prove the software is good. They do prove I was not typing every change myself.</p><p>The work moved when I stopped treating the agent as autocomplete and started treating the repository as an operating environment. Instructions became code-adjacent. Review rules became executable. Agents received narrow authority and had to return evidence. Failures fed changes back into the system instead of becoming one-off corrections in a chat window.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!QTlF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23349c88-6cec-400f-b03a-82d71918f52f_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!QTlF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23349c88-6cec-400f-b03a-82d71918f52f_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!QTlF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23349c88-6cec-400f-b03a-82d71918f52f_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!QTlF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23349c88-6cec-400f-b03a-82d71918f52f_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!QTlF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23349c88-6cec-400f-b03a-82d71918f52f_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!QTlF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23349c88-6cec-400f-b03a-82d71918f52f_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/23349c88-6cec-400f-b03a-82d71918f52f_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1203383,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/214072948?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23349c88-6cec-400f-b03a-82d71918f52f_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!QTlF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23349c88-6cec-400f-b03a-82d71918f52f_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!QTlF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23349c88-6cec-400f-b03a-82d71918f52f_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!QTlF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23349c88-6cec-400f-b03a-82d71918f52f_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!QTlF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F23349c88-6cec-400f-b03a-82d71918f52f_1024x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Delegation Has a Cost</h2><p>Good delegation takes time. That is a feature, not a bug. It is why strong engineers resist it.</p><p>Vincent Schmalbach&#8217;s June 2026 pilot study on <a href="https://arxiv.org/abs/2606.17099">software delegation contracts</a> compared ordinary issue-style prompts with explicit contracts covering the task, authority, returned work package, and acceptance context. Across 64 runs on ten small TypeScript tasks, every run passed the hidden acceptance tests. The contracts did not improve correctness because the tasks were already easy for the agents.</p><p>They improved reviewability. Evidence sufficiency improved in 22 of 30 paired comparisons and worsened in none. Reviewers saw more changed-file lists, known limitations, residual risks, and checklists. The contracts also used 13% more agent tokens and took 38% more wall-clock time.</p><p>That is a good trade only when reviewability is worth the cost. On a tiny task you will do once, it may not be. On repeated work, risky work, or work distributed across several actors, a reusable contract can pay for itself because the next handoff starts with an interface instead of another explanation.</p><p>Delegation also fails when it becomes abandonment. Giving somebody an outcome does not transfer away responsibility for the problem framing, the quality bar, or integration with the larger system. Autonomy needs a boundary and a feedback loop.</p><p>In one leadership model from my notes, exploratory work happened in a sandbox against broad targets. Useful results returned to an owner who could productionize them. The exploration was allowed to fail. The production handoff was not.</p><p>That distinction matters with agents. &#8220;Go figure it out&#8221; can be appropriate when the blast radius is small and the work is exploratory. It is negligence when the agent can change production state and nobody has defined how it returns evidence.</p><h2>Teach the Skill</h2><p>Delegation should appear in engineering development before somebody becomes a manager.</p><p>Code review already gives us the right setting. Ask an engineer to write the issue another engineer could implement without a private briefing. Ask what the assignee can decide, which constraints are real, which test would change the author&#8217;s mind, and when the work should come back. Then evaluate the handoff alongside the patch. Architecture exercises can do the same. A candidate should be able to design a subsystem and divide it into owned outcomes with interfaces between them.</p><p>Senior promotion rubrics should look for expertise that travels: working agreements, executable checks, useful documentation, and teams that can make decisions without routing every question back through the expert.</p><p>Engineering education is beginning to move this way. After two 2026 roundtables with roughly 30 to 40 academic and industry participants each, Sungmin Kang, Baishakhi Ray, and Abhik Roychoudhury argued that <a href="https://arxiv.org/abs/2606.21894">future engineers should learn to translate intent into machine-checkable specifications</a> and evaluate whether those specifications capture the intended result. They also put agent orchestration and verification into the proposed curriculum.</p><p>I would call the umbrella skill delegation.</p><p>The engineer who can do everything is not necessarily the engineer who can create the most throughput. Agentic development rewards the engineer who can make judgment portable without pretending judgment has disappeared.</p><p>We will still need people who can solve the hard problem themselves. They are the ones most able to recognize a wrong abstraction, an unsafe shortcut, or a test that proves less than it claims. But if all of that understanding remains inside one person&#8217;s implementation process, it can guide only one stream of work.</p><p><strong>The goal isn&#8217;t to do less engineering. It&#8217;s to move more of the engineering into a form that other capable actors can use.</strong></p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a>, a hospitality profit and revenue-management platform, and writes about AI-augmented engineering practice.  Previously, Bob has been CTO of Tripadvisor, Citymaps, and Runtime Technologies.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/harness-engineering">What Is Harness Engineering? (And Do You Need to Learn It?)</a>. Why the system around the model determines how much useful work it can do</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a>. Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a>. Strategic AI insights for CTOs and engineering leaders</p></li></ul>]]></content:encoded></item><item><title><![CDATA[The Word Problem: Why Developers Can’t Stand Talking to Opus 5]]></title><description><![CDATA[An genuinely honest, load-bearing, structural take worth reading]]></description><link>https://hyperdev.matsuoka.com/p/the-word-problem-why-developers-cant</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/the-word-problem-why-developers-cant</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Mon, 31 Aug 2026 11:31:51 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Mqay!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17e0b9a0-0583-4635-b0ec-db75e026f559_1152x864.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Mqay!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17e0b9a0-0583-4635-b0ec-db75e026f559_1152x864.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Mqay!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17e0b9a0-0583-4635-b0ec-db75e026f559_1152x864.png 424w, https://substackcdn.com/image/fetch/$s_!Mqay!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17e0b9a0-0583-4635-b0ec-db75e026f559_1152x864.png 848w, https://substackcdn.com/image/fetch/$s_!Mqay!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17e0b9a0-0583-4635-b0ec-db75e026f559_1152x864.png 1272w, https://substackcdn.com/image/fetch/$s_!Mqay!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17e0b9a0-0583-4635-b0ec-db75e026f559_1152x864.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Mqay!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17e0b9a0-0583-4635-b0ec-db75e026f559_1152x864.png" width="1152" height="864" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17e0b9a0-0583-4635-b0ec-db75e026f559_1152x864.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:864,&quot;width&quot;:1152,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:981456,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/213180440?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17e0b9a0-0583-4635-b0ec-db75e026f559_1152x864.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Mqay!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17e0b9a0-0583-4635-b0ec-db75e026f559_1152x864.png 424w, https://substackcdn.com/image/fetch/$s_!Mqay!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17e0b9a0-0583-4635-b0ec-db75e026f559_1152x864.png 848w, https://substackcdn.com/image/fetch/$s_!Mqay!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17e0b9a0-0583-4635-b0ec-db75e026f559_1152x864.png 1272w, https://substackcdn.com/image/fetch/$s_!Mqay!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17e0b9a0-0583-4635-b0ec-db75e026f559_1152x864.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>On August 10, after a phrase had shown up for the ninth time, I typed this into a Claude Code session running Opus 5:</p><blockquote><p>&#8220;the load-bearing section&#8221; this is too jargon-y. Do not use language like this (add it to tm output style)</p></blockquote><p>Forty-nine minutes later, the rule was live in trusty-tools:</p><blockquote><p>No borrowed-metaphor jargon. &#8220;Load-bearing&#8221; is the instance that prompted this rule... Wrong: &#8220;that section is load-bearing&#8221; / Right: &#8220;deleting that section breaks X&#8221;.</p></blockquote><p>I had seen it nine times in two days. Opus 5 used &#8220;load-bearing&#8221; for a constraint, then a crate, then a test suite, all in six sessions on August 9 and 10. I got tired of reading it, so I banned the word.</p><p>I remembered a much worse week. In my version, I fought with Opus 5 for days, switched to Fable 5, and everything got better. Then I checked my logs.</p><p>Across eight days, I typed 201 messages to the two models. The record is boring: no capability problems, no ALL-CAPS, no profanity, no &#8220;that&#8217;s not what I said.&#8221; I was calm and typo-heavy throughout (agents understand intent well enough that my typing has gotten worse). The switch to Fable happened in the middle of a session without a command or comment from me.</p><p>The window&#8217;s one documented correction appears secondhand, inside an auto-generated summary. It concerned version-bump authorization. My complaint about Opus 5&#8217;s prose came two days earlier. I remembered a fight, but the logs recorded something else.</p><p>Meanwhile, Opus 5 produced 1.37 million words of output on real engineering tasks, next to Fable 5&#8217;s 193,000 words. The capability argument was gone. I was left with the way it talked.</p><h2>TL;DR</h2><ul><li><p>My trusty-tools logs show zero capability problems with Opus 5 over the window that supposedly caused a switch to Fable 5. My original thesis was wrong.</p></li><li><p>Across the month-long sample, Opus 5 opens sentences with &#8220;worth noting/naming/knowing&#8221; 10.7x more often than Fable 5 (0.451 vs. 0.042 per 1,000 words).</p></li><li><p>A recurring metaphor (something &#8220;wearing&#8221; something else&#8217;s clothing) shows up 26 times across 12+ Opus 5 sessions and zero times anywhere else in the sample.</p></li><li><p>Eight named developers across Hacker News and X say they prefer GPT-5.6 Sol specifically for how it communicates. None turned up making that same claim on Reddit.</p></li><li><p>The voice predates Opus 5. A GitHub issue filed 11 days before it shipped describes the same tics, and Claude Code gives Opus 5 a system prompt a bit over a third the length of Sonnet 5&#8217;s.</p></li></ul><h2>How Opus 5 Talks</h2><p>The trusty-tools bans landed between August 10 and 12. From August 12 onward, the same rule applied to every model: no &#8220;worth noting,&#8221; no &#8220;load-bearing,&#8221; no borrowed-metaphor jargon, and none of the &#8220;honest/honestly&#8221; family. The month-long counts also include earlier output, so this is not a clean post-ban comparison. It is a record of how often the patterns appeared in the logs I have.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Aqo4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f0d33b2-a6f3-4795-95ad-f3ef95dbabfd_719x246.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Aqo4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f0d33b2-a6f3-4795-95ad-f3ef95dbabfd_719x246.png 424w, https://substackcdn.com/image/fetch/$s_!Aqo4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f0d33b2-a6f3-4795-95ad-f3ef95dbabfd_719x246.png 848w, https://substackcdn.com/image/fetch/$s_!Aqo4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f0d33b2-a6f3-4795-95ad-f3ef95dbabfd_719x246.png 1272w, https://substackcdn.com/image/fetch/$s_!Aqo4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f0d33b2-a6f3-4795-95ad-f3ef95dbabfd_719x246.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Aqo4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f0d33b2-a6f3-4795-95ad-f3ef95dbabfd_719x246.png" width="719" height="246" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2f0d33b2-a6f3-4795-95ad-f3ef95dbabfd_719x246.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:246,&quot;width&quot;:719,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:38568,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/213180440?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f0d33b2-a6f3-4795-95ad-f3ef95dbabfd_719x246.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Aqo4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f0d33b2-a6f3-4795-95ad-f3ef95dbabfd_719x246.png 424w, https://substackcdn.com/image/fetch/$s_!Aqo4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f0d33b2-a6f3-4795-95ad-f3ef95dbabfd_719x246.png 848w, https://substackcdn.com/image/fetch/$s_!Aqo4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f0d33b2-a6f3-4795-95ad-f3ef95dbabfd_719x246.png 1272w, https://substackcdn.com/image/fetch/$s_!Aqo4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2f0d33b2-a6f3-4795-95ad-f3ef95dbabfd_719x246.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The sample covers July 28 through August 27 in the same repository: 9,148 Opus 5 turns (1.37M words) and 2,732 Fable 5 turns (193K words). Opus 5 averages 149.5 words a turn on equivalent status-update and engineering tasks, more than double Fable 5&#8217;s 70.5.</p><p>On a routine GitHub API check, Opus 5 wrote:</p><blockquote><p>Worth noting how it verified: its first PATCH used <code>-f description=@file</code>, which isn&#8217;t a real <code>gh api</code> shorthand...</p></blockquote><p>Later, in a <code>SlotRegistry</code> design note:</p><blockquote><p>It&#8217;s load-bearing for assertion B, it&#8217;s about to be cited in an ADR, and right now nobody but you can retrieve it.</p></blockquote><p>The &#8220;wearing&#8221; metaphor got under my skin. Opus 5 used some version of it 26 times across at least 12 sessions that month. A conflict resolution became &#8220;unreviewed code wearing a rebase&#8217;s clothing.&#8221; A misattributed fact came back as &#8220;a hypothesis wearing a rule&#8217;s clothing.&#8221; A faster teardown was &#8220;a regression dressed as a performance win.&#8221; Fable 5 produced zero in 193,000 words. The thinner Opus 4.8 and Sonnet 5 samples produced zero too.</p><p>Fable 5, working the same repository, the same kind of task, under the same ban:</p><blockquote><p>Error arms proven load-bearing (swallowing provider errors turns four tests red)...</p><p>You&#8217;re right, and the honest accounting is that the day went to the plumbing that kept invalidating runs...</p><p>The pinned gate run is genuinely green.</p></blockquote><p>Fable used the same forbidden words, but less often. The &#8220;worth noting&#8221; family appears roughly a tenth as often.</p><p>&#8220;Honest&#8221; is the exception. Opus 5 sits at 0.193 per 1,000 words against Fable 5&#8217;s 0.176, close enough to call a wash. The other gaps run from roughly 2x to 10.7x.</p><h2>Developers Recognize the Voice</h2><p>On Hacker News alone, I found 29 items from April through August 2026 about Opus 5&#8217;s communication style. Roughly 55 people joined the complaint. The <a href="https://news.ycombinator.com/item?id=49296740">highest-engagement thread</a>, &#8220;Why does Opus 5 feel worse to work with?&#8221;, sits at 993 points and 873 comments. Two Ask HN threads pose nearly the same question: <a href="https://news.ycombinator.com/item?id=49045140">&#8220;Which is the least sloppy and claudeism free model you have used?&#8221;</a> (July 25) and <a href="https://news.ycombinator.com/item?id=49444272">&#8220;Why are Claude models so verbose?&#8221;</a> (August 26).</p><p>On X, 13 tweets with verifiable URLs carry the same complaint from 11 distinct authors, dated July 24 through August 19.</p><p>By comparison, eight named people across both platforms say they prefer GPT-5.6 Sol specifically for how it communicates: six on Hacker News (including <a href="https://news.ycombinator.com/item?id=49239262">matheusmoreira</a>: &#8220;Opus has a distinctive sentence structure and Fable somehow manages to be even more obtuse&#8221;), plus two on X: <a href="https://x.com/eshear/status/2081583932999700758">Emmett Shear</a> (&#8221;Talking to Opus makes me angry and depressed in a way that&#8217;s hard to articulate. It&#8217;s actually somehow worse than Sol&#8221;) and <a href="https://x.com/yacineMTB/status/2081217662323982724">yacineMTB</a> (&#8221;Meh. Switched back to sol. I am trying to get shit done not be hypnotized by a robot&#8221;).</p><p>The <a href="https://news.ycombinator.com/item?id=49412301">816-point thread</a> that gets cited most often in this argument, &#8220;Anthropic&#8217;s best AI model struggles to attract users as cheaper tools thrive,&#8221; is a price story. One top comment calls the output &#8220;peak verbosity vomit,&#8221; but the thread itself is about pricing pressure from cheaper models, not a referendum on tone.</p><p>These quotes come from Reddit aggregator blogs that include thread names, usernames, and vote counts. One step removed. The style complaint still shows up with its own vocabulary: &#8220;essay of slop,&#8221; &#8220;Claudeslop,&#8221; &#8220;benchslop,&#8221; and &#8220;load-bearing&#8221; appear independently across multiple write-ups as stock Opus 5 language. But I could not find a single Reddit thread where someone says they prefer Sol because of how it talks. The Sol arguments there are about benchmarks and price. Two threads praise Fable 5&#8217;s voice over Opus 4.8, which says nothing about Fable 5 versus Opus 5.</p><p>Counted across the Hacker News threads read for this piece: &#8220;load-bearing&#8221; appears 83 times, &#8220;verbose&#8221; 69, &#8220;honest/honestly&#8221; 66, &#8220;jargon&#8221; 34, &#8220;seam(s)&#8221; 36.</p><h2>The Voice Predates the Model</h2><p><a href="https://github.com/anthropics/claude-code/issues/77136">GitHub issue #77136</a> was filed July 13, 2026, eleven days before Opus 5 shipped. It documents &#8220;load bearing,&#8221; invented jargon, &#8220;hand-waving,&#8221; &#8220;reflexive hedging,&#8221; and &#8220;honest framing&#8221; in Opus 4.7, Opus 4.8, and Fable 5. The complaint is older than the model now taking the blame.</p><p>Claude Code also gives Opus 5 a <a href="https://lucadidomenico.studio/en/blog/opus-5-verbose-system-prompt-claude-code">system prompt of roughly 11,000 characters</a> with almost no anti-verbosity instructions. Sonnet 5 gets about 29,000 characters, much of the difference made up of rules that suppress this behavior. In the published test, restoring the longer prompt changed the behavior. The model may supply the prose, but the harness decides how much restraint to ask for.</p><p>And somebody made the opposite switch. <a href="https://homeless-entrepreneur.web.app/blog/opus-vs-sol">Manu Parasuraman</a> moved from Sol to Opus 5 because Sol was harder to read: &#8220;Sol answers questions nobody asked in language nobody speaks,&#8221; while Opus 5 &#8220;walks you through its reasoning and stays concrete.&#8221;</p><p>I started this piece assuming Opus 5&#8217;s personality was the whole problem. The system-prompt evidence complicates that: I may be hearing Claude Code&#8217;s prompt as much as the model underneath it, then blaming the name in the model picker.</p><h2>What This Comparison Can and Can&#8217;t Show</h2><p>The 10.7x gap is descriptive. It combines output from before and after the bans, and it does not tell me what either model would say on an empty prompt. Sonnet 5 contributed 13 turns from one session, too thin to rate. GPT-5.6 Sol never appears in the trusty-tools logs, which means every Sol comparison here comes from other people rather than my own measurement. The Reddit material is secondhand because the primary site was unreachable during research.</p><h2>Ten Days Later, Anthropic Shipped a Concise Style</h2><p>Anthropic reportedly shipped a built-in Concise output style for Claude Code on August 20, according to <a href="https://botmonster.com/ai/make-opus-5-less-verbose/">a third-party writeup</a>. That was ten days after my rule landed and six days after the 993-point Hacker News thread. I haven&#8217;t checked whether it helps. I prefer my own.</p><p>The logs explain why I wrote the rule. Opus 5 completed the engineering work, at scale, while repeatedly talking past an instruction written for it in plain English. I don&#8217;t call that a capability failure, but it is a product problem. The rule stays in trusty-tools.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a>, a hospitality revenue-management platform, and writes about AI-augmented engineering practice.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/when-sessions-talk">When Sessions Talk</a> &#8212; What 375 messages between my own Claude Code sessions said to each other</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul>]]></content:encoded></item><item><title><![CDATA[The Cheap Model Blinked First]]></title><description><![CDATA[Economics of AI Addiction]]></description><link>https://hyperdev.matsuoka.com/p/the-cheap-model-blinked-first</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/the-cheap-model-blinked-first</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Mon, 10 Aug 2026 11:31:36 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!CRg0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96817f0e-67da-4dcf-83c9-1b5aef5bd8b4_889x498.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!CRg0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96817f0e-67da-4dcf-83c9-1b5aef5bd8b4_889x498.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!CRg0!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96817f0e-67da-4dcf-83c9-1b5aef5bd8b4_889x498.png 424w, https://substackcdn.com/image/fetch/$s_!CRg0!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96817f0e-67da-4dcf-83c9-1b5aef5bd8b4_889x498.png 848w, https://substackcdn.com/image/fetch/$s_!CRg0!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96817f0e-67da-4dcf-83c9-1b5aef5bd8b4_889x498.png 1272w, https://substackcdn.com/image/fetch/$s_!CRg0!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96817f0e-67da-4dcf-83c9-1b5aef5bd8b4_889x498.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!CRg0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96817f0e-67da-4dcf-83c9-1b5aef5bd8b4_889x498.png" width="889" height="498" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/96817f0e-67da-4dcf-83c9-1b5aef5bd8b4_889x498.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:498,&quot;width&quot;:889,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1042157,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/210415157?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F37a386f5-14d8-441b-89b0-11aeb1ff9bc6_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!CRg0!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96817f0e-67da-4dcf-83c9-1b5aef5bd8b4_889x498.png 424w, https://substackcdn.com/image/fetch/$s_!CRg0!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96817f0e-67da-4dcf-83c9-1b5aef5bd8b4_889x498.png 848w, https://substackcdn.com/image/fetch/$s_!CRg0!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96817f0e-67da-4dcf-83c9-1b5aef5bd8b4_889x498.png 1272w, https://substackcdn.com/image/fetch/$s_!CRg0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96817f0e-67da-4dcf-83c9-1b5aef5bd8b4_889x498.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In June of 2025 I predicted that Chinese open weight models would set the LLM pricing floor. The piece pointed at DeepSeek-V3 and DeepSeek-R1 turning up in the cheap tiers of tools like Windsurf and said &#8220;these will likely become the new baseline for price-sensitive usage.&#8221; I wrote it a month after a $622 Anthropic bill (up from around $50 the month before) and after burning a $20 Zed plan&#8217;s entire monthly allocation in a few hours. My reasoning was this: the frontier providers were selling inference below cost, the subsidy would end, and the cheap models were waiting underneath.</p><p>I gave it two time windows. The tighter one sat with the claim that &#8220;Chinese competition will eventually drive costs down significantly,&#8221; following &#8220;a period&#8212;probably 6-12 months&#8212;where usage-based pricing hits hard.&#8221; That one ran out around June. The looser one sat with a broader claim, that Chinese pressure would &#8220;eventually drive down pricing across the board,&#8221; with the caveat that &#8220;eventually&#8221; might be 12-18 months out. That one runs to December.</p><p>The first window failed outright. The second is failing in a direction I didn&#8217;t anticipate: not that the floor held, but that the floor is about to go up.</p><p>DeepSeek&#8217;s <a href="https://api-docs.deepseek.com/quick_start/pricing/">own API pricing documentation</a> now carries this: &#8220;We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected.&#8221; No percentage. No effective date. No list of affected models. The <a href="https://www.scmp.com/tech/tech-trends/article/3363129/deepseek-signals-significant-price-hike-amid-surge-demand-low-cost-ai-models">South China Morning Post</a>, which reported the notice, has DeepSeek telling developers that specifics are &#8220;to be advised later&#8221; and that &#8220;users should plan their usage accordingly.&#8221; SCMP attributes the pressure to demand for DeepSeek-V4-Flash-0731, a 284-billion-parameter model released on July 31, a week before the notice went up.</p><p>The company I expected to set the floor has signaled it will raise it. The labs I expected to raise prices spent the last four months cutting prices or raising limits.</p><h2>TL;DR</h2><ul><li><p>In June 2025 I gave Chinese models two windows to set the price floor: 6-12 months, and 12-18 months. The first ran out around June 2026. The second runs to December, and DeepSeek is now warning of a &#8220;significant&#8221; price increase with no figure and no date attached, confirmed on its own pricing docs, not just in press coverage.</p></li><li><p>Over the same stretch the frontier moved the other way. Anthropic doubled Claude Code&#8217;s 5-hour limits and dropped peak-hour throttling (May 6). OpenAI split its $200 Pro tier into $100 and $200 tiers running identical models, differing only in usage volume (April 9). Google cut its top Gemini subscription from $249.99 to $199.99 (May 19).</p></li><li><p>On July 30 OpenAI cut GPT-5.6 Luna by 80%, to $0.20 input and $1.20 output per million tokens. Multiply DeepSeek V4-Flash by ten and it lands at $1.40/$2.80: 3.6x to 7x under Opus 5 and Fable 5 on input where today it is 36x to 71x under them, and more expensive than Luna.</p></li><li><p>OpenAI&#8217;s stated reason for the cut was efficiency: 20% lower serving costs and 15%+ better token-generation efficiency. That covers a 20% cut. Terra got 20%. Luna got 80%.</p></li><li><p>The rationing I predicted did happen, one layer down. Cursor converted flat billing to metered credits in June 2025. GitHub Copilot moved to token-metered AI Credits on June 1, 2026, at unchanged sticker prices.</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!uvKD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1315a340-9d36-4c46-8c32-d2187474abd1_1024x539.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!uvKD!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1315a340-9d36-4c46-8c32-d2187474abd1_1024x539.png 424w, https://substackcdn.com/image/fetch/$s_!uvKD!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1315a340-9d36-4c46-8c32-d2187474abd1_1024x539.png 848w, https://substackcdn.com/image/fetch/$s_!uvKD!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1315a340-9d36-4c46-8c32-d2187474abd1_1024x539.png 1272w, https://substackcdn.com/image/fetch/$s_!uvKD!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1315a340-9d36-4c46-8c32-d2187474abd1_1024x539.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!uvKD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1315a340-9d36-4c46-8c32-d2187474abd1_1024x539.png" width="1024" height="539" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1315a340-9d36-4c46-8c32-d2187474abd1_1024x539.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:539,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1294023,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/210415157?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0552316-002e-47ee-8b53-1a61c870279d_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!uvKD!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1315a340-9d36-4c46-8c32-d2187474abd1_1024x539.png 424w, https://substackcdn.com/image/fetch/$s_!uvKD!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1315a340-9d36-4c46-8c32-d2187474abd1_1024x539.png 848w, https://substackcdn.com/image/fetch/$s_!uvKD!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1315a340-9d36-4c46-8c32-d2187474abd1_1024x539.png 1272w, https://substackcdn.com/image/fetch/$s_!uvKD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1315a340-9d36-4c46-8c32-d2187474abd1_1024x539.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>What DeepSeek Disclosed</h2><p>V4-Flash lists at $0.14 per million input tokens and $0.28 per million output, with cache hits at $0.0028. V4-Pro lists at $0.435 and $0.87. Those are the pre-hike numbers. DeepSeek disclosed neither the size of the increase nor when it lands. Any number you see attached to this story is somebody&#8217;s guess.</p><p>A &#8220;significant&#8221; increase off a base that low still leaves room. Double V4-Flash and it sits at $0.28/$0.56, cheaper than most of what the frontier sold a year ago. The magnitude we don&#8217;t know yet. The direction we do, and that is the part I called backwards.</p><h2>The Frontier Went The Other Way</h2><p>Anthropic <a href="https://www.anthropic.com/news/higher-limits-spacex">doubled Claude Code&#8217;s 5-hour rate limits on May 6</a> for Pro, Max, Team, and seat-based Enterprise plans, and removed the peak-hours limit reduction for Pro and Max. Max 5x is still $100 a month. Max 20x is still $200. Both buy more than they did in the spring.</p><p>On April 9 of this year, OpenAI split its $200 Pro tier into $100 and $200 tiers running identical models and differing only in usage volume. That&#8217;s a second fixed-price plan at Claude Max&#8217;s number. Neither tier is uncapped. OpenAI doesn&#8217;t publish the Pro-model allowance on either, and running it out soft-degrades you to a smaller model rather than cutting you off.</p><p>Google restructured <a href="https://blog.google/products-and-platforms/products/google-one/google-ai-subscriptions/">AI Ultra at I/O on May 19</a>. The single $249.99 tier became $99.99 at 5x Pro limits and $199.99 at 20x, with daily prompt caps replaced by compute-weighted 5-hour refresh and weekly metering. The top tier alone dropped $50.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Dm-1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb95a2297-eaa6-48df-b8e1-a070230a2de1_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Dm-1!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb95a2297-eaa6-48df-b8e1-a070230a2de1_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!Dm-1!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb95a2297-eaa6-48df-b8e1-a070230a2de1_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!Dm-1!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb95a2297-eaa6-48df-b8e1-a070230a2de1_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!Dm-1!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb95a2297-eaa6-48df-b8e1-a070230a2de1_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Dm-1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb95a2297-eaa6-48df-b8e1-a070230a2de1_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b95a2297-eaa6-48df-b8e1-a070230a2de1_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1318172,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/210415157?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb95a2297-eaa6-48df-b8e1-a070230a2de1_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Dm-1!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb95a2297-eaa6-48df-b8e1-a070230a2de1_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!Dm-1!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb95a2297-eaa6-48df-b8e1-a070230a2de1_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!Dm-1!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb95a2297-eaa6-48df-b8e1-a070230a2de1_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!Dm-1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb95a2297-eaa6-48df-b8e1-a070230a2de1_1024x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Who GPT-5.6 Luna Is For</h2><p>Then came July 30. OpenAI cut GPT-5.6 Terra from $2.50/$15 to $2/$12, a 20% reduction. It cut GPT-5.6 Luna from $1/$6 to $0.20/$1.20, an 80% reduction. The stated reason, <a href="https://x.com/OpenAI/status/2082577277246972300">from OpenAI&#8217;s own account</a>: &#8220;20% lower serving costs from production GPU kernel improvements. 15%+ better token-generation efficiency from improved speculative decoding.&#8221;</p><p>That explains Terra. Twenty percent cheaper to serve, twenty percent off the price, a straight pass-through. It doesn&#8217;t explain Luna. The same kernels and the same speculative decoding produced a cut four times larger on the cheaper model. Efficiency gains don&#8217;t sort themselves by tier. Pricing decisions do.</p><p>Max plans holding steady is not by itself evidence that providers depend on volume. Falling serving costs cover that on their own. A lab building a tier at the bottom of the market and then cutting it 80% in a single move is harder to explain that way, because a $0.20 input tier is not where a company with a $5/$25 flagship model makes its money. It&#8217;s where a company goes when it wants a number that reads lower than DeepSeek&#8217;s.</p><p><a href="https://hyperdev.matsuoka.com/p/weve-turned-a-corner">Two months ago I wrote</a> that dealers give the first one away for a reason, and that the Max plan was my first bag. The habit runs in both directions. A dealer who keeps cutting the price of the first bag is a dealer who can&#8217;t afford to have you buy it somewhere else.</p><p>Run my June 2025 claim forward against a tenfold hike. DeepSeek disclosed nothing supporting that multiple or any other, so it&#8217;s a stress test, not a forecast. V4-Flash at ten times list: $1.40 input, $2.80 output. Against the models people think of when they say &#8220;frontier&#8221;, a multiple survives. Claude Opus 5 at $5/$25 is 3.6x the input price and 8.9x the output. Claude Fable 5 at $10/$50 is 7.1x and 17.9x. GPT-5.5 at $5/$30 is 3.6x and 10.7x. Those are reduced from much bigger numbers. At today&#8217;s list price Opus 5 costs 36x what V4-Flash does on input and 89x on output. A tenfold hike takes about ninety percent of that away. At 36x, price picks the model. At 3.6x, the model does.</p><p>Further down the market it reverses. Against Luna at $0.20/$1.20, a tenfold-hiked V4-Flash costs seven times more on input and more than twice as much on output. Against Gemini 3.6 Flash at $1.50/$7.50 it&#8217;s roughly a wash on input. Even Gemini 3.1 Pro at $2/$12 for prompts at or under 200k tokens (above that it goes to $4/$18) sits within 1.4x on input (4.3x on output). Run the stress test and the frontier ends up priced below the discounter, not for the flagship but for the tier built for exactly the work I said DeepSeek would take.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gU6k!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0af20b9a-f7a8-45fb-9a3e-e429287618d6_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gU6k!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0af20b9a-f7a8-45fb-9a3e-e429287618d6_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!gU6k!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0af20b9a-f7a8-45fb-9a3e-e429287618d6_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!gU6k!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0af20b9a-f7a8-45fb-9a3e-e429287618d6_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!gU6k!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0af20b9a-f7a8-45fb-9a3e-e429287618d6_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gU6k!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0af20b9a-f7a8-45fb-9a3e-e429287618d6_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0af20b9a-f7a8-45fb-9a3e-e429287618d6_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:958904,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/210415157?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0af20b9a-f7a8-45fb-9a3e-e429287618d6_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!gU6k!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0af20b9a-f7a8-45fb-9a3e-e429287618d6_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!gU6k!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0af20b9a-f7a8-45fb-9a3e-e429287618d6_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!gU6k!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0af20b9a-f7a8-45fb-9a3e-e429287618d6_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!gU6k!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0af20b9a-f7a8-45fb-9a3e-e429287618d6_1024x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The Case Against</h2><p>Three things complicate this reading.</p><p><strong>DeepSeek&#8217;s prices track chip supply, not just strategy.</strong> The company cited &#8220;constraints in high-end compute capacity&#8221; when it priced V4-Pro in April 2026, a constraint widely linked to US export controls on Nvidia H20 chips, though DeepSeek didn&#8217;t specify. In late May it made a 75% V4-Pro price cut permanent without saying why, on timing that lines up with easing Huawei Ascend 950 supply. The August notice gives no reason at all, though SCMP&#8217;s reporting points at demand for a model released the week before. I can&#8217;t tell a capacity bottleneck from a strategic retreat from outside the company.</p><p><strong>The rationing happened at the consumer level.</strong> Cursor replaced flat request-based Pro billing with usage-based credits pegged to API cost in June 2025, then ran a refund program from June 16 to July 4 after users got painful bills. GitHub Copilot moved from flat Premium Request Units to token-metered AI Credits on June 1, 2026, at unchanged sticker prices, with the code-review multiplier rising to 13x for legacy annual-plan holders. Anthropic added weekly caps on top of its 5-hour caps in August 2025, citing round-the-clock usage and account resale, and <a href="https://techcrunch.com/2025/07/28/anthropic-unveils-new-rate-limits-to-curb-claude-code-power-users/">estimated at announcement</a> that they would &#8220;apply to less than 5% of subscribers based on current usage.&#8221; Model-lab sticker prices held. The metering moved to heroin layer.</p><p><strong>The efficiency explanation may be the whole explanation.</strong> Serving costs have fallen steeply, and Anthropic&#8217;s gross margin has reportedly climbed off a deeply negative base, roughly -94% in 2024. <a href="https://www.investing.com/news/stock-market-news/anthropic-trims-profit-margin-outlook-as-ai-operating-costs-rise--the-information-4459316">The Information reported in January 2026</a> that Anthropic had trimmed its own projection to about 40% for 2025, still a steep climb off that base. If cost per token falls faster than price per token, generous Max plans need no dependency on volume to explain them. My reading rests on the 20/80 split between Terra and Luna, and that&#8217;s one data point on one day from one vendor.</p><h2>What Would Prove Me Wrong Again</h2><p>The June 2025 piece got the $200 line right, which was the easy part. What I got wrong was which end of the market was fragile. Any of these would tell me I&#8217;m getting it wrong a second time.</p><ul><li><p><strong>DeepSeek publishes numbers and the increase is under 2x on V4-Flash</strong>, leaving it under $0.30 input. That would be a capacity adjustment &#8212; the floor is where I thought it was.</p></li><li><p><strong>OpenAI raises Luna back toward $1 within two quarters.</strong> If that happens, July 30 was pass-through of a real engineering result, and I read a price war into it.</p></li><li><p><strong>Anthropic or Google restores throttling, or cuts 5-hour allowances at unchanged prices, before the end of 2026.</strong> That would mean the generosity was a promotional window, and June 2025 was early rather than wrong.</p></li><li><p><strong>The tool layer keeps metering while the model labs keep expanding:</strong> the pricing pressure was sitting at the tool layer the whole time, and both of my previous pieces were aimed at the wrong tier.</p></li><li><p><strong>DeepSeek&#8217;s hike lands and its API volume holds anyway</strong> &#8212; which would mean price wasn&#8217;t what price-sensitive buyers were choosing on.</p></li></ul><p>If you&#8217;re budgeting against a Chinese-model price floor for 2027, price in the possibility that the floor is now a frontier lab&#8217;s customer-acquisition tier.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a>, a hospitality revenue-management platform, and writes about AI-augmented engineering practice.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/p/the-other-shoe-will-drop">The Other Shoe Will Drop</a> &#8212; The June 2025 prediction this piece corrects.</p></li><li><p><a href="https://hyperdev.matsuoka.com/p/the-other-shoe-has-dropped">The Other Shoe Has Dropped</a> &#8212; Why per-token price cuts stopped reaching the invoice.</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners.</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders.</p></li></ul>]]></content:encoded></item><item><title><![CDATA[OpenAI's "Comeback"]]></title><description><![CDATA[Can it really be a comeback with they have so much money, so much exposure, and so little revenue?]]></description><link>https://hyperdev.matsuoka.com/p/openais-comeback</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/openais-comeback</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Fri, 31 Jul 2026 11:30:03 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!hQ4A!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!hQ4A!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!hQ4A!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png 424w, https://substackcdn.com/image/fetch/$s_!hQ4A!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png 848w, https://substackcdn.com/image/fetch/$s_!hQ4A!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png 1272w, https://substackcdn.com/image/fetch/$s_!hQ4A!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!hQ4A!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png" width="1456" height="874" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:874,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1378063,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/209163913?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!hQ4A!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png 424w, https://substackcdn.com/image/fetch/$s_!hQ4A!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png 848w, https://substackcdn.com/image/fetch/$s_!hQ4A!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png 1272w, https://substackcdn.com/image/fetch/$s_!hQ4A!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd36b4f06-c8d8-4a65-9808-b8f069539891_1600x960.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>GPT-5.6 Sol scores 80 on the <a href="https://artificialanalysis.ai/articles/gpt-5-6-has-landed">Artificial Analysis Coding Agent Index</a> against 77 for Claude Fable 5, at a lower cost per task (about $1.04). That post went up on July 9.</p><p>I called model parity back in April, in <a href="https://hyperdev.matsuoka.com/p/its-the-harness-stupid">a piece</a> with a section headed &#8220;GPT-5.4 Caught Up.&#8221; What I missed was everything around the model: that OpenAI would close the product gap, and buy a decade of compute while doing it.</p><p>Over the last few months I&#8217;ve been using ChatGPT more and more for work-like tasks. Not Duetto work, because we don&#8217;t have an enterprise plan and Duetto material stays out of it. Work-like things. Both the capabilities and the UX have improved to the point where I&#8217;d put the product at least on par with claude.ai, to my surprise. There are still gaps. It&#8217;s still good.</p><h2>TL;DR</h2><ul><li><p>GPT-5.6 Sol leads the Artificial Analysis Coding Agent Index 80 to 77 over Claude Fable 5 (July 9) and Terminal-Bench 2.1 85.77% to 84.64% over Claude Opus 5 (July 22), the second one with a methodology caveat that widens the gap.</p></li><li><p>Claude Opus 5 still leads the broader Artificial Analysis Intelligence Index at 60.7 against GPT-5.6 Sol&#8217;s 58.9 (July 29). Anthropic&#8217;s remaining model lead is general capability, not agentic coding.</p></li><li><p>Anthropic leads on revenue ($47B run-rate, May 28, 2026) and valuation ($965B), and leads enterprise API share 40% to 27% (Menlo, December 2025). OpenAI closed the capability gap and most of the product gap while trailing commercially.</p></li><li><p>OpenAI is far ahead on planned compute (Stargate at nearly 7GW toward a stated 10GW commitment, the Broadcom &#8220;Jalape&#241;o&#8221; inference ASIC). In July it moved its usage limits three times in seventeen days, in both directions, around a near-global outage on the 25th. Anthropic has been raising its limits since May.</p></li><li><p>No open-weight provider I found fields an app ecosystem within reach of ChatGPT or Claude Code, which keeps this a two-horse race even as the weights converge. Chinese domestic app numbers could complicate that.</p></li></ul><h2>Where the coding lead sits</h2><p>Two current measures put GPT ahead on agentic coding. The Coding Agent Index lists GPT-5.6&#8217;s three effort tiers at Sol 80, Terra 77.4 and Luna 74.6, with Claude Fable 5 at 77.2 (Terra and Fable 5 land 0.2 apart). Sol&#8217;s per-task cost runs below both Fable 5 and Opus 4.8. On <a href="https://www.vals.ai/benchmarks/terminal-bench-2-1">Terminal-Bench 2.1</a> (89 tasks, Terminus 2 harness, snapshot dated July 22), GPT-5.6 Sol leads Claude Opus 5 by about a point, 85.77% to 84.64%.</p><p>The Terminal-Bench result comes with a caveat from the publisher. Opus 5&#8217;s run used Claude Opus 4.8 as a refusal fallback. Count the nine affected passing results as failures instead and Opus 5 drops to 81.27%, which widens the GPT lead to roughly four and a half points. The lead is real. Its size depends on how you count.</p><p>Now the other direction. On the broader Artificial Analysis Intelligence Index, <a href="https://benchlm.ai/benchmarks/artificialAnalysis">as of July 29</a>, Claude Opus 5 sits first at 60.7 and Claude Fable 5 second at 59.9, with GPT-5.6 Sol third at 58.9. Anthropic holds the top two spots on general capability while losing the coding-agent tables. My &#8220;diminishing edge&#8221; framing was directionally right if imprecise. The edge moved. It didn&#8217;t disappear.</p><p>SWE-bench Pro, the number everybody reached for six months ago, has no usable current standing. Scale&#8217;s <a href="https://labs.scale.com/leaderboard/swe_bench_pro_public">public-set leaderboard</a> doesn&#8217;t list GPT-5.6, Claude Opus 5 or Claude Fable 5 at all, and the top of it reshuffled again this month. The aggregator sites quoting newer figures disagree with each other and cite at least one model name I can&#8217;t confirm exists. Skip it.</p><h2>The product story is a two-way race now</h2><p>This is the first half of what I missed. Codex spans CLI, IDE, web, and a hosted cloud mode for long-running tasks. The <a href="https://learn.chatgpt.com/docs/changelog?type=codex-cli">Codex CLI changelog</a> for the last two weeks of July includes resumable sessions with paginated thread history, session naming and branching, configurable sub-agents, audio input, and an <code>/import</code> command that migrates project-scoped memories out of Cursor and Claude Code. That last one reads like a decision made by somebody thinking hard about switching costs. On <a href="https://9to5mac.com/2026/07/09/openai-announcing-the-next-chapter-for-chatgpt-today-watch-here/">July 9</a> OpenAI also folded the standalone Codex desktop app into a single ChatGPT desktop app with Chat, Work, and Codex modes. Existing Codex users were moved across automatically, and Codex remains a full mode inside the merged app.</p><p>What I notice in use isn&#8217;t a benchmark, it&#8217;s flow. Codex and GPT live in the same product, and when I go looking for something I worked on weeks ago, ChatGPT finds it wherever I left it. Claude&#8217;s Projects still behave like sealed folders for me. That isolation was an advantage once, back when context was scarce and a hard wall around what the model could see was the safer design. With current context budgets it mostly means I go hunting. GPT is also the far better image renderer, which makes for a short comparison, since Claude doesn&#8217;t ship a first-party image model at all.</p><p>Anthropic didn&#8217;t stand still while any of this happened. <a href="https://www.anthropic.com/news/claude-design-anthropic-labs">Claude Design</a> shipped out of Anthropic Labs on April 17, and it&#8217;s excellent: you talk to it and get slides, wireframes, one-pagers and pitch decks, with export to PPTX or PDF, inline comments, adjustment sliders, and handoff into Claude Code. Codex has no equivalent that I&#8217;m aware of (and I&#8217;d bet real money they&#8217;re building one). Claude Cowork expanded to web and mobile on <a href="https://techcrunch.com/2026/07/07/the-coding-agent-wars-are-spilling-into-the-rest-of-the-office-claude-cowork/">July 7</a>. Both labs shipped serious B2B product work inside the same four months.</p><h2>Capacity: a much bigger future, a strained present</h2><p>The second half is compute. OpenAI is buying a much bigger future than Anthropic is, and is visibly straining in the present.</p><p>On committed buildout it isn&#8217;t close. OpenAI, Oracle and SoftBank put combined planned Stargate capacity at <a href="https://openai.com/index/five-new-stargate-sites/">nearly 7 gigawatts</a> across the Abilene flagship plus five more sites, with over $400B invested and the full commitment still stated as $500B and 10GW. Tomasz Tunguz&#8217;s aggregation of the announced vendor deals puts OpenAI&#8217;s 2025&#8211;2035 infrastructure commitment at <a href="https://tomtunguz.com/openai-hardware-spending-2025-2035/">roughly $1.15 trillion</a> across seven suppliers, which is derived rather than OpenAI-confirmed. On June 24, OpenAI and Broadcom unveiled <a href="https://openai.com/index/openai-broadcom-jalapeno-inference-chip/">Jalape&#241;o</a>, OpenAI&#8217;s first custom inference ASIC, designed to tape-out in about nine months, with engineering samples already running Codex workloads in the lab and first production deployment targeted at gigawatt scale in late 2026 (<a href="https://www.cnbc.com/2026/06/24/openai-and-broadcom-reveal-jalapeno-first-ai-chip-in-partnership.html">CNBC&#8217;s coverage</a> has the same details).</p><p>Anthropic&#8217;s disclosed commitments are significant but much smaller: <a href="https://www.anthropic.com/news/anthropic-amazon-compute">up to 5GW with AWS</a> announced April 20 (with nearly 1GW of Trainium online by year end and $100B+ committed over ten years), <a href="https://www.anthropic.com/news/google-broadcom-partnership-compute">well over 1GW of Google TPU capacity in 2026</a> with the Broadcom-built expansion arriving from 2027, and <a href="https://www.anthropic.com/news/higher-limits-spacex">more than 300MW from SpaceX&#8217;s Colossus 1</a>. Its custom silicon is at the conversation stage, <a href="https://techcrunch.com/2026/07/02/anthropic-is-discussing-a-new-custom-chip-with-samsung/">reportedly early talks with Samsung</a> on a 2nm accelerator with nothing shipping before late 2027.</p><p>The strain shows up in the usage limits, which OpenAI moved three times in seventeen days, in both directions. Sol&#8217;s launch week doubled traffic inside 48 hours, so on July 12 and 13 it lifted the five-hour cap on Plus, Pro and Business tiers to absorb the load. On July 25 came a <a href="https://status.openai.com/incidents/01KYC921K145JTR1JK7DYKGWH1">near-global outage</a>, 9:17 to 11:08 AM ET, elevated errors across ChatGPT, the API and Codex, resolved the same morning. On July 29 it reset banked limits for ChatGPT Work and Codex users, whose quota Sol was burning faster than the math anticipated, then put the five-hour cap back the next day. Anthropic has been moving the other way since spring. On May 6 it permanently doubled Claude Code&#8217;s five-hour limits for Pro, Max, Team and seat-Enterprise users and dropped peak-hour throttling for Pro and Max, on the back of that SpaceX capacity.</p><h2>Where Anthropic still leads</h2><p>Anthropic&#8217;s remaining, potentially durable, advantages are enterprise adoption and commercial scale. Anthropic is ahead on total revenue and on valuation, and well ahead on enterprise API share. Its <a href="https://www.anthropic.com/news/series-h">Series H</a>, a $65B round closed May 28, 2026, disclosed a $47B run-rate and a $965B post-money valuation. OpenAI&#8217;s most recent confirmed figures are roughly $25B annualized (about $2B a month) at the <a href="https://www.cnbc.com/2026/03/31/openai-funding-round-ipo.html">March 31, 2026 close</a> of its $122B round, at an <a href="https://www.bloomberg.com/news/articles/2026-03-31/openai-valued-at-852-billion-after-completing-122-billion-round">$852B post-money valuation</a>. Menlo Ventures&#8217; <a href="https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/">December 2025 enterprise report</a> put Anthropic at 40% of enterprise LLM API usage against OpenAI&#8217;s 27%, and at 54% in AI coding specifically. No 2026 Menlo report exists yet, so December is the current number.</p><p>So OpenAI has closed the capability gap and most of the product gap while trailing on money and on enterprise accounts. That&#8217;s the actual shape of the comeback. The lead Anthropic still holds on the Intelligence Index is general capability, and general capability is the harder thing to put in front of a CTO who&#8217;s buying a coding agent this quarter.</p><h2>Apple versus Microsoft, again?</h2><p>I reached for that analogy myself in May 2025, when OpenAI spent close to $10 billion in three weeks on Jony Ive&#8217;s io Products and on Windsurf, and I called it <a href="https://hyperdev.matsuoka.com/p/openais-apple-moment-building-a-walled">OpenAI&#8217;s Apple moment</a>: control the stack from silicon to screen, consumer-first, design-led. Anthropic in that mapping is Microsoft, enterprise and developer-led, winning the accounts while nobody writes headlines about it.</p><p>I don&#8217;t fully trust it. Apple versus Microsoft was never settled by who had the better quarter. It ran on which company owned the layer everybody else had to build on top of, and on switching costs measured in years. Nobody owns that layer here. Both companies rent it, from Amazon, Google, Microsoft, Nvidia, Broadcom and now SpaceX. For a developer, changing frontier models is a config change and an eval run. What costs real time is changing the tool you work in, which is a much smaller lock than owning the platform. The Apple/Microsoft dynamic assumed you couldn&#8217;t leave.</p><p>The mapping also keeps slipping. Anthropic is the enterprise player in it, which is right, but it&#8217;s also the one with the bigger revenue base and the higher valuation, which is not the role the analogy assigns it.</p><h2>Can either of them make money</h2><p>Neither is profitable and both are burning billions, so anything past that is projection.</p><p>OpenAI&#8217;s advertising business is <a href="https://247wallst.com/investing/2026/07/21/openai-is-on-pace-to-miss-its-own-ad-revenue-forecast-by-90-heres-what-it-means-for-the-ai-trade/">on pace to miss its own five-year forecast by about 90%</a>, OpenAI projected $2.5B in ad revenue this year, scaling to $100B by 2030. eMarketer puts the entire US chatbot-ad market, not just OpenAI, under $1B this year and around $5.4B by 2030, per 24/7 Wall St. on July 21. The widely circulated 2026 cash-burn estimates in the $25B range are analyst syntheses rather than disclosures, because OpenAI is private and doesn&#8217;t publish the figures that would settle it. Same caution applies to the gross-margin numbers people quote for both companies.</p><p>Anthropic&#8217;s topline looks better but is contested. Ed Zitron argues in <a href="https://www.wheresyoured.at/anthropics-profitability-swindle/">&#8221;Anthropic&#8217;s &#8216;Profitability&#8217; Swindle&#8221;</a> that the near-term EBITDA profitability claim is a one-quarter accounting artifact of temporarily discounted SpaceX compute during exactly the months profitability was claimed, and that costs still rise linearly with revenue. He also flags CFO Krishna Rao&#8217;s March 9 sworn filing citing cumulative revenue &#8220;exceeding $5 billion to date&#8221; against the $19B run-rate announced six days earlier. That second point is weaker than it reads, since cumulative revenue and annualized run-rate aren&#8217;t the same measure. The first point is the one I&#8217;d want answered, and it hasn&#8217;t been.</p><p>Hardware doesn&#8217;t resolve it. Chris Lehane said at Davos in January that OpenAI&#8217;s first consumer device would debut in the <a href="https://9to5mac.com/2026/01/19/openai-teases-hardware-unveil-this-year-as-jony-ives-team-hires-more-apple-alumni/">&#8221;latter part&#8221; of 2026</a>, which is an unveiling and explicitly not an on-sale date. MacRumors reported in <a href="https://www.macrumors.com/2026/02/20/jony-ive-openai-smart-speaker-2027/">February</a> that the thing is a smart speaker with a camera launching in 2027. Everything else circulating about it (screenless, voice-first, 360-degree camera, Foxconn in Vietnam) is unconfirmed. As of today nothing has shipped and no date is confirmed, so I&#8217;d file the device as a 2027 question and stop considering it this year.</p><h2>Open weights are a side argument</h2><p>The open-weight families are close to the frontier on raw capability. Kimi K3, the highest-scoring open model on the Artificial Analysis Intelligence Index, sits at 57, about four points behind Claude Opus 5&#8217;s 60.7 and under two behind GPT-5.6 Sol&#8217;s 58.9, the same index cited above. On <a href="https://arena.ai/leaderboard">LMArena</a> the best open-weight entry, Kimi K3-max at 1491, trails Claude Fable 5 at the top by 17 points.</p><p>What none of them have is the app ecosystem. No open-weight provider fields anything within reach of ChatGPT or Claude Code on integration depth or developer reach, at least none I found, though I didn&#8217;t dig into DeepSeek&#8217;s or Qwen&#8217;s domestic app numbers in China, which could complicate that. Weights converging matters much less than it sounds like it should when the surface people work in belongs to two companies.</p><p>Which is where I end up. The model tables are the smallest part of it. OpenAI put Codex and ChatGPT in one app with an import command pointed straight at Claude Code&#8217;s users, and it has nearly 7 gigawatts of planned capacity standing behind that. Claude Opus 5 still holds the top of the Intelligence Index at 60.7, and Anthropic still held 40% of enterprise API spend as of last December. OpenAI is knocking on the door. A year ago I&#8217;d have told you that door was closed.</p><p>A coda. Could we imagine a world in which Anthropic and OpenAI run products on other models? Seems very unlikely now. But the cost of developing new models is astronomical, the cost of derived models much less (even though both companies are <a href="https://www.completeskeptic.com/p/is-it-even-possible-for-the-chinese">fighting distillation</a> tooth and nail). So there is a world in which they use the model building work they&#8217;ve done as an audience and tool building exercise, and focus on expanding that share by expanding options, let the models fight for their own justification. Farfetched, as I mentioned, but if they start losing to application clones that use open weight models, they&#8217;re essentially inviting someone to compete. The existing assumption that the app ecosystem was just to drive captured inference falls apart if that inference is cost prohibitive. We shall see.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and also writes about AI business at <a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496">AI Power Ranking</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/p/openais-apple-moment-building-a-walled">OpenAI&#8217;s &#8216;Apple Moment&#8217;: Building a Walled-Garden AI Stack</a> &#8212; Where I first made the Apple analogy, three weeks and $10 billion into OpenAI&#8217;s acquisition run.</p></li><li><p><a href="https://hyperdev.matsuoka.com/p/its-the-harness-stupid">It&#8217;s the Harness, Stupid</a> &#8212; The April piece where I already conceded model parity, under a section header that says so. Everything I got wrong after that is in this article.</p></li><li><p><a href="https://hyperdev.matsuoka.com/p/i-switched-to-claudeai-from-chatgpt">I Switched to Claude.AI from ChatGPT As My Main AI Assistant</a> &#8212; The May 2025 call I&#8217;m revisiting here, with the context-management and stability reasons that drove it.</p></li><li><p><a href="https://aipowerranking.com">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners.</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders.</p></li></ul>]]></content:encoded></item><item><title><![CDATA[What A Difference A Year Makes]]></title><description><![CDATA[claude-mpm to trusty-tools/mpm]]></description><link>https://hyperdev.matsuoka.com/p/what-a-difference-a-year-makes</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/what-a-difference-a-year-makes</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Fri, 24 Jul 2026 12:30:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Ko1F!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3035844-db97-4357-a169-309d987dc3b0_1619x971.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Ko1F!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3035844-db97-4357-a169-309d987dc3b0_1619x971.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Ko1F!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3035844-db97-4357-a169-309d987dc3b0_1619x971.png 424w, https://substackcdn.com/image/fetch/$s_!Ko1F!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3035844-db97-4357-a169-309d987dc3b0_1619x971.png 848w, https://substackcdn.com/image/fetch/$s_!Ko1F!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3035844-db97-4357-a169-309d987dc3b0_1619x971.png 1272w, https://substackcdn.com/image/fetch/$s_!Ko1F!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3035844-db97-4357-a169-309d987dc3b0_1619x971.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Ko1F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3035844-db97-4357-a169-309d987dc3b0_1619x971.png" width="1456" height="873" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e3035844-db97-4357-a169-309d987dc3b0_1619x971.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:873,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2305672,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/208282371?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3035844-db97-4357-a169-309d987dc3b0_1619x971.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Ko1F!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3035844-db97-4357-a169-309d987dc3b0_1619x971.png 424w, https://substackcdn.com/image/fetch/$s_!Ko1F!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3035844-db97-4357-a169-309d987dc3b0_1619x971.png 848w, https://substackcdn.com/image/fetch/$s_!Ko1F!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3035844-db97-4357-a169-309d987dc3b0_1619x971.png 1272w, https://substackcdn.com/image/fetch/$s_!Ko1F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3035844-db97-4357-a169-309d987dc3b0_1619x971.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I made the first commit to my project, called <a href="https://github.com/bobmatnyc/claude-mpm">claude-mpm</a>, on July 24, 2025. It was a multi-agent project manager built on top of Claude Code: a Python package that gave you specialist agents, a ticket workflow, a thin memory layer, and a way to route work between them. Building the agents was most of the work, because at that time Claude Code had no notion of a subagent. You wrote your agents as Markdown, loaded them yourself, and orchestrated the handoffs by hand.</p><p><a href="https://code.claude.com/docs/en/sub-agents">Custom subagents</a> shipped in Claude Code on July 25, 2025, the day after that first commit.</p><p>I&#8217;m not telling this story because of that coincidence. I&#8217;m telling it because it marks a moment. In July 2025, multi-agent orchestration was something novel. A year later it&#8217;s commonplace. And that single move (capability migrating out of the code you write and into the tool you run) is a line connecting the past twelve months in agentic coding, measurable in two of my own repositories.</p><h2>TL;DR</h2><ul><li><p><strong>Same author, same subscription, two eras.</strong> claude-mpm (Python) took roughly twelve months and 4,552 commits, with memory and search either hand-rolled or bolted on from outside. <a href="https://github.com/bobmatnyc/trusty-tools">trusty-tools</a>, a 25-crate Rust monorepo with first-class memory, semantic search, worktree orchestration, and PR review, came together in about two months.</p></li><li><p><strong>The budget held. The developer improved, but the harness and the models leapt.</strong> I spent the year getting better at driving a team of agents. Even so, most of the difference sits with the tooling, not with me. Both projects ran on the same Claude Max subscription. What changed underneath was the frontier model and everything Claude Code learned to do on its own.</p></li><li><p><strong>A year ago you hand-built the PM layer.</strong> Agents-as-Markdown, an external vector-search dependency, a memory subsystem you maintained yourself. Working in tickets felt like an edge.</p></li><li><p><strong>Now the harness carries it.</strong> Subagents nest and run in the background, worktrees are a native flag, skills and memory and search are first-class. Ticket-driven development is table stakes. Running many worktrees against one repo is my working model now.</p></li><li><p><strong>The safe extrapolation:</strong> The observed path is tickets &#8594; worktrees &#8594; many concurrent sessions per repo. If the next year rhymes with the last, the harness manages more parallelism than a person can hold in their head.</p></li></ul><h2>A year ago: you built the orchestration yourself</h2><p>claude-mpm was shaped by what the harness couldn&#8217;t do at that time.</p><p>The frontier models in late July 2025 were Claude Opus 4 and Sonnet 4, which had <a href="https://www.anthropic.com/news/claude-4">reached general availability</a> on May 22, 2025. Claude Code itself was about two months into GA, capable but young. Subagents didn&#8217;t exist until the day after I started. <a href="https://www.anthropic.com/news/agent-skills">Agent Skills</a> wouldn&#8217;t launch until October 2025. There was no native worktree support. If you wanted isolated parallel work, you ran <code>git worktree add</code> yourself and wired it up by hand. MCP was maybe eight months old. Memory and semantic search were not primitives the harness handed you.</p><p>So we (claude-mpm and me) built all of it. The agents were Markdown templates the package loaded and orchestrated. The ticket workflow lived in the CLI. Memory was thin by necessity, a couple of files under <code>src/claude_mpm/memory/</code>, and even that was on its way out. The changelog shows the memory hooks being removed and handed off to an external successor rather than maintained in-tree. Semantic search wasn&#8217;t native either. It came in as an outside dependency, <code>mcp-vector-search</code>, referenced across dozens of files. Worktree awareness existed at the level of a consumer that knew the concept, not a subsystem that owned it.</p><p>That was the state of the art then, and it worked. Over about twelve months claude-mpm accumulated 4,552 commits (roughly 380 a month) and grew into a real system: multiple MCP channel servers built into the Python package, a plugin path exposing 56-odd skills, agent templates for a spread of specialist roles. Ticket-driven development, routing discrete units of work through a queue instead of narrating one long conversation, felt like an advance at the time. You had to construct the scaffolding to get there, and the scaffolding was the hard part of the project.</p><h2>Now: the harness carries it</h2><p>trusty-tools took its first commit on May 19, 2026. It was still under active development the day I pulled these numbers. In roughly two months it reached 1,726 commits on <code>main</code>, about 860 a month. Same author, same Claude Max subscription, a different era.</p><p>That first commit wasn&#8217;t a blank slate. It already carried claude-mpm&#8217;s PM scaffolding, and one absorbed component holds claude-mpm session logs dated May 11, a week before the new repo existed. The Python predecessor was building its Rust successor. Then the successor began building itself: on July 6 the commit attribution switched to &#8220;generated with trusty-mpm,&#8221; in a commit that was fixing trusty-mpm&#8217;s own guard for its own subagents. Three days later the instructions moved out of CLAUDE.md into trusty-mpm&#8217;s own convention. About seven weeks were built by its predecessor before it took over its own development.</p><p>A lot shipped in that window. The frontier moved to <a href="https://www.anthropic.com/news/claude-opus-4-8">Claude Opus 4.8</a> (late May 2026) and <a href="https://www.anthropic.com/news/claude-sonnet-5">Claude Sonnet 5</a> (end of June 2026), with a Mythos-class model, <a href="https://www.anthropic.com/news/claude-fable-5-mythos-5">Fable 5</a>, arriving in June at a million tokens of context. The million-token context window had <a href="https://claude.com/blog/1m-context-ga">gone GA at standard pricing</a> in early 2026. Inside Claude Code, subagents now nest several deep and run in the background by default. <code>/fork</code><a href="https://code.claude.com/docs/en/changelog"> and </a><code>/subtask</code><a href="https://code.claude.com/docs/en/changelog"> landed</a> in mid-2026, and a <a href="https://code.claude.com/docs/en/changelog">native </a><code>--worktree</code><a href="https://code.claude.com/docs/en/changelog"> flag</a> had arrived in 2026. The harness I was building on top of in mid-2026 was a different animal from the one I started claude-mpm against.</p><p>And trusty-tools reflects that, because it didn&#8217;t have to build the parts the harness now provides. It could spend that effort building capabilities further out. The result is a 25-crate workspace consisting of about 600K lines of Rust. The pieces claude-mpm hand-rolled or imported are now first-class crates of their own:</p><ul><li><p><strong>Memory</strong> is <a href="https://github.com/bobmatnyc/trusty-tools">trusty-memory</a>, a dedicated crate with knowledge-graph operations, &#8220;dream&#8221; consolidation that compacts and reorganizes stored facts, multiple namespaced memory &#8220;palaces,&#8221; and chat-session persistence. claude-mpm had no equivalent. Its two-file memory layer was being handed off precisely because maintaining that by hand no longer made sense.</p></li><li><p><strong>Search</strong> is a first-class crate (semantic, lexical, and knowledge-graph search, call-chain lookup, typeahead, indexing) rather than an external MCP dependency referenced across the codebase.</p></li><li><p><strong>Worktree orchestration</strong> is heavy and native to the design: <code>EnterWorktree</code>/<code>ExitWorktree</code> as real operations, and fifteen-plus simultaneous live worktrees as the ordinary way of working, not a party trick.</p></li><li><p><strong>Ticketing</strong> is a dedicated subsystem spanning GitHub issues and JIRA, not an example command.</p></li><li><p>And then the crates with no claude-mpm analogue at all: PR and diff review, git analytics, a code intelligence layer, a TUI, an embedding daemon.</p></li><li><p>It also has an original (not meta) harness called trusty-code (still a work in progress but early indications are that it will perform similarly to open-code), and a personal agents harness called trusty-agents. Both leverage many of the same orchestration tools that trusty-mpm uses, which is why I&#8217;m including them in the package.</p></li></ul><p>On top of that sits a catalog of 37 specialist agents over five foundation layers, and a live skill catalog surfacing more than 190 skill names, triple claude-mpm&#8217;s 56. The prompting style changed too. A year ago you spent your prompt budget teaching the model how to be an agent. Now you spend it telling a competent agent what you want, because the harness supplies the how (in the form of memory, search, specs, and tickets).</p><p>Speaking of, ticket-driven development, the edge a year ago, is now the assumed baseline. The frontier moved up a level: many worktrees against a single repo or monorepo, several sessions of work in flight at once.</p><h2>The measured contrast</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wZXl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa668e6d3-c3a0-4473-934c-49b405983fd7_1600x688.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wZXl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa668e6d3-c3a0-4473-934c-49b405983fd7_1600x688.png 424w, https://substackcdn.com/image/fetch/$s_!wZXl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa668e6d3-c3a0-4473-934c-49b405983fd7_1600x688.png 848w, https://substackcdn.com/image/fetch/$s_!wZXl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa668e6d3-c3a0-4473-934c-49b405983fd7_1600x688.png 1272w, https://substackcdn.com/image/fetch/$s_!wZXl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa668e6d3-c3a0-4473-934c-49b405983fd7_1600x688.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wZXl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa668e6d3-c3a0-4473-934c-49b405983fd7_1600x688.png" width="1456" height="626" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a668e6d3-c3a0-4473-934c-49b405983fd7_1600x688.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:626,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:91029,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/208282371?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa668e6d3-c3a0-4473-934c-49b405983fd7_1600x688.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!wZXl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa668e6d3-c3a0-4473-934c-49b405983fd7_1600x688.png 424w, https://substackcdn.com/image/fetch/$s_!wZXl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa668e6d3-c3a0-4473-934c-49b405983fd7_1600x688.png 848w, https://substackcdn.com/image/fetch/$s_!wZXl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa668e6d3-c3a0-4473-934c-49b405983fd7_1600x688.png 1272w, https://substackcdn.com/image/fetch/$s_!wZXl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa668e6d3-c3a0-4473-934c-49b405983fd7_1600x688.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Two projects, one author, one subscription. Here&#8217;s the comparison. Effort and wall-clock framing are estimates. Commit counts and dates are exact. The productivity story they imply is inference.</p><p><em>Time to build a comparable system: claude-mpm ~12 months versus trusty-tools ~2 months, roughly a sixth of the wall-clock time.</em></p><p>A more capable system (25 crates with first-class memory, search, worktree orchestration, PR review, and git analytics) came together in about a sixth of the calendar time.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!mZkp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffffec223-9ac6-409c-8a8e-eaa391cea2b4_1600x736.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!mZkp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffffec223-9ac6-409c-8a8e-eaa391cea2b4_1600x736.png 424w, https://substackcdn.com/image/fetch/$s_!mZkp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffffec223-9ac6-409c-8a8e-eaa391cea2b4_1600x736.png 848w, https://substackcdn.com/image/fetch/$s_!mZkp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffffec223-9ac6-409c-8a8e-eaa391cea2b4_1600x736.png 1272w, https://substackcdn.com/image/fetch/$s_!mZkp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffffec223-9ac6-409c-8a8e-eaa391cea2b4_1600x736.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!mZkp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffffec223-9ac6-409c-8a8e-eaa391cea2b4_1600x736.png" width="1456" height="670" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fffec223-9ac6-409c-8a8e-eaa391cea2b4_1600x736.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:670,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:97055,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/208282371?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffffec223-9ac6-409c-8a8e-eaa391cea2b4_1600x736.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!mZkp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffffec223-9ac6-409c-8a8e-eaa391cea2b4_1600x736.png 424w, https://substackcdn.com/image/fetch/$s_!mZkp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffffec223-9ac6-409c-8a8e-eaa391cea2b4_1600x736.png 848w, https://substackcdn.com/image/fetch/$s_!mZkp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffffec223-9ac6-409c-8a8e-eaa391cea2b4_1600x736.png 1272w, https://substackcdn.com/image/fetch/$s_!mZkp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffffec223-9ac6-409c-8a8e-eaa391cea2b4_1600x736.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Commit velocity: claude-mpm ~380/month versus trusty-tools ~860/month, roughly 2.3x, though squash-merges across many worktrees likely undercount trusty-tools&#8217; true activity.</em></p><p>Velocity roughly 2.3x per month. A year of tech-lead work sharpened the exact skills this way of working rewards, driving a team, now a team of agents, and holding SDLC discipline as the branches multiply. And it shows in the git record, not just in my own say-so. Three signals I can actually measure. Commit messages that name a driving issue or decision rose from roughly one commit in ten to about nine in ten. The code now logs <em>why</em>, not only <em>what</em>. The eight near-identical agent files I copy-pasted early in claude-mpm collapsed into one composed base definition, a fix I started mid-project and carried further into trusty-tools. And architecture decision records went from essentially none to eighteen numbered ADRs plus a per-crate decisions taxonomy, a habit that matured across both projects rather than one trusty-tools invented. Developer growth and harness capability compound. They push the same direction. But even with that, a person going from good to better doesn&#8217;t buy you a sixth of the calendar time. The impressive part is still the tooling.</p><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tj7D!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19050f4c-efa9-4513-9c5a-0c97c3860b5b_1650x1110.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tj7D!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19050f4c-efa9-4513-9c5a-0c97c3860b5b_1650x1110.png 424w, https://substackcdn.com/image/fetch/$s_!tj7D!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19050f4c-efa9-4513-9c5a-0c97c3860b5b_1650x1110.png 848w, https://substackcdn.com/image/fetch/$s_!tj7D!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19050f4c-efa9-4513-9c5a-0c97c3860b5b_1650x1110.png 1272w, https://substackcdn.com/image/fetch/$s_!tj7D!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19050f4c-efa9-4513-9c5a-0c97c3860b5b_1650x1110.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tj7D!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19050f4c-efa9-4513-9c5a-0c97c3860b5b_1650x1110.png" width="1456" height="979" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/19050f4c-efa9-4513-9c5a-0c97c3860b5b_1650x1110.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:979,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:180011,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/208282371?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19050f4c-efa9-4513-9c5a-0c97c3860b5b_1650x1110.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!tj7D!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19050f4c-efa9-4513-9c5a-0c97c3860b5b_1650x1110.png 424w, https://substackcdn.com/image/fetch/$s_!tj7D!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19050f4c-efa9-4513-9c5a-0c97c3860b5b_1650x1110.png 848w, https://substackcdn.com/image/fetch/$s_!tj7D!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19050f4c-efa9-4513-9c5a-0c97c3860b5b_1650x1110.png 1272w, https://substackcdn.com/image/fetch/$s_!tj7D!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19050f4c-efa9-4513-9c5a-0c97c3860b5b_1650x1110.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>Capability checklist across three columns, a year ago, now, and a speculative year from now, for orchestration, memory, search, worktrees, skills, ticketing, and PR review.</em></p><p>Every row that read &#8220;you build this&#8221; a year ago reads &#8220;the harness provides this&#8221; now. Multi-agent orchestration: hand-built then, native flag now. Memory: thin and hand-maintained then, a knowledge-graph crate now. Search: external dependency then, first-class now. Worktrees: DIY shell commands then, a <code>--worktree</code> flag and fifteen live trees now.</p><p>The harness now enforces the workflow, not just supplies the parts. Spec-linked documentation was a prompt hint in claude-mpm, off by default. In trusty-tools it&#8217;s a build-blocking lint gate (<code>trusty-sld-lint</code>) wired into CI and pre-commit. Documentation that fails the build, not documentation you&#8217;re reminded to write. Ticket discipline hardened the same way: that rise in issue-referenced commits now sits inside a codified chain (spec &#8594; issue &#8594; PR-linked-to-issue &#8594; review gate &#8594; squash-merge), process-enforced, not yet a hard check that rejects an unlinked PR. And claude-mpm&#8217;s main branch required zero approvals and no passing checks, protected in name only. trusty-tools&#8217; main requires a review approval and six passing CI checks, with an LLM review pass (<code>trusty-review</code>) as a named gate. A year ago a disciplined developer chose these. Now the harness refuses to merge without a review and passing checks.</p><h2>A word on the subscription plan</h2><p>Claude Max was announced in April 2025 at <a href="https://support.claude.com/en/articles/11049741-what-is-the-max-plan">$100/month (Max 5x) and $200/month (Max 20x)</a>, with usage shared across chat and Code. Those price points appear to have held from then through mid-2026. The most visible change over that whole span was the five-hour rate limits for Claude Code, which Max subscribers draw on, <a href="https://www.anthropic.com/news/higher-limits-spacex">being doubled</a> in mid-2026. More headroom, not a different product tier.</p><p>So the input that stayed roughly constant was the money and the person. The input that changed was the model quality and the harness capability. When you hold the developer and the budget still and the output grows by this much, the variable that moved most would be the model and the harness (mostly the model, though in my experience good memory and search make a huge difference), even granting the developer improved and the language changed.</p><h2>A year from now</h2><p>The trajectory has a clear shape: tickets &#8594; worktrees &#8594; multi-session parallelism. Working in tickets was novel in mid-2025 and is common sense now. Native worktrees arrived in early 2026 and, within months, running many of them at once became my working model rather than a demo. Each step took a workflow that used to live in the developer&#8217;s head (tracking the units of work, isolating the parallel branches) and moved it into the tool. There are no hard adoption statistics for this progression. I&#8217;m describing a tooling timeline and a direction, not a measured majority practice. The tooling timeline itself is real, though: worktree support arriving across editors and agents through late 2025 and into 2026 follows the same curve.</p><p>The next frontier isn&#8217;t hard to name. If tickets became common sense and worktrees became my working model, the thing after worktrees is many concurrent sessions against one repo or monorepo. Enough parallel work in flight so that no human is tracking all of it, and the harness keeps the branches, the memory, and the review gates coherent. The <code>/fork</code> and <code>/subtask</code> primitives that landed in mid-2026, and subagents running in the background by default, are early moves in exactly that direction. The developer&#8217;s job shifts further from writing the steps toward specifying the outcome and reviewing the merge.</p><p>That&#8217;s not a prediction of artificial general anything. It&#8217;s the same migration, run one more turn: complexity leaving the code you write by hand and entering the tool you run. A year ago the hard part of the project was the scaffolding. This year it&#8217;s the harness. Next year it&#8217;s the coordination of more parallel work than one person can follow. So that one person gets pushed further up into the spec and the design.</p><h2>What a difference a year makes</h2><p>I built claude-mpm, and I&#8217;m very proud of it. For its moment it was a good answer to a real constraint. That&#8217;s the ordinary fate of scaffolding once the platform grows the feature underneath it, the design working as intended.</p><p>The measure of the year isn&#8217;t that I got better, though I hope I did. It&#8217;s that the same person, on the same max plan, with the same working habits, could build a materially more capable system in a fraction of the time, because the models got sharper and the harness absorbed the work that used to be mine to do. Twelve months of hand-built PM layer on one side, two months of composing first-class parts on the other. Same author. Same subscription.</p><p>What a difference a year makes.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and also writes about AI business at <a href="https://aipowerranking.com/">AI Power Ranking</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/p/its-the-harness-stupid">It&#8217;s the Harness, Stupid</a> &#8212; Why orchestration, not raw model quality, drives the spread in outcomes &#8212; the mechanism underneath this whole comparison.</p></li><li><p><a href="https://hyperdev.matsuoka.com/p/what-is-harness-engineering">What Is Harness Engineering? (And Do You Need to Learn It?)</a> &#8212; The durable skill under the tooling: designing the loop, not just the scaffolding.</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners.</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders.</p></li></ul>]]></content:encoded></item><item><title><![CDATA[What Is Harness Engineering?]]></title><description><![CDATA[And Do You Need to Learn It?]]></description><link>https://hyperdev.matsuoka.com/p/what-is-harness-engineering</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/what-is-harness-engineering</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Fri, 10 Jul 2026 11:30:51 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!yCIE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!yCIE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yCIE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png 424w, https://substackcdn.com/image/fetch/$s_!yCIE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png 848w, https://substackcdn.com/image/fetch/$s_!yCIE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png 1272w, https://substackcdn.com/image/fetch/$s_!yCIE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yCIE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png" width="1456" height="1092" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1092,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2396213,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/206397665?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!yCIE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png 424w, https://substackcdn.com/image/fetch/$s_!yCIE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png 848w, https://substackcdn.com/image/fetch/$s_!yCIE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png 1272w, https://substackcdn.com/image/fetch/$s_!yCIE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F08ade5e0-a528-4d25-ab18-f607fbc7ada9_1536x1152.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Enhancing the Golden Egg</figcaption></figure></div><p>An engineer at work asked me a good question last week. &#8220;What&#8217;s harness engineering, and do I actually need to learn it &#8212; or is it going to be obsolete by the time I do?&#8221; Fair question. The term is about five months old, half the people using it mean different things by it, and the people who build the most capable coding agents around keep going on record to say the thing you&#8217;d build will get absorbed into the next model.</p><p>So here is my answer, stated plainly, because I think most engineers are getting it wrong: harness engineering is the most important skill you can build right now beyond coding and architecture themselves. And the strongest evidence that it matters is that most engineers don&#8217;t yet believe they need it.</p><p>I&#8217;ve been writing harnesses for over a year, so treat that as a disclosed bias rather than a neutral survey. What follows argues against the smartest version of the other side &#8212; the case that harness work is disposable scaffolding the models will eat for breakfast &#8212; because that case is largely correct, and it still doesn&#8217;t touch the skill I&#8217;m talking about.</p><h2>TL;DR</h2><ul><li><p><strong>Harness engineering is real but young.</strong> The phrase traces to <a href="https://mitchellh.com/writing/my-ai-adoption-journey">Mitchell Hashimoto in February 2026</a> and an <a href="https://openai.com/index/harness-engineering/">OpenAI Codex case study days later</a>; <a href="https://martinfowler.com/articles/harness-engineering.html">Fowler and B&#246;ckeler</a> formalized it as &#8220;Agent = Model + Harness.&#8221; Report it as an emerging frame, not a settled discipline &#8212; and no &#8220;harness engineer&#8221; job title exists yet.</p></li><li><p><strong>The labs are right that the crutch layer shrinks.</strong> Anthropic&#8217;s Cat Wu says <a href="https://www.lennysnewsletter.com/p/how-anthropics-product-team-moves">&#8220;the models will eat your harness for breakfast&#8221;</a>; Boris Cherny says scaffolding gets &#8220;pushed into the model itself.&#8221; No one at Anthropic said &#8220;don&#8217;t learn it.&#8221;</p></li><li><p><strong>The benchmark fight lives at the crutch layer.</strong> Same model, different scaffold moves scores by <a href="https://arxiv.org/abs/2606.08529">up to 28 points on GAIA</a> and <a href="https://www.tbench.ai/leaderboard/terminal-bench/2.0">18.6 points on Terminal-Bench 2.0</a> &#8212; but <a href="https://agents-last-exam.org/blogs/harness-matters">Agents&#8217; Last Exam</a> shows model choice drives roughly 3x the spread of harness choice. Scores were beside the point.</p></li><li><p><strong>The skill lives at a layer the benchmarks don&#8217;t measure.</strong> Not a better crutch &#8212; a different unit of work: research &#8594; spec &#8594; ticket &#8594; build &#8594; PR &#8594; review &#8594; merge &#8594; deploy, driven by the harness. <a href="https://openai.com/index/harness-engineering/">OpenAI ran that loop with 3 engineers to ~1M lines and 1,500 PRs</a> with zero human-written code.</p></li><li><p><strong>Harness engineering is not IDE engineering.</strong> If you treat the harness as an extension of your old editor workflow, you cap your ceiling. The durable move is letting it drive and jumping in when needed.</p></li><li><p><strong>Yes, you need to learn it.</strong> The specific syntax is throwaway. That&#8217;s exactly why the skill matters &#8212; you&#8217;re learning to operate a new unit of work, not one vendor&#8217;s config file.</p></li></ul><h2>The question, stated bluntly</h2><p>Start with what &#8220;harness engineering&#8221; is even asking you to do, because the word carries two arguments at once and people talk past each other constantly.</p><p>I already made the case that orchestration beats model quality in <a href="https://hyperdev.matsuoka.com/p/its-the-harness-stupid">It&#8217;s the Harness, Stupid</a> back in April &#8212; same model, large spread in outcomes, the competitive edge moving from model superiority to ecosystem superiority. That piece defined what a harness <em>is</em> and showed that it dominates results. I&#8217;m not going to re-argue it. This is the follow-up question that piece left open: if the harness matters that much, is <em>building</em> one a skill worth learning &#8212; or a treadmill that resets every model release?</p><p>That distinction is the whole article. Because the answer the evidence points to is: parts of it reset every release, and the part that doesn&#8217;t is the part almost nobody is naming. Most of the public argument is being had about the parts that reset.</p><h2>What a harness actually is</h2><p>The cleanest definition comes from Martin Fowler and <a href="https://martinfowler.com/articles/harness-engineering.html">Birgitta B&#246;ckeler</a>: &#8220;the harness&#8221; is everything in an AI agent except the model itself. Agent = Model + Harness. They split it into guides that push instructions forward and sensors that feed results back, and they frame the whole practice as a specific form of context engineering. <a href="https://simonwillison.net/guides/agentic-engineering-patterns/how-coding-agents-work/">Simon Willison</a> puts it the same way from the other direction: a coding agent is software that acts as a harness for an LLM.</p><p>When I say &#8220;using a harness,&#8221; here&#8217;s the concrete inventory I mean:</p><ul><li><p><strong>Agents</strong> &#8212; the loop that calls the model and routes its tool calls, plus any sub-agents you dispatch work to.</p></li><li><p><strong>Skills</strong> &#8212; reusable capabilities you can invoke by name instead of re-explaining every time.</p></li><li><p><strong>Hooks</strong> &#8212; deterministic gates that fire on events: run the tests, block a commit, reformat on save.</p></li><li><p><strong>Workflow</strong> &#8212; the orchestrated path from research to deploy, and who (or what) drives each step.</p></li><li><p><strong>Harness-specific instructions</strong> &#8212; how <em>this</em> harness should behave, kept distinct from project-specific instructions about <em>this</em> codebase. Conflating those two is one of the most common configuration mistakes I see.</p></li><li><p><strong>Memory and search</strong> &#8212; what persists across sessions, and how the agent retrieves it.</p></li></ul><p>Internalize that inventory before you touch any tool, because every product arranges these pieces differently and calls them different things. <a href="https://huggingface.co/blog/agent-glossary">Hugging Face</a> is the one source I&#8217;ve found that formally separates the <em>scaffolding</em> (the behavior layer &#8212; prompts, tool descriptions, memory) from the <em>harness</em> proper (the execution layer that calls the model and decides when to stop), then notes that most products just call the whole bundle a harness. That ambiguity is real, not something to paper over: the term is contested, <a href="https://haverin.substack.com/p/what-is-harness-engineering-ai-hype">some practitioners call it an old idea in new packaging</a>, and Latent Space literally ran a piece titled <a href="https://www.latent.space/p/ainews-is-harness-engineering-real">&#8220;Is Harness Engineering Real?&#8221;</a>. I don&#8217;t want to oversell a five-month-old buzzword. I want to separate the durable part from the disposable part, and to do that I have to give the skeptics their strongest swing first.</p><h2>&#8220;The models will eat your harness for breakfast&#8221;</h2><p>Here&#8217;s the skeptical case in its own words, and it&#8217;s a good case.</p><p>Cat Wu, who heads product for Claude Code, has a line for it: <a href="https://www.lennysnewsletter.com/p/how-anthropics-product-team-moves">the models will eat your harness for breakfast</a>. Her team does a system-prompt audit on every new model and deletes the reminders the model no longer needs. Her example: a to-do enforcement tool built to stop Claude Code from overclaiming that a refactor was finished became dead weight once newer models completed multi-step refactors on their own.</p><p>Boris Cherny, who created Claude Code, says the same thing from the architecture side. In an <a href="https://every.to/podcast/transcript-how-to-use-claude-code-like-the-people-who-built-it">Every.to interview</a>, he described the tool as &#8220;the thinnest possible wrapper over the model&#8221; and said that as models advance, &#8220;stuff that used to be scaffolding... gets pushed into the model itself.&#8221; His team builds harness features they expect to delete: &#8220;we build most things... even if that means we&#8217;ll have to get rid of it in three months. If anything, we hope that we will get rid of it in three months.&#8221; Disposable on purpose. And at Sequoia&#8217;s AI Ascent this spring he extended it forward &#8212; prompt-injection defenses, static command verification, permission modes, human-in-the-loop gates would all become less critical, he argued, &#8220;because models will do the right thing themselves.&#8221;</p><p>This isn&#8217;t only an Anthropic view. It&#8217;s the <a href="http://www.incompleteideas.net/IncIdeas/BitterLesson.html">bitter lesson</a> applied to agents: building in how we think the work should be structured tends to lose, over time, to raw capability and scale. Han Lee makes the practitioner version <a href="https://leehanchung.github.io/blogs/2026/05/08/hidden-technical-debt-agent-harness/">bluntly</a>: &#8220;Almost all of it is going to dissolve into the next generation of models... build each production harness like you mean to replace it.&#8221; Tool wrappers dissolve because models read OpenAPI specs directly now. Elaborate memory layers collapse into &#8220;plain text in progress.md plus git log.&#8221;</p><p>Two concessions I&#8217;ll make up front, because the fair version of this argument requires them. First: nobody at Anthropic said &#8220;don&#8217;t learn it.&#8221; The claim that they did is a paraphrase &#8212; a real cluster of &#8220;the harness shrinks&#8221; statements from Wu and Cherny, compressed by repetition into something stronger than anyone actually said. Second, and this is the one that stings: even the benchmark evidence <em>for</em> harnesses says model choice usually wins. On <a href="https://agents-last-exam.org/blogs/harness-matters">Agents&#8217; Last Exam</a>, swapping models with the harness fixed produced an 18-point pass-rate spread; swapping harnesses with the model fixed produced about 6. &#8220;The model accounts for about 3x the pass-rate spread of the harness.&#8221; If your goal is a higher number on the leaderboard, buy the better model before you tune the scaffold.</p><p>So the skeptics have a benchmark, a bitter lesson, and the people who build the reference implementation all pointing the same way. If I stopped here, the answer to &#8220;do you need to learn it&#8221; would be &#8220;not really &#8212; wait for the next model.&#8221; I don&#8217;t stop here, because all of that is arguing about a layer I don&#8217;t mean.</p><h2>Why both are right &#8212; and why it doesn&#8217;t touch the real skill</h2><p>The reconciliation is that &#8220;harness&#8221; names two different things, and the argument above is entirely about the first one.</p><p>The first layer is scaffolding-as-crutch. A hook that reminds the model to actually run the tests. A tool wrapper that translates an API the model can&#8217;t yet read. A permission gate that catches a mistake the model still makes. Anthropic&#8217;s own framing nails why this dissolves: <a href="https://www.anthropic.com/engineering/harness-design-long-running-apps">&#8220;every component in a harness encodes an assumption about what the model can&#8217;t do on its own.&#8221;</a> When the model can suddenly do that thing, the component becomes dead weight &#8212; exactly Wu&#8217;s deleted to-do tool. This layer is <em>supposed</em> to shrink. Anthropic builds it disposable on purpose. The benchmark gaps live here too: the <a href="https://arxiv.org/abs/2606.08529">28-point GAIA swing</a> and the <a href="https://www.tbench.ai/leaderboard/terminal-bench/2.0">18.6-point Terminal-Bench spread</a> measure how much a scaffold props up a fixed model&#8217;s score. Prop-up value falls as the model climbs. That&#8217;s the whole skeptical case, and it&#8217;s correct.</p><p>The second layer is workflow orchestration. Letting the harness drive the entire loop &#8212; research &#8594; spec &#8594; ticket &#8594; build &#8594; iterate &#8594; PR &#8594; review &#8594; merge &#8594; deploy &#8212; as one continuous operation instead of a sequence of prompts you babysit. This layer does not dissolve into a better model, because a better model doesn&#8217;t decide what&#8217;s safe to run unattended, what the blast radius of an autonomous change is, or where a human judgment call has to sit. Better models make the loop <em>run better</em>. They don&#8217;t make the loop <em>design itself</em>.</p><p>Blake Crosley draws the same line and I think it&#8217;s the sharpest version: harness <em><a href="https://blakecrosley.com/blog/loops-win-where-verification-is-cheap">syntax</a></em><a href="https://blakecrosley.com/blog/loops-win-where-verification-is-cheap"> is ephemeral and gets absorbed, but </a><em><a href="https://blakecrosley.com/blog/loops-win-where-verification-is-cheap">verification judgment</a></em><a href="https://blakecrosley.com/blog/loops-win-where-verification-is-cheap"> is durable</a> &#8212; knowing what&#8217;s safe to run without watching, what the acceptable failure modes are, where the loop needs a gate. The config file you write today is throwaway. The judgment about how to structure autonomous work is not.</p><p>The killer piece of evidence sits in the phrase&#8217;s own origin story. When <a href="https://openai.com/index/harness-engineering/">OpenAI published its harness-engineering case study</a>, the headline number was three engineers producing roughly a million lines of code across 1,500 pull requests, with zero human-written code, by engineering the harness around Codex. Read that carefully. That is not a better crutch bolted onto a fixed workflow. It&#8217;s a different unit of work &#8212; the engineers stopped writing lines and started operating a loop. No model upgrade alone produces that shape of output, because the shape is a workflow-design decision, not a capability. Ryan Lopopolo&#8217;s summary of what changed is the tell: the only scarce resource left was synchronous human attention. That&#8217;s an orchestration problem, and no amount of model progress makes it go away.</p><p>This is why I can concede the entire benchmark argument without losing anything. Model choice beats harness choice on scores &#8212; sure, roughly 3x on Agents&#8217; Last Exam. But the fight over scores is being had at the crutch layer, and the skill I mean lives at the orchestration layer, which those benchmarks don&#8217;t even measure. Even the strongest model still needs someone who knows how to hand it a whole workflow instead of a single task.</p><h2>Harness engineering vs. IDE engineering</h2><p>Here&#8217;s the crux, and it&#8217;s where I think most engineers cap their own ceiling without noticing.</p><p>The biggest difference between harness engineering and IDE engineering is what the unit of work is. In the editor era &#8212; including the AI-autocomplete-in-your-editor era &#8212; the unit is a change you make, assisted. You&#8217;re still driving. The tool suggests, you accept, you commit. A good harness inverts that. The unit becomes an <em>outcome you delegate</em>: research through deploy, handled by the harness, with you supervising the loop rather than typing inside it.</p><p>If you approach a harness as an extension of your existing editor workflow &#8212; a faster autocomplete, a smarter pair &#8212; you&#8217;ll get some lift and you&#8217;ll hit a ceiling fast, because you&#8217;re still the bottleneck on every step. The results I&#8217;ve gotten that actually surprised me came from letting the harness drive the whole loop and jumping in only where my judgment was needed: at the spec, at the review, at the &#8220;is this safe to merge&#8221; gate. That maps exactly onto Crosley&#8217;s durable skill. You&#8217;re not writing less carefully. You&#8217;re spending your attention on the decisions that don&#8217;t delegate, and letting the loop own the ones that do.</p><p>This is a materially different workflow from anything I did before agents, and I say that as someone who&#8217;s lived through several supposed paradigm shifts that turned out to be the same job with new keybindings. This one isn&#8217;t. The muscle you build isn&#8217;t &#8220;prompt the model well.&#8221; It&#8217;s &#8220;decompose an outcome into a loop a machine can run mostly unattended, and know precisely where to stand in it.&#8221; Harrison Chase, who runs LangChain, frames harness engineering as <a href="https://venturebeat.com/orchestration/langchains-ceo-argues-that-better-models-alone-wont-get-your-ai-agent-to">an extension of context engineering</a>, and that lineage is right &#8212; context engineering is about a year old and settled, harness engineering is the newer, contested layer on top. But the operative verb changed. You&#8217;re not composing a context window. You&#8217;re operating a workflow.</p><h2>Do you need to learn it? Yes &#8212; here&#8217;s how</h2><p>Yes. The fact that it still feels optional is the problem, not a reason to wait.</p><p>Here&#8217;s the path I&#8217;d suggest, and it&#8217;s roughly the one I took.</p><p><strong>Start by understanding what using a harness means</strong> &#8212; the inventory from earlier: agents, skills, hooks, workflow, harness-specific versus project-specific instructions, memory, search. Not as vocabulary. As the actual pieces you&#8217;ll arrange. If those seven words don&#8217;t map to concrete settings you can change, start there before you touch a workflow.</p><p><strong>Try several, and configure them well.</strong> Don&#8217;t judge the category from one tool on defaults. Learn to tune Claude Code yourself until it performs, rather than running it out of the box and concluding the harness &#8220;doesn&#8217;t matter.&#8221; Try <a href="https://openai.com/index/harness-engineering/">Codex</a>, Gemini, Auggie, OpenCode. I built <a href="https://github.com/bobmatnyc/claude-mpm">claude-mpm</a>, so weight my enthusiasm for the multi-agent approach accordingly &#8212; the point isn&#8217;t which one wins, it&#8217;s that you can&#8217;t feel the shape of the skill from a single vendor&#8217;s config.</p><p><strong>Lean into the differences instead of smoothing them over.</strong> The instinct is to find the tool that feels most like your old editor and stop. The results live in the opposite direction &#8212; in the workflows that feel least familiar, where the harness drives and you supervise. Best results come from leaning into what&#8217;s different, not translating it back into what you already knew.</p><p><strong>Let it drive, and know where to stand.</strong> This is the whole skill in one sentence. Hand the loop the outcome, supervise at the gates your judgment actually owns, jump in when the blast radius or the ambiguity demands it. That standing-in-the-right-place instinct is what transfers across every tool and survives every model release.</p><p>And that&#8217;s the reframe I&#8217;ll close on. Everything about the <em>specific</em> harness you learn this quarter is throwaway. The config syntax, the exact hooks, the tool wrappers &#8212; Han Lee is right, the next model eats away at it. But that&#8217;s precisely why the skill is worth building, not a reason to skip it. You&#8217;re not learning one vendor&#8217;s settings file. You&#8217;re learning to operate a new unit of work &#8212; an autonomous loop from research to deploy &#8212; and that competence is the thing the model upgrades keep <em>raising the value of</em>, not erasing. The syntax is disposable. The judgment about how to run the loop is what compounds.</p><p>Most engineers will figure this out eventually, when the workflow shift is obvious in hindsight. The ones who figure it out now get a head start measured in the gap between &#8220;my editor got smarter&#8221; and &#8220;my unit of work changed.&#8221; I&#8217;d rather be early on that one.</p><h2>One layer up: loop engineering</h2><p>There&#8217;s a move past the harness, and it picked up a name while I was writing this.</p><p>The stack is starting to read like a ladder: prompt engineering, then context engineering, then harness engineering &#8212; and now loop engineering on top. Each rung stops being where you spend your attention once the rung below it gets good enough to trust. You quit hand-tuning prompts when context engineering settled into a roughly year-old, mostly-solved practice. The bet underneath loop engineering is the same shape one level up: once your harness is solid, you stop prompting the agent and start designing the loops that prompt it for you.</p><p>The term isn&#8217;t mine. <a href="https://addyosmani.com/blog/loop-engineering/">Addy Osmani formalized it in June</a>, and his definition is the one I&#8217;d hand someone first: &#8220;Loop engineering is replacing yourself as the person who prompts the agent. You design the system that does it instead.&#8221; He also collected two lines that land the shift faster than I can. Peter Steinberger&#8217;s version: you should be designing the loops that prompt your agents. And Boris Cherny &#8212; the same Cherny from the &#8220;eat your harness for breakfast&#8221; section &#8212; put it flatly: &#8220;I don&#8217;t prompt Claude anymore&#8230; my job is to write loops.&#8221; <a href="https://www.langchain.com/blog/the-art-of-loop-engineering">LangChain picked it up a week later</a>, and their framing is the practical one: you don&#8217;t build a loop, you stack them &#8212; an agent loop inside a verification loop inside an event-driven loop inside a hill-climbing loop, each one checking the one below it.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!KNlZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3fd96b6-ed6f-43c1-be83-2956e9647562_1536x1152.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!KNlZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3fd96b6-ed6f-43c1-be83-2956e9647562_1536x1152.png 424w, https://substackcdn.com/image/fetch/$s_!KNlZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3fd96b6-ed6f-43c1-be83-2956e9647562_1536x1152.png 848w, https://substackcdn.com/image/fetch/$s_!KNlZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3fd96b6-ed6f-43c1-be83-2956e9647562_1536x1152.png 1272w, https://substackcdn.com/image/fetch/$s_!KNlZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3fd96b6-ed6f-43c1-be83-2956e9647562_1536x1152.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!KNlZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3fd96b6-ed6f-43c1-be83-2956e9647562_1536x1152.png" width="1456" height="1092" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d3fd96b6-ed6f-43c1-be83-2956e9647562_1536x1152.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1092,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2248716,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/206397665?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3fd96b6-ed6f-43c1-be83-2956e9647562_1536x1152.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!KNlZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3fd96b6-ed6f-43c1-be83-2956e9647562_1536x1152.png 424w, https://substackcdn.com/image/fetch/$s_!KNlZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3fd96b6-ed6f-43c1-be83-2956e9647562_1536x1152.png 848w, https://substackcdn.com/image/fetch/$s_!KNlZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3fd96b6-ed6f-43c1-be83-2956e9647562_1536x1152.png 1272w, https://substackcdn.com/image/fetch/$s_!KNlZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd3fd96b6-ed6f-43c1-be83-2956e9647562_1536x1152.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Loop Engineering</figcaption></figure></div><p>Here&#8217;s the seam between the two layers, drawn plainly. Harness engineering is the discipline of building the scaffolding &#8212; the agents, hooks, skills, gates, and workflow from the inventory up top. Loop engineering is what you do once that scaffolding is good enough that you stop typing prompts and start directing repeatable cycles. The harness is what you build. The loop is what you run on it, again and again, with your attention moved up to which loops to run and when to trust them unattended. That&#8217;s the same &#8220;let it drive, know where to stand&#8221; instinct from a few paragraphs back, pushed one rung higher: now you&#8217;re not standing inside the loop at all &#8212; you&#8217;re choosing which loops get to run.</p><p>I&#8217;m not planting a flag here. The people already naming it are out ahead of me, and that&#8217;s the reason to point at it &#8212; this is where the harness work goes next, not a term I&#8217;m coining. If harness engineering closes the gap between &#8220;my editor got smarter&#8221; and &#8220;my unit of work changed,&#8221; loop engineering is what shows up on the far side of that gap, once the unit of work is a cycle you supervise instead of a task you run. I&#8217;m watching that one closely. And I&#8217;d start learning it before it feels obvious.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and also writes about AI business at <a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">AI Power Ranking</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/p/its-the-harness-stupid">It&#8217;s the Harness, Stupid</a> &#8212; The predecessor to this piece: same model, wide spread in outcomes, and why the competitive edge moved from model quality to orchestration. It defined what a harness is; this piece argues that building one is a skill worth learning.</p></li><li><p><a href="https://hyperdev.matsuoka.com/p/hyperdevs-three-golden-rules">HyperDev&#8217;s Three Golden Rules</a> &#8212; The working rules I keep coming back to for professional AI work, and the discipline that keeps a driven-by-the-harness loop from running off the rails.</p></li><li><p><a href="https://hyperdev.matsuoka.com/p/claude-sonnet-5-takes-the-default">Claude Sonnet 5 Takes the Default Driver Slot</a> &#8212; A concrete example of the crutch layer shrinking: adaptive thinking folds interleaved reasoning into the model, removing work harness authors used to do by hand.</p></li><li><p><a href="https://addyosmani.com/blog/loop-engineering/">Loop Engineering</a> &#8212; Addy Osmani&#8217;s June 2026 piece that named the layer above the harness: once the scaffolding holds, you stop prompting the agent and design the loops that prompt it. The forward edge of the arc this article traces.</p></li><li><p><a href="https://www.langchain.com/blog/the-art-of-loop-engineering">The Art of Loop Engineering</a> &#8212; LangChain&#8217;s treatment of loop engineering as stacked loops &#8212; agent, verification, event-driven, hill-climbing &#8212; each one checking the one beneath. The practitioner&#8217;s map of where harness work heads next.</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners, including the coding-agent leaderboards this piece leans on.</p></li></ul>]]></content:encoded></item><item><title><![CDATA[Coding's Great Depresh ]]></title><description><![CDATA[&#8212; and How to Find Your Energesh]]></description><link>https://hyperdev.matsuoka.com/p/codings-great-depresh</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/codings-great-depresh</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Mon, 29 Jun 2026 12:16:41 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!4RhI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4RhI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4RhI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!4RhI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!4RhI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!4RhI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4RhI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1503107,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/203766515?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4RhI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!4RhI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!4RhI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!4RhI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcc3ff9ce-c117-4182-9b49-d918a3630ad5_1024x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In 2019, comedian Gary Gulman released an HBO special called <em><a href="https://www.imdb.com/title/tt10409666/">The Great Depresh</a></em>. The title puns on the 1930s collapse, but the special is about Gulman&#8217;s own clinical depression &#8212; the hospitalization, the electroconvulsive therapy, the long climb back. He calls it &#8220;depresh&#8221; throughout. The diminutive is the point: &#8220;depression&#8221; was too heavy to say out loud for years, so he found a smaller word that let him talk about it at all. The special ends on a line he delivers like a weather report: &#8220;My depresh is in remish.&#8221;</p><p>Something adjacent is moving through the developer community, and most people don&#8217;t have a name for it yet. It isn&#8217;t burnout, and it isn&#8217;t fear of replacement, though that gets blamed for it. It&#8217;s a specific malaise in senior engineers who, by outward measures, are doing fine &#8212; shipping more, faster, with tools that work. They feel worse anyway. Call it the Coding Depresh.</p><p>The Depresh is real, it&#8217;s documented, and there&#8217;s a path out that doesn&#8217;t require pretending the loss isn&#8217;t a loss. I write from an unusual position. I spent eight years as CTO of TripAdvisor managing more than 560 engineers, mostly not writing code, and I missed it. I came back to hands-on coding in March 2025, entirely through AI tools &#8212; my re-entry and AI-assisted development are inseparable. I never had to grieve a pre-AI coding identity, because I didn&#8217;t have one to defend. That gives me a strange vantage point on the engineers who did.</p><h2>TL;DR</h2><ul><li><p>The Coding Depresh is identity disconfirmation, not job-loss fear: when the thing you built your professional self around gets automated, &#8220;who am I as a developer?&#8221; stops having an easy answer.</p></li><li><p>It&#8217;s documented. A 15-year veteran describes shipping an AI-built pipeline and feeling grief &#8220;for an identity.&#8221; Stack Overflow&#8217;s 2025 survey shows distrust in AI accuracy (46%) outrunning trust (33%) for the first time, even as usage climbed to 84%.</p></li><li><p>The most-cited evidence that AI slows experienced developers &#8212; METR&#8217;s 2025 trial showing them ~19% slower &#8212; has been retired by the same researchers, whose redesigned data now points the other way. Even the skeptics&#8217; own number moved.</p></li><li><p>A different camp &#8212; the Energesh &#8212; reports feeling more capable, not less. The fault line isn&#8217;t seniority or skill. It&#8217;s your answer to &#8220;who are you as a developer.&#8221;</p></li><li><p>The most useful question I&#8217;ve seen comes from an Anthropic engineer: &#8220;I thought that I really enjoyed writing code, and I think instead I actually just enjoy what I <em>get</em> out of writing code.&#8221; Where you land tells you where you live.</p></li><li><p>For engineers in the Depresh, and the CTOs managing them: the craft instinct that makes the tools feel wrong is the most valuable thing you own. It doesn&#8217;t have to die for you to adopt the tools.</p></li></ul><h2>The Depresh is real, and it&#8217;s documented</h2><p>Start with the most precise account I&#8217;ve found. <a href="https://medium.com/codetodeploy/ai-existential-dread-and-developer-ego-death-aef8bfc93214">George Violaris</a>, a developer with fifteen years&#8217; experience, wrote in March 2026 about shipping a data pipeline with AI assistance in an afternoon &#8212; work that would have taken him three days by hand. He expected pride. He got this instead:</p><blockquote><p>&#8220;That evening, I felt something I can only describe as grief. Not for a person. For an identity.&#8221;</p></blockquote><p>He goes on: &#8220;Three years ago, writing code wasn&#8217;t just what I did. It was who I was.&#8221; Then: &#8220;The ego death was real. &#8216;I am a person who writes excellent code&#8217; had to die.&#8221;</p><p>This is not a man worried about his next paycheck. He shipped the thing. The pipeline is in production. What broke wasn&#8217;t his employment; it was the relationship between his sense of self and the work. Most of the public conversation treats developer anxiety as a labor-market story &#8212; real, severe for junior developers, but separate, and not the one I&#8217;m writing about.</p><p>The Depresh I&#8217;m describing hits a different person: the experienced developer whose answer to &#8220;who are you?&#8221; was &#8220;I am the person who writes excellent code.&#8221; For that person, the tools don&#8217;t threaten the job first. They threaten the identity first.</p><p>The numbers underneath the mood are strange. <a href="https://survey.stackoverflow.co/2025/ai/">Stack Overflow&#8217;s 2025 Developer Survey</a>, the largest census of working developers we have, shows two lines crossing. Favorable sentiment toward AI tools fell from 77% in 2023 to 60% in 2025, while usage or intent to use rose from 70% to 84%. People are adopting tools they feel worse about. Trust in AI accuracy dropped to 33% against 46% who distrust it &#8212; the first year distrust outran trust &#8212; and the top frustration, cited by 66%, was &#8220;AI solutions that are almost right, but not quite.&#8221; The most experienced developers are the most skeptical: only 2.5% &#8220;highly trust&#8221; the output, and use it all day anyway. Working with something you don&#8217;t trust, that demands verification you can&#8217;t delegate &#8212; that&#8217;s the exhaustion of vigilance without resolution.</p><p>Then there&#8217;s the METR study &#8212; less for the number most people quoted than for what happened to it. In 2025 the nonprofit METR <a href="https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/">ran a controlled trial</a> with experienced open-source developers and found them 19% slower with AI assistants while they believed they&#8217;d been roughly 20% faster. That became the most-cited evidence that the tools don&#8217;t pay off. Then in February 2026 the same team <a href="https://metr.org/blog/2026-02-24-uplift-update/">retired it</a>, writing that the original no longer reflects &#8220;the current impact of AI models on open-source developer productivity&#8221;; their redesigned measurement runs the other way, an estimated 18% speedup for the same returning developers. The intervals are wide, so the magnitude is soft &#8212; but the direction reversed, and why it had to be rebuilt is the tell: 30 to 50% of developers refused to submit tasks they expected AI to speed up, and many refused to work without AI at all. The measurement broke down because people won&#8217;t give up the tools. I read the original not as &#8220;the tools are bad&#8221; but as disorientation: when your instincts and the stopwatch disagree, then a year later the stopwatch reverses, your competence comes unmoored.</p><h2>There is another camp, and it&#8217;s also documented</h2><p>Here is what keeps the Depresh from being the whole story. A different group of engineers is having close to the opposite experience, and they&#8217;re not naive optimists or vendors.</p><p>Andrej Karpathy <a href="https://x.com/karpathy/status/1886192184808149383">named the moment in February 2025</a>: &#8220;a new kind of coding I call vibe coding, where you fully give in to the vibes, embrace exponentials, and forget that the code even exists.&#8221; He was describing a feeling, not a methodology. Simon Willison later coined <a href="https://simonwillison.net/2025/Oct/7/vibe-engineering/">&#8220;vibe engineering&#8221;</a> for the responsible professional version &#8212; engineers using these tools deliberately to accelerate real work.</p><p>The output side is loud. Pieter Levels built a <a href="https://levels.io/fly-pieter-com-vibecoded-flight-simulator">browser flight simulator</a> with no game-development background and reached around $1M in annualized revenue in 17 days. <a href="https://newsletter.pragmaticengineer.com/p/building-claude-code-with-boris-cherny">Boris Cherny</a>, who leads Claude Code at Anthropic, runs five parallel instances and ships 20 to 30 pull requests a day: &#8220;once there is a good plan, it will one-shot the implementation almost every time.&#8221;</p><p>This is where my own story sits. When I came back, the mechanics of typing code line by line weren&#8217;t part of my working identity &#8212; I&#8217;d been away from the keyboard for years. What I found waiting was the part I&#8217;d missed: building things, deciding what should exist, watching it take shape, fixing what&#8217;s wrong with it. The joy of building and the mechanics of writing code were never the same thing.</p><p>I can be concrete, because I&#8217;ve done it in the open. Since March 2025 I&#8217;ve shipped real, full-cycle software through these tools &#8212; not snippets, complete projects. The flagship is <a href="https://github.com/bobmatnyc/claude-mpm">claude-mpm</a>, a multi-agent orchestration platform for Claude with a real user base. Most of 2025 was Python; then in 2026 I taught myself Rust to build the trusty-* ecosystem &#8212; code search, memory, PR review, orchestration &#8212; something I wouldn&#8217;t have attempted by hand while running an org. Architecture, review, releases, the full loop, at a volume I couldn&#8217;t reach typing every line. Most of it is public on GitHub under <a href="https://github.com/bobmatnyc">bobmatnyc</a>, so the claim is checkable. I&#8217;m not theorizing about the Energesh. I&#8217;m living in it.</p><h2>What actually separates the two camps</h2><p>Not seniority. Not raw skill. Not the domain you work in, though each shades the experience. The fault line is the answer you&#8217;d have given, before any of this started, to a single question: who are you as a developer? There&#8217;s an older version of the same split.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xDfL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbba5a639-91a8-4072-939c-05ed95492b5b_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xDfL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbba5a639-91a8-4072-939c-05ed95492b5b_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!xDfL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbba5a639-91a8-4072-939c-05ed95492b5b_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!xDfL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbba5a639-91a8-4072-939c-05ed95492b5b_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!xDfL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbba5a639-91a8-4072-939c-05ed95492b5b_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xDfL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbba5a639-91a8-4072-939c-05ed95492b5b_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bba5a639-91a8-4072-939c-05ed95492b5b_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1319015,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/203766515?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbba5a639-91a8-4072-939c-05ed95492b5b_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!xDfL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbba5a639-91a8-4072-939c-05ed95492b5b_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!xDfL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbba5a639-91a8-4072-939c-05ed95492b5b_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!xDfL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbba5a639-91a8-4072-939c-05ed95492b5b_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!xDfL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbba5a639-91a8-4072-939c-05ed95492b5b_1024x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In the early nineteenth century the <em><a href="https://en.wikipedia.org/wiki/Canut_revolts">canuts</a></em> were the silk weavers of Lyon, clustered in the Croix-Rousse district, working intricate patterns by hand. It was master-craft work, identity-defining &#8212; the skill lived in the fingers. Then Joseph Marie Jacquard demonstrated a loom that wove those same complex patterns automatically, driven by punched cards, doing in a single pass what had taken the weaver and an assistant by hand. The canut whose sense of self lived in the act of weaving, in the tight and skilled handwork itself, was displaced at the identity layer. The canut whose relationship was with the silk itself, the finished cloth, still had somewhere to stand. The fabric still got made. More of it, in fact. If your identity comes from the tight weaving rather than from seeing the cloth made, you are a Canut.</p><p>Coding has the same shape. If your answer to &#8220;who are you&#8221; was inseparable from &#8220;I am the person who writes the code&#8221; &#8212; the line-by-line authorship, the mechanical understanding of every decision, the craft of it &#8212; then AI arrives at the identity layer, not the job layer. Violaris&#8217;s grief is the correct, proportional response; it scales with how much of himself he invested. If your answer was closer to &#8220;I am the person who builds the thing that exists at the end,&#8221; the same tools feel like a gain. You were always pointed at the cloth, not the weave. The result just got cheaper to make.</p><p>There&#8217;s a quiet irony in the comparison. Jacquard&#8217;s punched cards are a direct ancestor of the computer &#8212; they ran on into Babbage&#8217;s Analytical Engine and later Hollerith&#8217;s tabulators and the IBM punch card. The mechanism that displaced the weaver became the machine the coder works on, now automating a layer of the coder&#8217;s craft too. Same lineage, one more turn.</p><p>The cleanest articulation comes from an Anthropic engineer, quoted in <a href="https://www.anthropic.com/research/how-ai-is-transforming-work-at-anthropic">Anthropic&#8217;s own internal research on AI-assisted work</a>:</p><blockquote><p>&#8220;I thought that I really enjoyed writing code, and I think instead I actually just enjoy what I <em>get</em> out of writing code.&#8221;</p></blockquote><p>That is the skeleton key. Most developers in the Depresh believe they love writing code, and they&#8217;re not wrong, exactly &#8212; but the two halves of that sentence always came bundled. You couldn&#8217;t get the output without the authorship. AI unbundles them. Do you love the writing, or what the writing gets you? Where you land tells you which camp you&#8217;re in, and it&#8217;s not always where you assumed.</p><p><a href="https://newsletter.kentbeck.com/p/augmented-coding-beyond-the-vibes">Kent Beck&#8217;s vocabulary</a> dissolves a false worry. He distinguishes <em>augmented coding</em> from <em>vibe coding</em>. In vibe coding you don&#8217;t care about the code, only the behavior. In augmented coding you care deeply about the code, its complexity, its tests &#8212; &#8220;it&#8217;s just that I don&#8217;t type much of that code.&#8221; The Energesh isn&#8217;t about lowering your standards: the taste and architectural judgment that took decades to build are still doing the real work.</p><p>David Heinemeier Hansson gives the sharpest version of the craft objection. DHH spent most of 2025 resisting agent-first coding, and his reason was a craft reason. On <a href="https://lexfridman.com/dhh-david-heinemeier-hansson">Lex Fridman&#8217;s podcast</a> he said the joy is to type the code himself; he keeps AI in a separate window because letting it drive made him &#8220;feel competence draining out of [his] fingers.&#8221; He has since <a href="https://newsletter.pragmaticengineer.com/p/dhhs-new-way-of-writing-code">switched to an agent-first workflow</a>, on the logic that the tools finally met his standard, not that his standard moved. His values stayed put; what changed was his assessment of whether the tools honored them. That&#8217;s the template: the resister and the convert are the same man with the same principles.</p><p>Which is why the craft instinct deserves defending, not demolishing. The tools feel wrong partly because you&#8217;re judging them against that instinct, not just against output. Its firing isn&#8217;t a malfunction &#8212; it&#8217;s the sharpest instrument you own, calibrated over years, telling you when something is almost right but not quite, the same complaint 66% named in the Stack Overflow data. The mistake is concluding it has to be put down for the tools to be picked up. It doesn&#8217;t.</p><h2>The path from Depresh to Energesh</h2><p>This part is for the reader sitting in the Depresh right now. It isn&#8217;t a pep talk, and I&#8217;m not going to tell you the feeling is irrational, because it isn&#8217;t.</p><p><strong>Grieve first.</strong> The ego death Violaris describes is real, and you&#8217;re allowed to mourn it. The craft identity you spent years building had value &#8212; it shipped real systems and earned you a career. Skipping the grief doesn&#8217;t work. The developers I&#8217;ve watched leap straight to enthusiasm tend to adopt the tools resentfully and use them badly, half-hoping they&#8217;ll fail. Let the loss be a loss before you look for what&#8217;s on the other side.</p><p><strong>Then ask the real question.</strong> The Anthropic engineer&#8217;s version: do I love writing code, or what I <em>get</em> from writing code? You&#8217;ve probably never had to answer it, because the two were never separable before. They are now. The answer isn&#8217;t a verdict on your worth as an engineer. It&#8217;s a map of where you actually live, and either answer is fine &#8212; but you can&#8217;t find the path until you know which is true for you.</p><p><strong>Learn harness engineering.</strong> The on-ramp for a senior engineer goes up a level, not down. Vibe coding pulls you toward the model&#8217;s altitude; this pulls you above it. The harness is the tooling layer around the model &#8212; orchestration, context management, verification scaffolding, the agent configuration that decides what the model sees and whether you can trust what comes back. Engineering that layer rewards the judgment you spent years building: systems thinking, architecture, the discipline of making an unreliable component dependable. My own claude-mpm and trusty-* tools are harness engineering and nothing else. So skip the vibe-coded toy app. Pick something real, build the harness that drives it, and watch how it feels &#8212; whether <em>you</em> feel more like yourself, or less.</p><p><strong>Reframe from author of code to author of outcomes.</strong> Your craft instinct &#8212; what good architecture looks like, what breaks at scale, when a design is quietly wrong &#8212; isn&#8217;t going away, and the model doesn&#8217;t have it. The model is fast, capable, judgment-free. You are slow by comparison and you have taste. The job becomes directing the thing and knowing whether what came back is right. Not a demotion from engineer to button-pusher. It&#8217;s the part of the work that was always hardest to teach, now occupying most of your day.</p><p>A word for the CTOs reading this: you have a version of this problem you may not have named. The Depresh shows up in your metrics before your one-on-ones: review cycles stretching out, code-quality variance widening, your most experienced engineers going quiet in design reviews. The path for them is the path for an individual &#8212; don&#8217;t push adoption before you&#8217;ve made room for the grief. And understand who you&#8217;re dealing with: the engineers who feel the Depresh most sharply are frequently your best ones, who invested most in the craft. Their standards are the feature, not the bug. Burn that instinct down to force faster adoption and you&#8217;ll get the adoption and lose the standards.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1RiK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe76ca282-db68-4682-8a30-3d327f8a6e50_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1RiK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe76ca282-db68-4682-8a30-3d327f8a6e50_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!1RiK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe76ca282-db68-4682-8a30-3d327f8a6e50_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!1RiK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe76ca282-db68-4682-8a30-3d327f8a6e50_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!1RiK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe76ca282-db68-4682-8a30-3d327f8a6e50_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1RiK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe76ca282-db68-4682-8a30-3d327f8a6e50_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e76ca282-db68-4682-8a30-3d327f8a6e50_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1317916,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/203766515?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe76ca282-db68-4682-8a30-3d327f8a6e50_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!1RiK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe76ca282-db68-4682-8a30-3d327f8a6e50_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!1RiK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe76ca282-db68-4682-8a30-3d327f8a6e50_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!1RiK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe76ca282-db68-4682-8a30-3d327f8a6e50_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!1RiK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe76ca282-db68-4682-8a30-3d327f8a6e50_1024x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Where this leaves you</h2><p>Gulman didn&#8217;t end his special by announcing he was cured. He said his depresh was in remish &#8212; smaller word, smaller claim, a thing managed rather than defeated. That&#8217;s about the right register for where the industry is.</p><p>The craft you built was real, and so is the grief if you&#8217;re feeling it. But the thing you actually loved &#8212; if it turns out to be the building, the deciding, the watching something work that didn&#8217;t exist this morning &#8212; that part never depended on you typing every character yourself. The canut whose love was the cloth still had cloth to make. You can find out which one you are. Most developers go a whole career without having to ask. You get to.</p><p>Your depresh can be in remish. The energesh is on the other side of one unsparing question, and you already know how to ask those. You&#8217;ve been debugging your own assumptions for years.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and writes about AI business trends at <a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">AI Power Ranking</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/p/tide-has-turned-senior-devs">The Tide Has Turned: Senior Developers Are Finally Adopting AI Tools</a> &#8212; Why the holdouts changed their minds, and what shifted to make it happen</p></li><li><p><a href="https://hyperdev.matsuoka.com/p/era-of-the-leader-practitioner">The Era of the Leader/Practitioner</a> &#8212; How AI tools made the hybrid leader-who-builds role viable again</p></li><li><p><a href="https://hyperdev.matsuoka.com/p/weve-turned-a-corner">We&#8217;ve Turned a Corner</a> &#8212; On the shift from skepticism to working practice</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul>]]></content:encoded></item><item><title><![CDATA[The Math Doesn't Work (Yet): Inside the AI Profitability Problem]]></title><description><![CDATA[Why Scaling Doesn't Lead To Profitability]]></description><link>https://hyperdev.matsuoka.com/p/the-math-doesnt-work-yet-inside-the</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/the-math-doesnt-work-yet-inside-the</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Wed, 10 Jun 2026 11:31:58 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!XBpX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!XBpX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!XBpX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png 424w, https://substackcdn.com/image/fetch/$s_!XBpX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png 848w, https://substackcdn.com/image/fetch/$s_!XBpX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png 1272w, https://substackcdn.com/image/fetch/$s_!XBpX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!XBpX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png" width="1195" height="896" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:896,&quot;width&quot;:1195,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1847405,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/201357245?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!XBpX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png 424w, https://substackcdn.com/image/fetch/$s_!XBpX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png 848w, https://substackcdn.com/image/fetch/$s_!XBpX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png 1272w, https://substackcdn.com/image/fetch/$s_!XBpX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61162dd8-5fc3-4493-856f-338f6a569f95_1195x896.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h1>The Math Doesn&#8217;t Work (Yet): Inside the AI Profitability Problem</h1><p>OpenAI&#8217;s own projections show losses getting bigger as revenue gets bigger. Leaked investor documents reported across WSJ, Fortune, and The Information put the company at roughly <a href="https://fortune.com/2025/11/12/openai-cash-burn-rate-annual-losses-2028-profitable-2030-financial-documents/">$74 billion in operating losses in 2028</a> &#8212; on roughly $100 billion in projected revenue. That pairing is the headline. Not a smaller loss as scale arrives. A loss that grows faster than the top line.</p><p>That single relationship is the whole story. We tend to model AI companies as software businesses that will eventually grow into their cost structure the way SaaS companies did before them. The numbers say otherwise. These are capital-intensive infrastructure plays wearing software-company clothing, and the unit economics underneath them run in the opposite direction from the SaaS playbook most of us internalized over the last fifteen years.</p><p>A caveat before the numbers, because it matters for what&#8217;s below: neither OpenAI nor Anthropic publishes audited financials. Most figures here come from leaked investor decks, run-rate annualizations the companies announce in funding rounds, or SEC filings made by their cloud partners. The uncertainty is part of the analysis, not a footnote to it. When a specific timeline or quarterly figure couldn&#8217;t survive cross-checking against primary sources, I left it out.</p><h2>TL;DR</h2><ul><li><p>OpenAI&#8217;s leaked projections show operating losses <em>widening</em> as revenue grows &#8212; roughly <a href="https://fortune.com/2025/11/12/openai-cash-burn-rate-annual-losses-2028-profitable-2030-financial-documents/">$74B in losses on ~$100B revenue projected for 2028, with cumulative cash burn near $115B through 2029</a>.</p></li><li><p>In 2025 OpenAI spent about <a href="https://fortune.com/2025/11/12/openai-cash-burn-rate-annual-losses-2028-profitable-2030-financial-documents/">$1.69 for every dollar of revenue (~$9B net loss on ~$13B revenue)</a>, per leaked documents confirmed by multiple outlets.</p></li><li><p>Anthropic&#8217;s revenue trajectory is steep &#8212; roughly $1B annualized at the end of 2024 to a figure announced in the tens of billions by mid-2026 &#8212; with <a href="https://sacra.com/c/anthropic/">Claude Code alone reported at $2.5B annualized by February 2026</a>.</p></li><li><p><a href="https://www.investing.com/analysis/the-ai-token-pricing-crisis-behind-openai-and-anthropics-revenue-race-200680777">Inference token prices fell about 75% in a year</a>. Selling more AI makes the per-unit economics cheaper, which makes revenue growth harder, not easier.</p></li><li><p><a href="https://www.techtimes.com/articles/317542/20260601/ai-agent-economics-token-tax-locks-gross-margins-30-points-below-saas-baseline.htm">AI-native gross margins sit near 45% versus 75&#8211;85% for mature SaaS</a> &#8212; a structural gap of 23&#8211;33 points no company has yet closed.</p></li><li><p>No company has a verified path to profitability. Every specific breakeven-by-year claim I tried to confirm fell apart under scrutiny.</p></li></ul><h2>Two Companies, Two Shapes</h2><p>OpenAI ended 2025 at <a href="https://futurumgroup.com/insights/ai-capex-2026-the-690b-infrastructure-sprint/">roughly $20 billion in annualized revenue, a figure CFO Sarah Friar has stated directly</a>. That is a large business by any normal measure. It is also a business that, in the same year, spent about <a href="https://fortune.com/2025/11/12/openai-cash-burn-rate-annual-losses-2028-profitable-2030-financial-documents/">$1.69 for every dollar it took in &#8212; somewhere around a $9 billion net loss on roughly $13 billion in recognized revenue</a>, according to leaked documents that WSJ, Fortune, and The Information each reported. The company is majority funded by Microsoft, runs its compute primarily on Azure, and its strategy is scale-first: build the largest models, capture the most usage, and trust that revenue follows the curve.</p><p>Anthropic&#8217;s shape is different. Its revenue trajectory is steeper and more concentrated. The company grew from roughly $1 billion annualized at the end of 2024 to a figure it announced in the tens of billions by mid-2026 &#8212; <a href="https://sacra.com/c/anthropic/">the number it cited in its Series H materials</a>. I&#8217;m deliberately not pinning an exact figure to a month here, because the company has grown several-fold inside a single five-month window and any precise number is stale by the time you read it. What&#8217;s verifiable is the slope, and the slope is steep.</p><p>The more interesting detail is the concentration of value. Claude Code, one product, was reported at <a href="https://sacra.com/c/anthropic/">roughly $2.5 billion annualized by February 2026</a>. A single coding tool driving that much of a company&#8217;s run rate tells you something about where the margin-bearing demand actually lives. Anthropic is backed by Amazon (over $8 billion invested) and Google (over $2 billion), and it has <a href="https://techcrunch.com/2026/05/20/anthropic-will-pay-xai-1-25-billion-per-month-for-compute/">committed to spend more than $100 billion on AWS over ten years</a>, with roughly 1 GW of Trainium capacity targeted by the end of 2026. Two large companies funding it; one of them also selling it the silicon it runs on.</p><h2>Why Compute Is the Problem</h2><p>What separates these companies from every SaaS business you&#8217;ve evaluated: they don&#8217;t own their infrastructure. They rent it, at hyperscaler rates, from the same companies that fund them.</p><p>OpenAI&#8217;s Azure spend reportedly ran around <a href="https://www.theregister.com/2025/11/12/openai_spending_report/">$3.7 billion in 2024 and roughly $8.7 billion across the first three quarters of 2025</a>. Treat those numbers as medium-confidence &#8212; they come from leaked documents, and Microsoft pushed back that the figures &#8220;aren&#8217;t quite right.&#8221; But the direction is consistent with everything else: compute cost is the dominant line item, it&#8217;s largely fixed, and it grows with usage.</p><p>Anthropic&#8217;s arrangement produced one of the stranger details in modern enterprise finance I&#8217;ve read in some time. The company signed a deal for compute from Colossus 1 &#8212; Elon Musk&#8217;s Memphis data center, operated by xAI &#8212; at <a href="https://techcrunch.com/2026/05/20/anthropic-will-pay-xai-1-25-billion-per-month-for-compute/">roughly $1.25 billion per month for 300 MW of capacity, running through May 2029</a>. That&#8217;s not from a leak. It surfaced in SpaceX&#8217;s S-1 SEC filing and was confirmed by CNBC, Axios, and Data Center Dynamics, with a potential total value above $40 billion. There&#8217;s a 90-day mutual cancellation clause, so the headline total overstates the firm commitment. Still: Anthropic &#8212; funded by Google and Amazon &#8212; is paying Elon Musk&#8217;s company more than a billion dollars a month for compute. The AI capital world is stranger from the inside than the press releases suggest.</p><p>Zoom out and the renter problem gets sharper. Hyperscaler capex for 2026 is projected at <a href="https://futurumgroup.com/insights/ai-capex-2026-the-690b-infrastructure-sprint/">$660&#8211;690 billion</a>. Against that, OpenAI&#8217;s $20 billion ARR is roughly 3% of a single year&#8217;s data-center buildout by its suppliers. The companies selling AI applications are small tenants in an infrastructure market they don&#8217;t control and can&#8217;t currently price against.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!GBF-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88636760-5028-4c28-a47a-a9be5e14e2e7_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!GBF-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88636760-5028-4c28-a47a-a9be5e14e2e7_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!GBF-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88636760-5028-4c28-a47a-a9be5e14e2e7_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!GBF-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88636760-5028-4c28-a47a-a9be5e14e2e7_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!GBF-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88636760-5028-4c28-a47a-a9be5e14e2e7_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!GBF-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88636760-5028-4c28-a47a-a9be5e14e2e7_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/88636760-5028-4c28-a47a-a9be5e14e2e7_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1143520,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/201357245?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88636760-5028-4c28-a47a-a9be5e14e2e7_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!GBF-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88636760-5028-4c28-a47a-a9be5e14e2e7_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!GBF-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88636760-5028-4c28-a47a-a9be5e14e2e7_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!GBF-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88636760-5028-4c28-a47a-a9be5e14e2e7_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!GBF-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F88636760-5028-4c28-a47a-a9be5e14e2e7_1024x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The Unit Economics Trap</h2><p>This story inverts an instinct most of us trust.</p><p>In normal software, even in the Cloud, scale is your friend. Marginal cost trends toward zero, gross margin climbs as you grow, and a mature SaaS business lands at <a href="https://www.techtimes.com/articles/317542/20260601/ai-agent-economics-token-tax-locks-gross-margins-30-points-below-saas-baseline.htm">75&#8211;85% gross margin</a> because serving the millionth customer costs almost nothing. Volume is the cure.</p><p>Inference doesn&#8217;t behave that way. Every token generated costs compute &#8212; real, metered, non-zero compute &#8212; so the marginal cost of serving usage stays stubbornly positive. And the price you can charge for that token is collapsing. Enterprise transaction data from Ramp shows <a href="https://www.investing.com/analysis/the-ai-token-pricing-crisis-behind-openai-and-anthropics-revenue-race-200680777">inference prices falling roughly 75% in a single year, from around $10 per million tokens to around $2.50</a>. Capability per dollar is improving fast, which is good for buyers and brutal for sellers, because it means the revenue you booked at last year&#8217;s prices reprices downward while your compute bill does not.</p><p>Put the two forces together and you get a squeeze that worsens with success. The better you are at selling inference, the more usage you drive; the more usage you drive, the more the per-unit price falls; the more it falls, the harder it is to grow revenue against a compute bill that scales with that same usage. Volume isn&#8217;t the cure here. Under these dynamics it&#8217;s part of the disease.</p><p>The survey data puts a number on how far this world sits from SaaS. ICONIQ Capital polled about 300 software executives and <a href="https://www.techtimes.com/articles/317542/20260601/ai-agent-economics-token-tax-locks-gross-margins-30-points-below-saas-baseline.htm">pegged AI-native gross margins at 41% in 2024, 45% in 2025, and a projected 52% in 2026</a>. Improving &#8212; but starting from a base 30-plus points below mature SaaS, and closing the distance slowly. A 52% gross margin is a respectable hardware business. It is a structurally difficult software business, especially one still spending heavily to grow.</p><h2>Two Different Bets</h2><p>OpenAI and Anthropic are running different experiments on how you eventually close that gap. Neither has been validated.</p><p>OpenAI&#8217;s bet is scale and breadth. Build the broadest platform, capture consumer and enterprise and API demand simultaneously, and assume that at sufficient scale you gain pricing power over compute, model-efficiency gains compound, and the revenue base grows fast enough to absorb the fixed cost. The leaked projections embody the risk in this bet: they show <a href="https://fortune.com/2025/11/12/openai-cash-burn-rate-annual-losses-2028-profitable-2030-financial-documents/">losses </a><em><a href="https://fortune.com/2025/11/12/openai-cash-burn-rate-annual-losses-2028-profitable-2030-financial-documents/">widening</a></em><a href="https://fortune.com/2025/11/12/openai-cash-burn-rate-annual-losses-2028-profitable-2030-financial-documents/"> through 2028 even as revenue approaches $100 billion, with cumulative cash burn near $115 billion through 2029</a>. The theory requires the curve to bend after the window we can currently see.</p><p>Anthropic&#8217;s bet is narrower and more product-led. Find a wedge where the work is valuable enough that buyers tolerate real prices, prove the margin there, and expand outward. Claude Code is that wedge made concrete &#8212; <a href="https://sacra.com/c/anthropic/">$2.5 billion annualized from developers</a> who pay because the output is worth more than the inference under it. Coding, agents, and enterprise automation are higher-value work than chat, and higher-value work supports prices that don&#8217;t immediately erode under token deflation. The risk: revenue concentration in a single product line, and an infrastructure bill &#8212; AWS commitments plus the xAI deal &#8212; that&#8217;s enormous relative to a company still proving the model.</p><p>Two theories of the same problem. Scale your way past the margin gap, or find work valuable enough that the gap doesn&#8217;t bind. We don&#8217;t yet have the data to say either works.</p><h2>What Would It Actually Take</h2><p>I&#8217;ll skip the timeline speculation &#8212; every specific breakeven-by-year claim I tried to verify died on contact with the sources. The structural requirements are clearer than the dates.</p><p>Three things have to move. First, gross margins have to climb from the mid-40s toward something defensible &#8212; call it 60-plus &#8212; and stay there while volume grows. That means model-efficiency gains (cheaper inference per unit of capability) have to outrun price deflation, rather than getting passed straight through to buyers as lower prices.</p><p>Second, these companies need pricing power over compute, which today they don&#8217;t have. At current scale they&#8217;re tenants. The open question is at what ARR a vendor becomes large enough to negotiate compute like a partner instead of a customer &#8212; or to build its own. Anthropic&#8217;s Trainium commitment and OpenAI&#8217;s various infrastructure moves are bets that vertical integration eventually changes the cost equation. That&#8217;s unproven, and it&#8217;s expensive in the interim.</p><p>Third, the product mix has to keep shifting toward work that resists deflation &#8212; enterprise agents, coding tools, automation that&#8217;s measured against labor cost rather than against the falling price of a token. Claude Code is the cleanest evidence that this category exists and that buyers will pay. Whether it&#8217;s a large enough share of total volume to lift blended margins across a company is the question that decides the whole thing.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!I6FB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b6f837-7c33-4d8b-9fce-0ef475dfb701_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!I6FB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b6f837-7c33-4d8b-9fce-0ef475dfb701_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!I6FB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b6f837-7c33-4d8b-9fce-0ef475dfb701_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!I6FB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b6f837-7c33-4d8b-9fce-0ef475dfb701_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!I6FB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b6f837-7c33-4d8b-9fce-0ef475dfb701_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!I6FB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b6f837-7c33-4d8b-9fce-0ef475dfb701_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b6b6f837-7c33-4d8b-9fce-0ef475dfb701_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1409865,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/201357245?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b6f837-7c33-4d8b-9fce-0ef475dfb701_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!I6FB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b6f837-7c33-4d8b-9fce-0ef475dfb701_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!I6FB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b6f837-7c33-4d8b-9fce-0ef475dfb701_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!I6FB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b6f837-7c33-4d8b-9fce-0ef475dfb701_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!I6FB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6b6f837-7c33-4d8b-9fce-0ef475dfb701_1024x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The Open Questions</h2><p>What we don&#8217;t know outweighs what we do.</p><p>We don&#8217;t know the actual, current gross margins at either company &#8212; only the <a href="https://www.techtimes.com/articles/317542/20260601/ai-agent-economics-token-tax-locks-gross-margins-30-points-below-saas-baseline.htm">AI-native sector estimate of roughly 45%</a>. Neither company publishes the number that would settle the argument. We don&#8217;t know whether Anthropic&#8217;s stack of compute commitments, the decade-long AWS deal alongside the month-by-month xAI arrangement, creates structural tension or healthy redundancy. A company hedging across three infrastructure providers is either diversifying supply or revealing that no single supplier can meet its demand. We don&#8217;t know the ARR threshold at which compute pricing becomes negotiable, which is the hinge the entire margin story turns on.</p><p>And there&#8217;s the strategic risk that has no clean precedent: your infrastructure supplier is also your competitor. Microsoft ships Copilot. Amazon and Google both build models that compete with Anthropic&#8217;s. xAI builds Grok. Every dollar these companies pay for compute partly funds a rival&#8217;s model program. In normal software you don&#8217;t hand your gross margin to the company trying to beat you. Here it&#8217;s the default arrangement.</p><p>So the real question isn&#8217;t whether the AI labs are growing. They obviously are, faster than almost any companies in history. The question is whether revenue growth and margin improvement are the same trend or opposing ones. The SaaS era trained a generation of operators to believe that scale fixes economics. The leaked numbers describe a business where scale, so far, makes the loss bigger. Until one of these companies publishes a gross margin that shows the curve bending, that&#8217;s the math we have. And the math doesn&#8217;t work yet.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and writes about AI-powered engineering at HyperDev.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/p/the-first-70-era">The First 70% Era</a> &#8212; Where agentic AI delivers value and where it stops, and why the higher-value work resists token deflation</p></li><li><p><a href="https://hyperdev.matsuoka.com/p/ai-and-the-rise-of-the-hyperdev">AI and the Rise of the Hyperdev</a> &#8212; Why developers pay real money for AI tooling, the demand side of the margin story</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul>]]></content:encoded></item><item><title><![CDATA[What’s Old Is New Again]]></title><description><![CDATA[Nine classic SDLC practices that AI finally makes practical]]></description><link>https://hyperdev.matsuoka.com/p/whats-old-is-new-again</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/whats-old-is-new-again</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Wed, 03 Jun 2026 11:31:06 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!FRWK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b2c67b8-b815-4f06-a401-6d57b2a9d884_873x576.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!FRWK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b2c67b8-b815-4f06-a401-6d57b2a9d884_873x576.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!FRWK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b2c67b8-b815-4f06-a401-6d57b2a9d884_873x576.png 424w, https://substackcdn.com/image/fetch/$s_!FRWK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b2c67b8-b815-4f06-a401-6d57b2a9d884_873x576.png 848w, https://substackcdn.com/image/fetch/$s_!FRWK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b2c67b8-b815-4f06-a401-6d57b2a9d884_873x576.png 1272w, https://substackcdn.com/image/fetch/$s_!FRWK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b2c67b8-b815-4f06-a401-6d57b2a9d884_873x576.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!FRWK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b2c67b8-b815-4f06-a401-6d57b2a9d884_873x576.png" width="873" height="576" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0b2c67b8-b815-4f06-a401-6d57b2a9d884_873x576.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:576,&quot;width&quot;:873,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1435244,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/200236698?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F43b89068-cf8e-4789-95f1-e357c61076b0_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!FRWK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b2c67b8-b815-4f06-a401-6d57b2a9d884_873x576.png 424w, https://substackcdn.com/image/fetch/$s_!FRWK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b2c67b8-b815-4f06-a401-6d57b2a9d884_873x576.png 848w, https://substackcdn.com/image/fetch/$s_!FRWK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b2c67b8-b815-4f06-a401-6d57b2a9d884_873x576.png 1272w, https://substackcdn.com/image/fetch/$s_!FRWK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0b2c67b8-b815-4f06-a401-6d57b2a9d884_873x576.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Most of the best ideas in software engineering aren&#8217;t new. They&#8217;ve been written up in books, argued over at conferences, taught in every &#8220;best practices&#8221; deck since the late 1990s. And most teams quietly don&#8217;t do them.</p><p>Not because anyone thinks they&#8217;re wrong. Test-driven development, design by contract, architecture decision records, mutation testing &#8212; ask a room of senior engineers whether these are good ideas and you&#8217;ll get nods. Ask the same room who practices them consistently under deadline pressure and the hands stay down. I&#8217;ve been in that room for twenty-five years, on both sides of the question. I&#8217;ve also been the engineering leader who let those practices slip because shipping the feature mattered more this quarter.</p><p>There&#8217;s a single economic reason these practices lose. The upfront cost is high, the payoff is real but distant, and human attention is the binding constraint. Write the test before the code, document the decision, specify the invariant &#8212; every one of those is a tax you pay now against a benefit you collect later, maybe, if the project lives long enough. Under deadline pressure, that&#8217;s a losing trade for a human. So we skip it, ship, and pay the interest later in bugs and confusion. Call it the impatience tax.</p><p>Agents don&#8217;t pay that tax. They have infinite patience for upfront rigor and roughly zero marginal cost for the tedious work that rigor demands. Writing a thorough test suite for code that doesn&#8217;t exist yet is psychologically brutal for a person and completely fine for a model. That single shift &#8212; the cost of patience going to zero &#8212; quietly inverts the economics of a whole list of practices we knew were right and gave up on anyway.</p><p>This isn&#8217;t a piece about what AI makes <em>possible</em>. Lots of things are possible. It&#8217;s about a narrower, more useful question: which disciplines did we already agree were correct, fight about for decades, and abandon for reasons that no longer hold?</p><p>Here are nine.</p><h2>TL;DR</h2><ul><li><p>These nine practices share one structure: high upfront cost, distant payoff. Human attention is the constraint that kills them under deadline pressure.</p></li><li><p>AI removes the constraint. A failing test is the clearest prompt you can hand an agent; a spec is its input; an ADR is its context. The discipline becomes the interface.</p></li><li><p>TDD, design by contract, and property-based testing turn from &#8220;things we should do&#8221; into the most effective way to <em>constrain</em> agent behavior and prevent hallucinated correctness.</p></li><li><p>Documentation, ADRs, and living docs get a bilateral ROI: agents generate them from code, and they make agents far more effective in your codebase.</p></li><li><p>The catch is real. A 2025 METR randomized trial found experienced developers were about 19% <em>slower</em> with AI assistance. These practices pay off only when AI is used with discipline, not as autocomplete.</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!hmow!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff117aee0-1765-4975-a0c5-dd7bdddbe34d_1024x431.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!hmow!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff117aee0-1765-4975-a0c5-dd7bdddbe34d_1024x431.png 424w, https://substackcdn.com/image/fetch/$s_!hmow!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff117aee0-1765-4975-a0c5-dd7bdddbe34d_1024x431.png 848w, https://substackcdn.com/image/fetch/$s_!hmow!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff117aee0-1765-4975-a0c5-dd7bdddbe34d_1024x431.png 1272w, https://substackcdn.com/image/fetch/$s_!hmow!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff117aee0-1765-4975-a0c5-dd7bdddbe34d_1024x431.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!hmow!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff117aee0-1765-4975-a0c5-dd7bdddbe34d_1024x431.png" width="1024" height="431" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f117aee0-1765-4975-a0c5-dd7bdddbe34d_1024x431.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:431,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1083571,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/200236698?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5f4ba76c-18f5-413a-bf35-a56dc7861a00_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!hmow!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff117aee0-1765-4975-a0c5-dd7bdddbe34d_1024x431.png 424w, https://substackcdn.com/image/fetch/$s_!hmow!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff117aee0-1765-4975-a0c5-dd7bdddbe34d_1024x431.png 848w, https://substackcdn.com/image/fetch/$s_!hmow!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff117aee0-1765-4975-a0c5-dd7bdddbe34d_1024x431.png 1272w, https://substackcdn.com/image/fetch/$s_!hmow!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff117aee0-1765-4975-a0c5-dd7bdddbe34d_1024x431.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>1. Test-Driven Development</h2><p>Start with the practice most teams abandoned first.</p><p>Writing tests before code was always the theoretically superior move. It forces you to define the interface before you build behind it, catches bugs at the moment of definition instead of during integration, and leaves behind a living specification of what the code is supposed to do. Kent Beck made the case decades ago and the case held up.</p><p>Almost nobody did it consistently. The reason is psychological, not technical. Writing detailed tests for code that doesn&#8217;t exist yet, while a deadline breathes on your neck, feels like building scaffolding for a house you haven&#8217;t designed. Your brain screams at you to just write the function. So you write the function, promise yourself you&#8217;ll add tests after, and &#8212; well. You know how that goes.</p><p>Now flip the perspective. To an agent, a failing test isn&#8217;t scaffolding. It&#8217;s the clearest possible specification of intent you can provide. &#8220;Make this pass, don&#8217;t break anything else&#8221; is an unambiguous, machine-checkable instruction, which is exactly what a probabilistic system needs to stay honest. The test suite becomes a guardrail that prevents the most dangerous failure mode in AI-assisted coding: confident, plausible, wrong. Hallucinated correctness dies against a red bar.</p><p>TDD went from the discipline most teams couldn&#8217;t sustain to one of the best tools we have for bounding what an agent is allowed to claim it did. Same practice. Opposite economics.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!a841!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb60534e-05b5-4873-ab47-586bb7ec68bd_1024x670.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!a841!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb60534e-05b5-4873-ab47-586bb7ec68bd_1024x670.png 424w, https://substackcdn.com/image/fetch/$s_!a841!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb60534e-05b5-4873-ab47-586bb7ec68bd_1024x670.png 848w, https://substackcdn.com/image/fetch/$s_!a841!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb60534e-05b5-4873-ab47-586bb7ec68bd_1024x670.png 1272w, https://substackcdn.com/image/fetch/$s_!a841!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb60534e-05b5-4873-ab47-586bb7ec68bd_1024x670.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!a841!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb60534e-05b5-4873-ab47-586bb7ec68bd_1024x670.png" width="1024" height="670" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fb60534e-05b5-4873-ab47-586bb7ec68bd_1024x670.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:670,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1640038,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/200236698?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a4cda20-f557-4918-88d1-529fec00bdee_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!a841!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb60534e-05b5-4873-ab47-586bb7ec68bd_1024x670.png 424w, https://substackcdn.com/image/fetch/$s_!a841!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb60534e-05b5-4873-ab47-586bb7ec68bd_1024x670.png 848w, https://substackcdn.com/image/fetch/$s_!a841!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb60534e-05b5-4873-ab47-586bb7ec68bd_1024x670.png 1272w, https://substackcdn.com/image/fetch/$s_!a841!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb60534e-05b5-4873-ab47-586bb7ec68bd_1024x670.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>2. Spec-Driven Development and Design by Contract</h2><p>Bertrand Meyer formalized design by contract in the 1980s and built it into the Eiffel language: specify preconditions, postconditions, and invariants, then let the implementation follow from the contract.</p><p>The idea was sound and the adoption was thin, for one stubborn economic reason: the contract only pays off if someone <em>else</em> writes the implementation from it. If you&#8217;re writing both the spec and the code, the spec is overhead &#8212; you already know what you meant. The contract&#8217;s value lives in the handoff, and for most of software history there was no cheap handoff to hand it to.</p><p>Now there is. You write the contract; the agent writes the implementation from it. Spec-driven development stops being a documentation chore and becomes the actual control surface for delegation. The spec is the part requiring human judgment about <em>what</em> the system should do. The implementation &#8212; the part that used to eat the hours &#8212; is the part you delegate. Meyer&#8217;s economics finally close, forty years late, because the missing party in the transaction showed up.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!P-wa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5605bf-95f5-490e-b318-7334ff70ae87_1024x487.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!P-wa!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5605bf-95f5-490e-b318-7334ff70ae87_1024x487.png 424w, https://substackcdn.com/image/fetch/$s_!P-wa!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5605bf-95f5-490e-b318-7334ff70ae87_1024x487.png 848w, https://substackcdn.com/image/fetch/$s_!P-wa!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5605bf-95f5-490e-b318-7334ff70ae87_1024x487.png 1272w, https://substackcdn.com/image/fetch/$s_!P-wa!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5605bf-95f5-490e-b318-7334ff70ae87_1024x487.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!P-wa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5605bf-95f5-490e-b318-7334ff70ae87_1024x487.png" width="1024" height="487" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ee5605bf-95f5-490e-b318-7334ff70ae87_1024x487.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:487,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1370470,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/200236698?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F82d6c7c5-50c0-4b34-9a46-0700c4bf1e82_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!P-wa!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5605bf-95f5-490e-b318-7334ff70ae87_1024x487.png 424w, https://substackcdn.com/image/fetch/$s_!P-wa!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5605bf-95f5-490e-b318-7334ff70ae87_1024x487.png 848w, https://substackcdn.com/image/fetch/$s_!P-wa!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5605bf-95f5-490e-b318-7334ff70ae87_1024x487.png 1272w, https://substackcdn.com/image/fetch/$s_!P-wa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee5605bf-95f5-490e-b318-7334ff70ae87_1024x487.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>3. Architecture Decision Records</h2><p>Why does the codebase look like this? Why Postgres and not Dynamo, why this queue, why the weird module boundary that everyone trips over?</p><p>ADRs were the right answer to that question &#8212; a short dated record of each significant decision, the context, and the alternatives rejected. The discipline almost never held. Same shape as everything else here: the cost is immediate (stop, write the thing) and the value accrues slowly, mostly to some future engineer who isn&#8217;t in the room yet.</p><p>Two things flipped at once, which makes this one more interesting than the rest. First, agents can generate ADRs from an existing codebase &#8212; read the git history, the dependency choices, the structure, and reconstruct the decisions that produced them. The retroactive cost of documentation drops toward zero. Second, and this is the part people miss: existing ADRs dramatically improve what an agent can do <em>in</em> your codebase. An agent that can read why you chose eventual consistency won&#8217;t keep proposing changes that assume strong consistency.</p><p>So the ROI went bilateral. Agents help you write ADRs, and ADRs help agents help you. The practice that used to only cost now pays on both ends.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cfey!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F256e6ffa-ac04-401b-bd1d-276e0dc51f21_1024x445.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cfey!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F256e6ffa-ac04-401b-bd1d-276e0dc51f21_1024x445.png 424w, https://substackcdn.com/image/fetch/$s_!cfey!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F256e6ffa-ac04-401b-bd1d-276e0dc51f21_1024x445.png 848w, https://substackcdn.com/image/fetch/$s_!cfey!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F256e6ffa-ac04-401b-bd1d-276e0dc51f21_1024x445.png 1272w, https://substackcdn.com/image/fetch/$s_!cfey!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F256e6ffa-ac04-401b-bd1d-276e0dc51f21_1024x445.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cfey!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F256e6ffa-ac04-401b-bd1d-276e0dc51f21_1024x445.png" width="1024" height="445" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/256e6ffa-ac04-401b-bd1d-276e0dc51f21_1024x445.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:445,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1079171,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/200236698?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0eca5b97-c8ef-45f0-97de-eb23cc9969c2_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!cfey!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F256e6ffa-ac04-401b-bd1d-276e0dc51f21_1024x445.png 424w, https://substackcdn.com/image/fetch/$s_!cfey!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F256e6ffa-ac04-401b-bd1d-276e0dc51f21_1024x445.png 848w, https://substackcdn.com/image/fetch/$s_!cfey!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F256e6ffa-ac04-401b-bd1d-276e0dc51f21_1024x445.png 1272w, https://substackcdn.com/image/fetch/$s_!cfey!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F256e6ffa-ac04-401b-bd1d-276e0dc51f21_1024x445.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>4. Continuous Code Review</h2><p>&#8220;Catch issues early&#8221; has sat on every best-practices list since Extreme Programming put continuous review on the map. The advice was never controversial. The bottleneck was always the same: human reviewer attention is finite, expensive, and easily exhausted. So review collapsed into batch PR review &#8212; a tired engineer reading a 600-line diff on a Friday afternoon, approving most of it on faith.</p><p>AI review on every commit &#8212; not batched at the PR boundary, but running as code lands &#8212; is moving from aspirational toward baseline: increasingly the default on teams that have wired it in, though not yet universal. The marginal cost of a careful read went to nearly nothing, and the read happens while the context is still warm.</p><p>But this one comes with an emergent problem worth naming, because it bites teams that adopt the tooling without rethinking the model. PRs are getting larger and arriving faster under AI-assisted development. An agent can produce a 2,000-line change in an afternoon. If your review model still routes everything through a human approver at the end, that human is now the rate limiter, drowning in volume they didn&#8217;t generate and can&#8217;t realistically read. AI review on every commit is part of the answer. The harder part is restructuring <em>what</em> the human reviews &#8212; architecture, intent, the decisions a model shouldn&#8217;t make alone &#8212; and letting the machine handle line-level correctness continuously. Adopt the tool without rethinking the workflow and you&#8217;ve just built a faster way to overwhelm your best reviewer.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!SXYP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2eb6d6b9-de10-452f-8b64-7bb0bcba1041_1024x456.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!SXYP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2eb6d6b9-de10-452f-8b64-7bb0bcba1041_1024x456.png 424w, https://substackcdn.com/image/fetch/$s_!SXYP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2eb6d6b9-de10-452f-8b64-7bb0bcba1041_1024x456.png 848w, https://substackcdn.com/image/fetch/$s_!SXYP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2eb6d6b9-de10-452f-8b64-7bb0bcba1041_1024x456.png 1272w, https://substackcdn.com/image/fetch/$s_!SXYP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2eb6d6b9-de10-452f-8b64-7bb0bcba1041_1024x456.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!SXYP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2eb6d6b9-de10-452f-8b64-7bb0bcba1041_1024x456.png" width="1024" height="456" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2eb6d6b9-de10-452f-8b64-7bb0bcba1041_1024x456.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:456,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1098227,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/200236698?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ce64d7a-b849-4c7e-b523-2565960ee237_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!SXYP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2eb6d6b9-de10-452f-8b64-7bb0bcba1041_1024x456.png 424w, https://substackcdn.com/image/fetch/$s_!SXYP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2eb6d6b9-de10-452f-8b64-7bb0bcba1041_1024x456.png 848w, https://substackcdn.com/image/fetch/$s_!SXYP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2eb6d6b9-de10-452f-8b64-7bb0bcba1041_1024x456.png 1272w, https://substackcdn.com/image/fetch/$s_!SXYP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2eb6d6b9-de10-452f-8b64-7bb0bcba1041_1024x456.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>5. Pair Programming</h2><p>Pairing always looked expensive in the most obvious way: two engineers, one task, double the salary against a single unit of output. That intuition was wrong on the numbers &#8212; the measured overhead from the pair-programming studies was closer to 15%, often recovered through fewer defects &#8212; but the 2&#215; gut feeling is what drove the decisions. The benefits were real &#8212; knowledge transfer, real-time review, fewer dumb mistakes &#8212; but the perceived math meant most teams reserved it for critical paths, gnarly bugs, or onboarding a new hire. A luxury, rationed.</p><p>The pair is now a human and an agent, and it&#8217;s available to every engineer continuously, not rationed to the critical path. The knowledge-transfer benefit generalizes &#8212; the agent can explain unfamiliar parts of the codebase on demand. The real-time-review benefit generalizes &#8212; a second set of eyes on every line, every time, without scheduling two calendars. The economics that made pairing a rationed luxury simply don&#8217;t apply when one half of the pair has near-zero marginal cost.</p><p>Worth a caveat: agent-as-pair is genuinely good at the review and explanation half of pairing, and weaker at the part where a human partner pushes back on a bad <em>design</em> before you&#8217;ve written a line. You still need humans pairing with humans for that. But the day-to-day, line-by-line version of pairing just became free, and that&#8217;s most of what pairing was for.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!DHDP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e5b59c-a9a7-44ae-8725-1333f8e08687_1024x592.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!DHDP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e5b59c-a9a7-44ae-8725-1333f8e08687_1024x592.png 424w, https://substackcdn.com/image/fetch/$s_!DHDP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e5b59c-a9a7-44ae-8725-1333f8e08687_1024x592.png 848w, https://substackcdn.com/image/fetch/$s_!DHDP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e5b59c-a9a7-44ae-8725-1333f8e08687_1024x592.png 1272w, https://substackcdn.com/image/fetch/$s_!DHDP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e5b59c-a9a7-44ae-8725-1333f8e08687_1024x592.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!DHDP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e5b59c-a9a7-44ae-8725-1333f8e08687_1024x592.png" width="1024" height="592" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/18e5b59c-a9a7-44ae-8725-1333f8e08687_1024x592.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:592,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1439433,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/200236698?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F611b1613-4024-4d10-8e69-2b2a1936432d_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!DHDP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e5b59c-a9a7-44ae-8725-1333f8e08687_1024x592.png 424w, https://substackcdn.com/image/fetch/$s_!DHDP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e5b59c-a9a7-44ae-8725-1333f8e08687_1024x592.png 848w, https://substackcdn.com/image/fetch/$s_!DHDP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e5b59c-a9a7-44ae-8725-1333f8e08687_1024x592.png 1272w, https://substackcdn.com/image/fetch/$s_!DHDP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F18e5b59c-a9a7-44ae-8725-1333f8e08687_1024x592.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>6. Mutation Testing</h2><p>Coverage numbers lie, and most engineers know it. Eighty percent line coverage tells you eighty percent of your lines got executed by a test &#8212; not that any of those tests would <em>notice</em> if the behavior broke. Mutation testing is the honest measure: it deliberately introduces bugs (flip a comparison, drop a line, change a constant) and checks whether your test suite catches them. If a mutant survives, you have a test that runs code without actually validating it.</p><p>Mutation testing was always the gold standard and almost never run continuously, for one reason: it&#8217;s computationally expensive. You&#8217;re effectively running your whole suite many times over, once per mutation. On a real codebase that&#8217;s brutal. So it lived in research papers and the occasional heroic CI job that someone eventually disabled for being too slow.</p><p>That constraint is mostly gone &#8212; compute is cheap and parallel, and we got more comfortable spending it. And the practice arrived right when we suddenly need it most. AI-generated tests have a characteristic failure mode: they drift toward coverage metrics without meaningful assertions. The model writes a test that calls the function, exercises the path, and asserts almost nothing of substance &#8212; green checkmark, zero protection. Coverage looks great. Mutation testing is the thing that catches exactly that. It&#8217;s the verification layer for a verification layer, and it matters more now than when it was invented, because now a machine is writing the tests.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!EY9I!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70ad3dac-f7cb-48da-9d0a-a6098d28cbe6_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!EY9I!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70ad3dac-f7cb-48da-9d0a-a6098d28cbe6_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!EY9I!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70ad3dac-f7cb-48da-9d0a-a6098d28cbe6_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!EY9I!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70ad3dac-f7cb-48da-9d0a-a6098d28cbe6_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!EY9I!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70ad3dac-f7cb-48da-9d0a-a6098d28cbe6_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!EY9I!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70ad3dac-f7cb-48da-9d0a-a6098d28cbe6_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/70ad3dac-f7cb-48da-9d0a-a6098d28cbe6_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1482204,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/200236698?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70ad3dac-f7cb-48da-9d0a-a6098d28cbe6_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!EY9I!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70ad3dac-f7cb-48da-9d0a-a6098d28cbe6_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!EY9I!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70ad3dac-f7cb-48da-9d0a-a6098d28cbe6_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!EY9I!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70ad3dac-f7cb-48da-9d0a-a6098d28cbe6_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!EY9I!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F70ad3dac-f7cb-48da-9d0a-a6098d28cbe6_1024x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>7. Living Documentation</h2><p>Documentation was always supposed to be a first-class artifact. It almost never was, and the reason is by now familiar: writing docs is tedious, and the penalty for stale docs accrues slowly and lands on someone else. So docs rotted. Every team has a wiki that&#8217;s a graveyard of half-true pages from two reorgs ago.</p><p>AI changes both halves of the equation at once. It generates docs from code, so the writing cost drops. And it <em>consumes</em> docs as context, so the docs earn their keep immediately &#8212; a well-documented codebase is a measurably more useful codebase for an agent working in it. The ROI is immediate and bilateral, same structure as ADRs.</p><p>There&#8217;s a quiet shift hiding in there. Documentation used to be written for humans who&#8217;d mostly never read it. Now it&#8217;s also written for the agent that will read it on every task, which means stale docs don&#8217;t just confuse a future engineer &#8212; they actively degrade your tooling today. The feedback loop tightened from months to minutes. That&#8217;s the kind of change that actually moves behavior, because the cost of skipping it shows up now instead of later.</p><h2>8. Runbook Generation from Incidents</h2><p>On-call always leaned too hard on tribal knowledge. The person who knows why the payment service wedges at 3 a.m. is asleep, on vacation, or left the company last spring. Writing a runbook after each incident was obviously the right move and reliably the thing nobody did, because the incident was <em>over</em> and everyone wanted to go back to bed.</p><p>Incidents become runbooks automatically now. The agent has the incident timeline, the chat transcript, the commands that resolved it, the postmortem &#8212; and it can turn that into a structured runbook while the details are fresh, without asking an exhausted engineer to relive the night. The cost that used to fall right when motivation was lowest now falls on a system that doesn&#8217;t get tired or resentful.</p><p>I&#8217;d treat the generated runbook as a draft a human still signs off on, not gospel. But &#8220;imperfect draft, reviewed in five minutes&#8221; beats &#8220;blank page nobody ever fills in,&#8221; and that was always the real competition.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!dM5v!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c754a4-9b5b-424f-b008-30dbcb217887_1024x693.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!dM5v!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c754a4-9b5b-424f-b008-30dbcb217887_1024x693.png 424w, https://substackcdn.com/image/fetch/$s_!dM5v!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c754a4-9b5b-424f-b008-30dbcb217887_1024x693.png 848w, https://substackcdn.com/image/fetch/$s_!dM5v!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c754a4-9b5b-424f-b008-30dbcb217887_1024x693.png 1272w, https://substackcdn.com/image/fetch/$s_!dM5v!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c754a4-9b5b-424f-b008-30dbcb217887_1024x693.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!dM5v!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c754a4-9b5b-424f-b008-30dbcb217887_1024x693.png" width="1024" height="693" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d0c754a4-9b5b-424f-b008-30dbcb217887_1024x693.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:693,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1615996,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/200236698?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3f83774-28c6-4b46-addd-83efad33390b_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!dM5v!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c754a4-9b5b-424f-b008-30dbcb217887_1024x693.png 424w, https://substackcdn.com/image/fetch/$s_!dM5v!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c754a4-9b5b-424f-b008-30dbcb217887_1024x693.png 848w, https://substackcdn.com/image/fetch/$s_!dM5v!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c754a4-9b5b-424f-b008-30dbcb217887_1024x693.png 1272w, https://substackcdn.com/image/fetch/$s_!dM5v!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c754a4-9b5b-424f-b008-30dbcb217887_1024x693.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>9. Property-Based Testing</h2><p>Example-based tests check the cases you thought of. Property-based testing is stronger: you specify the <em>invariants</em> a system must always satisfy &#8212; reversing a list twice returns the original, a serialized-then-deserialized object equals the original, the account balance never goes negative &#8212; and the framework generates hundreds of adversarial inputs trying to break them. QuickCheck pioneered the approach; it finds the edge cases you&#8217;d never have written by hand.</p><p>It never went mainstream outside a few communities, and the bottleneck wasn&#8217;t tooling &#8212; good property-based libraries exist for most languages. The bottleneck was writing good property specifications. Identifying the right invariants requires deep domain reasoning: you have to understand the system well enough to state what must <em>always</em> be true, which is harder than writing a few example cases. Most engineers, under pressure, defaulted to the easier thing.</p><p>This is where AI helps in a way that&#8217;s less obvious than &#8220;it writes the code.&#8221; A model can generate property suites from a spec, and &#8212; more usefully &#8212; it can reason about <em>what invariants a system should satisfy</em> in the first place, surfacing properties you hadn&#8217;t articulated. That&#8217;s the expensive, judgment-heavy part it actually offloads. Combined with mutation testing to keep the generated properties honest, you get a testing approach that was always more powerful than example-based testing and was always too expensive in human reasoning to adopt widely.</p><h2>The Catch</h2><p>I&#8217;d be selling you something if I stopped there, and the data won&#8217;t let me.</p><p>In July 2025, METR ran a randomized controlled trial with experienced open-source developers working on real tasks in repositories they knew well. The developers expected AI assistance to speed them up. It slowed them down &#8212; by roughly 19%. METR&#8217;s February 2026 follow-up found that gap narrowing, and reversing on some measures, as the same kind of developers gained real experience with the tools &#8212; which is to say the 19% was a snapshot of the unfamiliar, undisciplined path, and it closes precisely as people pick up the habits this piece is about.</p><p>That finding is real and it isn&#8217;t a contradiction of everything above. It&#8217;s the missing condition. Every practice in this piece works <em>because</em> it imposes structure on the agent &#8212; TDD as a guardrail, the spec as input, the ADR as context, mutation testing as the check on the check. Used that way, with discipline, AI is constrained toward correctness. Used the other way &#8212; as autocomplete, as a vibe-coding partner you don&#8217;t supervise &#8212; you get more code, faster, with less correctness and a slower path to done once you account for the cleanup. The METR developers, working in code they already understood deeply, may well have been paying exactly that tax: accepting plausible suggestions that took longer to vet and fix than writing it themselves would have.</p><p>So the inversion isn&#8217;t automatic. The cost of patience dropped to zero, which makes the rigorous path finally affordable. It does not make the undisciplined path good. If anything it makes discipline more important, because a tool that produces plausible output at high volume is precisely the tool that most needs a guardrail you can&#8217;t talk your way past. A red test bar doesn&#8217;t care how confident the model sounds.</p><h2>What To Do With This</h2><p>The interesting question was never &#8220;what does AI make possible.&#8221; That list is enormous, mostly speculative, and not very actionable. The better question is the one this whole piece is built on: which practices did we already know were right, argue about for decades, and quietly give up on?</p><p>That list is short, specific, and yours to write. Go pull your own team&#8217;s &#8220;we should really do this but we don&#8217;t&#8221; backlog &#8212; the standing items in retros that everyone agrees with and nobody owns. I&#8217;d bet most of them have the same economic shape: high upfront cost, distant payoff, killed by human impatience under deadline. Test coverage on the legacy module. The runbooks. The ADRs for the three decisions everyone keeps re-litigating. The integration tests that would&#8217;ve caught last quarter&#8217;s outage.</p><p>Run each one through a single question: was this abandoned because it was <em>wrong</em>, or because it was <em>expensive in human patience</em>? The wrong ones, leave abandoned. The expensive-in-patience ones just got cheap. Those are the ones to pick back up first.</p><p>The impatience tax got repealed. The disciplines it used to make unaffordable are sitting right there, mostly unchanged, waiting for someone to notice the price changed.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and writes about AI-powered engineering at <a href="https://hyperdev.substack.com/">HyperDev</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p>The Other Shoe Has Dropped &#8212; Why enterprise AI bills don&#8217;t match the per-token price collapse</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul>]]></content:encoded></item><item><title><![CDATA[The Era of the Leader/Practitioner]]></title><description><![CDATA[Putting "Do" back into "Lead"]]></description><link>https://hyperdev.matsuoka.com/p/the-era-of-the-leaderpractitioner</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/the-era-of-the-leaderpractitioner</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Mon, 01 Jun 2026 11:31:28 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!E_Vr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3731b710-763c-4f36-860f-560f95b11d2b_1024x700.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!E_Vr!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3731b710-763c-4f36-860f-560f95b11d2b_1024x700.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!E_Vr!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3731b710-763c-4f36-860f-560f95b11d2b_1024x700.png 424w, https://substackcdn.com/image/fetch/$s_!E_Vr!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3731b710-763c-4f36-860f-560f95b11d2b_1024x700.png 848w, https://substackcdn.com/image/fetch/$s_!E_Vr!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3731b710-763c-4f36-860f-560f95b11d2b_1024x700.png 1272w, https://substackcdn.com/image/fetch/$s_!E_Vr!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3731b710-763c-4f36-860f-560f95b11d2b_1024x700.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!E_Vr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3731b710-763c-4f36-860f-560f95b11d2b_1024x700.png" width="1024" height="700" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3731b710-763c-4f36-860f-560f95b11d2b_1024x700.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:700,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1378805,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/200067020?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb6656d57-be88-4993-a972-b7c0c5fd743d_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!E_Vr!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3731b710-763c-4f36-860f-560f95b11d2b_1024x700.png 424w, https://substackcdn.com/image/fetch/$s_!E_Vr!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3731b710-763c-4f36-860f-560f95b11d2b_1024x700.png 848w, https://substackcdn.com/image/fetch/$s_!E_Vr!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3731b710-763c-4f36-860f-560f95b11d2b_1024x700.png 1272w, https://substackcdn.com/image/fetch/$s_!E_Vr!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3731b710-763c-4f36-860f-560f95b11d2b_1024x700.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Something shifted in the last year, and it took me a while to name it.</p><p>A growing number of people are running organizations while still doing real hands-on technical work. Not as a hobby, not on weekends, not as a vanity exercise to keep their commit graph green. They are building things their teams depend on &#8212; and they are doing it as a deliberate part of the job, not in the cracks between meetings. The work has a specific shape. It is rarely a production feature. It is the layer underneath: developer productivity tooling, internal services, agentic harnesses, MCP connectors, the infrastructure that unblocks everyone else.</p><p>For most of my career this combination didn&#8217;t really hold together. You could code or you could lead, and the moment you tried to do both seriously, one of them rotted. I&#8217;ve watched plenty of technical executives keep a foot in the codebase and slowly become the bottleneck everyone routed around politely. The pattern was familiar enough to be a warning.</p><p>What changed is not that leaders suddenly got more disciplined. It&#8217;s that the time cost of meaningful technical contribution collapsed. Agentic coding made a hybrid role viable that wasn&#8217;t viable before &#8212; and the more I look at it, the more I think this isn&#8217;t a new invention at all. It&#8217;s the recovery of a very old idea that modern specialization interrupted.</p><p>I&#8217;m writing this as someone living in the middle of it. I spend roughly 30% of my time coding. And when I say coding, I mean directing a team of agents &#8212; much closer to that than hands-on work, which is probably the whole point. That&#8217;s not a full-time IC&#8217;s week, and it isn&#8217;t meant to be. It&#8217;s enough to stay close to the work that matters and to build the enabling infrastructure I think is worth my own hands on the keyboard.</p><h2>TL;DR</h2><ul><li><p>A distinct role is emerging: leaders who run organizations and still do deep technical work &#8212; specifically enabling/institutional work (tooling, harnesses, internal services), not critical-path product features.</p></li><li><p>This satisfies Charity Majors&#8217; actual advice. Her line was never &#8220;stop coding.&#8221; It was &#8220;stop writing code in the critical path.&#8221; Enabling work fits that exactly.</p></li><li><p>The integration of strategist and practitioner has deep cross-cultural precedent &#8212; Japan&#8217;s <em>bunbu-ry&#333;d&#333;</em>, Rome&#8217;s Marcus Aurelius, China&#8217;s <em>wen-wu</em>, the Renaissance polymath, Mattis&#8217;s &#8220;warrior monk.&#8221; The modern role is a recovery, not a novelty.</p></li><li><p>Agentic coding is what makes it newly viable: focused sessions now deliver output that once required sustained, uninterrupted immersion. The Anthropic 2026 data shows ~27% of AI-assisted work is work that &#8220;wouldn&#8217;t have been done otherwise.&#8221;</p></li><li><p>Directing agents feels like delegation &#8212; the same skill leaders already use with human reports. Which is why experienced leaders adapt to it more naturally than juniors do.</p></li></ul><h2>The pattern, named</h2><p>The difficulty with the coding executive was never philosophical. It was attentional.</p><p>Charity Majors mapped this years ago in <a href="https://charity.wtf/2017/05/11/the-engineer-manager-pendulum/">The Engineer/Manager Pendulum</a>, and the piece holds up because she was precise about the mechanism. Management is interruptive by design &#8212; your job is to be available, to unblock, to absorb the chaos so your team doesn&#8217;t have to. Serious engineering is the opposite. It requires blocking interruptions for long enough to hold a complex system in your head. Two incompatible attention modes. Try to run both at once and you do neither well.</p><p>But here is the part people skip when they quote her. Majors never said managers should stop coding. Her actual advice was sharper: <em>don&#8217;t write code in the critical path.</em> Don&#8217;t be the person others are waiting on. Stay technical, stay sharp, just don&#8217;t make yourself a dependency that blocks shipping. &#8220;The best frontline eng managers in the world,&#8221; she wrote, &#8220;are the ones that are never more than 2-3 years removed from hands-on work.&#8221;</p><p>That distinction carries the whole argument. Because there is a category of technical work that is consequential without being critical-path, and it turns out to be exactly the work senior people are best positioned to do.</p><p>Call it enabling work. Internal tooling. Developer productivity infrastructure. Agentic harnesses. MCP services that other teams plug into. Architectural prototypes that prove a direction before anyone commits to it. None of this is what blocks a release on Thursday. All of it multiplies whoever comes after. It tolerates interruption &#8212; you can pick it up Tuesday afternoon and put it down when a real fire starts &#8212; precisely because nobody is standing at your desk waiting for it.</p><p>This is also where the industry is putting its money. Gartner named platform engineering a top strategic trend for two years running and projects that 80% of large engineering organizations will run dedicated platform teams by 2026. The structural reason a leader can work here without becoming the bottleneck is built into the definition of the domain: its output unblocks others rather than blocking them.</p><p>Will Larson&#8217;s <a href="https://lethain.com/staff-engineer-archetypes/">Staff Engineer archetypes</a> circle the same territory without quite landing on it. His &#8220;Architect&#8221; sits in a permanent argument &#8212; some organizations demand the Architect stay deep in the code, others forbid it. The leader/practitioner resolves that argument by relocating it: deep in the enabling and infrastructure work, absent from production product code. Not the pendulum, not the staff IC. A real hybrid, and a newly coherent one.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!jHB6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd3763-7ad5-4119-96d4-4e0c1d01f0cd_1024x684.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!jHB6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd3763-7ad5-4119-96d4-4e0c1d01f0cd_1024x684.png 424w, https://substackcdn.com/image/fetch/$s_!jHB6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd3763-7ad5-4119-96d4-4e0c1d01f0cd_1024x684.png 848w, https://substackcdn.com/image/fetch/$s_!jHB6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd3763-7ad5-4119-96d4-4e0c1d01f0cd_1024x684.png 1272w, https://substackcdn.com/image/fetch/$s_!jHB6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd3763-7ad5-4119-96d4-4e0c1d01f0cd_1024x684.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!jHB6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd3763-7ad5-4119-96d4-4e0c1d01f0cd_1024x684.png" width="1024" height="684" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dffd3763-7ad5-4119-96d4-4e0c1d01f0cd_1024x684.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:684,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1540586,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/200067020?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa8e7bf7a-504a-4f0c-b75f-1832889fc598_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!jHB6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd3763-7ad5-4119-96d4-4e0c1d01f0cd_1024x684.png 424w, https://substackcdn.com/image/fetch/$s_!jHB6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd3763-7ad5-4119-96d4-4e0c1d01f0cd_1024x684.png 848w, https://substackcdn.com/image/fetch/$s_!jHB6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd3763-7ad5-4119-96d4-4e0c1d01f0cd_1024x684.png 1272w, https://substackcdn.com/image/fetch/$s_!jHB6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdffd3763-7ad5-4119-96d4-4e0c1d01f0cd_1024x684.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>This is not new</h2><p>Here is where I want to slow down, because the most interesting thing about this role is how old it is.</p><p>The idea that a leader should be both a strategist and a practitioner &#8212; not separate modes to alternate between but a single integrated way of operating &#8212; shows up independently across at least five civilizations. That kind of convergence usually means a culture has found a durable answer to a real problem.</p><p>Japan gave it a name: <em>bunbu-ry&#333;d&#333;</em> (&#25991;&#27494;&#20001;&#36947;), the way of both the literary and the martial. <em>Bun</em> is letters, cultivation, strategy. <em>Bu</em> is the martial, the active, the practitioner&#8217;s hand. <em>Ry&#333;d&#333;</em> means both ways, together &#8212; not balanced, not traded off, but held at once. By the mid-fourteenth century the dual-talented warrior was already established as the model leader, and during the Edo period the Tokugawa shogunate made it official policy for the samurai class. The phrase that survives captures the stakes: culture without power is ineffective, and power without culture is barbarous.</p><p>The archetype is Miyamoto Musashi. Undefeated in more than sixty duels, often fighting with a wooden sword against live steel. He founded a two-sword school, and in the last months of his life he wrote <em>The Book of Five Rings</em> in a cave. He was also a recognized master of ink painting and calligraphy &#8212; his <em>Shrike on a Withered Branch</em> survives as a designated Important Cultural Property of Japan. The same hands that won sixty duels produced fine art a nation still protects. He didn&#8217;t oscillate between the sword and the brush. He held both, and each sharpened the other. &#8220;When I apply the principle of strategy to the ways of different arts and crafts,&#8221; he wrote, &#8220;I no longer have need for a teacher in any domain.&#8221; Mastery in one discipline illuminating all the others.</p><p>Rome had Marcus Aurelius, the philosopher-king made historical rather than theoretical. He ran the empire and commanded its armies on the Danube, and he wrote the <em>Meditations</em> in the war camps &#8212; fragmentary notes to himself, composed in the middle of campaigning and administration. That book was never published philosophy. It was a working journal, the most powerful man in the ancient world writing to stay grounded while doing the job.</p><p>China institutionalized the same ideal as <em>wen</em> and <em>wu</em> &#8212; civil cultivation and martial capability &#8212; and ran it for roughly thirteen centuries through the scholar-official class. Zeng Guofan is the canonical case: he rose through the imperial examinations to high Confucian office, then built and commanded an army of more than a hundred thousand, reportedly keeping a diary on Neo-Confucian ethics even as he directed the campaigns. The integration wasn&#8217;t left to personal taste. It was built into the examination and career structure.</p><p>And the archetype still lands today. James Mattis earned the nickname &#8220;Warrior Monk&#8221; &#8212; battlefield commander and devoted reader, a 7,000-book library, the <em>Meditations</em> carried into combat. The chain from Aurelius to Mattis is literal: the same book, eighteen centuries apart. That we still reach for &#8220;warrior monk&#8221; as a compliment for a leader tells you the integration never stopped resonating.</p><p>Across all of it, the answer is the same. The contemplative and the active were not specializations to assign to different people. They were a single discipline, each half informing the other. Modernity &#8212; with its org charts, its clean role boundaries, its professional specialization &#8212; interrupted that. The leader/practitioner is not a tech-industry novelty. It&#8217;s an old integration becoming feasible again.</p><h2>Why now</h2><p>So what actually changed? Not the wisdom. The economics.</p><p>The thing that made the coding executive a bad idea was the attention math. Serious technical work demanded long, unbroken stretches of focus &#8212; the exact resource a leadership schedule cannot reliably provide. You cannot design a system in the fifteen minutes between a board prep and a one-on-one. The pendulum was a real constraint, not a failure of will.</p><p>Agentic coding changes that math directly. The unit of work moved up a level. Instead of holding every implementation detail in working memory across a four-hour session, you specify intent, direct an agent, review what comes back, correct course, and direct again. A focused 30-minute session now produces what used to require an afternoon of immersion &#8212; not because the thinking got easier, but because the implementation cost collapsed.</p><p>The Anthropic <a href="https://resources.anthropic.com/2026-agentic-coding-trends-report">2026 Agentic Coding Trends Report</a> puts numbers on the shift. Average session length has climbed to 23 minutes in the agentic era, up from about 4 in the autocomplete era &#8212; the work got denser, not just faster. 78% of Claude Code sessions now involve multi-file edits, up from 34% a year earlier. Teams running multi-agent workflows report 2&#8211;4x faster delivery from task creation to deployment. And the figure that matters most for this argument: roughly 27% of AI-assisted work consists of tasks that &#8220;wouldn&#8217;t have been done otherwise&#8221; &#8212; the scaling projects, the nice-to-have tools, the exploratory infrastructure that was never quite worth the manual hours.</p><p>That 27% is the enabling work. It is the category that lives or dies on time cost, and it&#8217;s the category a leader/practitioner is best placed to take on.</p><p>The arithmetic is what makes 30% credible. I documented a 6&#8211;10x multiplier on focused technical sessions in <a href="https://hyperdev.matsuoka.com/the-irreducibles-what-a-pattern-master-does">The Irreducibles</a> earlier this year &#8212; a project I estimated at 150&#8211;200 billable hours compressed into roughly 50&#8211;70 hours of wall-clock time, most of which wasn&#8217;t coding at all. If directed work runs several times faster than hand-coding, then a day and a half a week can produce what once consumed a full-time engineer&#8217;s week. That&#8217;s not a marginal gain. It&#8217;s a change in what&#8217;s structurally possible.</p><p>There&#8217;s another dimension that doesn&#8217;t show up in the productivity numbers: the work itself is unstable. Agentic coding patterns are still shaking out. There aren&#8217;t many experienced practitioners, the field is moving fast, and we don&#8217;t yet have good consensus on which patterns are load-bearing and which are fashion. A manager who&#8217;s only reading about it can&#8217;t make that distinction on behalf of a team. You have to be in it to know.</p><p>There&#8217;s a counterintuitive wrinkle worth naming: the people best positioned to exploit this are the senior ones. A University of Chicago working paper from late 2025 found experienced developers were 5&#8211;6% more likely to succeed with AI agents for every standard deviation of work experience, largely because they worked plan-first &#8212; laying out objectives, alternatives, and steps before invoking the tool. That&#8217;s the opposite of the assumption that AI flattens the seniority curve. Expertise improves your ability to delegate to a model for the same reason it improves your ability to delegate to a person. AI doesn&#8217;t change what senior engineering is. It reveals what it always was.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!EYNR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F727ed32e-2441-4dc7-bf6b-4ea9a19a2cb7_1024x634.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!EYNR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F727ed32e-2441-4dc7-bf6b-4ea9a19a2cb7_1024x634.png 424w, https://substackcdn.com/image/fetch/$s_!EYNR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F727ed32e-2441-4dc7-bf6b-4ea9a19a2cb7_1024x634.png 848w, https://substackcdn.com/image/fetch/$s_!EYNR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F727ed32e-2441-4dc7-bf6b-4ea9a19a2cb7_1024x634.png 1272w, https://substackcdn.com/image/fetch/$s_!EYNR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F727ed32e-2441-4dc7-bf6b-4ea9a19a2cb7_1024x634.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!EYNR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F727ed32e-2441-4dc7-bf6b-4ea9a19a2cb7_1024x634.png" width="1024" height="634" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/727ed32e-2441-4dc7-bf6b-4ea9a19a2cb7_1024x634.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:634,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1514078,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/200067020?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1745d9f4-d955-45cb-b61a-c12956852f96_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!EYNR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F727ed32e-2441-4dc7-bf6b-4ea9a19a2cb7_1024x634.png 424w, https://substackcdn.com/image/fetch/$s_!EYNR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F727ed32e-2441-4dc7-bf6b-4ea9a19a2cb7_1024x634.png 848w, https://substackcdn.com/image/fetch/$s_!EYNR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F727ed32e-2441-4dc7-bf6b-4ea9a19a2cb7_1024x634.png 1272w, https://substackcdn.com/image/fetch/$s_!EYNR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F727ed32e-2441-4dc7-bf6b-4ea9a19a2cb7_1024x634.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Directing agents is delegation</h2><p>This is the part that interested me the most, and it&#8217;s the bridge between the leadership job and the technical one.</p><p>A year or so ago, working with Claude Code felt like coding. Now it feels like delegating. I <a href="https://hyperdev.matsuoka.com/coding-to-delegation-shift">wrote about that shift</a> when it first became undeniable &#8212; the move from being a programmer who uses AI to being something closer to a technical project manager who directs it. Anthropic&#8217;s report uses the same vocabulary, describing engineers moving &#8220;from writing code to orchestrating the systems that write it.&#8221;</p><p>What I didn&#8217;t fully appreciate at the time is how directly that maps onto the muscle leaders already have. Delegating to an agent feels, in practice, just like delegating to a human engineer. You frame the problem, set the constraints, hand it off, and come back to assess the result. Give a directive, walk away, return to completed work. That loop is the daily reality of management. Leaders developed it because they had to, and it transfers to agents almost without friction.</p><p>So the leader/practitioner doesn&#8217;t have to become a coder again in the old sense. The skill in demand is judgment plus delegation, and that&#8217;s the skill leadership has been building all along. The hands-on knowledge tells you what to ask for and whether the answer is any good. The delegation instinct does the rest.</p><p>This is why the enabling work and the agentic tools fit together so cleanly. Enabling work tends to be well-defined, non-user-facing, and long-horizon &#8212; exactly the profile agents handle well and exactly the profile that tolerates a leader&#8217;s interrupted schedule. The hands-on contribution mostly takes the form of specifying constraints and patterns, which is what I&#8217;ve called <a href="https://hyperdev.matsuoka.com/what-does-a-pattern-master-do">pattern mastery</a>: when you write the pattern down, you&#8217;ve written the spec, and the spec multiplies everyone else&#8217;s output.</p><h2>What it actually looks like</h2><p>Let me ground this without turning it into a war story.</p><p>The concrete examples from my own work are the kind of thing I mean. I built an agentic harness &#8212; the orchestration layer I <a href="https://hyperdev.matsuoka.com/its-the-harness-stupid">argued is the real determinant of AI coding outcomes</a>, where the same model can swing more than a quality point depending on the scaffolding around it. I built MCP services, the <a href="https://hyperdev.matsuoka.com/is-this-the-era-of-the-connector">org-specific connectors</a> that replaced a handful of standalone tools in a few hours of directed work each. None of that was a production feature. All of it was infrastructure other people now depend on.</p><p>Here&#8217;s a detail that may make the point. I now have an &#8220;AI architect&#8221; on my leadership team helping maintain the very infrastructure I originally built &#8212; not just the harness, but our inference relationships, our training program, office hours, the real human work I no longer have the time, or the right, to be doing myself. And I expect to hand off more over time. The enabling work I do today partly becomes the system that does tomorrow&#8217;s enabling work. That handoff is the role in miniature: you build the thing that multiplies the team, then you put someone in place to build the next version.</p><p>The proportion matters. Around 30% hands-on keeps judgment fresh without putting me in the critical path. Even full-time senior ICs aren&#8217;t full-time coders &#8212; Bain&#8217;s Jue Wang, quoted in MIT Technology Review last December, put developer coding time at 20&#8211;40%, with the rest going to analysis, strategy, and the surrounding work. A leader at 30% isn&#8217;t doing something exotic. They&#8217;re spending their technical budget on the layer where it compounds.</p><p>The decision is not &#8220;how do I find time to code.&#8221; It&#8217;s &#8220;what enabling work is worth my own hands?&#8221; Those are different questions. The first leads to the bottleneck I watched so many executives become. The second leads somewhere useful.</p><h2>The choice</h2><p>I&#8217;ll resist overselling this, because it isn&#8217;t for everyone and it isn&#8217;t automatic.</p><p>This is a deliberate role, not a default. Staying technically current costs ongoing investment, and the work is often invisible &#8212; enabling infrastructure rarely shows up in a quarterly review the way a shipped feature does. The role is easy to misread, too. From the outside, a CTO who codes can look like a CTO who hasn&#8217;t let go. The defense against that reading is the discipline Majors named: stay out of the critical path. Build the multipliers, not the blockers.</p><p>The returns are real, though. Fresh judgment &#8212; the kind that lets you evaluate not just what and why but how. Trust from engineers who see you in the work rather than above it. And institutional infrastructure that makes the whole team faster, built by the person with both the technical depth and the positional authority to prioritize it.</p><p>There&#8217;s a closing note in the history worth keeping. <em>Bunbu-ry&#333;d&#333;</em> wasn&#8217;t only a personal aspiration. The Tokugawa shogunate institutionalized it &#8212; built career and class structures around the assumption that a leader should be both. China did the same with its examination system. We&#8217;re not there yet. For now, the leader/practitioner is an individual choice, made one person at a time, made viable by tools that finally collapsed the cost of staying hands-on.</p><p>But the precedent suggests where this could go. When a way of working proves durable, organizations eventually build structures around it. The era of the leader/practitioner is early. It is also, I&#8217;d argue, a return &#8212; to an integration we knew was valuable long before we had the means to make it practical again.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and writes about AI-powered engineering at <a href="https://hyperdev.substack.com/">HyperDev</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/coding-to-delegation-shift">From Coding with AI to Managing AI</a> &#8212; When agentic coding starts to feel like delegation</p></li><li><p><a href="https://hyperdev.matsuoka.com/its-the-harness-stupid">It&#8217;s The Harness, Stupid!</a> &#8212; Why orchestration quality dominates AI coding outcomes</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul>]]></content:encoded></item><item><title><![CDATA[The Other Shoe Has Dropped]]></title><description><![CDATA[The Economics of Enterprise Inference Usage]]></description><link>https://hyperdev.matsuoka.com/p/the-other-shoe-has-dropped</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/the-other-shoe-has-dropped</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Fri, 29 May 2026 11:31:28 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!-Qpp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-Qpp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-Qpp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!-Qpp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!-Qpp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!-Qpp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-Qpp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1174362,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/199673840?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!-Qpp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!-Qpp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!-Qpp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!-Qpp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41ae5cbb-29e9-4a44-96a2-fb1724f0bf79_1024x768.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Two stories from the last two weeks. Uber <a href="https://fortune.com/2026/05/26/uber-coo-ai-spending-tokens-claude-code/">burned through its entire 2026 AI budget in four months</a> on Claude Code, with COO Andrew Macdonald telling the <em>Rapid Response</em> podcast that the link between that spend and shipped consumer features &#8220;is not there yet.&#8221; <a href="https://www.theinformation.com/newsletters/applied-ai/uber-cto-shows-claude-code-can-blow-ai-budgets">The Information had the underlying numbers a few weeks earlier</a>: engineer adoption from 32% to 84% between December and March, heavy users running $500&#8211;$2,000/month in tokens, and CTO Praveen Neppalli Naga torching $1,200 in a two-hour demo. Same week, Microsoft told thousands of engineers in its Experiences + Devices division that their Claude Code access is going away. <a href="https://www.windowscentral.com/microsoft/microsoft-cancels-claude-code-licenses-shifting-developers-to-github-copilot-cli-a-move-likely-driven-by-financial-motives">Windows Central, summarizing The Verge&#8217;s Notepad scoop</a>, has the cutoff at June 30 &#8212; end of fiscal year &#8212; with cost as the actual driver even though EVP Rajesh Jha framed it publicly as convergence on Copilot CLI.</p><p>Two of the most AI-forward enterprises on the planet, same tool, same week. The &#8220;AI is failing&#8221; takes were live within hours.</p><p>I don&#8217;t buy that framing.</p><p>The headlines are getting it wrong. Uber didn&#8217;t cancel anything &#8212; adoption ran ahead of the budget and the company blew its annual spend keeping up. That&#8217;s a planning failure, not a verdict on the tool. Microsoft didn&#8217;t divorce Anthropic either; they&#8217;re still consuming Claude through Azure Foundry and M365 Copilot. What they cancelled is a specific license &#8212; Claude Code at the engineer-seat level &#8212; because engineers preferred it over GitHub Copilot CLI and the division was paying for that preference.</p><p>What both stories show: AI is a new tool and we haven&#8217;t learned to use it well yet. The teams over budget pointed it at problems it wasn&#8217;t the cheapest way to solve, then let it decide for itself how much work to do per task.</p><p>I&#8217;ve made <a href="https://hyperdev.matsuoka.com/p/what-the-other-shoe-sounds-like-when">the cloud parallel here before</a>. Early cloud was expensive and misused. Lift-and-shift workloads routinely ran two or three times their on-prem cost &#8212; I watched that play out across teams I ran, and it took years to correct through architecture. Then the industry learned: right-sizing, reserved instances, autoscaling, serverless where it fit, on-prem where it didn&#8217;t. The bills came down. Not because compute got dramatically cheaper, but because we got more careful about what we asked the cloud to do. AI is in the same phase. Cheap per-token, expensive per-task, and the gap is architectural.</p><p>A few weeks ago I ran controlled head-to-head tests on Opus 4.6 and Opus 4.7 against identical coding tasks. Both models passed every test. Opus 4.7 cost 3.6&#215; more to do it. Same outcomes, same rate card, dramatically more tokens.</p><p>Finout&#8217;s analysis of production deployments <a href="https://www.finout.io/blog/claude-opus-4.7-pricing-the-real-cost-story-behind-the-unchanged-price-tag">tells the same story at scale</a>: up to a 35% cost increase overnight, driven by tokenizer changes that don&#8217;t show up on the per-token rate card. Not one team&#8217;s bad luck &#8212; the shape of the bill across the enterprise AI buyer base right now. The second of two shoes on AI economics.</p><p>I wrote about that <a href="https://hyperdev.matsuoka.com/p/opus-46-vs-47-the-real-cost-of-incremental">version-to-version cost drift in detail</a>. Providers can collapse per-token prices in public while the per-task bill drifts upward in private. The first shoe was the per-token price collapse that made everyone optimistic. The second is the behavioral and architectural cost overhang now landing on quarterly P&amp;Ls.</p><p><strong>TL;DR</strong></p><ul><li><p>Per-token costs at GPT-3.5-equivalent performance are down roughly 280&#215; since late 2022, per <a href="https://aiindex.stanford.edu/report/">Stanford&#8217;s AI Index 2025</a>. Vendor revenue tells the opposite story: Anthropic&#8217;s annualized revenue went from <a href="https://www.pymnts.com/artificial-intelligence-2/2026/anthropic-hits-30-billion-run-rate-as-enterprise-demand-accelerates/">$1B in January 2025 to $30B by April 2026</a> &#8212; a 30&#215; move in 15 months, coming from enterprise inference, not consumer subscriptions.</p></li><li><p>Gartner&#8217;s April 2026 survey: just 28% of AI use cases fully meet ROI expectations, 78% of IT leaders report material AI charges that didn&#8217;t show up in any procurement model.</p></li><li><p>Gartner estimates agentic workflows consume 5&#8211;30&#215; more tokens than equivalent chatbot interactions; Stanford&#8217;s Digital Economy Lab puts the upper bound for coding agents at 1,000&#215;. The cost driver isn&#8217;t the model &#8212; it&#8217;s the workflow architecture wrapped around it.</p></li><li><p>Two patterns hold the line in production. Search-first architectures put inference at the end of a deterministic pipeline. Consolidated single-shot designs replace multi-call chains.</p></li><li><p>Inference is a power tool, not a default. Use it with specific ROI goals per call, apply it to <em>code</em> solutions rather than to directly solve problems, and bound it with deterministic structure on both ends.</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!TPPJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F491e72b0-16e1-49fe-a551-391ba05461bc_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!TPPJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F491e72b0-16e1-49fe-a551-391ba05461bc_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!TPPJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F491e72b0-16e1-49fe-a551-391ba05461bc_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!TPPJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F491e72b0-16e1-49fe-a551-391ba05461bc_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!TPPJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F491e72b0-16e1-49fe-a551-391ba05461bc_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!TPPJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F491e72b0-16e1-49fe-a551-391ba05461bc_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/491e72b0-16e1-49fe-a551-391ba05461bc_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1163417,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/199673840?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F491e72b0-16e1-49fe-a551-391ba05461bc_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!TPPJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F491e72b0-16e1-49fe-a551-391ba05461bc_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!TPPJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F491e72b0-16e1-49fe-a551-391ba05461bc_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!TPPJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F491e72b0-16e1-49fe-a551-391ba05461bc_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!TPPJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F491e72b0-16e1-49fe-a551-391ba05461bc_1024x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Why the Per-Token Savings Didn&#8217;t Reach the Invoice</h2><p>Per-token economics of frontier models have been collapsing for two years. <a href="https://aiindex.stanford.edu/report/">Stanford&#8217;s AI Index 2025</a> puts the decline at roughly 280&#215; from a late-2022 baseline at GPT-3.5-equivalent performance. Most enterprise budget conversations in 2024 started from that headline. The implicit assumption: bills should be going <em>down</em>.</p><p>They&#8217;re not. The clearest read comes from the vendor side. Anthropic&#8217;s annualized revenue <a href="https://www.anthropic.com/news/anthropic-raises-30-billion-series-g-funding-380-billion-post-money-valuation">went from $1B in January 2025 to $30B by April 2026</a>, with roughly 80% from enterprise and API usage rather than consumer subscriptions. Anthropic <a href="https://www.saastr.com/anthropic-just-passed-openai-in-revenue-while-spending-4x-less-to-train-their-models/">now discloses 1,000+ customers spending more than $1M per year</a> &#8212; a cohort that doubled in under two months &#8212; alongside roughly 300,000 business customers. The mid-tier ($100K&#8211;$1M/year) grew 7&#215; year over year.</p><p>Menlo Ventures&#8217; <a href="https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/">2025 State of Generative AI report</a> cross-checks at the market level: enterprise GenAI spend tripled to $37B in 2025, with LLM API consumption alone at $8.4B by mid-year. The high tier shows up in named deals: <a href="https://sacra.com/c/anthropic/">Snowflake&#8217;s $200M multi-year partnership</a> implies a $50&#8211;70M annual run rate from one customer, and <a href="https://sacra.com/c/anthropic/">Deloitte is deploying Claude across 470,000 employees</a>.</p><p>For a typical enterprise running multiple production AI workloads, $500K&#8211;$2M per year is now the realistic floor. Fortune 100 is running $10M&#8211;$50M+, the most AI-intensive past $100M. The Gartner numbers point the same direction: <a href="https://www.gartner.com/en/newsroom/press-releases/2026-04-07-gartner-says-artificial-intelligence-projects-in-infrastructure-and-operations-stall-ahead-of-meaningful-roi-returns">just 28% of AI use cases fully meet ROI expectations and 20% fail outright</a>, and <a href="https://zylo.com/blog/saas-management-index/">78% of IT leaders report material AI charges</a> that didn&#8217;t show up in any procurement model.</p><p>The per-token chart is real. The invoice is also real. What closes the gap is <em>behavior</em>. Three behaviors specifically.</p><p><strong>Models do more work per task.</strong> Reasoning models reason. Agentic loops loop. The prompt that used to consume 4K tokens now consumes 40K because the assistant explores, plans, second-guesses, and verifies. Some of that is valuable. Much of it is the model performing thoroughness in a way that costs you money. The Opus 4.6-to-4.7 jump I documented earlier: same task, same outcome, 2.9&#215; more output tokens and 4.8&#215; more cache reads.</p><p><strong>Workflows fan out.</strong> A &#8220;single&#8221; task in a modern agentic system might trigger a planner, researcher, coder, reviewer, and summarizer. Each makes its own LLM calls over overlapping context. Gartner&#8217;s March 2026 analysis puts agentic workflows at 5&#8211;30&#215; the token consumption of an equivalent chatbot ask. Stanford Digital Economy Lab&#8217;s April 2026 arXiv paper goes further: coding agents can consume 1,000&#215; more tokens than equivalent chat completions. The agent isn&#8217;t more expensive because it&#8217;s smarter. It&#8217;s more expensive because it&#8217;s louder.</p><p><strong>Context windows fill themselves.</strong> Long context is a feature in marketing and a bill in practice. In our own enterprise Claude.AI usage &#8212; 82,852 messages from 329 employees over 3.5 months, audited via the Anthropic Compliance API &#8212; the average request carried 366,000 input tokens, mostly from 10-turn conversations dragging accumulated history forward into every new turn. Most production systems I&#8217;ve audited show the same fingerprint: pipelines paying for context they aren&#8217;t actually using.</p><p>None of this is fraud and none of it is mysterious. It&#8217;s the natural consequence of letting probabilistic systems decide how much work to do on every call. The savings from cheaper tokens were real. They just got consumed by an order of magnitude more tokens per task.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!T_ba!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7441c01-b144-4b0c-83a2-113c2bb730fd_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!T_ba!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7441c01-b144-4b0c-83a2-113c2bb730fd_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!T_ba!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7441c01-b144-4b0c-83a2-113c2bb730fd_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!T_ba!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7441c01-b144-4b0c-83a2-113c2bb730fd_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!T_ba!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7441c01-b144-4b0c-83a2-113c2bb730fd_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!T_ba!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7441c01-b144-4b0c-83a2-113c2bb730fd_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c7441c01-b144-4b0c-83a2-113c2bb730fd_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1253809,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/199673840?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7441c01-b144-4b0c-83a2-113c2bb730fd_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!T_ba!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7441c01-b144-4b0c-83a2-113c2bb730fd_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!T_ba!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7441c01-b144-4b0c-83a2-113c2bb730fd_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!T_ba!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7441c01-b144-4b0c-83a2-113c2bb730fd_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!T_ba!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc7441c01-b144-4b0c-83a2-113c2bb730fd_1024x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>What I&#8217;ve Found Shipping These Systems</h2><p>The teams handling this well aren&#8217;t the ones cutting AI usage. They&#8217;re changing the <em>shape</em> of how they use it.</p><p>The pattern I keep coming back to: treat inference as the expensive step at the end of a mostly deterministic pipeline. Do the cheap, structured work in code. Reserve the model call for the part that actually requires judgment. Then bound the call hard &#8212; context budget, output budget, quality gate on whether the call even runs.</p><p>Two examples from systems I&#8217;ve been building illustrate this from different angles.</p><h3>Example 1: Code-Intelligence &#8212; Search-First Architecture</h3><p>The naive version of a code review tool is obvious: dump the changed files into Claude, ask for a review. That works. It also costs roughly $0.03 per operation, scales linearly with repo size, and produces a lot of review output you didn&#8217;t need. Claude.AI offers a code review service &#8212; it ended up costing us thousands a month for just a few repos. Augment Code offers a well-regarded one as a GitHub app, but charges a platform fee (a meaningful fraction of our Anthropic spend) <em>just</em> to connect.</p><p>So we built our own. It leverages a multimodal search/RAG/KG engine I&#8217;d already built, so this wasn&#8217;t from scratch.</p><p>The version I actually ship uses a multi-tier search pipeline with the LLM call at the very end:</p><pre><code><code>Stage 1: Vector Search    (~$0.0002 per query, semantic similarity)
Stage 2: BM25 Reranking   (~$0.0001 per query, lexical relevance)
Stage 3: Static Analysis  (~$0.0001 per query, AST + symbol resolution)
Stage 4: Quality Gate     (free, deterministic threshold check)
Stage 5: Single LLM Call  (~$0.03 per call, only if Stages 1-4 passed)</code></code></pre><p>The first four stages cost about $0.0004 combined. They do the bulk of the work: deciding <em>what code is actually relevant</em>, ranking it, pulling structural relationships, and deciding whether the result is even worth asking an LLM about.</p><p>Hard budget controls run through the whole pipeline:</p><pre><code><code># Budget enforcement, not aspiration
MAX_CONTEXT_FILES = 6          # cap on what we send to the model
MAX_REVIEW_WORDS = 500         # cap on what the model returns
RELEVANCE_FLOOR = 0.005        # quality gate before calling the model

if combined_relevance_score &lt; RELEVANCE_FLOOR:
    # no point spending $0.03 to get a review of weakly-related code
    return SkipReason("below relevance floor")

context = select_top_n(ranked_results, MAX_CONTEXT_FILES)
review = llm.review(context, max_output_tokens=MAX_REVIEW_WORDS * 1.4)</code></code></pre><p>The <code>RELEVANCE_FLOOR</code> check is the part to underline. A meaningful percentage of review requests in real codebases don&#8217;t justify an LLM call at all &#8212; the changes are mechanical, the related code trivial, or the search signal weak enough that whatever the model says will be hallucinated context. Refusing to spend $0.03 on those cases is where most of the savings come from.</p><p>Rough economics across a quarter of usage:</p><p>Approach Cost per operation LLM calls per 1K operations Direct &#8220;review the diff&#8221; ~$0.030 1,000 Search-first with gates ~$0.0034 average ~430</p><p>About 80% of the workflow logic ends up deterministic: search, ranking, static analysis, gating. The model handles the last 20% &#8212; judgment on curated context. Same outcome from the user&#8217;s perspective, roughly an order of magnitude cheaper, with more predictable failure modes because most of the pipeline is debuggable code rather than prompt behavior.</p><p>The limitation: this is more work than wiring up a single LLM call, and the gates are only as good as your search infrastructure. The payoff is on the cost and determinism side, not on speed of initial implementation.</p><h3>Example 2: duetto-intelligence &#8212; Context Injection Instead of Replacement</h3><p>The second pattern comes from duetto-intelligence, internal tooling I&#8217;ve been building against that same enterprise Claude.AI usage &#8212; 82,852 real employee messages over 3.5 months, not a thought experiment. The problem here isn&#8217;t &#8220;should we call the LLM at all.&#8221; It&#8217;s: given that our people are already routing structured-data questions through a $0.274-per-request multi-turn Sonnet conversation, what&#8217;s a cheaper path that doesn&#8217;t degrade the answer?</p><p>The audit data made the gap concrete. Average request: 366K input tokens, ten-turn conversation, $0.274 to Anthropic. The same query answered through a Haiku single-pass against deterministically-retrieved internal data: $0.0009. A 300:1 cost ratio on the slice of traffic about structured product knowledge, CRM/account prep, JIRA, people and org lookups.</p><p>Not all traffic. Somewhere in the 35&#8211;40% range based on classified samples. About half of remaining queries genuinely need full Sonnet or Opus reasoning &#8212; writing, debugging, free-form analysis &#8212; and shouldn&#8217;t be intercepted at all.</p><p>The framing matters, because it&#8217;s easy to mis-read this as &#8220;replace Claude with a smaller model.&#8221; It isn&#8217;t. duetto-intelligence acts as a <strong>context injection layer</strong> in front of the user-facing model. When a query has structured-data intent, we route a sub-query to DI, get back a bounded structured result, and inject that into the prompt the larger model sees. The expensive model still does the reasoning &#8212; it just stops being responsible for the deterministic data retrieval it&#8217;s bad at and expensive for.</p><p>The naive design for the routing layer looks like this:</p><pre><code><code>1. Classify the user's intent             &#8594; LLM call (~150 tokens)
2. Plan which subsystems to query         &#8594; LLM call (~250 tokens)
3. Call subsystem A, summarize response   &#8594; LLM call (~300 tokens)
4. Call subsystem B, summarize response   &#8594; LLM call (~300 tokens)
5. Synthesize a final answer              &#8594; LLM call (~250 tokens)
                                          Total: ~1,250 tokens, 5 calls</code></code></pre><p>Each call is plausible on its own. Together they&#8217;re a tax on every user interaction. Latency stacks linearly with calls, and any one of the five can hallucinate in a way that corrupts the rest of the chain.</p><p>The consolidated design replaces three steps with deterministic code:</p><pre><code><code>1. Classify intent                  &#8594; LLM call  (~80 tokens, tight tagger prompt)
2. Fan-out to subsystems            &#8594; code      (0 tokens, intent &#8594; call map)
3. Consolidated synthesis           &#8594; LLM call  (~200 tokens, structured input)
                                       Total: ~280 tokens, 2 calls</code></code></pre><p>The trick is the intent classifier. It produces a tag from a fixed vocabulary of 87 tags &#8212; <code>revenue.query.ytd</code>, <code>forecast.compare.year_over_year</code>, <code>account.lookup.contact</code>, and so on. Each tag maps deterministically to a set of downstream calls in plain Python. No LLM in the routing step. The model isn&#8217;t asked &#8220;what should we do?&#8221; It&#8217;s asked &#8220;what is the user asking about?&#8221; &#8212; a much smaller, bounded question.</p><p>We validated against a 50-query test corpus drawn directly from the compliance data &#8212; real questions people had asked the model in production. After tuning, 100% of those queries land on the fast path with no LLM call required for routing. That&#8217;s the proof the deterministic-discipline part holds at the boundary; the routing isn&#8217;t quietly falling back to a second model call to bail itself out.</p><p>Budget enforcement is explicit in every prompt template:</p><pre><code><code>SOURCE_CHAR_BUDGET = 600    # per data source pulled into context
OUTPUT_TOKEN_BUDGET = 200   # cap on synthesis response
INTENT_TAG_VOCAB = load_intent_taxonomy()  # 87 tags, versioned

def synthesize(intent_tag: str, sources: list[Source]) -&gt; str:
    trimmed = [s.truncate(SOURCE_CHAR_BUDGET) for s in sources]
    return llm.complete(
        prompt=template(intent_tag, trimmed),
        max_tokens=OUTPUT_TOKEN_BUDGET,
    )</code></code></pre><p>The economics, against measured baselines rather than estimates: current Claude.AI Chat spend across the 329-user population runs $15,264 over 3.5 months. Roughly $4,360/month, driven by that $0.274-per-request multi-turn average. If DI intercepts the 35&#8211;40% of traffic that&#8217;s structured-retrieval underneath, projected savings come in around $1,500&#8211;1,700/month, or $18&#8211;20K/year on this single user population. The leverage isn&#8217;t from picking a cheaper model. It&#8217;s from refusing to pay Sonnet rates to answer questions a deterministic system already has the data for.</p><p>The intent vocabulary is the contract. New capability means a new tag, a new downstream mapping, a new prompt template. The model never has to invent structure on the fly. This is what people mean by &#8220;use the LLM to code solutions, not to solve problems directly&#8221; &#8212; the routing logic lives in code, the tagger is a thin call, the synthesis is bounded.</p><p>The source-character budget matters more than the output budget. The compliance audit confirmed it: production overspend is on the <em>input</em> side. 366K input tokens against a few hundred output. Models will happily consume whatever context you hand them. Trimming at the source &#8212; 600 characters per source, no exceptions &#8212; is how you keep per-call cost from drifting upward as the system gets more capable.</p><p>The limitation: this only works on the routable slice. The pattern isn&#8217;t &#8220;eliminate inference.&#8221; It&#8217;s &#8220;stop spending $0.274 to answer questions that have a structured answer at $0.0009.&#8221;</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-uHZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe0775c6-8743-46e4-908a-ee4e2a8db5d6_1024x768.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-uHZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe0775c6-8743-46e4-908a-ee4e2a8db5d6_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!-uHZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe0775c6-8743-46e4-908a-ee4e2a8db5d6_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!-uHZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe0775c6-8743-46e4-908a-ee4e2a8db5d6_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!-uHZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe0775c6-8743-46e4-908a-ee4e2a8db5d6_1024x768.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-uHZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe0775c6-8743-46e4-908a-ee4e2a8db5d6_1024x768.png" width="1024" height="768" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fe0775c6-8743-46e4-908a-ee4e2a8db5d6_1024x768.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:768,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1351360,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/199673840?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe0775c6-8743-46e4-908a-ee4e2a8db5d6_1024x768.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!-uHZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe0775c6-8743-46e4-908a-ee4e2a8db5d6_1024x768.png 424w, https://substackcdn.com/image/fetch/$s_!-uHZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe0775c6-8743-46e4-908a-ee4e2a8db5d6_1024x768.png 848w, https://substackcdn.com/image/fetch/$s_!-uHZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe0775c6-8743-46e4-908a-ee4e2a8db5d6_1024x768.png 1272w, https://substackcdn.com/image/fetch/$s_!-uHZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffe0775c6-8743-46e4-908a-ee4e2a8db5d6_1024x768.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Why &#8220;Just Use a Cheaper Model&#8221; Doesn&#8217;t Save You</h2><p>A reasonable objection: aren&#8217;t the cheaper, smaller models supposed to handle this? Why not route everything to Haiku or an open model and call it solved?</p><p>The pricing math seems to support it. The behavioral math doesn&#8217;t.</p><p>Two things go wrong when you swap a cheaper model into an unstructured workflow. Cheaper models are usually less efficient <em>per task</em> &#8212; more turns to converge, more exploration, more hallucination, which means more retries and verification calls. A model 5&#215; cheaper per token can run 1.5&#8211;2&#215; more expensive per completed task if your workflow lets it spin.</p><p>And the workflow itself is where most of the cost lives. The 5&#8211;30&#215; multiplier is structural, not modal &#8212; it exists regardless of which model you point at it. Switching from Sonnet to Haiku inside an unbounded agent loop changes the per-token cost. It doesn&#8217;t change the loop.</p><p>Model choice is a 2&#8211;5&#215; lever. Architecture choice is closer to an order-of-magnitude lever in the systems I&#8217;ve shipped &#8212; consistently larger than what model swaps deliver. Most teams are over-tuning the model selection and under-tuning the structure around it.</p><p>The default assumption &#8212; including from vendors with strong incentives to sell you more tokens &#8212; is that the answer to AI cost is buying inference more cleverly. The actual answer is using inference less, more deliberately, with hard bounds on what each call is allowed to do.</p><h2>How I Think About Inference Now</h2><p><strong>Inference is a power tool.</strong> Not a default. You don&#8217;t reach for it when a search query, a regex, or a <code>switch</code> statement would do. You reach for it when you need probabilistic judgment over unstructured input. Every call you don&#8217;t make is the cheapest call.</p><p><strong>Use it to code solutions, not to solve problems.</strong> The highest-leverage use of LLMs in my workflow is generating the deterministic code that then handles the workflow without further LLM calls. A model that writes you a 50-line classifier is more valuable than a model that <em>acts as</em> the classifier on every request forever. The first costs tokens once. The second costs tokens every transaction for the life of the system.</p><p><strong>Wrap every call in a budget.</strong> Context budget on the input, token budget on the output, quality gate on whether the call runs at all. Treat the LLM call as you&#8217;d treat a paid API with rate limits and SLA penalties. Because it is.</p><p><strong>Set specific ROI targets per call.</strong> &#8220;AI-assisted code review&#8221; is too coarse to optimize. &#8220;Reviewing files with relevance score &gt; 0.005, capped at 6 files, returning 500 words&#8221; is something you can measure cost-per-outcome on. Even loose ROI math at the call level surfaces where you&#8217;re paying for theater.</p><p><strong>Treat behavioral cost as the primary risk.</strong> Model rate cards will keep coming down. They are not your problem. Your problem is what your pipeline asks of the model and what the model decides to do once asked. That&#8217;s the line item that grew while the unit cost dropped 280&#215;. That&#8217;s the shoe that just dropped.</p><h2>What This Means If You&#8217;re Running an AI Budget</h2><p>Three things to look at, in order of how much they&#8217;ll move the line item.</p><p><strong>Audit the call graph, not the rate card.</strong> Pull a representative day of production traffic and trace the actual LLM calls per user task. Count them. Most teams find a handful of workflows producing the majority of cost, and most of those have 2&#8211;4 LLM calls that could be replaced by deterministic code. That&#8217;s the consolidated-design pattern from the duetto-intelligence example. 50&#8211;80% reductions are common when you actually look.</p><p><strong>Put quality gates in front of inference.</strong> For any workflow where the LLM call is expensive and the input quality is variable, add a deterministic check that decides whether the call is worth making. That&#8217;s the search-first pattern from the code-intelligence example. The savings come from the calls you <em>don&#8217;t</em> make, which never show up on the invoice.</p><p><strong>Set hard context budgets and enforce them in code.</strong> Per-source character limits, per-call token caps, no &#8220;just in case&#8221; context stuffing. The output budget gets attention because it&#8217;s visible. The input budget is usually where the actual money goes.</p><p>None of this requires changing models, switching providers, or making bets on the next frontier release. It&#8217;s architectural work inside the pipeline you already have &#8212; work the per-token price chart has been letting people defer.</p><p>The teams that do it over the next two quarters will look like they got a 5&#8211;10&#215; cost improvement from &#8220;AI getting cheaper.&#8221; The teams that don&#8217;t will look like AI got 3&#8211;4&#215; more expensive while everyone else&#8217;s costs fell. Same providers, same models, same rate cards. Different shoe.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and writes about AI-powered engineering at <a href="https://hyperdev.substack.com/">HyperDev</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/p/opus-46-vs-47-the-real-cost-of-incremental">Opus 4.6 vs 4.7: The Real Cost of Incremental AI Improvements</a> &#8212; The first shoe, on per-task cost drift between model versions</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul><p></p>]]></content:encoded></item><item><title><![CDATA[Is Your Digital Brain the Light Saber of the AI Era?]]></title><description><![CDATA[Jedi Knight Tools for the Knowledge Worker]]></description><link>https://hyperdev.matsuoka.com/p/is-your-digital-brain-the-light-saber</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/is-your-digital-brain-the-light-saber</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Wed, 20 May 2026 12:02:03 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!90qs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!90qs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!90qs!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png 424w, https://substackcdn.com/image/fetch/$s_!90qs!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png 848w, https://substackcdn.com/image/fetch/$s_!90qs!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png 1272w, https://substackcdn.com/image/fetch/$s_!90qs!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!90qs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png" width="1456" height="1007" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1007,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5326200,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/198515384?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!90qs!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png 424w, https://substackcdn.com/image/fetch/$s_!90qs!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png 848w, https://substackcdn.com/image/fetch/$s_!90qs!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png 1272w, https://substackcdn.com/image/fetch/$s_!90qs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6533e48e-54c3-4232-aa80-b2fe447f0136_1993x1378.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Your Digital Lightsaber is your AI Memory</figcaption></figure></div><p>My Duetto colleague Jake Becker is sharp. He&#8217;s been ahead of AI adoption on our team &#8212; experimenting early, pushing for new tools, staying current. Last week he messaged me on Slack: &#8220;I wish I had my own CTO Assistant. Like what you have.&#8221;</p><p>I paused.</p><p>I have one. I&#8217;ve had several, in different forms, going back months. But I hadn&#8217;t said anything about it publicly.</p><p>That gap &#8212; between an AI-forward person who knows the tools and a practitioner who actually has the thing &#8212; is what this piece is about.</p><p><strong>TL;DR</strong></p><ul><li><p>Off-the-shelf AI tools are capable but contextually blind. You still have to leave your domain to get help.</p></li><li><p>Serious knowledge workers &#8212; writers, researchers, historians &#8212; have always built their own knowledge systems. Never trusted a vendor to hold their material.</p></li><li><p>Coding is shifting toward what writing has always been: directing, curating, maintaining a living body of knowledge.</p></li><li><p>Building your own AI memory and search layer is the new rite of passage. The tool you build IS the skill being developed.</p></li><li><p>The latest versions of my own stack: trusty-memory and trusty-search &#8212; both open source, both installable today.</p></li></ul><h2>Everyone Wants One. Few Have One.</h2><p>Jake&#8217;s comment revealed something I&#8217;d been taking for granted.</p><p>The major vendors are actively working on this &#8212; memory features, RAG pipelines, personalization layers. But those solutions are optimized for breadth, not for one person&#8217;s actual work across functions.</p><p>ChatGPT, Claude, Copilot &#8212; these are all capable. They&#8217;re also still contextually blind to <em>you</em>. You have to leave your environment, paste in context, explain your situation from scratch, then interpret the output back into your actual work. Vendors are working on it. But the solutions remain generic, and every session still starts from zero.</p><p>I wrote about this in March &#8212; <a href="https://hyperdev.matsuoka.com/personal-bots-abomination">Everyone Blamed Clawd Bot&#8217;s Execution. The Concept Was the Problem.</a> The structural flaw of universal assistants isn&#8217;t fixable. They require you to leave your context to get help. What actually works is the opposite: your tools get assistant capabilities, and assistance comes to where your context lives.</p><p>Off-the-shelf tools haven&#8217;t solved this. They&#8217;ve gotten more powerful &#8212; better reasoning, longer context, faster inference &#8212; but they still don&#8217;t know your codebase, your decisions, your institutional history, your current sprint. They know a lot about the world, and very little about you.</p><p>Jake wanted <em>my</em> assistant. But what he actually wants is <em>his</em> assistant. The one that knows what he knows.</p><p>That&#8217;s a different problem entirely.</p><p>The job changed first.</p><h2>Coding Is Becoming Writing</h2><p>Practitioners feel it before analysts name it.</p><p>A few years ago, being a strong engineer meant writing a lot of code quickly and correctly. Today, with agentic AI coders at their disposal, the best engineers I watch spend their time directing, reviewing, specifying, and curating. The unit of work has moved up a level. Implementation is increasingly delegated. Judgment &#8212; about architecture, trade-offs, what to build and why &#8212; is the differentiator.</p><p>This is not what happens when automation replaces a skill. It&#8217;s what happens when a new discipline appears.</p><p>Writers have always worked this way. A novelist doesn&#8217;t produce words per minute as a primary metric. They produce decisions &#8212; what to say, in what order, with what emphasis. The words are the output of the decisions, not the work itself. What makes a writer productive over a career isn&#8217;t typing speed. It&#8217;s having a system: notes, research, accumulated material, patterns of thought that compound over years.</p><p>Some (typically senior) engineers are arriving at the same realization. Your value isn&#8217;t the code. It&#8217;s the judgment, the accumulated context, the knowledge of what was tried and why it failed. The question is whether that accumulates in your head alone &#8212; which doesn&#8217;t scale, and doesn&#8217;t survive a context switch &#8212; or whether it lives in a system.</p><p>Good engineers who learn to use their AI tools effectively generate better code &#8212; the stack amplifies judgment and accumulated knowledge. The gap isn&#8217;t between skilled and unskilled engineers in isolation; it&#8217;s between engineers who&#8217;ve wired their knowledge into their tools and those who haven&#8217;t. The knowledge is real in both cases. Only in one case does it compound.</p><p>The writers I&#8217;ve observed who sustain serious output over decades all have the same property: they know where things are. Their research is retrievable. Their earlier thinking is available to their current thinking. The system makes the person bigger than their working memory.</p><p>That&#8217;s the gap.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Ynth!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd8f91e4-5bab-49cc-8f75-bb43db12e9e7_1999x1227.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Ynth!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd8f91e4-5bab-49cc-8f75-bb43db12e9e7_1999x1227.png 424w, https://substackcdn.com/image/fetch/$s_!Ynth!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd8f91e4-5bab-49cc-8f75-bb43db12e9e7_1999x1227.png 848w, https://substackcdn.com/image/fetch/$s_!Ynth!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd8f91e4-5bab-49cc-8f75-bb43db12e9e7_1999x1227.png 1272w, https://substackcdn.com/image/fetch/$s_!Ynth!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd8f91e4-5bab-49cc-8f75-bb43db12e9e7_1999x1227.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Ynth!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd8f91e4-5bab-49cc-8f75-bb43db12e9e7_1999x1227.png" width="1456" height="894" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dd8f91e4-5bab-49cc-8f75-bb43db12e9e7_1999x1227.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:894,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:5603002,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/198515384?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd8f91e4-5bab-49cc-8f75-bb43db12e9e7_1999x1227.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Ynth!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd8f91e4-5bab-49cc-8f75-bb43db12e9e7_1999x1227.png 424w, https://substackcdn.com/image/fetch/$s_!Ynth!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd8f91e4-5bab-49cc-8f75-bb43db12e9e7_1999x1227.png 848w, https://substackcdn.com/image/fetch/$s_!Ynth!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd8f91e4-5bab-49cc-8f75-bb43db12e9e7_1999x1227.png 1272w, https://substackcdn.com/image/fetch/$s_!Ynth!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd8f91e4-5bab-49cc-8f75-bb43db12e9e7_1999x1227.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"> The AI Zettelkasten</figcaption></figure></div><h2>Writers Don&#8217;t Trust Vendors With Their Material</h2><p>Niklas Luhmann, a German sociologist working in the 1950s, produced 70 books and nearly 400 articles over his career. I haven&#8217;t written a book yet &#8212; but I have published over 200 articles at HyperDev. Worth naming the parallel. He credits his output to his Zettelkasten &#8212; a slip-box system of 90,000 interconnected index cards, each with a unique identifier linking it to related thoughts. Not a filing cabinet. A network of ideas that got richer with every addition.</p><p>And it&#8217;s worth asking: is that so different from what an AI knowledge system does today? The Zettelkasten was an analog precursor to what trusty-memory and trusty-search do programmatically &#8212; indexing ideas, linking related thoughts, surfacing connections that wouldn&#8217;t otherwise be visible. Luhmann was doing manually what these tools do automatically. Same architecture. New substrate.</p><p>He didn&#8217;t use a vendor product. He built a system that reflected how he thought. The architecture of his Zettelkasten was itself an expression of his intellectual method.</p><p>This isn&#8217;t a historical quirk. It&#8217;s a pattern. Serious knowledge workers have always built their own systems &#8212; commonplace books, research archives, private wikis. The reason is structural: the schema you design reflects how you think. That&#8217;s not something a product gives you. A product gives everyone the same schema.</p><p>Andrej Karpathy pointed at something similar last month with his LLM Wiki gist. His framing: use LLMs not just to write code, but to build and maintain a personal knowledge base. &#8220;Obsidian is the IDE, the LLM is the programmer, the wiki is the codebase.&#8221; Three folders, structured Markdown, a large context window. He concluded: &#8220;I think there is room here for an incredible new product.&#8221;</p><p>He&#8217;s right there&#8217;s room. I wrote about his framing in <a href="https://hyperdev.matsuoka.com/whats-in-your-second-brain">What&#8217;s In Your Second Brain?</a> The product comment is where I&#8217;d push back. You can build tooling around the pattern. You can&#8217;t productize the schema. The schema is the moat &#8212; because it reflects how <em>you</em> think, not how a product manager thinks you think. The ones who get it aren&#8217;t waiting for a product.</p><h2>The Lightsaber Rite of Passage</h2><p>In Star Wars canon, a Padawan doesn&#8217;t receive a lightsaber. They build one.</p><p>The ritual is called the Gathering. Initiates travel alone to the Crystal Caves of Ilum. They have to find their kyber crystal &#8212; the crystal that&#8217;s attuned to them through the Force. The caves are shaped by the initiate&#8217;s own fears and insecurities. The crystal doesn&#8217;t go to the strongest or the fastest. It bonds with the person who confronts what&#8217;s in the way.</p><p>Then they build it themselves, guided by Professor Huyang.</p><p>You can&#8217;t buy this. You can&#8217;t inherit it. The construction is the training. The tool reflects the builder.</p><p>I&#8217;m not the first to reach for this metaphor in tech. But I think it lands differently now. Building your own AI memory and search layer isn&#8217;t just useful. It&#8217;s diagnostic. You can&#8217;t do it without confronting what you actually know, how you actually think, what deserves to persist and what doesn&#8217;t. The schema you design for your knowledge base is a statement about your mind.</p><p>The engineers I know who are operating at the highest level right now &#8212; CTOs, senior architects, tech leads at places moving fast &#8212; they all quietly roll their own. They don&#8217;t announce it. They just have it.</p><h2>My Own Lineage</h2><p>I&#8217;ve been building versions of this for months.</p><p>trusty-izzie was the first &#8212; a simple wrapper. Then ai-commander, a more structured approach to context management. Then open-mpm and claude-mpm, which was where I started thinking seriously about multi-agent orchestration. Then kuzu-memory, a graph-backed memory layer. Then mcp-vector-search, semantic search over my entire codebase.</p><p>Each iteration taught me something about what I actually needed. Not what I thought I needed. What the practice revealed.</p><p>This piece was drafted with a configured writing assistant &#8212; <a href="https://github.com/bobmatnyc/claude-mpm">claude-mpm</a> loaded with my publication style guide, my voice patterns, my article archive. That&#8217;s a saber too. Not a generic chat interface. A tool shaped around how I think and write, producing work I can actually publish rather than work I have to fix. The saber list keeps growing.</p><p>The latest two are the most capable tools I&#8217;ve built.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cyse!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fc209a6-05f2-4b6d-8950-09c7562bd1ac_2005x1419.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cyse!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fc209a6-05f2-4b6d-8950-09c7562bd1ac_2005x1419.png 424w, https://substackcdn.com/image/fetch/$s_!cyse!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fc209a6-05f2-4b6d-8950-09c7562bd1ac_2005x1419.png 848w, https://substackcdn.com/image/fetch/$s_!cyse!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fc209a6-05f2-4b6d-8950-09c7562bd1ac_2005x1419.png 1272w, https://substackcdn.com/image/fetch/$s_!cyse!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fc209a6-05f2-4b6d-8950-09c7562bd1ac_2005x1419.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cyse!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fc209a6-05f2-4b6d-8950-09c7562bd1ac_2005x1419.png" width="1456" height="1030" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0fc209a6-05f2-4b6d-8950-09c7562bd1ac_2005x1419.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1030,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:6229361,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/198515384?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fc209a6-05f2-4b6d-8950-09c7562bd1ac_2005x1419.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!cyse!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fc209a6-05f2-4b6d-8950-09c7562bd1ac_2005x1419.png 424w, https://substackcdn.com/image/fetch/$s_!cyse!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fc209a6-05f2-4b6d-8950-09c7562bd1ac_2005x1419.png 848w, https://substackcdn.com/image/fetch/$s_!cyse!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fc209a6-05f2-4b6d-8950-09c7562bd1ac_2005x1419.png 1272w, https://substackcdn.com/image/fetch/$s_!cyse!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0fc209a6-05f2-4b6d-8950-09c7562bd1ac_2005x1419.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The Latest Sabers</h2><p><strong>trusty-memory</strong> is a machine-wide AI memory daemon written in Rust. It uses what I call the Memory Palace architecture &#8212; multiple named palaces, each for a different domain. Sub-5ms baseline retrieval on Apple Silicon. It runs as an MCP server for Claude Code, which means my assistant stores and retrieves memories automatically, across sessions, without me managing any of it explicitly.</p><pre><code><code>cargo install trusty-memory</code></code></pre><p>Available at <a href="https://crates.io/crates/trusty-memory">crates.io/crates/trusty-memory</a>.</p><p><strong>trusty-search</strong> is a machine-wide hybrid code search daemon, also in Rust. Always-on, one install per machine. It combines BM25 lexical search with HNSW vector search (all-MiniLM-L6-v2 INT8) and a Knowledge Graph with 1-2 hop expansion, fused via Reciprocal Rank Fusion. It exposes an MCP server with 11 tools. Stdio and HTTP/SSE transports drop straight into Claude Code.</p><pre><code><code>cargo install trusty-search</code></code></pre><p>Available at <a href="https://crates.io/crates/trusty-search">crates.io/crates/trusty-search</a>.</p><p>These tools, along a set of custom reporting pythons apps along with a custom Slack Bot I use to access the data remotely, comprise my digital brain.</p><p>These aren&#8217;t products I bought. These are tools I built, iterated, and use daily. They know my codebase the way a Zettelkasten knows a scholar&#8217;s intellectual territory &#8212; not because a vendor configured them, but because I did.</p><p>To be precise: trusty-memory and trusty-search are infrastructure utilities &#8212; the memory layer and the search layer. Building the actual assistant that uses them is a separate act of customization. That&#8217;s where the lightsaber metaphor completes: the kyber crystal is only part of it. The construction &#8212; what you build with the crystal &#8212; is the saber.</p><p>When Jake said he wished he had a CTO Assistant, this is what he was gesturing at. Not a prompt template. Not a workflow. A living knowledge layer that compounds.</p><h2>The Right Question</h2><p>Jake asked: &#8220;Can I get a CTO Assistant?&#8221;</p><p>That&#8217;s the wrong question. It assumes the thing is available off the shelf, and the task is finding and configuring it.</p><p>The right question is: &#8220;What would it take to build one that knows what I know?&#8221;</p><p>That question is harder. It requires confronting the shape of your knowledge, what&#8217;s worth persisting, how to structure retrieval. It&#8217;s uncomfortable in the same way the Crystal Caves are uncomfortable &#8212; not because the work is technically difficult, but because you have to be honest about what you actually have.</p><p>Not everyone needs to write Rust. The specific technology isn&#8217;t the point. The point is that the engineers asking the right question are already operating differently. They&#8217;re working like writers &#8212; maintaining a living body of knowledge, building systems that compound, treating their accumulated context as an asset rather than a liability.</p><p>Writers&#8217; discipline has been creeping into engineering for a while. AI made it urgent.</p><p>If you&#8217;re waiting for a product to hand you the thing, you&#8217;re waiting for someone to build your lightsaber. It won&#8217;t be yours.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and writes about AI-powered engineering at <a href="https://hyperdev.substack.com/">HyperDev</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/whats-in-your-second-brain">What&#8217;s In Your Second Brain?</a> &#8212; Karpathy&#8217;s LLM Wiki and the case for a compounding knowledge layer</p></li><li><p><a href="https://hyperdev.matsuoka.com/personal-bots-abomination">Everyone Blamed Clawd Bot&#8217;s Execution. The Concept Was the Problem.</a> &#8212; Why universal assistants are architecturally broken</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul>]]></content:encoded></item><item><title><![CDATA[What’s In Your Second Brain?]]></title><description><![CDATA[Tooling for the modern CTO]]></description><link>https://hyperdev.matsuoka.com/p/whats-in-your-second-brain</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/whats-in-your-second-brain</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Mon, 04 May 2026 13:26:26 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!oH6q!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30309841-bfd3-4d55-a5d2-03d8d872cc9b_1024x649.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!oH6q!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30309841-bfd3-4d55-a5d2-03d8d872cc9b_1024x649.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!oH6q!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30309841-bfd3-4d55-a5d2-03d8d872cc9b_1024x649.png 424w, https://substackcdn.com/image/fetch/$s_!oH6q!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30309841-bfd3-4d55-a5d2-03d8d872cc9b_1024x649.png 848w, https://substackcdn.com/image/fetch/$s_!oH6q!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30309841-bfd3-4d55-a5d2-03d8d872cc9b_1024x649.png 1272w, https://substackcdn.com/image/fetch/$s_!oH6q!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30309841-bfd3-4d55-a5d2-03d8d872cc9b_1024x649.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!oH6q!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30309841-bfd3-4d55-a5d2-03d8d872cc9b_1024x649.png" width="1024" height="649" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/30309841-bfd3-4d55-a5d2-03d8d872cc9b_1024x649.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:649,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1458693,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/196419582?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa25c8af5-4b58-4f5d-b775-532bd485770e_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!oH6q!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30309841-bfd3-4d55-a5d2-03d8d872cc9b_1024x649.png 424w, https://substackcdn.com/image/fetch/$s_!oH6q!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30309841-bfd3-4d55-a5d2-03d8d872cc9b_1024x649.png 848w, https://substackcdn.com/image/fetch/$s_!oH6q!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30309841-bfd3-4d55-a5d2-03d8d872cc9b_1024x649.png 1272w, https://substackcdn.com/image/fetch/$s_!oH6q!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F30309841-bfd3-4d55-a5d2-03d8d872cc9b_1024x649.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The modern CTO toolkit isn&#8217;t just apps and coding tools. The real differentiator is a custom knowledge layer &#8212; databases, search indices, memory graphs, behavioral instructions that compound over time. No product gives you this. You build it.</p><p><a href="https://karpathy.ai/">Andrej Karpathy</a> gestured at something similar last month when he posted a <a href="https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f">GitHub Gist</a> he called &#8220;LLM Wiki.&#8221; His framing: stop using LLMs <em>just</em> to write code, use them to build and maintain a personal knowledge base instead. <em>&#8220;Obsidian is the IDE, the LLM is the programmer, the wiki is the codebase.&#8221;</em> Three folders, structured Markdown, a large context window, a few Python scripts. No RAG, no vector database. He concluded with: <em>&#8220;I think there is room here for an incredible new product.&#8221;</em></p><p>He&#8217;s right that there&#8217;s room. But the product comment is where I&#8217;d push back, and I&#8217;ll get to that. What Karpathy is describing isn&#8217;t a note-taking system. It&#8217;s a personal operational knowledge layer. For CTOs specifically, that layer needs to be broader than a personal wiki &#8212; it needs live organizational data, agent-connected search, and context that persists across months of decisions. No app hands you that.</p><h2>TL;DR</h2><ul><li><p>Karpathy&#8217;s LLM Wiki shows the direction: LLMs as knowledge compilers, not just code generators</p></li><li><p>A modern CTO&#8217;s &#8220;second brain&#8221; is more than PKM &#8212; it&#8217;s live databases, custom agents, and contextual search across organizational data</p></li><li><p>When I joined Duetto as CTO, my custom toolkit let me synthesize a 150-person R&amp;D org in weeks instead of months</p></li><li><p>The power isn&#8217;t Obsidian. It&#8217;s what you connect to it &#8212; MCP servers, search indices, knowledge graphs</p></li><li><p>Productizing this is theoretically possible and practically very hard, because the schema is the moat</p></li></ul><h2>The toolkit article got it half right</h2><p>In <a href="https://hyperdev.matsuoka.com/p/whats-in-my-claude-code-toolkit">What&#8217;s In My Toolkit: Claude Code and Family</a>, I wrote about vanilla Claude Code&#8217;s core limitations: context evaporates, code search is keyword-based, memory doesn&#8217;t persist, execution is single-threaded. The tools I built &#8212; <a href="https://github.com/bobmatnyc/claude-mpm">Claude MPM</a>, <a href="https://github.com/bobmatnyc/mcp-vector-search">mcp-vector-search</a>, <a href="https://github.com/bobmatnyc/kuzu-memory">kuzu-memory</a> &#8212; address each of those gaps.</p><p>But that article was about coding workflows. The real story is broader.</p><p>The same architecture that makes a coding session more effective &#8212; persistent memory, semantic search, specialized agents pulling from structured data &#8212; turns out to be extraordinarily useful for executive work. Understanding an organization, tracking decisions over time, querying data across systems, maintaining context across months of meetings and analysis. The toolkit I built for software development became the toolkit I used to onboard as a CTO.</p><p><a href="https://hyperdev.matsuoka.com/p/i-built-a-coding-tool-then-i-used">That onboarding story</a> is documented in detail elsewhere. Short version: I pointed a multi-agent framework at GitHub, JIRA, Slack, Confluence, and budget spreadsheets, and synthesized a 150-person R&amp;D organization in the weeks before my start date. The difference between doing that with a chat interface versus a CLI-based orchestration layer with parallel agents and persistent memory wasn&#8217;t 2x or 5x. It was closer to 10x.</p><p>But the onboarding was just the starting gun. The second brain I assembled keeps compounding.</p><h2>What&#8217;s actually in my second brain</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ICJf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd86cffcd-de7a-4db2-8ab3-ec6f86df65db_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ICJf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd86cffcd-de7a-4db2-8ab3-ec6f86df65db_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!ICJf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd86cffcd-de7a-4db2-8ab3-ec6f86df65db_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!ICJf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd86cffcd-de7a-4db2-8ab3-ec6f86df65db_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!ICJf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd86cffcd-de7a-4db2-8ab3-ec6f86df65db_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ICJf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd86cffcd-de7a-4db2-8ab3-ec6f86df65db_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d86cffcd-de7a-4db2-8ab3-ec6f86df65db_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1082407,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/196419582?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd86cffcd-de7a-4db2-8ab3-ec6f86df65db_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ICJf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd86cffcd-de7a-4db2-8ab3-ec6f86df65db_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!ICJf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd86cffcd-de7a-4db2-8ab3-ec6f86df65db_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!ICJf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd86cffcd-de7a-4db2-8ab3-ec6f86df65db_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!ICJf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd86cffcd-de7a-4db2-8ab3-ec6f86df65db_1024x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Let me be specific. Because when people (now) hear &#8220;second brain&#8221; they usually think Obsidian vaults with color-coded tags and pretty Markdown files. That&#8217;s part of it. It&#8217;s the surface layer.</p><p>The actual power comes from what&#8217;s underneath.</p><h3>The memory layer</h3><p><a href="https://github.com/bobmatnyc/kuzu-memory">kuzu-memory</a> is a KuzuDB-backed knowledge graph that persists across every AI session. It stores learnings from conversations, code commits, decisions, patterns. When I start a new Claude Code session on a problem I&#8217;ve touched before, the context isn&#8217;t blank &#8212; it&#8217;s enriched with what was learned the last time.</p><p>This is the thing people underestimate. A project-specific memory that accumulates over months of work develops a kind of organizational intelligence you can&#8217;t replicate in a single conversation. It knows why a particular architectural decision was made. It knows that a vendor was evaluated and found lacking. It knows the terminology your team uses internally that differs from industry standard.</p><p>KuzuDB isn&#8217;t a product choice for its own sake &#8212; it&#8217;s graph-native, which means it handles relationships well. The connections between people, systems, decisions, and code are as important as the facts themselves.</p><h3>The search layer</h3><p><a href="https://github.com/bobmatnyc/mcp-vector-search">mcp-vector-search</a> provides semantic search across all project files. Not keyword search &#8212; semantic search with AST parsing. When I ask &#8220;where is the analysis I did on contractor productivity last quarter,&#8221; it finds it even if the document never uses those exact words.</p><p>At Duetto, this covers everything in my CTO project: architecture records, meeting notes pulled from Granola, emails I&#8217;ve synthesized, analysis documents, planning artifacts. Months of accumulated context, all searchable in seconds. The underlying code intelligence for the engineering organization runs as a separate service &#8212; mcp-vector-search is for my working knowledge, not the codebase itself.</p><h3>The databases</h3><p>My CTO project has three:</p><ul><li><p><strong>cto.db</strong> &#8212; SQLite. Work classification, people analysis, contributor data, commit history. The operational database for running analyses and reports.</p></li><li><p><strong>analytics.duckdb</strong> &#8212; DuckDB. OLAP queries and analytics. When I need to slice engineering output data in different ways or run something that would be painful in SQLite, it goes here.</p></li><li><p><strong>duetto_knowledge.db</strong> &#8212; The RAG-queryable knowledge base backing a Flask web app for interactive exploration.</p></li></ul><p>These aren&#8217;t a product I bought. They&#8217;re a schema I designed, built incrementally, and own completely. The schema reflects how I think about the organization, which is precisely why it&#8217;s useful.</p><h3>The connectors</h3><p><a href="https://github.com/bobmatnyc/gworkspace-mcp">gworkspace-mcp</a> handles Drive, Docs, Sheets, Gmail, Calendar, and more. I wrote my own rather than using the off-the-shelf options &#8212; Google&#8217;s first-party integration and Anthropic&#8217;s default both have significant tool coverage gaps. Mine exposes substantially more of the Workspace API surface and integrates transparently with Claude MPM, so agents can use Google Workspace tools without any special configuration at the call site.</p><p>Beyond Workspace: Notion API for product specs and planning documents. Extraction scripts for JIRA, Confluence, Slack, Datadog, and AWS. Each system outputs to as raw data, which feeds analysis pipelines that generate reports stored in a project directory.</p><p>For company-wide memory, two more tools: duetto-memory and duetto-directory. These handle shared organizational context &#8212; information that needs to flow between tools and across team members rather than staying in a single session. Memory persists within our VPC, encrypted to individual users&#8217; OAuth keys. Not even our own IT has access to it. Context shared from Claude Code shows up in Claude.ai, and vice versa, without any manual sync.</p><p>The entire flow is queryable. From a single Claude session, I can ask about budget trends, team velocity, specific architectural decisions, or what a particular engineer has been working on for the last three months. Because it&#8217;s all in the same context-addressable system.</p><h3>Obsidian as the front door</h3><p>Yes, I use Obsidian. But it&#8217;s a front door, not the building. The vault holds my personal notes, research captures, and synthesized analysis. The Obsidian Web Clipper feeds raw material into the knowledge pipeline. Templates enforce consistent structure.</p><p>Karpathy&#8217;s insight about Obsidian as IDE is right in the narrow sense: it&#8217;s the interface you use to read and organize. But the interesting work happens outside it &#8212; in the databases, the agents, the search indices, the custom scripts.</p><h2>CLAUDE.md files everywhere</h2><p>The context layer isn&#8217;t just data. It&#8217;s also behavioral instructions.</p><p>Every major directory in my project has a CLAUDE.md. The root CTO project one is 400 lines of conventions, routing logic, document lifecycle rules, and architectural decisions. Every subdirectory has a more focused version. Every specialized agent has its own constraints.</p><p>These files are my second brain&#8217;s schema, expressed as instructions rather than data. A single routing rule &#8212; &#8220;if the prompt mentions meeting notes, save to <code>projects/meetings/2026-W##/</code>&#8220; &#8212; sounds trivial. But it means twelve months of meeting notes accumulate in consistent, queryable locations rather than wherever an agent happened to save them. Multiply that by forty routing rules across fifteen subdirectories, and the entire corpus becomes navigable. The CLAUDE.md files are what make the databases useful. Without them, the data is just data.</p><p>Karpathy put it well: &#8220;You share the schema, not the code.&#8221; The schema is the valuable part. The schema is what compounds.</p><p>My schema took months to build. It will keep getting better. No product ships with the right schema for my organization, because no product knows what I know about how Duetto&#8217;s R&amp;D works.</p><h2>The productization question</h2><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!-gqt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87551032-97c3-4a89-8697-729269302cfa_1024x708.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!-gqt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87551032-97c3-4a89-8697-729269302cfa_1024x708.png 424w, https://substackcdn.com/image/fetch/$s_!-gqt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87551032-97c3-4a89-8697-729269302cfa_1024x708.png 848w, https://substackcdn.com/image/fetch/$s_!-gqt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87551032-97c3-4a89-8697-729269302cfa_1024x708.png 1272w, https://substackcdn.com/image/fetch/$s_!-gqt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87551032-97c3-4a89-8697-729269302cfa_1024x708.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!-gqt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87551032-97c3-4a89-8697-729269302cfa_1024x708.png" width="1024" height="708" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/87551032-97c3-4a89-8697-729269302cfa_1024x708.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:708,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1454118,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/196419582?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F85bbeb78-8523-459e-9cb0-e65b1d5bce18_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!-gqt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87551032-97c3-4a89-8697-729269302cfa_1024x708.png 424w, https://substackcdn.com/image/fetch/$s_!-gqt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87551032-97c3-4a89-8697-729269302cfa_1024x708.png 848w, https://substackcdn.com/image/fetch/$s_!-gqt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87551032-97c3-4a89-8697-729269302cfa_1024x708.png 1272w, https://substackcdn.com/image/fetch/$s_!-gqt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87551032-97c3-4a89-8697-729269302cfa_1024x708.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Karpathy said there&#8217;s room for an incredible product. He&#8217;s not wrong about the gap. He might be wrong about the solution.</p><p>The structural problems with productizing a second brain:</p><p><strong>Context compounds, products don&#8217;t.</strong> My system gets smarter with every commit, meeting, and conversation. A SaaS product serves thousands of customers and maintains no one&#8217;s specific context. The more I use my system, the wider the gap between it and any off-the-shelf alternative.</p><p><strong>The schema is the moat.</strong> My knowledge architecture reflects how I think about engineering organizations. Someone else&#8217;s knowledge architecture would be different. Products that force their schema on you &#8212; and every product does &#8212; are imposing someone else&#8217;s way of thinking on your problem. That friction is small at first and grows over time.</p><p><strong>Privacy is structural, not incidental.</strong> My databases contain org structures, salary data, performance patterns, vendor negotiations. Routing that through third-party infrastructure creates risk that&#8217;s practically impossible to contain. When I built duetto-memory for enterprise use, the entire stack stays within our VPC, with memories encrypted to individual users&#8217; OAuth keys. Not even IT can read them. That level of isolation is nearly impossible to provide as a multi-tenant SaaS.</p><p>Some layers could be productized &#8212; the infrastructure, not the intelligence. A well-designed memory MCP with sensible defaults. Semantic search that works without configuration. Privacy-preserving graph storage you don&#8217;t have to host yourself. The plumbing.</p><p>The schema, the decisions, and the accumulated context can&#8217;t be productized. Those are yours. That&#8217;s the point &#8212; and it&#8217;s also why the product gap Karpathy sees will remain open even after someone tries to fill it.</p><h2>Can anyone do this?</h2><p>There&#8217;s an access problem here, and I&#8217;d be dishonest not to acknowledge it.</p><p>Building what I&#8217;ve described requires knowing Python well enough to write extraction scripts, understanding enough about graph databases to design a schema, and being comfortable with CLI-based tooling and MCP server configuration. Not every CTO has that background. Not every technical leader wants to spend weekends building personal infrastructure.</p><p>The irony is that the people who most need better organizational intelligence &#8212; executives without deep engineering backgrounds &#8212; are least equipped to build these systems. And the people who are most capable of building them are often less interested in the executive problems the systems could solve.</p><p>Tiago Forte, who wrote <em><a href="https://www.buildingasecondbrain.com/">Building a Second Brain</a></em>, has been making this point for years. His PARA method and CODE framework are accessibility layers &#8212; ways to make the underlying ideas approachable without requiring you to build a graph database. The methodology is sound. But it was designed for knowledge workers, not for CTOs running engineering organizations who need live data pipelines, not filing systems. Well-designed for whom?</p><p>Karpathy&#8217;s LLM Wiki is explicitly a system for someone comfortable writing Python and working with file systems. His Gist has code in it. That&#8217;s a feature for his audience and a barrier for everyone else.</p><h2>What I&#8217;d watch for</h2><p>A few trends that will determine whether this remains a DIY space or gets productized:</p><p><strong>MCP as infrastructure.</strong> The <a href="https://hyperdev.matsuoka.com/p/the-mcp-cat-is-out-of-the-bag">Model Context Protocol</a> creates a standard interface for exactly this kind of knowledge infrastructure. Memory servers, search servers, database connectors &#8212; they all expose the same interface to any compatible AI client. The ecosystem is growing fast. As more MCP servers mature, the configuration burden drops.</p><p><strong>Searchable context beats raw window size.</strong> Karpathy argues for plain Markdown because ~400K words fit in a modern context window. That&#8217;s true, and the window is getting larger. But the more important shift is that structured, searchable context doesn&#8217;t have a ceiling. A well-organized knowledge base that spans years of meetings, decisions, and analysis delivers more than any single context window can hold &#8212; and the value scales with the quality of the organization, not the size of the model.</p><p><strong>Local model quality.</strong> Karpathy runs Anthropic agents via Claude Code. But local model quality is improving fast. A system that uses the cloud API for synthesis and queries but runs a local model for routine indexing tasks would be significantly cheaper and more private. Not ready yet. Getting closer.</p><p>The product Karpathy thinks exists &#8212; if it gets built &#8212; probably looks like a well-designed local MCP server with clean configuration, sensible defaults, and a plugin ecosystem for connectors. Not a SaaS. Not a cloud database. Something you install and own.</p><p>The people who need it most will have already built their own before any product ships. And in the process of building it, they&#8217;ll have accumulated the one thing no product can give them: months of their own operational context, organized the way their own mind works.</p><p>That&#8217;s not a consolation prize. That&#8217;s the whole point.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and writes about AI-powered engineering at <a href="https://hyperdev.substack.com/">HyperDev</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/p/whats-in-my-claude-code-toolkit">What&#8217;s In My Toolkit: Claude Code and Family</a> &#8212; The coding layer of the stack</p></li><li><p><a href="https://hyperdev.matsuoka.com/p/i-built-a-coding-tool-then-i-used">I Built a Coding Tool. Then I Used It to Onboard as CTO</a> &#8212; Applying agent orchestration to organizational analysis</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul>]]></content:encoded></item><item><title><![CDATA[It’s The Harness, Stupid!]]></title><description><![CDATA[Why AI tool orchestration now matters more than foundation model quality]]></description><link>https://hyperdev.matsuoka.com/p/its-the-harness-stupid</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/its-the-harness-stupid</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Mon, 13 Apr 2026 17:17:16 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!376u!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c1f4dab-a192-4b50-bbbb-5c889e29af13_1024x523.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!376u!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c1f4dab-a192-4b50-bbbb-5c889e29af13_1024x523.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!376u!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c1f4dab-a192-4b50-bbbb-5c889e29af13_1024x523.png 424w, https://substackcdn.com/image/fetch/$s_!376u!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c1f4dab-a192-4b50-bbbb-5c889e29af13_1024x523.png 848w, https://substackcdn.com/image/fetch/$s_!376u!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c1f4dab-a192-4b50-bbbb-5c889e29af13_1024x523.png 1272w, https://substackcdn.com/image/fetch/$s_!376u!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c1f4dab-a192-4b50-bbbb-5c889e29af13_1024x523.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!376u!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c1f4dab-a192-4b50-bbbb-5c889e29af13_1024x523.png" width="1024" height="523" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5c1f4dab-a192-4b50-bbbb-5c889e29af13_1024x523.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:523,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1201765,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/193459844?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4af78c41-cc16-45d4-abb7-f0e0ee92722b_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!376u!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c1f4dab-a192-4b50-bbbb-5c889e29af13_1024x523.png 424w, https://substackcdn.com/image/fetch/$s_!376u!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c1f4dab-a192-4b50-bbbb-5c889e29af13_1024x523.png 848w, https://substackcdn.com/image/fetch/$s_!376u!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c1f4dab-a192-4b50-bbbb-5c889e29af13_1024x523.png 1272w, https://substackcdn.com/image/fetch/$s_!376u!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5c1f4dab-a192-4b50-bbbb-5c889e29af13_1024x523.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>It&#8217;s The Harness, Stupid!</h2><p><strong>Why AI tool orchestration now matters more than foundation model quality</strong></p><p><em>Author: Bob Matsuoka, CTO @ Duetto Research</em> <br><em>April 6, 2026</em></p><h2>TL;DR</h2><ul><li><p>Same-model testing reveals 0.82-point quality spread (3.93 to 4.75) and 7x efficiency differences&#8212;orchestration dominates outcomes</p></li><li><p>Market validation: Claude maintains 70% developer preference despite GPT-5.4 achieving model parity through superior harness quality</p></li><li><p>Reddit analysis confirms Codex efficiency gains come from orchestration improvements, not just model upgrades</p></li><li><p>Competitive advantage has shifted permanently from model superiority to ecosystem superiority</p></li></ul><p><strong>Bottom line: The harness era has begun. Choose tools based on workflow fit, not benchmark claims.</strong></p><h2>The $50B Model Myth</h2><p>The AI industry has a fixation problem. Every week brings breathless announcements about parameter counts, training costs, and benchmark scores. &#8220;GPT-6 has 50 trillion parameters!&#8221; &#8220;Our model scored 94.7% on SWE-bench!&#8221; &#8220;We spent $2 billion on compute!&#8221;</p><p>Three converging pieces of evidence prove this approach is fundamentally wrong.</p><div class="callout-block" data-callout="true"><p><strong>Evidence #1:</strong> I tested eight AI coding agents across five programming challenges. Four agents used identical Claude Sonnet 4.6 models. Quality scores ranged from 3.93 to 4.75&#8212;a 0.82-point spread on the same foundation model.</p></div><div class="callout-block" data-callout="true"><p><strong>Evidence #2:</strong> GPT-5.4 achieved parity with Claude Sonnet 4.6 on coding benchmarks. Yet Claude maintains 70% developer preference through superior ecosystem quality.</p></div><div class="callout-block" data-callout="true"><p><strong>Evidence #3:</strong> Reddit developer communities confirm Codex&#8217;s efficiency improvements come from orchestration architecture changes, not just model upgrades.</p></div><p><strong>The harness matters more than the model.</strong> Choosing an AI coding tool is now primarily an engineering decision, not a model selection decision. The next competitive advantage isn&#8217;t bigger models&#8212;it&#8217;s better orchestration.</p><h2>Evidence Pillar #1: The Smoking Gun Laboratory Data</h2><h3>The Bake-Off Setup</h3><p>I designed five programming challenges ranging from 30-minute tasks to 8-hour full-stack builds:</p><ul><li><p><strong>Level 1-2:</strong> Simple scripts and basic applications</p></li><li><p><strong>Level 3:</strong> API integration with Docker containerization</p></li><li><p><strong>Level 4:</strong> Extensible data processing pipeline (architecture test)</p></li><li><p><strong>Level 5:</strong> Full-stack web application with authentication</p></li></ul><p>Eight agents competed: Claude Code, Claude MPM, Codex, Gemini CLI, Auggie, Qwen+Aider, DeepSeek+Aider, and Warp AI. Each received identical prompts. A panel of expert developers blind-reviewed all submissions across eight criteria: functionality, correctness, best practices, architecture, code reuse, testing, error handling, and documentation.</p><h3>The Harness Advantage Data</h3><p><strong>Table 1: Same Model, Different Worlds</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!TErH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47d85d71-6503-4452-ad8c-957585385133_867x409.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!TErH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47d85d71-6503-4452-ad8c-957585385133_867x409.png 424w, https://substackcdn.com/image/fetch/$s_!TErH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47d85d71-6503-4452-ad8c-957585385133_867x409.png 848w, https://substackcdn.com/image/fetch/$s_!TErH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47d85d71-6503-4452-ad8c-957585385133_867x409.png 1272w, https://substackcdn.com/image/fetch/$s_!TErH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47d85d71-6503-4452-ad8c-957585385133_867x409.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!TErH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47d85d71-6503-4452-ad8c-957585385133_867x409.png" width="867" height="409" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/47d85d71-6503-4452-ad8c-957585385133_867x409.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:409,&quot;width&quot;:867,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:74882,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/193459844?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47d85d71-6503-4452-ad8c-957585385133_867x409.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!TErH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47d85d71-6503-4452-ad8c-957585385133_867x409.png 424w, https://substackcdn.com/image/fetch/$s_!TErH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47d85d71-6503-4452-ad8c-957585385133_867x409.png 848w, https://substackcdn.com/image/fetch/$s_!TErH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47d85d71-6503-4452-ad8c-957585385133_867x409.png 1272w, https://substackcdn.com/image/fetch/$s_!TErH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F47d85d71-6503-4452-ad8c-957585385133_867x409.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Four agents using identical Claude Sonnet 4.6 models. Quality scores from 3.93 to 4.75&#8212;a 0.82-point spread. <a href="https://github.com/bobmatnyc/claude-mpm">claude-mpm</a> finished in 45 minutes while warp took 313 minutes. Almost <strong>7x longer for lower quality results</strong>.</p><h3>The Scaling Pattern</h3><p>The harness advantage compounds with complexity:</p><ul><li><p><strong>Levels 1-2:</strong> All agents performed similarly. Simple tasks don&#8217;t reveal orchestration differences.</p></li><li><p><strong>Level 3:</strong> API integration and Docker setup separated agents that plan from those that code-and-fix. Clear gaps emerged.</p></li><li><p><strong>Levels 4-5:</strong> Architecture and full-stack challenges broke most agents. Only well-orchestrated systems completed the complex workflows.</p></li></ul><p>The pattern is clear: as complexity increases, harness quality becomes the primary determinant of success.</p><h2>Evidence Pillar #2: Market Validation &#8212; GPT-5.4 Caught Up</h2><h3>Model Parity Achievement</h3><p>February-April 2026 benchmarks confirm <strong>GPT-5.4 has achieved parity with Claude Sonnet 4.6</strong>:</p><p><strong>Core Benchmarks:</strong></p><ul><li><p><strong>SWE-bench Verified</strong>: GPT-5.4 ~80% vs Claude 79.6% (statistical tie)</p></li><li><p><strong>SWE-bench Pro</strong>: GPT-5.4 57.7% vs Claude 43.6% (GPT leads complex problems)</p></li><li><p><strong>Terminal-Bench</strong>: GPT-5.4 75.1% vs Claude ~65% (DevOps advantage)</p></li><li><p><strong>Context handling</strong>: Both models feature 1M token windows</p></li></ul><h3>Yet Claude Still Dominates Through Harness Advantages</h3><p>Despite achieving model parity, the competitive landscape tells the harness story:</p><p><strong>Market Reality:</strong></p><ul><li><p><strong>Developer preference</strong>: Claude 70% (superior workflow integration)</p></li><li><p><strong>Enterprise share</strong>: Anthropic +4.9% MoM growth, OpenAI -1.5% decline</p></li><li><p><strong>Revenue</strong>: Claude Code $2B ARR in 6 months</p></li></ul><p><strong>Even when models reach parity, harness quality determines adoption.</strong></p><h3>The Multi-Model Strategic Reality</h3><p>Leading organizations aren&#8217;t choosing between models anymore&#8212;they&#8217;re deploying <strong>three-tier strategic architectures</strong> based on cost-performance optimization:</p><p><strong>Tier 1: Daily Workhorse (60-70% of requests)</strong></p><ul><li><p><strong>Claude Sonnet 4.6</strong>: <a href="https://medium.com/@mkteam/gpt-5-4-vs-claude-sonnet-4-6-2026-the-ultimate-ai-model-comparison-49526cac8b14">$3/$15 per million tokens</a></p></li><li><p>High-volume development, routine coding tasks</p></li><li><p><a href="https://www.nxcode.io/resources/news/claude-sonnet-4-6-vs-gpt-5-4-coding-comparison-2026">95%+ of premium model quality at half the cost</a></p></li><li><p>Default choice for most enterprise development work</p></li></ul><p><strong>Tier 2: Specialized Operations (20-30% of requests)</strong></p><ul><li><p><strong>GPT-5.4</strong>: <a href="https://medium.com/@mkteam/gpt-5-4-vs-claude-sonnet-4-6-2026-the-ultimate-ai-model-comparison-49526cac8b14">$2.50/$15 per million tokens</a></p></li><li><p>Terminal operations, DevOps workflows, CI/CD debugging</p></li><li><p><a href="https://www.morphllm.com/best-ai-model-for-coding">75.1% Terminal-Bench score (10-point lead over competitors)</a></p></li><li><p><a href="https://medium.com/@ricardomsgarces/openai-codex-vs-github-copilot-why-codex-is-winning-the-future-of-coding-f9a2767695b0">Inherited Codex&#8217;s terminal operation dominance</a></p></li></ul><p><strong>Tier 3: Premium Analysis (10-20% of requests)</strong></p><ul><li><p><strong>Claude Opus 4.6</strong>: <a href="https://medium.com/@mkteam/gpt-5-4-vs-claude-sonnet-4-6-2026-the-ultimate-ai-model-comparison-49526cac8b14">$5/$25 per million tokens</a></p></li><li><p>Complex reasoning, architectural decisions, high-stakes analysis</p></li><li><p><a href="https://help.apiyi.com/en/gpt-5-4-vs-claude-opus-4-6-comparison-2026-en.html">World leader in abstract reasoning (87.4% vs GPT-5.4&#8217;s 83.9%)</a></p></li><li><p>When cost justifies maximum capability</p></li></ul><p>This confirms the core thesis: when models are &#8220;good enough,&#8221; teams optimize for <strong>strategic cost-performance fit</strong>, not raw capability or marketing claims.</p><h2>Evidence Pillar #3: Community Validation &#8212; The Codex Orchestration Story</h2><h3>Reddit Confirms Orchestration Improvements</h3><p>Reddit research explains Codex&#8217;s impressive efficiency results (42 minutes, 4.49 quality score). The evidence confirms improvements come from orchestration, not just model upgrades.</p><p><strong>Architectural Evolution Evidence:</strong></p><ul><li><p><a href="https://medium.com/@aliazimidarmian/openai-codex-from-2021-code-model-to-a-2025-autonomous-coding-agent-85ef0c48730a">Codex evolved from &#8220;embedded assistant&#8221; &#8594; &#8220;independent agent with multi-agent orchestration&#8221;</a></p></li><li><p><a href="https://www.digitalapplied.com/blog/gpt-5-2-codex-openai-model-guide-2026">GPT-5.2-Codex (Jan 2026) with 192K context + MCP tool orchestration</a></p></li><li><p><a href="https://developers.openai.com/blog/openai-for-developers-2025">&#8220;Command center for agents&#8221; interface launched Feb 2026</a></p></li></ul><p><strong>Workflow Efficiency Improvements:</strong></p><ul><li><p><a href="https://reelmind.ai/blog/openai-codex-code-generation-features-reddit-developer-insights">Developers report queuing &#8220;4-5 Codex tasks before diving into manual work&#8221;</a></p></li><li><p><a href="https://reelmind.ai/blog/openai-codex-code-generation-features-reddit-developer-insights">&#8220;2-3 completed PRs waiting for review&#8221; after a coffee break</a></p></li><li><p><a href="https://www.nxcode.io/resources/news/openai-codex-app-review-2026">P99 response time 45ms vs Copilot&#8217;s 55ms through better context management</a></p></li><li><p><strong>Parallel processing capabilities</strong> that enable true background orchestration</p></li></ul><p><strong>Enterprise Orchestration Benefits:</strong></p><ul><li><p><strong><a href="https://www.quantumrun.com/consulting/openai-codex-statistics/">70% more pull requests</a></strong><a href="https://www.quantumrun.com/consulting/openai-codex-statistics/"> merged weekly at OpenAI</a></p></li><li><p><strong><a href="https://www.quantumrun.com/consulting/openai-codex-statistics/">50% reduction</a></strong><a href="https://www.quantumrun.com/consulting/openai-codex-statistics/"> in code review times at Cisco</a></p></li><li><p><strong><a href="https://www.quantumrun.com/consulting/openai-codex-statistics/">67% reduction</a></strong><a href="https://www.quantumrun.com/consulting/openai-codex-statistics/"> in median turnaround time at Duolingo</a></p></li><li><p><strong><a href="https://www.quantumrun.com/consulting/openai-codex-statistics/">90% Fortune 100 adoption</a></strong><a href="https://www.quantumrun.com/consulting/openai-codex-statistics/"> validates orchestration value at scale</a></p></li></ul><h3>The Community Strategic Deployment Pattern</h3><p>Reddit developers now recommend <strong>different tools for different purposes</strong>:</p><ul><li><p><strong>Claude Code</strong>: Code quality and reasoning</p></li><li><p><strong>Cursor</strong>: Daily coding integration</p></li><li><p><strong>OpenAI Codex</strong>: Complex multi-agent workflows and long-horizon autonomy</p></li></ul><p>This matches exactly what the market data predicted: teams use orchestrated tools strategically rather than seeking one universal solution.</p><h2>The Harness Quality Ladder</h2><p>Based on all three evidence pillars, I see four tiers of orchestration quality emerging:</p><p><strong>Tier 1: Basic Wrappers</strong></p><ul><li><p>Simple API access, minimal context management</p></li><li><p>Examples: Raw ChatGPT interface, basic API wrappers</p></li><li><p>Limitation: No file coordination, poor context retention</p></li></ul><p><strong>Tier 2: Workflow Tools</strong></p><ul><li><p>File awareness, some context management</p></li><li><p>Examples: GitHub Copilot, basic IDE extensions</p></li><li><p>Capability: Single-file optimization, limited cross-file understanding</p></li></ul><p><strong>Tier 3: Orchestrated Systems</strong></p><ul><li><p>Multi-file coordination, workflow integration</p></li><li><p>Examples: Cursor, Claude Code, well-configured aider</p></li><li><p>Advantage: Understands project structure, handles complex tasks</p></li></ul><p><strong>Tier 4: Agentic Frameworks</strong></p><ul><li><p>Multi-agent coordination, planning, verification</p></li><li><p>Examples: claude-mpm, advanced orchestration systems</p></li><li><p>Power: Full project lifecycle, quality assurance, architectural thinking</p></li></ul><p>The performance cliff between tiers is exponential, not linear. Bad orchestration can make great models perform poorly; great orchestration can make good models perform excellently.</p><h2>Academic and Industry Validation</h2><p>This isn&#8217;t just empirical observation. Multiple 2026 research papers and industry studies support the harness thesis:</p><p><strong>Academic Consensus:</strong><br>The arXiv paper <a href="https://arxiv.org/html/2511.14136v1">&#8220;Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems&#8221;</a> shows that domain-tuned models with better orchestration achieve superior cost-normalized accuracy despite using smaller base models.</p><p><a href="https://pricepertoken.com/leaderboards/benchmark/humaneval">SWE-bench data</a> reveals the same pattern. Cursor, Claude Code, and Auggie all use similar base models yet score between 50.2% and 55.4%, while the raw model score is only 45.9%. The 5.9-point improvement comes entirely from better context retrieval and agent design.</p><p><strong>Business Reality Check:</strong><br><a href="https://claude5.com/news/enterprise-ai-adoption-2026-how-businesses-deploy-claude-gpt">Enterprise adoption surveys</a> show a clear shift in CTO priorities. &#8220;Model performance&#8221; is dropping in tool evaluation criteria, replaced by governance, integration quality, and workflow fit. As one 2026 McKinsey report put it: &#8220;CTOs are realizing their biggest bottleneck isn&#8217;t model performance&#8212;it&#8217;s governance.&#8221;</p><h2>What This Means for Engineering Leaders</h2><h3>Stop Optimizing for Benchmarks</h3><p>The old procurement mindset was model-first: &#8220;We need access to GPT-6 for competitive advantage.&#8221; The new reality is that benchmark performance doesn&#8217;t predict practical utility. SWE-bench scores don&#8217;t tell you whether a tool will integrate with your existing workflow, handle your codebase size, or recover gracefully from errors.</p><p>Start evaluating harness quality:</p><ul><li><p><strong>Context management:</strong> How well does it understand your project structure?</p></li><li><p><strong>File coordination:</strong> Can it work intelligently across multiple files?</p></li><li><p><strong>Error recovery:</strong> Does it handle failures gracefully or require constant babysitting?</p></li><li><p><strong>Workflow integration:</strong> How does it fit with your team&#8217;s existing development process?</p></li></ul><h3>Budget for Orchestration Quality</h3><p>The three evidence pillars show that investing in better orchestration yields measurable returns:</p><ul><li><p><strong>Quality per minute:</strong> claude-mpm achieved 4.75 quality in 45 minutes; warp achieved 3.94 in 313 minutes</p></li><li><p><strong>Market validation:</strong> Claude maintains dominance despite model parity through superior developer experience</p></li><li><p><strong>Enterprise results:</strong> 70% more PRs, 50% faster code review, 67% faster turnaround</p></li></ul><p>The ROI case for harness investment is clear and quantifiable.</p><h3>Team Productivity Focus</h3><p>Tool choice impacts your entire development pipeline. The 7x speed difference between well and poorly orchestrated tools using the same model means tool selection is a productivity multiplier, not just a capability decision.</p><p>Better tools also reduce onboarding time and increase adoption rates. A tool that works reliably gets used; one that requires constant troubleshooting gets abandoned.</p><h2>The Competitive Landscape Evolution</h2><h3>Codex Deserves Recognition</h3><p>Codex&#8217;s performance has significantly improved. At 42 minutes for all five levels with a 4.49 quality score, it achieved by far the best efficiency in my study. GPT-5.4+ combined with the orchestration improvements OpenAI made represents a compelling package. The Reddit research confirms this wasn&#8217;t just a model upgrade&#8212;it was an architectural evolution toward multi-agent orchestration.</p><h3>Claude Code&#8217;s Harness Moat</h3><p>While Claude Code performed well (4.53 quality score), the market validation shows its true strength: <strong>ecosystem superiority</strong>. Despite GPT-5.4 achieving model parity, Claude maintains 70% developer preference through superior harness quality. This is exactly what sustainable competitive advantage looks like in the post-parity era.</p><h3>The Multi-Model Future</h3><p>All evidence points to the same conclusion: the era of picking one model is over. Leading organizations deploy <strong>three-tier cost-performance architectures</strong>, optimizing for specific strengths rather than seeking universal solutions.</p><p>Real enterprise case studies validate this pattern:</p><ul><li><p><strong><a href="https://www.datastudios.org/post/claude-in-the-enterprise-case-studies-of-ai-deployments-and-real-world-results">TELUS (57,000 employees)</a></strong><a href="https://www.datastudios.org/post/claude-in-the-enterprise-case-studies-of-ai-deployments-and-real-world-results">: Uses Sonnet as core engine across developer teams</a></p></li><li><p><strong><a href="https://www.datastudios.org/post/claude-in-the-enterprise-case-studies-of-ai-deployments-and-real-world-results">Zapier</a></strong><a href="https://www.datastudios.org/post/claude-in-the-enterprise-case-studies-of-ai-deployments-and-real-world-results">: 800+ internal agents using strategic model selection</a></p></li><li><p><strong><a href="https://devtk.ai/en/blog/claude-api-pricing-guide-2026/">Financial Services</a></strong><a href="https://devtk.ai/en/blog/claude-api-pricing-guide-2026/">: Monthly costs ~$80 at massive scale through optimized routing</a></p></li></ul><p>The successful pattern: <strong>Sonnet for volume, GPT-5.4 for DevOps, Opus for complexity</strong>.</p><h2>The Token Economics Reality</h2><p>claude-mpm achieved the highest quality score (4.75) but used 87 million tokens versus codex&#8217;s 120K. This looks expensive until you consider the output: 262 comprehensive tests (vs codex&#8217;s 32), complete documentation, 100% verification rates, and multi-file coordination (note: this was also a wake-up call to me to focus on token optimization, current version is much stingier)</p><p>The 700x token multiplier isn&#8217;t overhead&#8212;it&#8217;s the cost of work a solo agent skips. <strong>Orchestration doesn&#8217;t waste tokens&#8212;it spends them on comprehensive deliverables.</strong></p><p>The optimization question: Could you achieve 80% of the quality benefits at 30% of the token cost? The opportunity isn&#8217;t eliminating orchestration&#8212;it&#8217;s finding the minimal viable team size for maximum impact.</p><h3>The Vendor Bias Problem: &#8220;Opus for Everything&#8221;</h3><p>Boris Cherny, the Claude Code lead, recently advocated for using &#8220;Opus for everything.&#8221; This perfectly illustrates the disconnect between vendor recommendations and practical deployment reality.</p><p><strong>Only someone working for Anthropic can say that.</strong></p><p>When your employer provides unlimited access to premium models, of course you&#8217;d recommend the most expensive option for every task. But real organizations operating with P&amp;L responsibility make strategic decisions about when premium capability justifies premium cost.</p><p>This vendor bias actually <strong>validates the multi-model thesis</strong>:</p><ul><li><p><strong>Vendors say:</strong> &#8220;Use our premium model for everything&#8221;</p></li><li><p><strong>Users do:</strong> Strategic model selection based on task complexity and budget constraints</p></li><li><p><strong>Market reality:</strong> 70% prefer Claude for daily coding (cost/speed), GPT-5.4 for complex reasoning (quality ceiling)</p></li></ul><p>Cherny&#8217;s comment inadvertently proves that <strong>cost-conscious orchestration</strong> is the real competitive battleground. Companies that figure out optimal model routing&#8212;not maximal model usage&#8212;will have sustainable advantages.</p><p>The vendors push premium. The market chooses strategically. <strong>The harness makes both possible.</strong></p><h2>The Future: Welcome to the Harness Era</h2><h3>What Changes for Developers</h3><p>Tool selection framework:</p><ol><li><p><strong>Workflow fit:</strong> Does it match how your team works?</p></li><li><p><strong>Integration quality:</strong> Plays well with existing tools?</p></li><li><p><strong>Reliability:</strong> Can you trust it with production code?</p></li><li><p><strong>Model quality:</strong> Fourth priority</p></li></ol><h3>What Changes for the Industry</h3><p>Foundation models are becoming commodities. Differentiation shifts to integration, context management, and user experience. The next unicorns will be harness companies, not model companies.</p><p>Major funding flows to orchestration companies. Enterprise procurement evaluates integration first, model second.</p><h3>The Competitive Moat Shift</h3><p>The old game was: train bigger models, claim benchmark superiority. The new game is: build better orchestration, solve real workflow problems. Model access becomes a utility; workflow mastery becomes the moat.</p><h2>Practical Recommendations</h2><h3>For CTOs and Engineering Leaders</h3><ul><li><p><strong>Audit orchestration quality</strong>: Test tools with your actual codebase for 2-week trials</p></li><li><p><strong>Budget 60/40</strong>: Spend more on harness development than model subscription fees</p></li><li><p><strong>Measure real metrics</strong>: Track pull request velocity and code review time, not benchmark scores</p></li><li><p><strong>Evaluate integration first</strong>: How well does it fit your existing CI/CD pipeline?</p></li></ul><h3>For Developers</h3><ul><li><p><strong>Test with real projects</strong>: Spend 2 days with each tool on actual work before deciding</p></li><li><p><strong>Learn orchestration patterns</strong>: Context management and file coordination matter more than prompts</p></li><li><p><strong>Invest in mastery</strong>: The 7x efficiency difference justifies significant learning time</p></li><li><p><strong>Ignore marketing claims</strong>: Model access means nothing without good orchestration</p></li></ul><h3>For the AI Industry</h3><ul><li><p><strong>Build for workflow integration</strong>: Solve real development pipeline problems</p></li><li><p><strong>Measure practical utility</strong>: Developer retention and task completion rates beat benchmarks</p></li><li><p><strong>Focus on context management</strong>: Multi-file coordination is the real competitive moat</p></li></ul><h2>Conclusion: The Questions That Matter Now</h2><p>The old question was: &#8220;What&#8217;s the best model?&#8221;</p><p>The new question is: &#8220;What&#8217;s the best harness for my team&#8217;s workflow?&#8221;</p><p>Three evidence sources prove we&#8217;ve crossed a threshold: foundation models are &#8220;good enough,&#8221; and orchestration quality now dominates outcomes. Laboratory testing, market validation, and community confirmation point to the same reality.</p><p>The foundation model is the engine. The harness is the car. The best engine in the world won&#8217;t get you anywhere without wheels.</p><p><strong>The harness era has begun. Drive accordingly.</strong></p><div><hr></div><p><em>Bob Matsuoka is CTO at <a href="https://www.duettocloud.com/">Duetto Research</a> and creator of <a href="https://github.com/bobmatnyc/claude-mpm">Claude MPM</a>, one of the agents evaluated in this study. All evaluation data and methodology are available at <a href="https://github.com/bobmatnyc/ai-coding-bake-off">github.com/bobmatnyc/ai-coding-bake-off</a> for reproducibility.</em></p><div><hr></div><h2>Appendix: Complete Results Data</h2><h3>Quality Scores by Criterion</h3><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!dSX7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96cbd80c-434e-456b-ae62-1fb565e1ec0d_981x368.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!dSX7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96cbd80c-434e-456b-ae62-1fb565e1ec0d_981x368.png 424w, https://substackcdn.com/image/fetch/$s_!dSX7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96cbd80c-434e-456b-ae62-1fb565e1ec0d_981x368.png 848w, https://substackcdn.com/image/fetch/$s_!dSX7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96cbd80c-434e-456b-ae62-1fb565e1ec0d_981x368.png 1272w, https://substackcdn.com/image/fetch/$s_!dSX7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96cbd80c-434e-456b-ae62-1fb565e1ec0d_981x368.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!dSX7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96cbd80c-434e-456b-ae62-1fb565e1ec0d_981x368.png" width="981" height="368" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/96cbd80c-434e-456b-ae62-1fb565e1ec0d_981x368.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:368,&quot;width&quot;:981,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:69754,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/193459844?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96cbd80c-434e-456b-ae62-1fb565e1ec0d_981x368.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!dSX7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96cbd80c-434e-456b-ae62-1fb565e1ec0d_981x368.png 424w, https://substackcdn.com/image/fetch/$s_!dSX7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96cbd80c-434e-456b-ae62-1fb565e1ec0d_981x368.png 848w, https://substackcdn.com/image/fetch/$s_!dSX7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96cbd80c-434e-456b-ae62-1fb565e1ec0d_981x368.png 1272w, https://substackcdn.com/image/fetch/$s_!dSX7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F96cbd80c-434e-456b-ae62-1fb565e1ec0d_981x368.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>GPT-5.4 vs Claude Sonnet 4.6 Market Data</h3><p><strong>SWE-bench Performance:</strong></p><ul><li><p>SWE-bench Verified: GPT-5.4 ~80% vs Claude 79.6% (statistical tie)</p></li><li><p>SWE-bench Pro: GPT-5.4 57.7% vs Claude 43.6% (GPT advantage on complex problems)</p></li><li><p>Terminal-Bench: GPT-5.4 75.1% vs Claude ~65% (GPT DevOps advantage)</p></li></ul><p><strong>Market Metrics:</strong></p><ul><li><p>Developer preference (daily coding): Claude 70%</p></li><li><p>Enterprise market share: Anthropic +4.9% MoM, OpenAI -1.5% MoM</p></li><li><p>Claude Code revenue: $2B ARR in 6 months</p></li></ul><h3>Methodology Notes</h3><ul><li><p><strong>Laboratory data:</strong> Single run evaluation with disclosed author bias</p></li><li><p><strong>Market data:</strong> Cross-validated across 15+ authoritative sources</p></li><li><p><strong>Community research:</strong> Reddit analysis across 8+ developer subreddits</p></li><li><p><strong>Statistical confidence:</strong> Mean inter-reviewer deviation of 0.216 points</p></li><li><p><strong>Reproducible:</strong> All data and prompts available in public repository</p></li></ul>]]></content:encoded></item><item><title><![CDATA[I Met a Movie Star Mila Jovovich — As a Coder]]></title><description><![CDATA[More evidence of the democratization of software]]></description><link>https://hyperdev.matsuoka.com/p/i-met-a-movie-star-mila-jovovich</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/i-met-a-movie-star-mila-jovovich</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Sat, 11 Apr 2026 12:31:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_YC7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d0e03e6-ff0f-4040-ab56-8239bf91a20d_1024x659.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_YC7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d0e03e6-ff0f-4040-ab56-8239bf91a20d_1024x659.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_YC7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d0e03e6-ff0f-4040-ab56-8239bf91a20d_1024x659.png 424w, https://substackcdn.com/image/fetch/$s_!_YC7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d0e03e6-ff0f-4040-ab56-8239bf91a20d_1024x659.png 848w, https://substackcdn.com/image/fetch/$s_!_YC7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d0e03e6-ff0f-4040-ab56-8239bf91a20d_1024x659.png 1272w, https://substackcdn.com/image/fetch/$s_!_YC7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d0e03e6-ff0f-4040-ab56-8239bf91a20d_1024x659.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_YC7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d0e03e6-ff0f-4040-ab56-8239bf91a20d_1024x659.png" width="1024" height="659" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6d0e03e6-ff0f-4040-ab56-8239bf91a20d_1024x659.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:659,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1463075,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/193848267?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F188090d8-280b-48b0-81af-3a11dec4dac3_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!_YC7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d0e03e6-ff0f-4040-ab56-8239bf91a20d_1024x659.png 424w, https://substackcdn.com/image/fetch/$s_!_YC7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d0e03e6-ff0f-4040-ab56-8239bf91a20d_1024x659.png 848w, https://substackcdn.com/image/fetch/$s_!_YC7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d0e03e6-ff0f-4040-ab56-8239bf91a20d_1024x659.png 1272w, https://substackcdn.com/image/fetch/$s_!_YC7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6d0e03e6-ff0f-4040-ab56-8239bf91a20d_1024x659.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I didn&#8217;t expect to meet Mila Jovovich through a GitHub issue.</p><p>But there I was last week, deep-diving into her AI memory framework called <a href="https://github.com/milla-jovovich/mempalace">MemPalace</a>, when I discovered something remarkable: the &#8220;Resident Evil&#8221; and &#8220;Fifth Element&#8221; star had created one of the most talked-about AI memory systems of 2026. And she&#8217;d done it using Claude Code, the same AI-assisted development environment I use daily.</p><p>More remarkably, when I found critical bugs in her benchmark methodology, she responded directly through her Claude Code workflow, acknowledging the issues and implementing fixes. Not through a PR team or engineering intermediaries &#8212; Mila herself, using AI-assisted development to debug complex memory retrieval algorithms at 9 AM on a Thursday.</p><p>This isn&#8217;t a story about a celebrity coding stunt. It&#8217;s about something much more profound: we&#8217;ve entered an era where outcomes and features drive development, not the technical limitations of writing code.</p><h2>The MemPalace Phenomenon</h2><p>In April 2026, Mila Jovovich and developer Ben Sigman released MemPalace, an open-source AI memory system that immediately went viral. Within 48 hours, it had <a href="https://github.com/milla-jovovich/mempalace">over 23,000 GitHub stars</a>. The system claimed to achieve the first perfect score on the LongMemEval benchmark, scoring 96.6% raw recall.</p><p>The project represents something unprecedented: a free, locally-running memory system that rivals expensive cloud alternatives like Mem0 ($19-249/month) and Zep ($25+/month). It uses the &#8220;memory palace&#8221; technique &#8212; a classical memory method dating back to ancient Greece &#8212; implemented through ChromaDB and SQLite, with zero ongoing API costs.</p><p>The technical architecture includes basic Claude Code integration (save hooks every 15 messages and before context compression) and 24 tools via the Model Context Protocol (MCP), making it compatible across multiple AI platforms.</p><p>The duo had spent months building it using Claude Code&#8217;s AI-assisted development environment. As Sigman noted, he provided &#8220;the engineering chops&#8221; while Jovovich drove the architectural vision.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://github.com/milla-jovovich" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!93ZZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d7e3197-469d-4e18-86f1-3033d8bd4a27_459x460.jpeg 424w, https://substackcdn.com/image/fetch/$s_!93ZZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d7e3197-469d-4e18-86f1-3033d8bd4a27_459x460.jpeg 848w, https://substackcdn.com/image/fetch/$s_!93ZZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d7e3197-469d-4e18-86f1-3033d8bd4a27_459x460.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!93ZZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d7e3197-469d-4e18-86f1-3033d8bd4a27_459x460.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!93ZZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d7e3197-469d-4e18-86f1-3033d8bd4a27_459x460.jpeg" width="459" height="460" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7d7e3197-469d-4e18-86f1-3033d8bd4a27_459x460.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:460,&quot;width&quot;:459,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:52413,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:&quot;https://github.com/milla-jovovich&quot;,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/193848267?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d7e3197-469d-4e18-86f1-3033d8bd4a27_459x460.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!93ZZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d7e3197-469d-4e18-86f1-3033d8bd4a27_459x460.jpeg 424w, https://substackcdn.com/image/fetch/$s_!93ZZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d7e3197-469d-4e18-86f1-3033d8bd4a27_459x460.jpeg 848w, https://substackcdn.com/image/fetch/$s_!93ZZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d7e3197-469d-4e18-86f1-3033d8bd4a27_459x460.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!93ZZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7d7e3197-469d-4e18-86f1-3033d8bd4a27_459x460.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>When Audits Meet AI-Generated Code</h2><p>That&#8217;s when things got interesting.</p><p>As someone who works extensively with AI memory systems &#8212; I maintain <a href="https://github.com/bobmatnyc/kuzu-memory">KuzuMemory</a>, a graph-based memory framework &#8212; I was naturally curious about MemPalace&#8217;s benchmark methodology. The claimed 96.6% recall rate was extraordinary, especially for a system running entirely locally.</p><p>So I dove in.</p><p>What I found were several methodological issues that fundamentally undermined the headline numbers. The benchmark adapter was discarding assistant turns in conversation history, causing systematic under-recall on certain question types. More critically, the benchmark wasn&#8217;t actually testing MemPalace&#8217;s core functionality &#8212; it was primarily testing ChromaDB&#8217;s raw vector search capabilities.</p><p>I filed <a href="https://github.com/milla-jovovich/mempalace/issues/242">Issue #242</a> documenting the assistant turn bug, and <a href="https://github.com/milla-jovovich/mempalace/issues/214">Issue #214</a> showing that the 96.6% score was essentially a ChromaDB score, not a MemPalace score.</p><p>Mila&#8217;s response was immediate and technically sophisticated:</p><blockquote><p>&#8220;Hey <a href="https://github.com/bobmatnyc">@bobmatnyc</a> &#8212; I&#8217;ve taken a look and ran it through CLI. This is a real bug and it&#8217;s urgent. You caught that <code>benchmarks/longmemeval_bench.py</code> at lines 189-190 builds each session&#8217;s indexed document by concatenating <em>only</em> <code>user</code> role turns... <strong>Fix priority: this must land before any public benchmark re-run.</strong>&#8220;</p></blockquote><p>She didn&#8217;t deflect or dismiss. She debugged the issue herself, identified the exact lines of code causing the problem, explained the downstream impact on other benchmarks, and outlined a detailed fix plan including regression tests.</p><p>This wasn&#8217;t PR speak. This was an AI-assisted developer engaging seriously with technical criticism.</p><h2>The Democratization Shift</h2><p>This interaction crystallized something profound about our current moment in software development.</p><p>We&#8217;re witnessing the emergence of a new class of builders: technically-minded individuals who understand software conceptually but may not have traditional coding backgrounds. AI-assisted development tools like Claude Code, GitHub Copilot, and Cursor have lowered the implementation barrier to the point where vision and domain expertise matter more than syntax mastery.</p><p>Mila Jovovich exemplifies this shift perfectly. Without formal technical education (she left school in 7th grade for modeling), she spent months intensively learning AI-assisted development through Claude Code starting in late 2025. She understood the conceptual framework of memory palaces deeply enough to architect a sophisticated system. Her collaboration with Ben Sigman &#8212; CEO of Bitcoin lending platform Libre Labs, who provided the engineering expertise while she drove architectural vision &#8212; represents a new model of software development where domain knowledge and AI tool fluency can substitute for traditional programming backgrounds.</p><p>The fact that a movie star can release a technically competent, widely-adopted memory framework isn&#8217;t a commentary on coding getting easier (though it has). It&#8217;s about software development becoming more accessible to domain experts and visionaries who previously couldn&#8217;t bridge the implementation gap.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!f3FK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b861799-f605-4e03-bec5-f88ec1387a42_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!f3FK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b861799-f605-4e03-bec5-f88ec1387a42_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!f3FK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b861799-f605-4e03-bec5-f88ec1387a42_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!f3FK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b861799-f605-4e03-bec5-f88ec1387a42_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!f3FK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b861799-f605-4e03-bec5-f88ec1387a42_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!f3FK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b861799-f605-4e03-bec5-f88ec1387a42_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8b861799-f605-4e03-bec5-f88ec1387a42_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1971642,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/193848267?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b861799-f605-4e03-bec5-f88ec1387a42_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!f3FK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b861799-f605-4e03-bec5-f88ec1387a42_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!f3FK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b861799-f605-4e03-bec5-f88ec1387a42_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!f3FK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b861799-f605-4e03-bec5-f88ec1387a42_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!f3FK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b861799-f605-4e03-bec5-f88ec1387a42_1024x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>What MemPalace Gets Right</h2><p>Despite the benchmark issues I uncovered, MemPalace demonstrates genuine technical sophistication. The memory palace metaphor isn&#8217;t just marketing &#8212; it&#8217;s a thoughtful architectural choice that makes AI memory systems more intuitive and debuggable.</p><p>The system includes elegant features like per-agent memory &#8220;wings&#8221; that prevent cross-contamination between different AI assistants. The Claude Code integration hooks are well-designed, automatically triggering memory saves at logical conversation boundaries. The MCP implementation is clean and follows established patterns.</p><p>Most importantly, the project tackles a real problem: most AI memory systems are either expensive cloud services or complex local installations. MemPalace provides a middle path that&#8217;s both free and relatively easy to deploy.</p><p>Through my testing and integration experiments, I learned techniques that improved my own KuzuMemory system. The competitive analysis forced me to think more carefully about memory organization patterns and retrieval strategies. This kind of cross-pollination benefits the entire ecosystem.</p><h2>The Validation Requirement</h2><p>But the benchmark controversy highlights a crucial point: democratized software development still requires traditional validation methods.</p><p>AI-assisted coding tools excel at implementation but can perpetuate subtle conceptual errors throughout a codebase. The MemPalace benchmark issues weren&#8217;t obvious bugs &#8212; they were methodological problems that required domain expertise to identify.</p><p>This creates an interesting dynamic: AI tools enable rapid development by non-traditional developers, but peer review by experienced practitioners becomes even more critical. The community response to MemPalace&#8217;s inflated benchmarks wasn&#8217;t hostile &#8212; it was collaborative debugging at scale.</p><p>Mila&#8217;s willingness to engage directly with technical criticism and implement fixes demonstrates the right approach. The democratization of software development doesn&#8217;t eliminate the need for technical rigor; it distributes that rigor across a broader community.</p><h2>The Harness Thesis Validated</h2><p>This story perfectly validates what I call the &#8220;harness thesis&#8221; &#8212; that we&#8217;ve entered an era where AI tool ecosystems matter more than underlying model capabilities.</p><p>MemPalace succeeded not because Mila wrote perfect code from scratch, but because she effectively orchestrated Claude Code to implement her vision. The system&#8217;s value comes from its architectural choices, integration quality, and user experience &#8212; not from novel algorithmic breakthroughs.</p><p>Similarly, my ability to audit and improve the system came not from superior coding skills, but from having developed complementary expertise with memory systems and benchmark methodology. The collaboration that emerged &#8212; distributed across GitHub issues, with contributors from multiple backgrounds &#8212; represents the new model of software development.</p><p>We&#8217;re not just building different software; we&#8217;re building software differently.</p><h2>Meeting Mila Through Code</h2><p>In the end, I did meet Mila Jovovich &#8212; through our AI Agents, lines of Python code, GitHub issues, and technical discussions about memory retrieval algorithms, mediated by our respective Claude Code workflows. Not the meeting I would have predicted, but somehow more meaningful than a typical celebrity encounter.</p><p>She embodies a new archetype: the technical visionary who uses AI tools to implement sophisticated ideas without traditional programming backgrounds. Her willingness to engage with criticism and continuously improve the system demonstrates the collaborative spirit that makes this new era of development possible.</p><p>The future of software isn&#8217;t just about better AI models or more powerful tools. It&#8217;s about enabling more people with domain expertise and creative vision to participate in building the systems that shape our digital world.</p><p>And sometimes, that means meeting your childhood movie star idol in a GitHub issue thread, debugging memory palace algorithms together.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and writes about AI-powered engineering at <a href="https://hyperdev.substack.com/">HyperDev</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/its-the-harness-stupid">It&#8217;s The Harness Stupid</a> &#8212; Why AI tool ecosystems matter more than model capabilities</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul><p><strong>Referenced Links:</strong></p><ul><li><p><a href="https://github.com/milla-jovovich/mempalace">MemPalace GitHub Repository</a></p></li><li><p><a href="https://github.com/bobmatnyc/kuzu-memory">KuzuMemory GitHub Repository</a></p></li><li><p><a href="https://github.com/milla-jovovich/mempalace/issues/242">Issue #242: Benchmark adapter bug</a></p></li><li><p><a href="https://github.com/milla-jovovich/mempalace/issues/214">Issue #214: ChromaDB vs MemPalace scoring</a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[The Software Factory is the Next Big Challenge]]></title><description><![CDATA[Many enterprises are rolling their own]]></description><link>https://hyperdev.matsuoka.com/p/the-software-factory-is-the-next</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/the-software-factory-is-the-next</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Wed, 08 Apr 2026 12:30:16 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!tsf8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58714619-d828-47fe-aa27-7777b26b3b11_1024x825.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tsf8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58714619-d828-47fe-aa27-7777b26b3b11_1024x825.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tsf8!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58714619-d828-47fe-aa27-7777b26b3b11_1024x825.png 424w, https://substackcdn.com/image/fetch/$s_!tsf8!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58714619-d828-47fe-aa27-7777b26b3b11_1024x825.png 848w, https://substackcdn.com/image/fetch/$s_!tsf8!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58714619-d828-47fe-aa27-7777b26b3b11_1024x825.png 1272w, https://substackcdn.com/image/fetch/$s_!tsf8!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58714619-d828-47fe-aa27-7777b26b3b11_1024x825.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tsf8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58714619-d828-47fe-aa27-7777b26b3b11_1024x825.png" width="1024" height="825" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/58714619-d828-47fe-aa27-7777b26b3b11_1024x825.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:825,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1823642,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/193118243?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F90d9b63c-9a04-4eb4-895e-c659c19b1b3b_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!tsf8!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58714619-d828-47fe-aa27-7777b26b3b11_1024x825.png 424w, https://substackcdn.com/image/fetch/$s_!tsf8!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58714619-d828-47fe-aa27-7777b26b3b11_1024x825.png 848w, https://substackcdn.com/image/fetch/$s_!tsf8!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58714619-d828-47fe-aa27-7777b26b3b11_1024x825.png 1272w, https://substackcdn.com/image/fetch/$s_!tsf8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F58714619-d828-47fe-aa27-7777b26b3b11_1024x825.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">The Software Factory is the future of software development</figcaption></figure></div><p>Stripe engineers send Slack messages that automatically become production code. Not suggestions. Not drafts. Production code merged into their main branch, supporting over a trillion dollars in annual payment processing.</p><p>Their <a href="https://stripe.dev/blog/minions-stripes-one-shot-end-to-end-coding-agents">&#8220;Minions&#8221; system</a> generates <a href="https://www.infoq.com/news/2026/03/stripe-autonomous-coding-agents/">1,300 pull requests per week</a> with zero human-written code. Fire-and-forget automation from conversation to deployment. While the rest of us debate whether AI can write good code, Stripe has built a software factory that produces enterprise-grade applications at scale.</p><p>The software factory isn&#8217;t a future concept. It&#8217;s a present reality, and it represents the next fundamental challenge for engineering organizations.</p><h2>What We&#8217;re Building at Duetto</h2><p>I&#8217;ve been thinking about this a lot lately. At Duetto, we&#8217;re exploring what a software factory could look like for hospitality technology. Not because we want to eliminate developers, but because we&#8217;re hitting the limits of traditional development approaches for domain-specific applications.</p><p>Our challenge isn&#8217;t just writing code&#8212;it&#8217;s translating complex hotel revenue management requirements into software that works reliably across thousands of properties with different systems, data formats, and business rules. The cognitive load of keeping all these variations in mind while building features is becoming unsustainable.</p><p>What if we could describe what we need in something like our APEX specifications, and have the system generate not just code, but complete deployments? Kubernetes instances running Claude Code agents, database migrations, monitoring setup, the whole stack configured for that specific use case.</p><p>The goal isn&#8217;t replacing our engineering team. Our developers should be solving revenue optimization algorithms and building domain-specific integrations, not configuring YAML files for the hundredth deployment variation.</p><h2>The Stripe Blueprint</h2><p>Stripe&#8217;s Minions reveal what a production software factory actually looks like when you strip away the hype and focus on what works.</p><p><strong>Five-Layer Pipeline</strong>: Their system transforms Slack messages into production-ready pull requests through a structured pipeline. Not magic&#8212;engineering discipline applied to automation.</p><p><strong>Sandboxed Execution</strong>: Every agent runs in isolated containers with codebase checkouts. They can&#8217;t access production systems, can&#8217;t cause cascading failures, can&#8217;t break things outside their designated scope. <a href="https://www.anup.io/stripes-coding-agents-the-walls-matter-more-than-the-model/">The walls matter more than the model</a>.</p><p><strong>Surgical Tool Selection</strong>: Their <a href="https://www.mindstudio.ai/blog/what-is-ai-agent-harness-stripe-minions">Model Context Protocol</a> provides access to hundreds of internal tools, but agents get intelligently prefetched access to only the ~15 tools relevant to their specific task. Not everything available&#8212;the right things available.</p><p><strong>One-Shot Optimization</strong>: Instead of conversational back-and-forth, their agents are <a href="https://www.sitepoint.com/stripe-minions-architecture-explained/">optimized for well-defined work</a> that completes in a single execution. Better latency, lower costs, more predictable outcomes.</p><p>The results speak for themselves: 1,300 PRs weekly, zero human-written code in merged changes, supporting their entire payment infrastructure. This isn&#8217;t a pilot program. This is their production development workflow.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!18qG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F405e3b24-3997-4ecc-89e7-bcd78eb0c218_1024x674.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!18qG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F405e3b24-3997-4ecc-89e7-bcd78eb0c218_1024x674.png 424w, https://substackcdn.com/image/fetch/$s_!18qG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F405e3b24-3997-4ecc-89e7-bcd78eb0c218_1024x674.png 848w, https://substackcdn.com/image/fetch/$s_!18qG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F405e3b24-3997-4ecc-89e7-bcd78eb0c218_1024x674.png 1272w, https://substackcdn.com/image/fetch/$s_!18qG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F405e3b24-3997-4ecc-89e7-bcd78eb0c218_1024x674.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!18qG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F405e3b24-3997-4ecc-89e7-bcd78eb0c218_1024x674.png" width="1024" height="674" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/405e3b24-3997-4ecc-89e7-bcd78eb0c218_1024x674.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:674,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1235022,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/193118243?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1885512e-da12-485c-9bc1-eaa9bd2512e6_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!18qG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F405e3b24-3997-4ecc-89e7-bcd78eb0c218_1024x674.png 424w, https://substackcdn.com/image/fetch/$s_!18qG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F405e3b24-3997-4ecc-89e7-bcd78eb0c218_1024x674.png 848w, https://substackcdn.com/image/fetch/$s_!18qG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F405e3b24-3997-4ecc-89e7-bcd78eb0c218_1024x674.png 1272w, https://substackcdn.com/image/fetch/$s_!18qG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F405e3b24-3997-4ecc-89e7-bcd78eb0c218_1024x674.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The Broader Software Factory Landscape</h2><p>Stripe isn&#8217;t alone in building these systems, just the most public about their approach.</p><p>Netflix has their federated developer console integrating dozens of tools into a single unified experience. <a href="https://engineering.atspotify.com/2024/4/supercharged-developer-portals">Spotify&#8217;s Backstage</a> holds 89% market share among internal developer platforms, reducing time-to-tenth-pull-request by 55% for new developers.</p><p>The open source ecosystem is catching up quickly. <a href="https://openhands.dev/">OpenHands</a> provides a model-agnostic platform for cloud coding agents with $18.8M in Series A funding. <a href="https://www.turing.com/blog/top-5-ai-code-generation-tools-in-2024">CodeT5</a> handles multi-language code generation. <a href="https://github.com/features/copilot/enterprise">GitHub Copilot Enterprise</a> is expanding beyond code completion into full workflow automation.</p><p>Major cloud providers are also building comprehensive platforms. Microsoft&#8217;s GitHub Copilot Workspace, Google&#8217;s Duet AI for developers, and Amazon&#8217;s Q Developer all represent enterprise-grade attempts at software factory capabilities.</p><p>According to Gartner, <a href="https://calmops.com/devops/internal-developer-platform-idp-2026-complete-guide/">80% of large engineering organizations</a> now have dedicated platform teams. The question isn&#8217;t whether software factories are coming&#8212;it&#8217;s whether your organization will build one or buy one.</p><h2>What a Proper Software Factory Requires</h2><p>Building a software factory isn&#8217;t just about connecting AI tools to deployment pipelines. Based on what&#8217;s working at Stripe and emerging patterns across the industry, here are the essential components:</p><h3>Artifact Response Systems</h3><p>Your factory needs to respond to structured specifications and generate complete deployments. At Duetto, this might mean taking an APEX specification for a new revenue optimization feature and producing:</p><ul><li><p>Kubernetes deployment configurations</p></li><li><p>Database migration scripts</p></li><li><p>Monitoring and alerting setup</p></li><li><p>Load testing scenarios</p></li><li><p>Documentation</p></li></ul><p>The system should handle the entire deployment lifecycle from specification to running production service, not just generate code that someone has to manually deploy.</p><h3>Strategic Human Review Checkpoints</h3><p>Notice I said strategic, not comprehensive. Stripe&#8217;s fire-and-forget model works because they&#8217;ve identified the specific points where human judgment adds value without blocking automation.</p><p>For enterprise applications, you need checkpoints at:</p><ul><li><p><strong>Specification validation</strong>: Do the requirements make business sense?</p></li><li><p><strong>Security review</strong>: Are access patterns and data handling appropriate?</p></li><li><p><strong>Integration testing</strong>: Does this work with existing systems?</p></li><li><p><strong>Production readiness</strong>: Are monitoring and rollback capabilities sufficient?</p></li></ul><p>The key is making these gates fast and decisive, not bureaucratic approval processes that defeat the purpose of automation.</p><h3>Scaffolding for Error Detection</h3><p>Your factory will produce broken code. That&#8217;s not a bug&#8212;that&#8217;s reality. The difference between a prototype and a production system is sophisticated error detection and recovery.</p><p>This means:</p><ul><li><p><strong>Isolated execution environments</strong> where failures can&#8217;t cause broader damage</p></li><li><p><strong>Automated testing and iteration</strong> when initial attempts fail</p></li><li><p><strong>Multi-layer validation</strong> before anything reaches production</p></li><li><p><strong>Comprehensive rollback capabilities</strong> for when something gets through anyway</p></li></ul><p>Stripe&#8217;s sandbox architecture is brilliant because it lets agents fail safely while learning from those failures to improve future attempts.</p><h3>Success Criteria Parameters</h3><p>Your factory needs to know what success looks like for each type of work. Not just &#8220;the code compiles,&#8221; but measurable business outcomes.</p><p>For a hospitality feature, success might mean:</p><ul><li><p>Performance benchmarks met under load</p></li><li><p>Integration tests pass with five different PMS systems</p></li><li><p>Revenue impact measurable within 30 days</p></li><li><p>Zero customer-facing errors in the first week</p></li></ul><p>Define these criteria upfront, build them into your validation pipeline, and let the factory optimize for actual business value rather than technical metrics alone.</p><h3>Cost Tracking and Optimization</h3><p>AI-powered development isn&#8217;t free. You need visibility into the computational costs, tool usage, and human review time for each generated system.</p><p>Stripe optimizes for this explicitly&#8212;their one-shot agents cost less than conversational approaches, their surgical tool selection reduces context costs, their automated testing prevents expensive human debugging cycles.</p><p>Track these metrics from day one. The difference between a cost-effective software factory and an expensive experiment is usually found in the operational details.</p><h3>Deployment Models</h3><p>Your factory needs sophisticated understanding of how to deploy different types of applications. Golden Path workflows that codify best practices, environment promotion strategies that reduce risk, and rollback procedures that restore service quickly when things go wrong.</p><p>This is where domain expertise becomes critical. A generic software factory might know how to deploy a web service, but does it understand the specific requirements for hospitality payment processing, guest data privacy, and integration with property management systems?</p><h2>The Duetto Context</h2><p>At Duetto, we&#8217;re thinking about how a software factory could handle the complexity of hospitality technology. Our domain has unique challenges:</p><p><strong>Data Integration Complexity</strong>: Every hotel uses different systems with different data formats. A software factory needs to understand these variations and generate appropriate integration code.</p><p><strong>Regulatory Requirements</strong>: Guest privacy, payment processing, accessibility compliance. The factory needs to embed these requirements into everything it produces.</p><p><strong>Performance Characteristics</strong>: Revenue management systems need to process pricing updates in near real-time across thousands of rooms and rate plans. The factory needs to optimize for these specific performance patterns.</p><p><strong>Operational Constraints</strong>: Hotels can&#8217;t afford downtime during peak booking periods. Deployment strategies need to account for hospitality business cycles.</p><p>We&#8217;re not trying to build a general-purpose software factory. We&#8217;re exploring how to build one that deeply understands our domain and can produce applications that work reliably in hospitality environments.</p><h2>The Reality Check</h2><p>Building a software factory is hard. Not because the technology doesn&#8217;t exist&#8212;Stripe proves it does&#8212;but because the organizational challenges are substantial.</p><p><strong>ROI Demonstration</strong>: You need to show measurable productivity improvements and cost savings. &#8220;The AI is impressive&#8221; isn&#8217;t sufficient justification for the investment required.</p><p><strong>Security and Compliance</strong>: Automated code generation that touches customer data or payment systems requires additional security layers and audit capabilities.</p><p><strong>Developer Workflow Changes</strong>: Your engineering team needs to learn new ways of working. Some will embrace it, others will resist. Change management is as important as the technical implementation.</p><p><strong>Quality Assurance Evolution</strong>: Your QA processes need to evolve from testing human-written code to validating AI-generated systems. Different failure modes, different testing strategies.</p><p><strong>Integration Complexity</strong>: Your factory needs to work with existing systems, databases, APIs, and workflows. The harder the integration challenge, the longer the implementation timeline.</p><p>These aren&#8217;t reasons to avoid building a software factory. They&#8217;re reasons to approach the project with realistic expectations and proper preparation.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bVFQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931edca-6b79-4588-bae5-ab6a883a7b66_1024x678.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bVFQ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931edca-6b79-4588-bae5-ab6a883a7b66_1024x678.png 424w, https://substackcdn.com/image/fetch/$s_!bVFQ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931edca-6b79-4588-bae5-ab6a883a7b66_1024x678.png 848w, https://substackcdn.com/image/fetch/$s_!bVFQ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931edca-6b79-4588-bae5-ab6a883a7b66_1024x678.png 1272w, https://substackcdn.com/image/fetch/$s_!bVFQ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931edca-6b79-4588-bae5-ab6a883a7b66_1024x678.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bVFQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931edca-6b79-4588-bae5-ab6a883a7b66_1024x678.png" width="1024" height="678" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5931edca-6b79-4588-bae5-ab6a883a7b66_1024x678.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:678,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1446253,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/193118243?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F71a61e37-d209-4927-8ab4-44a35ed22bb9_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!bVFQ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931edca-6b79-4588-bae5-ab6a883a7b66_1024x678.png 424w, https://substackcdn.com/image/fetch/$s_!bVFQ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931edca-6b79-4588-bae5-ab6a883a7b66_1024x678.png 848w, https://substackcdn.com/image/fetch/$s_!bVFQ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931edca-6b79-4588-bae5-ab6a883a7b66_1024x678.png 1272w, https://substackcdn.com/image/fetch/$s_!bVFQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5931edca-6b79-4588-bae5-ab6a883a7b66_1024x678.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Looking Forward</h2><p>The trajectory is clear. <a href="https://leanopstech.com/blog/platform-engineering-in-2025-the-future-of-developer-productivity/">Software factories are moving from experimental to mainstream</a>, with proven systems operating at enterprise scale and standardized architecture patterns emerging across the industry.</p><p>The question for engineering leaders isn&#8217;t whether this transformation will happen. It&#8217;s whether your organization will be an early adopter that shapes how software factories work in your domain, or a later adopter that implements patterns developed by others.</p><p>At Duetto, we&#8217;re betting on being early. Not because we want to be on the cutting edge for its own sake, but because the companies that figure out domain-specific software factories first will have a significant competitive advantage in application development speed and quality.</p><p>The software factory represents the next evolution of platform engineering. The organizations that master it will build better software faster than those that don&#8217;t.</p><p>The challenge isn&#8217;t technical anymore. It&#8217;s organizational, strategic, and operational.</p><p>The question is: Are you ready to build one?</p><div><hr></div><p><em>About this analysis: This piece draws from comprehensive research on production software factory implementations, including detailed analysis of Stripe&#8217;s Minions architecture, enterprise platform engineering initiatives, and emerging open source solutions. The author is exploring software factory applications for hospitality technology at Duetto.</em></p><p><em>About the author: Bob Matsuoka is Chief Technology Officer at Duetto and creator of Claude MPM (Multi-agent Project Manager). He has implemented AI-assisted development workflows across enterprise engineering teams and writes about the practical realities of AI integration in software development at <a href="https://hyperdev.substack.com/">HyperDev</a>.</em></p>]]></content:encoded></item><item><title><![CDATA[Is The Claude Code Team Moving Too Quickly?]]></title><description><![CDATA[What To Think of the Source Leak]]></description><link>https://hyperdev.matsuoka.com/p/is-the-claude-code-team-moving-too</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/is-the-claude-code-team-moving-too</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Mon, 06 Apr 2026 12:30:41 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!xLBj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd07636e2-b451-4839-853e-91fc8ca0b4b3_1024x595.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xLBj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd07636e2-b451-4839-853e-91fc8ca0b4b3_1024x595.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xLBj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd07636e2-b451-4839-853e-91fc8ca0b4b3_1024x595.png 424w, https://substackcdn.com/image/fetch/$s_!xLBj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd07636e2-b451-4839-853e-91fc8ca0b4b3_1024x595.png 848w, https://substackcdn.com/image/fetch/$s_!xLBj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd07636e2-b451-4839-853e-91fc8ca0b4b3_1024x595.png 1272w, https://substackcdn.com/image/fetch/$s_!xLBj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd07636e2-b451-4839-853e-91fc8ca0b4b3_1024x595.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xLBj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd07636e2-b451-4839-853e-91fc8ca0b4b3_1024x595.png" width="1024" height="595" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d07636e2-b451-4839-853e-91fc8ca0b4b3_1024x595.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:595,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1151803,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/193101232?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d9c1f20-243f-4cf5-b1c1-b708e63ffae8_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!xLBj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd07636e2-b451-4839-853e-91fc8ca0b4b3_1024x595.png 424w, https://substackcdn.com/image/fetch/$s_!xLBj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd07636e2-b451-4839-853e-91fc8ca0b4b3_1024x595.png 848w, https://substackcdn.com/image/fetch/$s_!xLBj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd07636e2-b451-4839-853e-91fc8ca0b4b3_1024x595.png 1272w, https://substackcdn.com/image/fetch/$s_!xLBj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd07636e2-b451-4839-853e-91fc8ca0b4b3_1024x595.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>On March 31, 2026, Anthropic accidentally shipped their entire Claude Code source&#8212;512,000 lines of TypeScript&#8212;in an npm package. What followed was perhaps the most intense technical autopsy in AI history. The verdict? Mixed, and revealing.</p><p>The criticism has been swift and pointed. A 5,594-line file with a single 3,167-line function sporting 12 levels of nesting. Regex-based frustration detection looking for &#8220;wtf&#8221; and &#8220;shit&#8221;. A quarter million wasted API calls per day from a three-line bug. As one critic put it: &#8220;A multi-billion-dollar AI company is detecting user frustration with a regex.&#8221;</p><p>But before we pile on, we need to ask: <strong>What does &#8220;good code&#8221; even mean when you&#8217;re building client-side LLM applications?</strong></p><h2>The Unprecedented Challenge</h2><p>Claude Code isn&#8217;t your typical software. It&#8217;s a client-side application that orchestrates conversations with large language models, manages context across sessions, and attempts to maintain coherent state while working with fundamentally non-deterministic systems.</p><p>This creates problems that traditional software engineering practices weren&#8217;t designed for:</p><ul><li><p><strong>Context management</strong>: Handling arbitrarily long conversations that exceed model limits</p></li><li><p><strong>Failure recovery</strong>: When your core computation is a 20% failure-rate API call</p></li><li><p><strong>State synchronization</strong>: Keeping UI, conversation history, and model context aligned</p></li><li><p><strong>Dynamic adaptation</strong>: Code that needs to adapt to changing model capabilities</p></li></ul><p>The leaked source reveals sophisticated solutions to these problems: a three-layer memory architecture, anti-distillation mechanisms, dual parser systems for safety. The engineering is <s>genuinely</s> impressive, even if the implementation is sometimes ugly.</p><h2>The Meta-Problem: AI Writing AI</h2><p>Claude Code was partially written by Claude Code. This represents the first documented case of a large-scale AI tool generating significant portions of its own source code&#8212;not just incremental improvement, but a categorical change in development methodology that creates unprecedented quality control challenges when AI-generated code scales beyond human review capacity.</p><p>When AI generates code at scales that exceed human review capacity, traditional quality control breaks down. That 3,167-line function? Probably not written by a human. The 12 levels of nesting? Algorithmic patterns, not human design choices.</p><p><strong>This is the real story</strong>: We&#8217;re witnessing the first major autopsy of self-bootstrapping AI tooling.</p><h2>Deterministic vs. LLM Code: Different Standards Apply</h2><p>I&#8217;ve been thinking about this distinction a lot lately in my work with <a href="https://github.com/bobmatnyc/claude-mpm">Claude MPM</a>, an open-source multi-agent code generation framework built on Claude Code that coordinates specialized AI agents for software development workflows. When you&#8217;re building traditional, deterministic software, all the usual rules apply. Clean functions, clear abstractions, maintainable architecture. Use your normal code analysis tools.</p><p>But when you&#8217;re building LLM-integrated systems, the rules change:</p><ol><li><p><strong>Failure is the default</strong>: Your core operations fail 20% of the time</p></li><li><p><strong>Context is expensive</strong>: Every token counts toward limits</p></li><li><p><strong>Behavior is emergent</strong>: The system does things you didn&#8217;t explicitly program</p></li><li><p><strong>Adaptation is constant</strong>: Model capabilities change monthly</p></li></ol><p>In this world, a 5,594-line file might be ugly, but if it successfully manages complex failure recovery across multiple conversation threads, it might also be <em>correct</em>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!aIuf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd26aa417-4783-42d6-8105-488131dfe518_1024x790.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!aIuf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd26aa417-4783-42d6-8105-488131dfe518_1024x790.png 424w, https://substackcdn.com/image/fetch/$s_!aIuf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd26aa417-4783-42d6-8105-488131dfe518_1024x790.png 848w, https://substackcdn.com/image/fetch/$s_!aIuf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd26aa417-4783-42d6-8105-488131dfe518_1024x790.png 1272w, https://substackcdn.com/image/fetch/$s_!aIuf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd26aa417-4783-42d6-8105-488131dfe518_1024x790.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!aIuf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd26aa417-4783-42d6-8105-488131dfe518_1024x790.png" width="1024" height="790" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d26aa417-4783-42d6-8105-488131dfe518_1024x790.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:790,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1466836,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/193101232?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd33fb826-655a-4d8b-87cc-bfade8984326_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!aIuf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd26aa417-4783-42d6-8105-488131dfe518_1024x790.png 424w, https://substackcdn.com/image/fetch/$s_!aIuf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd26aa417-4783-42d6-8105-488131dfe518_1024x790.png 848w, https://substackcdn.com/image/fetch/$s_!aIuf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd26aa417-4783-42d6-8105-488131dfe518_1024x790.png 1272w, https://substackcdn.com/image/fetch/$s_!aIuf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd26aa417-4783-42d6-8105-488131dfe518_1024x790.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The Code Analysis Checkpoint Strategy</h2><p>This is where I&#8217;ve found success with my recent updates to the code analyzer in Claude MPM. The analyzer utilizes <a href="https://github.com/modelcontextprotocol/servers/tree/main/src/mcp-vector-search">mcp-vector-search</a> for comprehensive codebase analysis, providing AST-based semantic search, full-text search capabilities, and knowledge graph construction for architectural pattern detection. Instead of trying to prevent AI from generating messy code (impossible), I focus on <strong>regular refactoring and analysis checkpoints</strong>.</p><p>The analyzer has gotten very good at catching two specific issues:</p><ol><li><p><strong>Drift</strong>: When AI-generated code slowly diverges from intended architecture</p></li><li><p><strong>Bloat</strong>: When generated solutions become unnecessarily complex over time</p></li></ol><p>I make a point to run these checkpoints regularly, treating them as essential maintenance rather than optional cleanup. It&#8217;s like running <code>cargo clippy</code> or <code>eslint</code>, but for AI-generated architectural decisions.</p><p>The key insight: <strong>AI code needs different kinds of maintenance than human code</strong>.</p><h2>Outcome-Based Generation: Does It Work?</h2><p>Here&#8217;s my perhaps controversial take: If Claude Code successfully helps developers ship better software faster, then the messy internals might not matter as much as we think.</p><p>The leaked code reveals a system that:</p><ul><li><p>Handles millions of conversations per day</p></li><li><p>Maintains context across arbitrarily long sessions</p></li><li><p>Provides sophisticated memory management</p></li><li><p>Implements multiple safety layers</p></li><li><p>Delivers a $2.5 billion ARR product experience</p></li></ul><p>Is the implementation elegant? No. Does it work? Apparently, yes. Because we can observe/measure what it&#8217;s building completely independently of what built it.</p><p>This doesn&#8217;t excuse basic engineering failures (that <code>.npmignore</code> mistake was embarrassing). But it does suggest we need new frameworks for evaluating AI-generated systems.</p><h2>The Scaffolding Solution</h2><p>Rather than trying to make AI generate perfect code, we can scaffold around the inevitable messiness:</p><p><strong>Automated refactoring checkpoints</strong>: Regular cleanup of AI-generated bloat<br><strong>Architectural constraints</strong>: Guard rails that prevent the worst patterns<br><strong>Outcome validation</strong>: Testing that focuses on behavior over implementation<br><strong>Human oversight</strong>: Strategic points where humans validate AI decisions</p><p>This is the approach I&#8217;ve been taking with Claude MPM, and it&#8217;s proven remarkably effective. Let the AI generate messy-but-functional code, then use tooling to clean it up systematically.</p><h2>What This Means for the Industry</h2><p>The Claude Code leak represents a watershed moment. It&#8217;s our first real look at what happens when AI tools build themselves at scale.</p><p>The criticism is valid&#8212;basic engineering discipline matters, even in AI systems. A missing <code>.npmignore</code> file is inexcusable for a billion-dollar product.</p><p>But the deeper question is whether we&#8217;re applying the right standards. Traditional code quality metrics may not capture what actually matters for AI-integrated systems.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NrDl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87df3326-02c3-4838-b302-ab23ab5d5e19_1024x825.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NrDl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87df3326-02c3-4838-b302-ab23ab5d5e19_1024x825.png 424w, https://substackcdn.com/image/fetch/$s_!NrDl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87df3326-02c3-4838-b302-ab23ab5d5e19_1024x825.png 848w, https://substackcdn.com/image/fetch/$s_!NrDl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87df3326-02c3-4838-b302-ab23ab5d5e19_1024x825.png 1272w, https://substackcdn.com/image/fetch/$s_!NrDl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87df3326-02c3-4838-b302-ab23ab5d5e19_1024x825.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NrDl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87df3326-02c3-4838-b302-ab23ab5d5e19_1024x825.png" width="1024" height="825" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/87df3326-02c3-4838-b302-ab23ab5d5e19_1024x825.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:825,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1937908,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/193101232?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1273dbb0-a196-49bc-b0ca-17f437545fd8_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!NrDl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87df3326-02c3-4838-b302-ab23ab5d5e19_1024x825.png 424w, https://substackcdn.com/image/fetch/$s_!NrDl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87df3326-02c3-4838-b302-ab23ab5d5e19_1024x825.png 848w, https://substackcdn.com/image/fetch/$s_!NrDl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87df3326-02c3-4838-b302-ab23ab5d5e19_1024x825.png 1272w, https://substackcdn.com/image/fetch/$s_!NrDl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F87df3326-02c3-4838-b302-ab23ab5d5e19_1024x825.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Moving Forward</h2><p>Anthropic probably <em>is</em> moving too quickly in some ways. The leak revealed security vulnerabilities, competitive intelligence losses, and quality control failures that suggest inadequate human oversight.</p><p>But they&#8217;re also pioneering entirely new categories of software. The problems they&#8217;re solving&#8212;context management, failure recovery, human-AI collaboration&#8212;don&#8217;t have established best practices yet.</p><p>The real lesson isn&#8217;t that AI-generated code is inherently bad. It&#8217;s that we need new practices for building, reviewing, and maintaining systems that exceed human comprehension scales.</p><p><strong>The question isn&#8217;t whether Claude Code&#8217;s internals are messy. It&#8217;s whether we can build better scaffolding around AI-generated systems to catch the problems that matter while accepting the messiness we can&#8217;t avoid.</strong></p><p>The Claude Code team probably needs to slow down on the basics&#8212;security, testing, deployment hygiene. But they&#8217;re moving fast on problems that genuinely require speed to solve before competitors do.</p><p>That&#8217;s a nuanced position in an industry that loves simple takes. But nuance is what the moment requires.</p><p><em>What do you think? Are we being too hard on AI-generated code, or not hard enough? Share your thoughts in the comments.</em></p><div><hr></div><p><em>About this analysis: This piece draws from extensive technical analysis of the March 31, 2026 Claude Code source leak, including community responses, security assessments, and business impact analysis. The author maintains active development projects using AI-assisted coding tools and has direct experience with the challenges discussed.</em></p><p><em>About the author: Bob Matsuoka is Chief Technology Officer at <a href="https://www.duettocloud.com/">Duetto</a> and creator of Claude MPM (Multi-agent Project Manager). He has implemented AI-assisted development workflows across enterprise engineering teams and writes about the practical realities of AI integration in software development at <a href="https://hyperdev.substack.com">HyperDev</a>.</em></p>]]></content:encoded></item><item><title><![CDATA[Moving Past the 10-Tab Workflow]]></title><description><![CDATA[Autonomous Orchestration Management Is Next]]></description><link>https://hyperdev.matsuoka.com/p/moving-past-the-10-tab-workflow</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/moving-past-the-10-tab-workflow</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Wed, 01 Apr 2026 12:32:10 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!wz2X!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5c54d57-2f21-4893-a87d-88523deb8ae7_1024x690.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wz2X!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5c54d57-2f21-4893-a87d-88523deb8ae7_1024x690.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wz2X!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5c54d57-2f21-4893-a87d-88523deb8ae7_1024x690.png 424w, https://substackcdn.com/image/fetch/$s_!wz2X!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5c54d57-2f21-4893-a87d-88523deb8ae7_1024x690.png 848w, https://substackcdn.com/image/fetch/$s_!wz2X!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5c54d57-2f21-4893-a87d-88523deb8ae7_1024x690.png 1272w, https://substackcdn.com/image/fetch/$s_!wz2X!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5c54d57-2f21-4893-a87d-88523deb8ae7_1024x690.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wz2X!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5c54d57-2f21-4893-a87d-88523deb8ae7_1024x690.png" width="1024" height="690" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a5c54d57-2f21-4893-a87d-88523deb8ae7_1024x690.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:690,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1272299,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/192741144?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb87ec05a-31c6-4037-a9e0-db7c8c8fcba4_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!wz2X!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5c54d57-2f21-4893-a87d-88523deb8ae7_1024x690.png 424w, https://substackcdn.com/image/fetch/$s_!wz2X!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5c54d57-2f21-4893-a87d-88523deb8ae7_1024x690.png 848w, https://substackcdn.com/image/fetch/$s_!wz2X!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5c54d57-2f21-4893-a87d-88523deb8ae7_1024x690.png 1272w, https://substackcdn.com/image/fetch/$s_!wz2X!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa5c54d57-2f21-4893-a87d-88523deb8ae7_1024x690.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">From Tab Chaos to Autonomic Orchestration</figcaption></figure></div><p>I&#8217;m looking at my (iTerm) terminal right now. Ten tmux sessions. Each session holds a different project context&#8212;one monitoring CI failures, another handling a code review, a third debugging a production issue.</p><p>This is the reality of modern agent development work.</p><h2>TL;DR</h2><ul><li><p><strong>Multi-session reality</strong>: Power users average 8-12 terminal sessions; most work involves modification, bug response, and PR handling&#8212;not new code generation</p></li><li><p><strong>Natural workflow origination</strong>: Future systems trigger from product team actions, CI failures, and automated events rather than human prompts </p></li><li><p><strong>Orchestration evolution</strong>: From human-orchestrated agents to orchestration-of-orchestrators where prime coordinators are non-human</p></li><li><p><strong>Production examples</strong>: Stripe&#8217;s Minions (1,300 PRs/week), GitLab&#8217;s Duo Agent Platform, Meta&#8217;s REA demonstrate hierarchical agent orchestration</p></li><li><p><strong>Architecture shift</strong>: Claude Code&#8217;s SDK model enables workflow-driven development through persistent, context-aware agent orchestration</p></li></ul><h2>The 10-Tab Reality</h2><p>According to recent developer workflow studies, <a href="https://www.heyuan110.com/posts/ai/2026-03-03-tmux-guide-ai-development/">tmux has become the standard for AI-assisted development</a>, with <a href="https://dev.to/_d7eb1c1703182e3ce1782/tmux-tutorial-the-complete-developer-workflow-guide-2026-33b3">persistent sessions solving the context-switching tax</a>. The productivity advantage isn&#8217;t the multiplexing&#8212;it&#8217;s the persistence. Projects become environments you step in and out of rather than things you open and close.</p><p>But here&#8217;s what the productivity tutorials miss: most of those tabs aren&#8217;t generating software.</p><p><strong>My current session breakdown:</strong></p><ul><li><p>3 sessions: non-coding -- my CTO knowledge base (currently analyzing our Sumo use), a writing assistant, and our Duetto product management framework</p></li><li><p>4 sessions: coding - various internal tools and MCP connectors</p></li><li><p>2 sessions: coding - new projects</p></li><li><p>1 session: code review</p></li></ul><p>The 8:2 ratio holds across most senior developers I&#8217;ve observed. Most development work involves responding to existing systems, not creating new ones.</p><p>This distribution points toward something significant: <strong>the future of development orchestration isn&#8217;t human-initiated.</strong></p><h2>Beyond Prompt-Driven Development</h2><p>Claude Code&#8217;s new SDK architecture reflects this reality. Instead of starting with human prompts, work originates from natural workflow events:</p><ul><li><p>Product team creates ticket &#8594; Implementation specification generated</p></li><li><p>CI pipeline fails &#8594; Diagnostic agent analyzes failure, proposes fix</p></li><li><p>PR submitted &#8594; Review agent examines code, suggests improvements</p></li><li><p>Production alert triggered &#8594; Incident response agent investigates, documents findings</p></li><li><p>Security scan detects vulnerability &#8594; Remediation agent generates patch</p></li></ul><p>The pattern: <strong>Event &#8594; Agent Response &#8594; Human Review &#8594; Autonomous Resolution</strong>.</p><p>Humans remain in the loop, but as orchestrators and validators rather than initiators. The shift from &#8220;What should I build?&#8221; to &#8220;How should this system respond?&#8221;</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!G7Uu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20f98126-397c-40f4-bf7d-6fc025ab018d_1024x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!G7Uu!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20f98126-397c-40f4-bf7d-6fc025ab018d_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!G7Uu!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20f98126-397c-40f4-bf7d-6fc025ab018d_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!G7Uu!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20f98126-397c-40f4-bf7d-6fc025ab018d_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!G7Uu!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20f98126-397c-40f4-bf7d-6fc025ab018d_1024x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!G7Uu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20f98126-397c-40f4-bf7d-6fc025ab018d_1024x1024.png" width="1024" height="1024" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/20f98126-397c-40f4-bf7d-6fc025ab018d_1024x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1024,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1418038,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/192741144?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20f98126-397c-40f4-bf7d-6fc025ab018d_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!G7Uu!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20f98126-397c-40f4-bf7d-6fc025ab018d_1024x1024.png 424w, https://substackcdn.com/image/fetch/$s_!G7Uu!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20f98126-397c-40f4-bf7d-6fc025ab018d_1024x1024.png 848w, https://substackcdn.com/image/fetch/$s_!G7Uu!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20f98126-397c-40f4-bf7d-6fc025ab018d_1024x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!G7Uu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20f98126-397c-40f4-bf7d-6fc025ab018d_1024x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Yes, Dall-e is Still Atrocious With Spelling.  Don&#8217;t @ me!</figcaption></figure></div><h2>Orchestration of Orchestrators: Production Examples</h2><h3>Stripe&#8217;s Blueprint Architecture</h3><p><a href="https://medium.com/@oracle_43885/how-stripe-built-secure-unattended-ai-agents-merging-1-000-pull-requests-weekly-1ff42f3fe550">Stripe&#8217;s Minions system demonstrates mature orchestration-of-orchestrators</a>. Their &#8220;blueprint&#8221; pattern alternates between deterministic code nodes and agentic reasoning loops, generating <a href="https://blog.bytebytego.com/p/how-stripes-minions-ship-1300-prs">1,300+ pull requests weekly</a>.</p><p><strong>Architecture insight</strong>: <a href="https://www.mindstudio.ai/blog/stripe-minions-blueprint-architecture-deterministic-agentic-nodes">Each blueprint functions as a strict contract between orchestration and execution</a>. Task definitions specify input requirements, output formats, constraints, and success criteria. The orchestrator manages workflow, agents handle implementation.</p><p><strong>Security model</strong>: <a href="https://www.sitepoint.com/stripe-minions-architecture-explained/">Every Minion execution runs in isolated VMs</a> with no internet or production access. The system has submission authority but not merge authority&#8212;all changes require human review.</p><h3>GitLab&#8217;s Intelligent Orchestration</h3><p><a href="https://about.gitlab.com/blog/agentic-sdlc-gitlab-and-tcs-deliver-intelligent-orchestration-across-the-enterprise/">GitLab&#8217;s Duo Agent Platform treats agents as durable actors</a> that plan, modify code, fix pipelines, and enforce security with traceability. Multiple AI agents handle parallel tasks&#8212;code generation, testing, CI/CD fixes&#8212;while developers maintain oversight through defined rules.</p><p><strong>Orchestration insight</strong>: <a href="https://docs.gitlab.com/user/duo_agent_platform/">GitLab positions itself as an AI orchestration plane</a> where humans and agents share delivery responsibility. The platform coordinates multi-agent workflows across the entire software lifecycle rather than providing isolated AI tools.</p><h3>Meta&#8217;s Hierarchical Agent Systems</h3><p><a href="https://engineering.fb.com/2026/03/17/developer-tools/ranking-engineer-agent-rea-autonomous-ai-system-accelerating-meta-ads-ranking-innovation/">Meta&#8217;s Ranking Engineer Agent (REA) demonstrates autonomous ML lifecycle management</a>. REA Planner and REA Executor components, supported by shared skill and knowledge systems, autonomously evolve ads ranking models at scale.</p><p><strong>Acquisition significance</strong>: <a href="https://venturebeat.com/orchestration/why-meta-bought-manus-and-what-it-means-for-your-enterprise-ai-agent">Meta&#8217;s $2B Manus acquisition</a> focused on orchestration infrastructure rather than foundation models. Manus&#8217;s achievement was engineering an execution layer enabling models to browse, code, manipulate files, and complete multi-step workflows autonomously.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!21xU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0980b901-da6a-408f-a4c7-2dd2188be40c_1024x663.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!21xU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0980b901-da6a-408f-a4c7-2dd2188be40c_1024x663.png 424w, https://substackcdn.com/image/fetch/$s_!21xU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0980b901-da6a-408f-a4c7-2dd2188be40c_1024x663.png 848w, https://substackcdn.com/image/fetch/$s_!21xU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0980b901-da6a-408f-a4c7-2dd2188be40c_1024x663.png 1272w, https://substackcdn.com/image/fetch/$s_!21xU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0980b901-da6a-408f-a4c7-2dd2188be40c_1024x663.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!21xU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0980b901-da6a-408f-a4c7-2dd2188be40c_1024x663.png" width="1024" height="663" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0980b901-da6a-408f-a4c7-2dd2188be40c_1024x663.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:663,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1147097,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/192741144?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fef64881a-c7f7-4bab-8d8d-2393ee5bf8b0_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!21xU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0980b901-da6a-408f-a4c7-2dd2188be40c_1024x663.png 424w, https://substackcdn.com/image/fetch/$s_!21xU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0980b901-da6a-408f-a4c7-2dd2188be40c_1024x663.png 848w, https://substackcdn.com/image/fetch/$s_!21xU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0980b901-da6a-408f-a4c7-2dd2188be40c_1024x663.png 1272w, https://substackcdn.com/image/fetch/$s_!21xU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0980b901-da6a-408f-a4c7-2dd2188be40c_1024x663.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The Architecture Implications</h2><h3>Beyond the Single-Agent Model</h3><p>The production examples reveal a consistent pattern: successful autonomous development requires <strong>hierarchical orchestration</strong> rather than monolithic AI assistants.</p><p><strong>Traditional approach</strong>: Human &#8594; Single Agent &#8594; Code<br><strong>Emerging pattern</strong>: Event &#8594; Orchestrator &#8594; Specialized Agents &#8594; Validation &#8594; Resolution</p><h3>Context Preservation at Scale</h3><p><a href="https://medium.com/@gveloper/using-iterm2s-built-in-integration-with-tmux-d5d0ef55ec30">The tmux paradigm</a> of persistent sessions maps directly to agent orchestration. Instead of recreating context for each interaction, systems maintain ongoing project understanding across multiple concurrent workflows.</p><p><strong>Implementation insight</strong>: <a href="https://iterm2.com/documentation-tmux-integration.html">iTerm2&#8217;s tmux integration (-CC mode)</a> provides the UI pattern for agent orchestration&#8212;persistent remote workspaces with native interface feel. The same architecture principles apply to agent coordination.</p><h2>Where This Leads</h2><h3>Non-Human Prime Orchestrators</h3><p>The logical endpoint isn&#8217;t humans managing multiple agents&#8212;it&#8217;s orchestrating systems that manage agent ecosystems. <a href="https://www.deloitte.com/us/en/insights/industry/technology/technology-media-and-telecom-predictions/2026/ai-agent-orchestration.html">According to Gartner&#8217;s 2025 Agentic AI research</a>, nearly 50% of surveyed vendors identified AI orchestration as their primary differentiator.</p><p><strong>Pattern emergence</strong>: Meta-agents or orchestrator-generalists will control specialized agents, assign tasks, interpret results, and revise goals in real-time. <a href="https://arxiv.org/pdf/2601.13671">Hierarchical orchestration becomes essential for enterprise-scale implementations</a>.</p><h3>The Developer Role Evolution</h3><p>Instead of managing 10 terminal sessions, a framework orchestrates autonomous workflows. Each workflow maintains its own context, responds to its own triggers, and escalates to human attention when required. Some of those will be human/experimentation/new development driven, the majority will be responding to the automated lifecycle.</p><p><strong>Skills that matter</strong>:</p><ul><li><p><strong>Workflow boundary definition</strong>: Which autonomous streams can operate independently?</p></li><li><p><strong>Escalation criteria design</strong>: When do workflows require human intervention?</p></li><li><p><strong>Cross-workflow dependency management</strong>: How do autonomous streams coordinate?</p></li><li><p><strong>Quality gate enforcement</strong>: What validation must occur before autonomous resolution?</p></li></ul><h2>Implementation Considerations</h2><p>Teams experimenting with orchestrated autonomous development should consider:</p><ol><li><p><strong>Event-driven architecture</strong>: Which existing workflows could trigger autonomous responses?</p></li><li><p><strong>Context preservation systems</strong>: How will agent workflows maintain project understanding?</p></li><li><p><strong>Isolation and security</strong>: What boundaries prevent autonomous agents from causing damage?</p></li><li><p><strong>Human oversight integration</strong>: Where do human validation points occur in autonomous workflows?</p></li><li><p><strong>Cross-workflow coordination</strong>: How do parallel autonomous streams avoid conflicts?</p></li></ol><p>The transition from 10-tab manual orchestration to autonomous lifecycle orchestration isn&#8217;t theoretical. Stripe, GitLab, and Meta demonstrate production implementations. The question becomes implementation timeline and organizational readiness.</p><p>Early adopters are discovering that the competitive advantage comes not from having the smartest individual AI agents, but from orchestrating networks of specialized agents that collaborate effectively at scale.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and writes about AI-powered engineering at <a href="https://hyperdev.substack.com/">HyperDev</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://www.coddykit.com/pages/blog-detail?id=512757">Stripe&#8217;s Minions: Inside Their Enterprise AI Coding Agent Strategy</a> &#8212; Blueprint orchestration architecture and production metrics</p></li><li><p><a href="https://docs.gitlab.com/user/duo_agent_platform/">GitLab Duo Agent Platform</a> &#8212; Intelligent orchestration across software lifecycle</p></li><li><p><a href="https://www.heyuan110.com/posts/ai/2026-03-03-tmux-guide-ai-development/">Tmux Complete Guide: AI-Powered Multi-Agent Workflows</a> &#8212; Terminal multiplexing for autonomous development</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul>]]></content:encoded></item><item><title><![CDATA[Everyone Blamed Clawd Bot’s Execution. The Concept Was the Problem.]]></title><description><![CDATA[Is A Personal Assistant Bot Really Helpful?]]></description><link>https://hyperdev.matsuoka.com/p/everyone-blamed-clawd-bots-execution</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/everyone-blamed-clawd-bots-execution</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Thu, 12 Mar 2026 11:32:09 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!51wj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F878e8350-5206-465d-9bf0-b600661c22ed_1024x774.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!51wj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F878e8350-5206-465d-9bf0-b600661c22ed_1024x774.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!51wj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F878e8350-5206-465d-9bf0-b600661c22ed_1024x774.png 424w, https://substackcdn.com/image/fetch/$s_!51wj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F878e8350-5206-465d-9bf0-b600661c22ed_1024x774.png 848w, https://substackcdn.com/image/fetch/$s_!51wj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F878e8350-5206-465d-9bf0-b600661c22ed_1024x774.png 1272w, https://substackcdn.com/image/fetch/$s_!51wj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F878e8350-5206-465d-9bf0-b600661c22ed_1024x774.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!51wj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F878e8350-5206-465d-9bf0-b600661c22ed_1024x774.png" width="1024" height="774" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/878e8350-5206-465d-9bf0-b600661c22ed_1024x774.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:774,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1722868,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/190672571?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb05bbe76-836e-475e-b360-793755bf1927_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!51wj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F878e8350-5206-465d-9bf0-b600661c22ed_1024x774.png 424w, https://substackcdn.com/image/fetch/$s_!51wj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F878e8350-5206-465d-9bf0-b600661c22ed_1024x774.png 848w, https://substackcdn.com/image/fetch/$s_!51wj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F878e8350-5206-465d-9bf0-b600661c22ed_1024x774.png 1272w, https://substackcdn.com/image/fetch/$s_!51wj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F878e8350-5206-465d-9bf0-b600661c22ed_1024x774.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The story everyone told about Clawd Bot missed the point entirely. Austrian developer Peter Steinberger built an open-source AI assistant that went viral &#8212; 145,000 GitHub stars, 2 million visitors in a week. Then Anthropic forced a trademark-based name change because &#8220;Clawd&#8221; was too similar to &#8220;Claude.&#8221; The community called it petty. DHH called Anthropic &#8220;customer hostile.&#8221; The irony: Clawd Bot users were actually buying more Claude subscriptions, providing free marketing to Anthropic, yet they still demanded the shutdown.</p><p>But everyone focused on the wrong drama. The trademark dispute was noise. The real problem was deeper: Clawd Bot was built because someone could, not because anyone needed it.</p><p>I tested Clawd Bot for about a week. The interface was clean, the onboarding smooth, the responses capable. But it required permissions I wouldn&#8217;t give to any tool &#8212; access to email, calendars, messaging, sensitive services. The execution had real problems. But even if those were fixed, it would still be solving the wrong problem.</p><p>Here&#8217;s where I should admit: I tried building a digital assistant, izzie, when I started experimenting with AI agents. I never got it to a point I found useful. Not because of technical limitations &#8212; because the entire concept of a universal assistant doesn&#8217;t match how work actually happens.</p><h2>TL;DR</h2><ul><li><p>Clawd Bot was successful open-source project by Peter Steinberger that Anthropic forced to rename; execution wasn&#8217;t the problem</p></li><li><p>The real question: when do you need an &#8220;assistant&#8221;? Most execs won&#8217;t trust AI scheduling; the value is intelligent data movement between services</p></li><li><p>Context switching is a symptom, not the root issue &#8212; the issue is what assistants should be doing at all</p></li><li><p>The product management sessions: Granola meeting notes, calendar checks, Slack updates, Notion sync &#8212; all from within one tool, data flowing intelligently between services</p></li><li><p>The commercial evidence: Cursor, Notion AI, Linear&#8217;s AI triage &#8212; the winners embedded AI in tools as infrastructure, not interface</p></li><li><p>trusty-izzie&#8217;s highest value isn&#8217;t the chat interface &#8212; it&#8217;s as a local MCP service exporting personal context to every other tool</p></li><li><p>The universal assistant category isn&#8217;t going to produce a winner. It&#8217;s going to dissolve.</p></li></ul><h2>What the Universal Assistant Model Gets Wrong (And It&#8217;s Not Just Execution)</h2><p>Clawd Bot had serious execution problems &#8212; it&#8217;s a security nightmare requiring broad permissions across email, calendars, messaging platforms, and sensitive services. You can&#8217;t ignore that. But even if the security issues were solved, universal assistants face a deeper structural problem: they assume people need an assistant in the traditional sense.</p><p>Walk through what even a well-executed version of the same product model looks like.</p><p>Smooth onboarding. Crystal-clear use cases. High-quality AI responses. Clean interface design. Users know exactly what to ask and how to ask it.</p><p>You still have to leave whatever you&#8217;re working on to use it. And when you do, the context you were carrying &#8212; the code you were reviewing, the initiative you were drafting, the design decision you were working through &#8212; is no longer present. You&#8217;ve moved somewhere that knows nothing about any of that.</p><p>So you explain. &#8220;I&#8217;m working on the &#8216;YYY&#8217; data ingestion initiative, and I need to check whether the points Mark raised in Tuesday&#8217;s meeting are addressed in the current design.&#8221; The assistant doesn&#8217;t know what &#8216;YYY&#8217; is. Doesn&#8217;t have Tuesday&#8217;s meeting. Doesn&#8217;t know Mark, the current design, or the organizational context that makes &#8220;addressed&#8221; mean something specific. You load all of it by hand.</p><p>In demos, this overhead is invisible. Demo tasks are self-contained by design &#8212; the context fits in a sentence or two. In practice, your working context isn&#8217;t self-contained. It&#8217;s weeks of accumulated decisions, relationships, dependencies, and constraints that live distributed across your tools. You can&#8217;t paste it into a chat window. You can&#8217;t even fully articulate it. It&#8217;s partially tacit, partially in documents, partially in the history of the tool you&#8217;re using.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!skTt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a50911a-4ef2-491e-8d2e-01234b0fec77_1024x641.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!skTt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a50911a-4ef2-491e-8d2e-01234b0fec77_1024x641.png 424w, https://substackcdn.com/image/fetch/$s_!skTt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a50911a-4ef2-491e-8d2e-01234b0fec77_1024x641.png 848w, https://substackcdn.com/image/fetch/$s_!skTt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a50911a-4ef2-491e-8d2e-01234b0fec77_1024x641.png 1272w, https://substackcdn.com/image/fetch/$s_!skTt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a50911a-4ef2-491e-8d2e-01234b0fec77_1024x641.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!skTt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a50911a-4ef2-491e-8d2e-01234b0fec77_1024x641.png" width="1024" height="641" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8a50911a-4ef2-491e-8d2e-01234b0fec77_1024x641.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:641,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1522333,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/190672571?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff92dde62-56c6-4fa2-ab9a-1147e8c99362_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!skTt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a50911a-4ef2-491e-8d2e-01234b0fec77_1024x641.png 424w, https://substackcdn.com/image/fetch/$s_!skTt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a50911a-4ef2-491e-8d2e-01234b0fec77_1024x641.png 848w, https://substackcdn.com/image/fetch/$s_!skTt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a50911a-4ef2-491e-8d2e-01234b0fec77_1024x641.png 1272w, https://substackcdn.com/image/fetch/$s_!skTt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8a50911a-4ef2-491e-8d2e-01234b0fec77_1024x641.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The Session That Clarified It</h2><p>A few weeks into the new role at Duetto, I was doing product management work in a <a href="https://github.com/bobmatnyc/claude-mpm">claude-mpm</a> session &#8212; reviewing open initiatives, managing the PR queue, creating proposals. Standard operational work for a new CTO getting oriented.</p><p>I wanted to add an infrastructure initiative. Cloud Dev Sleds &#8212; dedicated cloud development machines for the engineering team. The context was in a meeting I&#8217;d had the day before. In the old workflow, this would mean: switch to Granola, find the right meeting, read the transcript, extract the relevant points, switch back, and then write the initiative with that context now loaded in my head rather than in the tool.</p><p>Instead I just asked: &#8220;Review my meeting with Mark yesterday in Granola to get context. I want to create the initiative as a feasibility, cost, and LOE assessment.&#8221;</p><p>The tool pulled the notes. I created the initiative. The product context &#8212; what other infrastructure work was in flight, what the team structure looked like, what the related architectural decisions were &#8212; never left. The Granola content landed inside that context rather than requiring me to carry it manually between tools.</p><p>Same session: needed to check whether I had a conflict for an upcoming demo. Calendar check, without opening Google Calendar.</p><p>Same session: the team needed a status update. Posted directly to the engineering Slack channel, with proper <code>&lt;@USERID&gt;</code> mentions so people actually got notified. The message reflected the same initiatives I&#8217;d been working on all session &#8212; not because I copy-pasted anything, but because the tool already knew what was in flight.</p><p>Later: set up a Notion sync &#8212; initiative statuses with links to the docs, updated automatically.</p><p>The efficiency argument is real but secondary. The more important thing is that the product context never left. The tool knew what initiatives existed, who owned what, what the architectural decisions were, which PRs were waiting on which engineers. When I pulled Granola notes, they arrived inside that context. When I posted to Slack, the message was informed by that context. A universal assistant would have required me to reconstruct and transport that context manually every time I needed to cross a tool boundary.</p><p>No universal assistant is going to have that work knowledge. Not because the AI isn&#8217;t capable. Because the knowledge lives in the tool, accumulated over months &#8212; PRDs, design decisions, initiative history, team assignments, the proposals that got approved and the ones that didn&#8217;t. You don&#8217;t recreate that in a chat window.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!AP4m!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94610dbb-00aa-45e1-93a6-e96a14419f5c_1024x647.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!AP4m!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94610dbb-00aa-45e1-93a6-e96a14419f5c_1024x647.png 424w, https://substackcdn.com/image/fetch/$s_!AP4m!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94610dbb-00aa-45e1-93a6-e96a14419f5c_1024x647.png 848w, https://substackcdn.com/image/fetch/$s_!AP4m!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94610dbb-00aa-45e1-93a6-e96a14419f5c_1024x647.png 1272w, https://substackcdn.com/image/fetch/$s_!AP4m!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94610dbb-00aa-45e1-93a6-e96a14419f5c_1024x647.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!AP4m!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94610dbb-00aa-45e1-93a6-e96a14419f5c_1024x647.png" width="1024" height="647" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/94610dbb-00aa-45e1-93a6-e96a14419f5c_1024x647.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:647,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1372751,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/190672571?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F257a2fb4-ff26-4bb3-8f3e-82d9973a7a60_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!AP4m!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94610dbb-00aa-45e1-93a6-e96a14419f5c_1024x647.png 424w, https://substackcdn.com/image/fetch/$s_!AP4m!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94610dbb-00aa-45e1-93a6-e96a14419f5c_1024x647.png 848w, https://substackcdn.com/image/fetch/$s_!AP4m!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94610dbb-00aa-45e1-93a6-e96a14419f5c_1024x647.png 1272w, https://substackcdn.com/image/fetch/$s_!AP4m!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94610dbb-00aa-45e1-93a6-e96a14419f5c_1024x647.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>The Deep Context Problem</h2><p>The thing that makes domain tools irreplaceable isn&#8217;t AI capability. It&#8217;s accumulated context.</p><p>A product management tool carries months of initiative history. The CTO knowledge base carries organizational decisions, vendor relationships, strategic context that builds over time. These aren&#8217;t things you can summarize in a system prompt. They&#8217;re queryable, interconnected, grounded in real artifacts. The tool has developed something like institutional memory &#8212; and that memory is what makes AI assistance inside the tool qualitatively different from AI assistance outside it.</p><p>Universal assistants are built for breadth. Any question, any domain, any task. That breadth is the pitch and also the structural weakness. The model that&#8217;s ready for anything is primed for nothing specifically. It has no idea that &#8220;the YYY initiative&#8221; refers to a specific ingestion redesign with a particular set of constraints, a particular set of people involved, and three months of design decisions behind it.</p><p>The inversion worth stating plainly: the tools you work in every day already have more relevant context than any assistant will. The right move is surfacing AI capabilities inside those tools, not pulling people out of those tools into a separate assistant layer.</p><p>But here&#8217;s what&#8217;s happening at the executive level. I&#8217;m finding more and more technical executives using Claude Code as knowledge assistance &#8212; not because they&#8217;re universal assistants, but because the amount of data and complexity they can manage far exceeds what standard off-the-shelf tools provide. The deep context problem can&#8217;t be solved with generic solutions.</p><p>For MPM, I built specific connectors: gworkspace-mcp, slack-mpm, notion-mpm, granola-mcp (the last from Granola, the others myself because <a href="https://hyperdev.substack.com/p/mcp-was-a-brilliant-idea-but-it-needs">mcp has limitations</a>). That became as much of an &#8220;assistant&#8221; as I needed, besides izzie. No universal chat interface. Just targeted data bridges that let Claude access specific services when I&#8217;m working on something that needs their context.</p><p>The commercial evidence points the same direction. The AI tooling products with real adoption aren&#8217;t universal assistants. Cursor put AI in the editor. Notion AI put AI in the documents. Linear&#8217;s triage put AI in the issue tracker. Each works because the AI operates inside existing context. The pattern is consistent enough that it&#8217;s probably not coincidence.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tqnj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc968550-faa1-4415-a59d-39051323dc48_1024x751.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tqnj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc968550-faa1-4415-a59d-39051323dc48_1024x751.png 424w, https://substackcdn.com/image/fetch/$s_!tqnj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc968550-faa1-4415-a59d-39051323dc48_1024x751.png 848w, https://substackcdn.com/image/fetch/$s_!tqnj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc968550-faa1-4415-a59d-39051323dc48_1024x751.png 1272w, https://substackcdn.com/image/fetch/$s_!tqnj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc968550-faa1-4415-a59d-39051323dc48_1024x751.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tqnj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc968550-faa1-4415-a59d-39051323dc48_1024x751.png" width="1024" height="751" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fc968550-faa1-4415-a59d-39051323dc48_1024x751.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:751,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1457550,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/190672571?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26977664-13e1-43a1-b5a5-d1bf0f4eb443_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!tqnj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc968550-faa1-4415-a59d-39051323dc48_1024x751.png 424w, https://substackcdn.com/image/fetch/$s_!tqnj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc968550-faa1-4415-a59d-39051323dc48_1024x751.png 848w, https://substackcdn.com/image/fetch/$s_!tqnj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc968550-faa1-4415-a59d-39051323dc48_1024x751.png 1272w, https://substackcdn.com/image/fetch/$s_!tqnj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffc968550-faa1-4415-a59d-39051323dc48_1024x751.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Interface vs Infrastrccture</figcaption></figure></div><h2>What I Got Wrong About My Own Bot</h2><p>I (re)built trusty-izzie as a personal assistant &#8212; natural language queries over my email and calendar history, local graph database, vector embeddings, stays on my machine. It works. But &#8220;personal assistant&#8221; was the wrong frame for where the value lies.</p><p>The thing izzie has is a grounded, real-time, locally-stored representation of my professional life &#8212; people, relationships, projects, scheduling, communications history. That&#8217;s a context store. Every tool I use should have access to it without me switching to izzie to ask.</p><p>The right version of izzie isn&#8217;t the one you talk to. It&#8217;s the one that runs as a local MCP service &#8212; always on, queryable by anything that needs personal context. The product management tool asks it about scheduling. The writing environment surfaces relevant prior conversations. The coding environment knows who owns what system before I have to explain it. None of that requires me to open izzie. It requires izzie to be infrastructure rather than interface.</p><p>If you want to try izzie yourself: <a href="https://izzie.bot/">izzie.bot</a> has the details, and the full source is at <a href="https://github.com/bobmatnyc/trusty-izzie">github.com/bobmatnyc/trusty-izzie</a>. I strongly recommend building from source using an agentic coder to verify the code is safe &#8212; never trust AI tooling with your personal data without auditing it first.</p><p>Not there yet. But the frame shift changes what to build next.</p><h2>What the Architecture Looks Like</h2><p>If you&#8217;re building a personal AI tool, the question isn&#8217;t &#8220;what will users ask the assistant?&#8221; It&#8217;s &#8220;where do users have context, and how do you bring assistance there without making them leave?&#8221;</p><p>The test is simple. Does using your tool require leaving the context where the relevant information lives? If yes, you&#8217;re fighting the architecture. Users will use it occasionally, for low-friction tasks. They won&#8217;t build their workflow around it.</p><p>The tools that pass the test: Claude Code (your codebase is the context), Cursor (you stay in the editor), Notion AI (you stay in the document), Linear AI triage (you stay in the issue tracker). The tools that fail it: every standalone AI assistant that requires opening a new interface and re-explaining what you&#8217;re working on.</p><p>For domain tools with real depth &#8212; months of accumulated decisions, relationships, history &#8212; the connectors are the product. The LLM orchestration is the interface layer. The accumulated context is what no competitor can replicate by building a better general assistant. The moat isn&#8217;t the AI. It&#8217;s what the AI is operating inside.</p><p>For personal infrastructure like izzie: build the MCP service before the chat UI. The chat UI is useful and I use it. The MCP service is what makes the tool true infrastructure rather than one more thing to switch to.</p><p>The universal assistant category isn&#8217;t going to produce a winner because the category is structured wrong. The capabilities will get absorbed by the tools where the relevant context lives &#8212; because that&#8217;s where the value is, and users will figure that out even if product teams don&#8217;t.  The infrastructure driving this &#8212; entity and relationship detection, email, calendar, and task management (all built for Izzie) &#8212; will likely be delivered by the personal productivity tool providers (hello Google).</p><p>Clawd Bot wasn&#8217;t a failed product. It was wildly popular, but I suspect will have been a flash in the pan once the shininess wears off and the liabilities outweigh the usefulness. That distinction matters, because if you think it&#8217;s an execution problem, you go looking for a better universal assistant. If you understand it&#8217;s a conceptual problem &#8212; that most &#8220;assistant&#8221; work is intelligent data movement &#8212; you build infrastructure instead of interfaces.</p><div><hr></div><p><em>Bob Matsuoka is CTO of <a href="https://www.duettocloud.com/">Duetto</a> and writes about AI-powered engineering at <a href="https://hyperdev.substack.com/">HyperDev</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/p/what-does-a-pattern-master-actually">What Does A Pattern Master Do</a>? &#8212; The role of expertise in AI development</p></li><li><p><a href="https://aipowerranking.com/">AI Power Ranking</a> &#8212; Tool comparisons and benchmarks for AI practitioners</p></li><li><p><a href="https://www.linkedin.com/newsletters/ai-power-ranking-7345782916301418496/">LinkedIn Newsletter</a> &#8212; Strategic AI insights for CTOs and engineering leaders</p></li></ul>]]></content:encoded></item><item><title><![CDATA[Are We Heading to a World Where We Only Pay Inference Providers?]]></title><description><![CDATA[A future where you only pay for complexity and scale]]></description><link>https://hyperdev.matsuoka.com/p/are-we-heading-to-a-world-where-we</link><guid isPermaLink="false">https://hyperdev.matsuoka.com/p/are-we-heading-to-a-world-where-we</guid><dc:creator><![CDATA[Robert Matsuoka]]></dc:creator><pubDate>Thu, 05 Mar 2026 12:31:30 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!kbON!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee438ea7-5463-4374-99f1-b53d482ea31b_1024x616.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kbON!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee438ea7-5463-4374-99f1-b53d482ea31b_1024x616.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kbON!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee438ea7-5463-4374-99f1-b53d482ea31b_1024x616.png 424w, https://substackcdn.com/image/fetch/$s_!kbON!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee438ea7-5463-4374-99f1-b53d482ea31b_1024x616.png 848w, https://substackcdn.com/image/fetch/$s_!kbON!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee438ea7-5463-4374-99f1-b53d482ea31b_1024x616.png 1272w, https://substackcdn.com/image/fetch/$s_!kbON!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee438ea7-5463-4374-99f1-b53d482ea31b_1024x616.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kbON!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee438ea7-5463-4374-99f1-b53d482ea31b_1024x616.png" width="1024" height="616" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ee438ea7-5463-4374-99f1-b53d482ea31b_1024x616.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:616,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1174195,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/189515100?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd0c1480a-00f2-4fda-824e-0eb2a4a026fa_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!kbON!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee438ea7-5463-4374-99f1-b53d482ea31b_1024x616.png 424w, https://substackcdn.com/image/fetch/$s_!kbON!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee438ea7-5463-4374-99f1-b53d482ea31b_1024x616.png 848w, https://substackcdn.com/image/fetch/$s_!kbON!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee438ea7-5463-4374-99f1-b53d482ea31b_1024x616.png 1272w, https://substackcdn.com/image/fetch/$s_!kbON!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee438ea7-5463-4374-99f1-b53d482ea31b_1024x616.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Inference vs SASS</figcaption></figure></div><p>Starting Monday morning at 6am, my leadership team gets a Slack notification with our weekly engineering metrics. Commit patterns by developer, product area breakdowns, DORA approximations, behavioral insights like &#8220;Large batch changes&#8221; and &#8220;Afternoon developer&#8221; for each team member. The kind of data that DX or Jellyfish charge thousands per month to provide.</p><p>Total cost to generate that report: $0.005.</p><p>Not $5. Half a penny.</p><p>I just replaced enterprise developer productivity tooling with inference costs. And the replacement isn&#8217;t a compromise&#8212;it&#8217;s better. Custom reports sent directly to leadership through our existing Slack channels and email. Data correlated to our specific product initiatives. Insights that matter to how we actually work.</p><p>The realization landed differently because I&#8217;d been here before. Three times.</p><h2>TL;DR</h2><ul><li><p>Built GitFlow Analytics (GFA) to replace enterprise tools like DX/Jellyfish for $0.005 per weekly report vs. $36K-92K annually</p></li><li><p>Analyzes 154 developers across 100+ repositories, generates automated leadership reports with behavioral insights</p></li><li><p>Total build cost: ~$20K one-time investment vs. $36K-92K annually for enterprise alternatives</p></li><li><p>Requires data engineering skills&#8212;not accessible to every organization, but economics favor custom builds when possible</p></li><li><p>Enterprise analytics bifurcating: zero-footprint vendors (15-min integrations) vs. inference-only custom solutions</p></li><li><p>Pattern recognition becoming commodity; vendors survive on convenience, not intelligence</p></li></ul><p>Before we go further: I&#8217;m talking about enterprise tools that aggregate and analyze data&#8212;developer productivity platforms, team analytics, reporting dashboards. B2B software with complex domain logic or proprietary computation still has enormous value. But enterprise analytics are becoming inference costs.<br><br>I should also point out that DX is a great tool, simple interface, capable.  But expensive and basically unused months after it was installed.  I championed Jellyfish at Tripadvisor, I&#8217;d bet it&#8217;s barely being used there.  The former is due to the effort to personalize, the latter got bogged down in integration costs/time (to be fair, we were running a locally hosted version of JIRA that was a nightmare).</p><h2>The GitFlow Analytics Story</h2><p>GitFlow Analytics started as an internal project at a former client. We needed to understand engineering productivity across our distributed team, but existing solutions were either too expensive or too generic. So I built <a href="https://github.com/bobmatnyc/gitflow-analytics">Gitflow Analytics</a> (GFA)&#8212;a CLI tool that walks git repositories, classifies every commit by work type using inference, handles canonicalization of committers (a surprisingly complex problem), and generates structured reports.</p><p>The system is intentionally simple. Runs on a MacBook Pro with AWS credentials. No GPU, no training data, no vector databases. Just Python scripts that process git logs and make Bedrock API calls to classify commits into Feature, Bug Fix, KTLO, Refactoring, Infrastructure, etc.</p><p>At Duetto, I expanded GFA to analyze over 100 repositories across both Duetto and HotStats. 154 developers tracked. thousands of commits classified. The system handles identity resolution (because different git configs on developer machines create split identities), maps commits to product areas, and generates narrative reports about developer behavior patterns.</p><p>Every morning at 5am, GFA runs automatically. By 6am, my SELT (Senior Engineering Leadership Team) gets a Slack post with the weekly metrics. Individual team leads get personalized HTML reports by email with their direct reports&#8217; patterns.</p><p>The intelligence that interprets raw git data into actionable insights? That&#8217;s Claude Haiku at $0.25 per million input tokens.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!eubl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0c59b41-151e-4cb1-8c44-a554b9cfeeb8_1024x799.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!eubl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0c59b41-151e-4cb1-8c44-a554b9cfeeb8_1024x799.png 424w, https://substackcdn.com/image/fetch/$s_!eubl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0c59b41-151e-4cb1-8c44-a554b9cfeeb8_1024x799.png 848w, https://substackcdn.com/image/fetch/$s_!eubl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0c59b41-151e-4cb1-8c44-a554b9cfeeb8_1024x799.png 1272w, https://substackcdn.com/image/fetch/$s_!eubl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0c59b41-151e-4cb1-8c44-a554b9cfeeb8_1024x799.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!eubl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0c59b41-151e-4cb1-8c44-a554b9cfeeb8_1024x799.png" width="1024" height="799" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a0c59b41-151e-4cb1-8c44-a554b9cfeeb8_1024x799.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:799,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1594282,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/189515100?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F41d37fae-c632-4dd9-af9e-147962b43942_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!eubl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0c59b41-151e-4cb1-8c44-a554b9cfeeb8_1024x799.png 424w, https://substackcdn.com/image/fetch/$s_!eubl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0c59b41-151e-4cb1-8c44-a554b9cfeeb8_1024x799.png 848w, https://substackcdn.com/image/fetch/$s_!eubl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0c59b41-151e-4cb1-8c44-a554b9cfeeb8_1024x799.png 1272w, https://substackcdn.com/image/fetch/$s_!eubl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa0c59b41-151e-4cb1-8c44-a554b9cfeeb8_1024x799.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Economics</figcaption></figure></div><h2>The Economics Are Absurd</h2><p>Here&#8217;s what $0.26 bought me: classification of commits across our entire engineering organization. Every commit message analyzed, categorized, and understood in business context. The corpus represents months of engineering work across multiple product teams.</p><p>Weekly incremental runs cost $0.005-0.01. Maybe 200-400 new commits get classified, the reports regenerate, and leadership gets fresh data. The bottleneck isn&#8217;t inference cost&#8212;it&#8217;s data collection. Git logs are free. The interpretation layer costs pennies.</p><p>Compare this to DX or Jellyfish pricing. DX doesn&#8217;t publish pricing, but industry estimates put enterprise developer productivity tools at $20-50 per developer per month. For our 154-person engineering team, that&#8217;s $36,000-92,000 annually. Jellyfish is reportedly similar.</p><p>My system handles the same workload for about 50 cents per year in inference costs.</p><p>You&#8217;re not paying for intelligence when you buy enterprise analytics. Intelligence costs nothing. You&#8217;re paying for data collection infrastructure, UI development, customer support, sales teams, compliance certifications. All the overhead of running a SaaS business.</p><p>We obsess over the high costs of advanced coding agents and reasoning models. But the cost of capable text inference&#8212;the kind that powers pattern recognition and reporting&#8212;is dropping through the floor. Haiku handles commit classification as well as Opus would. For enterprise analytics, you don&#8217;t need the flagship models. You need reliable categorization and natural language generation, which commodity models deliver for pennies.</p><p>But if you know how to build? You don&#8217;t need any of that. You just need (more expensive) inference.</p><h2>What GFA Actually Delivers</h2><p>The data flowing to my leadership team isn&#8217;t generic dashboard noise. It&#8217;s intelligence designed around how we actually operate.</p><p>Every developer gets a behavioral profile: &#8220;Large batch changes,&#8221; &#8220;Afternoon developer,&#8221; &#8220;Exceptional performer (Top 20%).&#8221; These aren&#8217;t arbitrary labels&#8212;they&#8217;re LLM interpretations of quantitative patterns. Commit size distributions, time-of-day histograms, percentile rankings converted into readable insights.</p><p>Product area attribution happens automatically. The system maps our repositories to eight business areas: Frontend, Core Product, Integrations, Data Platform, Intelligence/ML, Infrastructure, QA/Testing, Developer Tools. When we see commit patterns shifting from Core Product to Infrastructure, that signals architectural decisions playing out in code.</p><p>DORA metrics approximate from git data. Deployment frequency tracks through release tags. Lead time measures first commit to merge. We don&#8217;t get the full Four Keys implementation, but we get enough signal to spot trends and outliers.</p><p>The identity resolution was the hardest part. Not the LLM calls&#8212;those work fine. But knowing that different developer machines create split git identities required human judgment to build the canonical mapping. Once you solve identity, everything else flows from structured data.</p><p>Most importantly, the reports answer questions executives actually ask. &#8220;Which teams are handling the most KTLO work?&#8221; &#8220;Are we seeing more bug fixes or new features this quarter?&#8221; &#8220;Who&#8217;s working weekends and why?&#8221; These aren&#8217;t metrics you find in generic productivity dashboards. They&#8217;re insights that matter for our specific business context.</p><h2>We&#8217;ve Been Here Before</h2><p>Enterprise software was expensive because intelligence was scarce. In the 1990s, turning raw data into insights required Oracle licenses, dedicated servers, and consultants. The SaaS revolution changed delivery but not fundamentals&#8212;you still paid massive recurring costs for pattern recognition and reporting.</p><p>Intelligence is now a commodity API call. You don&#8217;t need Tableau because you can generate charts and send them through Slack. You don&#8217;t need Looker because Claude can summarize SQL results. The bottleneck was never data storage&#8212;it was interpretation. When interpretation costs pennies, everything else becomes optional.</p><h2>The Development Cost Reality</h2><p>Building GFA required skills not every organization has&#8212;data modeling, Python scripting, API integration, identity resolution patterns. Conservative total development cost: around $20,000 including engineering time, infrastructure setup, testing, and iteration cycles.</p><p>For organizations with engineering talent, the economics have shifted dramatically. A $20,000 one-time investment delivers exactly what we needed versus $36,000-92,000 annually for enterprise alternatives. As one SELT member said: &#8220;This is exactly what we wanted.&#8221;</p><p>You own the data model. Custom breakdowns take SQL queries, not feature requests. When we needed behavioral insights, I added prompts interpreting commit patterns. When leadership wanted trends, I added 12-week windows. Each enhancement took hours, not months&#8212;because intelligence was delegated to inference, and data processing was just Python.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!G6dJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e43d720-88bf-459c-b372-4568c71af682_1024x608.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!G6dJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e43d720-88bf-459c-b372-4568c71af682_1024x608.png 424w, https://substackcdn.com/image/fetch/$s_!G6dJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e43d720-88bf-459c-b372-4568c71af682_1024x608.png 848w, https://substackcdn.com/image/fetch/$s_!G6dJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e43d720-88bf-459c-b372-4568c71af682_1024x608.png 1272w, https://substackcdn.com/image/fetch/$s_!G6dJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e43d720-88bf-459c-b372-4568c71af682_1024x608.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!G6dJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e43d720-88bf-459c-b372-4568c71af682_1024x608.png" width="1024" height="608" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6e43d720-88bf-459c-b372-4568c71af682_1024x608.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:608,&quot;width&quot;:1024,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1228168,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://hyperdev.matsuoka.com/i/189515100?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdd660a81-5330-41e9-87ac-77edd9141f8f_1024x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!G6dJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e43d720-88bf-459c-b372-4568c71af682_1024x608.png 424w, https://substackcdn.com/image/fetch/$s_!G6dJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e43d720-88bf-459c-b372-4568c71af682_1024x608.png 848w, https://substackcdn.com/image/fetch/$s_!G6dJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e43d720-88bf-459c-b372-4568c71af682_1024x608.png 1272w, https://substackcdn.com/image/fetch/$s_!G6dJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e43d720-88bf-459c-b372-4568c71af682_1024x608.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Simplicity</figcaption></figure></div><h2>The Pattern Playing Out</h2><p>Enterprise analytics tools are pattern-matching at scale across thousands of customers. But every organization&#8217;s data is different&#8212;your repositories, team structure, product areas, business priorities. Generic dashboards force you to map specific context onto generic data models. Insights lose precision in translation.</p><p>When custom analytics cost pennies, the calculation flips. Instead of generic insights that sort of fit, you build specific insights that exactly fit. Customization drops from &#8220;feature request, wait six months&#8221; to &#8220;write prompt, test output.&#8221;</p><p>Enterprise vendors aren&#8217;t solving technical problems&#8212;they&#8217;re solving procurement, compliance, and integration problems. Important work, but not work justifying massive recurring costs when intelligence is commodity inference.</p><h2>The Zero-Footprint Exception</h2><p>Not every vendor gets replaced. Some survive by making integration frictionless enough that convenience beats custom builds.</p><p>We&#8217;re evaluating <a href="https://www.augmentcode.com/">Augment Code</a>&#8217;s code review service for Duetto. Their value proposition isn&#8217;t features&#8212;it&#8217;s zero-footprint integration. Fifteen-minute setup call, quick estimate, running production code reviews with minimal configuration. When customers can build equivalent functionality for pennies, your value proposition becomes the path to value, not the functionality itself.  This is an important lesson for us.  We do handle massive complexity and data, hard for smaller customers to manage themselves, but need to do better at simplifying integration.</p><p>The intelligence is commodity; the packaging is differentiated.</p><h2>Where This Goes</h2><p>The unbundling accelerates. Enterprise analytics face two paths: become zero-footprint integration plays or get replaced by inference-only custom builds.</p><p>What survives:</p><p><strong>Genuinely complex software.</strong> Revenue management algorithms, fraud detection engines, supply chain optimizers&#8212;systems requiring proprietary computation, not pattern recognition.</p><p><strong>Zero-footprint integrations.</strong> Fifteen-minute setups with immediate value. When alternatives cost pennies but require engineering skills, convenience must be measured in minutes.</p><p><strong>Proprietary data advantages.</strong> GitHub&#8217;s intelligence benefits from every public repository. LinkedIn draws from member networks. Data moats protect against inference-only competition.</p><p>Everything else becomes vulnerable to custom builds powered by inference calls.</p><p>The market bifurcates. Companies with engineering teams build custom analytics for pennies and get better insights than generic dashboards. Companies without those skills pay for zero-friction integrations.</p><p>Are we heading to a world where we only pay inference providers? For organizations with the skills to build, we&#8217;re already there. For everyone else, vendors survive by making paying them simpler than learning to build alternatives.</p><div><hr></div><p><em>I&#8217;m Bob Matsuoka, CTO at Duetto and writer on AI development tools and software economics at <a href="https://hyperdev.substack.com/">HyperDev</a>.</em></p><p><strong>Related reading:</strong></p><ul><li><p><a href="https://hyperdev.matsuoka.com/p/ai-job-transformation">The AI Job Transformation: Pattern Masters, Not Coders</a> - The pattern-matching analogy and why reasoning matters</p></li><li><p><a href="https://hyperdev.matsuoka.com/p/fix-feat-ratio">The Fix:Feat Ratio - The Metric That Actually Matters</a> - Quality metrics in AI-assisted development</p></li><li><p><a href="https://hyperdev.matsuoka.com/p/opus-vs-sonnet-quality">Claude Opus 4.5 vs Sonnet 4.5: When Quality Beats Speed</a> - Choosing the right model for the task</p></li></ul>]]></content:encoded></item></channel></rss>