We’re starting to see what I’m calling “Billion Token” hires. Engineers capable of consuming a billion tokens every month, and increasingly other AI-native employees who use harnesses for their daily work. Coding makes that scale easier to reach because implementation agents, testing, and code review can keep running together. I’ll focus on engineers here.
I’ve been living this for over a year. My harness sessions run coding agents in parallel: implementing features, investigating bugs, writing tests, and reviewing changes. These agents are my shadow army. The harness organizes their work and checks what comes back. I remain responsible for what I accept and ship.
Let’s define a billion-token engineer (BTE) as one that consumes (let’s say) a billion tokens a month excluding cache reads. Repeatedly reading cached context still costs money, but it doesn’t count toward that billion. My year-end “wrapped” listed 70B tokens for 2025, but that included cache reads and inflated the number relative to this definition. Last week I used about 0.8B tokens excluding cache reads. A billion a month is hardly outlandish when I consumed most of that in a week. Token volume alone tells us little about how much useful software came out.
My September usage analysis put one BTE at roughly $17,000–$20,000 a month at the API list prices for the models I used, including the associated cache reads. That estimate uses incomplete transcript counts and my model mix. It is not my paid bill or a universal token price. For the thought experiment, use a $20,000 monthly inference budget per engineer, before salary and shared infrastructure.
Now hire twenty of them.
Today’s article discusses what happens when you do, and why I believe personal harness building may be the defining skill of this era.
TL;DR
Twenty BTEs would bring twenty shadow armies into the organization: agents driven by each engineer’s harnesses. A twenty-person team behind each engineer is an assumption for this thought experiment, not a measured productivity multiplier.
The workload has to fit through review, integration, and user acceptance. Increasing generation without expanding those capacities creates unfinished work.
Personal harnesses need shared rules for ownership and acceptance when their engineers work on the same repository.
I think we will increasingly hire engineers for the output of their harnesses and their willingness to stand behind it. Insurance and certification are possible consequences of that responsibility.
One Person, One Week
For September 14–20, I counted the activity in my local harness logs and matched it with the GitHub records I could access. These are aggregate figures across my work, with the project names left out.
One person directing the work, including sessions started or resumed by automation. The session counts cover the week, not simultaneous agents. Tokens include fresh input, cache creation, and output, with cache reads excluded.
The GitHub figures cover my account’s work in repositories represented in the logs. Access gaps leave some work uncounted. Builds include repeats and separate platform targets, but exclude unverified local builds. The 10,438 automated workflow runs overall included notifications and other checks, so I haven’t counted them all as builds.
Opened and merged PRs are different sets, as are created and closed issues. Neither a merge nor a closed ticket establishes user acceptance.
What I’m Delegating
I’m largely delegating this work: setting the goal, supplying context, and evaluating what comes back. Much of it goes into internal applications, a hackathon site, companion tools, and CTO reporting. Other work involves analyzing knowledge for my CTO responsibilities, running classification experiments, or building the harnesses themselves.
That lets us build optional internal tools without hiring or allocating engineers specifically for them. Some wouldn’t justify a dedicated hire, so part of the value is doing work that might otherwise stay on the wish list. The activity counts don’t translate into salaries saved (or jobs lost), but these projects become possible without that staffing commitment. You can also make the argument that this work makes ALL engineers more effective (I spent 8 years at a company struggling to make product work and code visible/searchable, now I can trace any code or PRD and make reasonable estimations on any effort quickly).
I can also delegate quick projects that make a proposed direction concrete. A working example gives people something to examine and helps us decide whether the idea deserves more investment, even if the example doesn’t become a production product.
Twenty Engineers and Their Shadow Armies
But let’s expand this notion in the workplace. Suppose each BTE can direct something resembling a team of twenty implementation workers. Twenty hires would bring four hundred workers’ worth of assumed implementation capacity. Their token budget would be $400,000 a month, or $4.8 million a year, plus compensation and shared infrastructure.
Twenty is a thought-experiment assumption. I don’t have a comparable human team completing the same assignments to the same acceptance standard, and agents spend much of their time checking or discarding other agents’ work. Use five and the company still has a hundred workers’ worth of assumed capacity arriving through twenty people. A useful multiplier would need to account for accepted changes, human review time, and subsequent defects and rework.
Each engineer runs a small software factory with its own practices and assumptions (presumably with alignment on core best practices) about when work is done. Someone has to divide useful work among them and settle disagreements about the result. Adding factories can increase the code available for review much faster than the organization’s ability to accept it. They still have to deliver a coherent product.
The Issue Debris Field
Mac, my colleague at work, described what happens to the follow-up issues our harnesses generate. He could keep track of five to seven parallel sessions working on substantial problems. What overwhelmed him was the “issue debris field” accumulating beneath those efforts, across several projects. A well-defined plan breaks into implementation issues. Working those issues produces follow-ups, some critical and some scope creep. Left unattended, they go stale as the code changes. Picking up a month-old ticket then requires someone to establish whether the problem still exists, whether later work partly solved it, and whether it deserves attention now. That consumes model time and human attention before implementation can begin.
Mac had tried limiting follow-up creation. A ban could hide critical problems, while a critical-only rule still left ownership and tracking unresolved. The direction we discussed was to attach those findings to a parent tracker, keep its state current, and resolve or reassess them before they became stale. Research would live in committed documents linked from the issues, so another contributor could recover the context without needing the original session.
Even working on the fix exposed the coordination problem. I’d opened an issue-management effort while Mac was researching a related approach. He spotted the overlap and warned that we were heading for a collision. We agreed to put his work into the existing effort. Generating another plan was easy. Recognizing the overlap required a conversation.
Think of a commander directing several units toward agreed objectives. Every advance uncovers side missions, blocked routes, and reports that may already be obsolete. The units can keep moving while the commander loses track of what their movements mean for the campaign. A shared map needs to connect each finding to the objective, its owner, and a decision about what happens next. That is the role we were proposing for the parent tracker.
Generating more work is no longer the limiting factor in this part of my practice. Organizing it is. A newly filed issue needs a place in the plan and a decision about whether to pursue it. Otherwise, the shadow army can fill the tracker faster than its owner can determine what still needs doing.
Three BTEs in the Same Repository
Mac and I mostly work in different repositories, which gives us room to operate independently. What happens when three BTEs work on the same repository? Separate branches protect unfinished edits, but three sensible assignments can still conflict. A shared interface changes beneath another army’s implementation. Its tests may pass against assumptions that are no longer true.
In When Sessions Talk..., I documented my own sessions negotiating ownership and merge windows, with some messages never appearing in the receiving model’s context. Even one human needed a durable place for decisions. Three humans bring separate priorities and authority.
I’d want work claimed before dispatch, naming the responsible engineer, the subsystem, and any shared contracts affected. The claim must change when scope expands. Teams also need limits on concurrent changes to shared components and agreement on interfaces before several armies implement against them. Agents can investigate alternatives while a decision is pending without all modifying the same subsystem.
A merge queue can test proposed changes against the target branch and earlier queued changes. It can’t settle an architectural disagreement the tests don’t express. A named owner still needs authority to hold the change until that decision is made.
The Work Waiting to Be Accepted
A tracker that says “coding complete” hides much of the remaining work. Implementation, independent review, integration, and user acceptance need distinct states, with evidence and someone responsible for the next decision.
Suppose twenty BTEs each submit five changes a day that need a human decision. At fifteen minutes apiece, those hundred decisions require twenty-five reviewer-hours a day, before rework or user acceptance testing. These are illustrative assumptions. The staffing problem gets harder if many changes need the same architect or product owner. A dashboard makes that queue visible without increasing the person’s capacity, while continued implementation changes the code underneath the waiting work.
I would manage the factories by accepted output and the age of their queues. How long does review take? How often does integration invalidate earlier checks? How much accepted work returns as a defect? A shadow army whose output waits unreviewed for days needs fewer new assignments and more effort clearing work in progress.
BTEs will need help tracking what is running, what is waiting, and what needs their attention. I have a possible solution in the concept of a supervisor. I’ll write more about that in a week or so.
User acceptance testing (UAT) needs its own capacity plan. Agents can prepare environments and test data, then replay agreed workflows. But passing criteria that misunderstand the customer’s task repeats the misunderstanding. A product owner or domain expert must decide whether the behavior is useful. They need a reviewable release candidate, a description of the affected workflow, and evidence tied to that version. If the implementation changes afterward, the team needs to know which acceptance results remain valid. Asking a user to approve a moving target makes the approval difficult to defend.
What the Harness Has to Establish
Writing an implementation gave me repeated opportunities to discover mistakes in my reasoning. When agents do more of that construction, the harness needs other ways to expose them. Types and static analysis help, alongside tests that exercise the surrounding system. Production observations need to feed back into the process. Agreement between agents provides weak evidence when both follow the same mistaken specification.
I also want to challenge the validation itself. Would the tests catch a deliberately broken implementation? Does a check fail when the behavior it protects disappears? A passing suite is more useful when I understand its limits, though that still leaves the question of whether we specified the right behavior.
The harness should retain what was requested, which version was evaluated, the checks performed, and the unresolved risks. A reviewer needs to reconstruct the acceptance decision without replaying the agent’s entire conversation. Building that ability consumes some of the capacity the shadow army appears to provide for free.
The Engineer Comes With a Factory
In Is Your Digital Brain the Light Saber of the AI Era?, I compared building a personal knowledge system to a Jedi building a lightsaber. The construction reflects how its owner thinks. I now think the engineer and harness relationship is a more apt analogy.
The harness is a tool of the trade, shaped by its driver’s style. Judgment appears in how the engineer decomposes problems, the failures their tests anticipate, and the conditions that make an agent stop and ask. It accumulates in instructions and executable checks that the engineer maintains.
That changes what I would look for in a hire. I want to see what their harness produces and how they decide to accept it. Show me a failure it caught. Show me a failure it missed, and what changed afterward. Show me how another engineer can investigate the result when you’re unavailable. A large token budget tells me that the factory runs. It doesn’t tell me whether I should trust its output. And I don’t particularly care which harness they use, just that they know how to adapt it for their own and my company’s needs: workflow, PR style, issue management, etc.
The company still needs common acceptance requirements and production-access rules. Personalization shouldn’t leave colleagues dependent on a private system they cannot inspect or operate. Hiring a team with shadow armies means judging whether their harnesses can work together and their owners can resolve conflicts.
You Can’t Fire an Agent for Negligence
An agent can be stopped, replaced, or given different instructions. None of those actions answers a customer asking who accepted the defective work. The engineer who directs the factory needs to stand behind what it ships. So does the organization that authorized the work and decided where to deploy it.
I think the professional definition of an engineer moves toward accountability for code produced under their direction, including code they didn’t write. I want to take that as far as professional liability. Will BTEs carry liability insurance? Will customers ask for certification of the harnesses producing their software? Those are possibilities to explore, not existing legal requirements. Delegated implementation should leave a recognizable person responsible for the engineering judgment.
A useful certification would examine the operating process, including how it handles failure. Changes to models, tools, or permissions could require reevaluation. An insurer, if this market develops, would have reason to examine those controls and the losses they failed to prevent. The harness supports the engineer’s assurance of quality by making decisions inspectable and providing a way to improve after failure. That is more persuasive to me than insisting they read every line, though neither promises defect-free code.
There is an apprenticeship problem inside this, too. Engineers need enough understanding to recognize bad output and challenge a weak test. Junior engineers will need practice investigating failures and operating the systems their agents help build. Giving them a finished harness without teaching them how to question it would leave the accountability mostly ceremonial.
I believe this is the engineering model for the next few years. You hire people whose factories produce solid, tested, defensible code, then build an organization capable of accepting that work. Their personal methods can differ. The shared product still needs owners who can explain why it was safe to ship and take responsibility when they were wrong.
Bob Matsuoka is CTO of Duetto, a hospitality profit and revenue-management platform, and writes about AI-augmented engineering practice. Previously, Bob has been CTO of Tripadvisor, Citymaps, and Runtime Technologies.
Related reading:
Is Your Digital Brain the Light Saber of the AI Era? — Building a personal knowledge system as part of developing your practice.
When Sessions Talk... — What my agent sessions coordinated, and where their messages failed to arrive.
We Need to Teach Delegation as an Engineering Skill — Defining outcomes and authority so another actor can do useful work.
AI Power Ranking — Tool comparisons and benchmarks for AI practitioners.
LinkedIn Newsletter — Strategic AI insights for CTOs and engineering leaders.






