OpenAI's "Comeback"
Can it really be a comeback with they have so much money, so much exposure, and so little revenue?
GPT-5.6 Sol scores 80 on the Artificial Analysis Coding Agent Index against 77 for Claude Fable 5, at a lower cost per task (about $1.04). That post went up on July 9.
I called model parity back in April, in a piece with a section headed “GPT-5.4 Caught Up.” What I missed was everything around the model: that OpenAI would close the product gap, and buy a decade of compute while doing it.
Over the last few months I’ve been using ChatGPT more and more for work-like tasks. Not Duetto work, because we don’t have an enterprise plan and Duetto material stays out of it. Work-like things. Both the capabilities and the UX have improved to the point where I’d put the product at least on par with claude.ai, to my surprise. There are still gaps. It’s still good.
TL;DR
GPT-5.6 Sol leads the Artificial Analysis Coding Agent Index 80 to 77 over Claude Fable 5 (July 9) and Terminal-Bench 2.1 85.77% to 84.64% over Claude Opus 5 (July 22), the second one with a methodology caveat that widens the gap.
Claude Opus 5 still leads the broader Artificial Analysis Intelligence Index at 60.7 against GPT-5.6 Sol’s 58.9 (July 29). Anthropic’s remaining model lead is general capability, not agentic coding.
Anthropic leads on revenue ($47B run-rate, May 28, 2026) and valuation ($965B), and leads enterprise API share 40% to 27% (Menlo, December 2025). OpenAI closed the capability gap and most of the product gap while trailing commercially.
OpenAI is far ahead on planned compute (Stargate at nearly 7GW toward a stated 10GW commitment, the Broadcom “Jalapeño” inference ASIC). In July it moved its usage limits three times in seventeen days, in both directions, around a near-global outage on the 25th. Anthropic has been raising its limits since May.
No open-weight provider I found fields an app ecosystem within reach of ChatGPT or Claude Code, which keeps this a two-horse race even as the weights converge. Chinese domestic app numbers could complicate that.
Where the coding lead sits
Two current measures put GPT ahead on agentic coding. The Coding Agent Index lists GPT-5.6’s three effort tiers at Sol 80, Terra 77.4 and Luna 74.6, with Claude Fable 5 at 77.2 (Terra and Fable 5 land 0.2 apart). Sol’s per-task cost runs below both Fable 5 and Opus 4.8. On Terminal-Bench 2.1 (89 tasks, Terminus 2 harness, snapshot dated July 22), GPT-5.6 Sol leads Claude Opus 5 by about a point, 85.77% to 84.64%.
The Terminal-Bench result comes with a caveat from the publisher. Opus 5’s run used Claude Opus 4.8 as a refusal fallback. Count the nine affected passing results as failures instead and Opus 5 drops to 81.27%, which widens the GPT lead to roughly four and a half points. The lead is real. Its size depends on how you count.
Now the other direction. On the broader Artificial Analysis Intelligence Index, as of July 29, Claude Opus 5 sits first at 60.7 and Claude Fable 5 second at 59.9, with GPT-5.6 Sol third at 58.9. Anthropic holds the top two spots on general capability while losing the coding-agent tables. My “diminishing edge” framing was directionally right if imprecise. The edge moved. It didn’t disappear.
SWE-bench Pro, the number everybody reached for six months ago, has no usable current standing. Scale’s public-set leaderboard doesn’t list GPT-5.6, Claude Opus 5 or Claude Fable 5 at all, and the top of it reshuffled again this month. The aggregator sites quoting newer figures disagree with each other and cite at least one model name I can’t confirm exists. Skip it.
The product story is a two-way race now
This is the first half of what I missed. Codex spans CLI, IDE, web, and a hosted cloud mode for long-running tasks. The Codex CLI changelog for the last two weeks of July includes resumable sessions with paginated thread history, session naming and branching, configurable sub-agents, audio input, and an /import command that migrates project-scoped memories out of Cursor and Claude Code. That last one reads like a decision made by somebody thinking hard about switching costs. On July 9 OpenAI also folded the standalone Codex desktop app into a single ChatGPT desktop app with Chat, Work, and Codex modes. Existing Codex users were moved across automatically, and Codex remains a full mode inside the merged app.
What I notice in use isn’t a benchmark, it’s flow. Codex and GPT live in the same product, and when I go looking for something I worked on weeks ago, ChatGPT finds it wherever I left it. Claude’s Projects still behave like sealed folders for me. That isolation was an advantage once, back when context was scarce and a hard wall around what the model could see was the safer design. With current context budgets it mostly means I go hunting. GPT is also the far better image renderer, which makes for a short comparison, since Claude doesn’t ship a first-party image model at all.
Anthropic didn’t stand still while any of this happened. Claude Design shipped out of Anthropic Labs on April 17, and it’s excellent: you talk to it and get slides, wireframes, one-pagers and pitch decks, with export to PPTX or PDF, inline comments, adjustment sliders, and handoff into Claude Code. Codex has no equivalent that I’m aware of (and I’d bet real money they’re building one). Claude Cowork expanded to web and mobile on July 7. Both labs shipped serious B2B product work inside the same four months.
Capacity: a much bigger future, a strained present
The second half is compute. OpenAI is buying a much bigger future than Anthropic is, and is visibly straining in the present.
On committed buildout it isn’t close. OpenAI, Oracle and SoftBank put combined planned Stargate capacity at nearly 7 gigawatts across the Abilene flagship plus five more sites, with over $400B invested and the full commitment still stated as $500B and 10GW. Tomasz Tunguz’s aggregation of the announced vendor deals puts OpenAI’s 2025–2035 infrastructure commitment at roughly $1.15 trillion across seven suppliers, which is derived rather than OpenAI-confirmed. On June 24, OpenAI and Broadcom unveiled Jalapeño, OpenAI’s first custom inference ASIC, designed to tape-out in about nine months, with engineering samples already running Codex workloads in the lab and first production deployment targeted at gigawatt scale in late 2026 (CNBC’s coverage has the same details).
Anthropic’s disclosed commitments are significant but much smaller: up to 5GW with AWS announced April 20 (with nearly 1GW of Trainium online by year end and $100B+ committed over ten years), well over 1GW of Google TPU capacity in 2026 with the Broadcom-built expansion arriving from 2027, and more than 300MW from SpaceX’s Colossus 1. Its custom silicon is at the conversation stage, reportedly early talks with Samsung on a 2nm accelerator with nothing shipping before late 2027.
The strain shows up in the usage limits, which OpenAI moved three times in seventeen days, in both directions. Sol’s launch week doubled traffic inside 48 hours, so on July 12 and 13 it lifted the five-hour cap on Plus, Pro and Business tiers to absorb the load. On July 25 came a near-global outage, 9:17 to 11:08 AM ET, elevated errors across ChatGPT, the API and Codex, resolved the same morning. On July 29 it reset banked limits for ChatGPT Work and Codex users, whose quota Sol was burning faster than the math anticipated, then put the five-hour cap back the next day. Anthropic has been moving the other way since spring. On May 6 it permanently doubled Claude Code’s five-hour limits for Pro, Max, Team and seat-Enterprise users and dropped peak-hour throttling for Pro and Max, on the back of that SpaceX capacity.
Where Anthropic still leads
Anthropic’s remaining, potentially durable, advantages are enterprise adoption and commercial scale. Anthropic is ahead on total revenue and on valuation, and well ahead on enterprise API share. Its Series H, a $65B round closed May 28, 2026, disclosed a $47B run-rate and a $965B post-money valuation. OpenAI’s most recent confirmed figures are roughly $25B annualized (about $2B a month) at the March 31, 2026 close of its $122B round, at an $852B post-money valuation. Menlo Ventures’ December 2025 enterprise report put Anthropic at 40% of enterprise LLM API usage against OpenAI’s 27%, and at 54% in AI coding specifically. No 2026 Menlo report exists yet, so December is the current number.
So OpenAI has closed the capability gap and most of the product gap while trailing on money and on enterprise accounts. That’s the actual shape of the comeback. The lead Anthropic still holds on the Intelligence Index is general capability, and general capability is the harder thing to put in front of a CTO who’s buying a coding agent this quarter.
Apple versus Microsoft, again?
I reached for that analogy myself in May 2025, when OpenAI spent close to $10 billion in three weeks on Jony Ive’s io Products and on Windsurf, and I called it OpenAI’s Apple moment: control the stack from silicon to screen, consumer-first, design-led. Anthropic in that mapping is Microsoft, enterprise and developer-led, winning the accounts while nobody writes headlines about it.
I don’t fully trust it. Apple versus Microsoft was never settled by who had the better quarter. It ran on which company owned the layer everybody else had to build on top of, and on switching costs measured in years. Nobody owns that layer here. Both companies rent it, from Amazon, Google, Microsoft, Nvidia, Broadcom and now SpaceX. For a developer, changing frontier models is a config change and an eval run. What costs real time is changing the tool you work in, which is a much smaller lock than owning the platform. The Apple/Microsoft dynamic assumed you couldn’t leave.
The mapping also keeps slipping. Anthropic is the enterprise player in it, which is right, but it’s also the one with the bigger revenue base and the higher valuation, which is not the role the analogy assigns it.
Can either of them make money
Neither is profitable and both are burning billions, so anything past that is projection.
OpenAI’s advertising business is on pace to miss its own five-year forecast by about 90%, OpenAI projected $2.5B in ad revenue this year, scaling to $100B by 2030. eMarketer puts the entire US chatbot-ad market, not just OpenAI, under $1B this year and around $5.4B by 2030, per 24/7 Wall St. on July 21. The widely circulated 2026 cash-burn estimates in the $25B range are analyst syntheses rather than disclosures, because OpenAI is private and doesn’t publish the figures that would settle it. Same caution applies to the gross-margin numbers people quote for both companies.
Anthropic’s topline looks better but is contested. Ed Zitron argues in ”Anthropic’s ‘Profitability’ Swindle” that the near-term EBITDA profitability claim is a one-quarter accounting artifact of temporarily discounted SpaceX compute during exactly the months profitability was claimed, and that costs still rise linearly with revenue. He also flags CFO Krishna Rao’s March 9 sworn filing citing cumulative revenue “exceeding $5 billion to date” against the $19B run-rate announced six days earlier. That second point is weaker than it reads, since cumulative revenue and annualized run-rate aren’t the same measure. The first point is the one I’d want answered, and it hasn’t been.
Hardware doesn’t resolve it. Chris Lehane said at Davos in January that OpenAI’s first consumer device would debut in the ”latter part” of 2026, which is an unveiling and explicitly not an on-sale date. MacRumors reported in February that the thing is a smart speaker with a camera launching in 2027. Everything else circulating about it (screenless, voice-first, 360-degree camera, Foxconn in Vietnam) is unconfirmed. As of today nothing has shipped and no date is confirmed, so I’d file the device as a 2027 question and stop considering it this year.
Open weights are a side argument
The open-weight families are close to the frontier on raw capability. Kimi K3, the highest-scoring open model on the Artificial Analysis Intelligence Index, sits at 57, about four points behind Claude Opus 5’s 60.7 and under two behind GPT-5.6 Sol’s 58.9, the same index cited above. On LMArena the best open-weight entry, Kimi K3-max at 1491, trails Claude Fable 5 at the top by 17 points.
What none of them have is the app ecosystem. No open-weight provider fields anything within reach of ChatGPT or Claude Code on integration depth or developer reach, at least none I found, though I didn’t dig into DeepSeek’s or Qwen’s domestic app numbers in China, which could complicate that. Weights converging matters much less than it sounds like it should when the surface people work in belongs to two companies.
Which is where I end up. The model tables are the smallest part of it. OpenAI put Codex and ChatGPT in one app with an import command pointed straight at Claude Code’s users, and it has nearly 7 gigawatts of planned capacity standing behind that. Claude Opus 5 still holds the top of the Intelligence Index at 60.7, and Anthropic still held 40% of enterprise API spend as of last December. OpenAI is knocking on the door. A year ago I’d have told you that door was closed.
A coda. Could we imagine a world in which Anthropic and OpenAI run products on other models? Seems very unlikely now. But the cost of developing new models is astronomical, the cost of derived models much less (even though both companies are fighting distillation tooth and nail). So there is a world in which they use the model building work they’ve done as an audience and tool building exercise, and focus on expanding that share by expanding options, let the models fight for their own justification. Farfetched, as I mentioned, but if they start losing to application clones that use open weight models, they’re essentially inviting someone to compete. The existing assumption that the app ecosystem was just to drive captured inference falls apart if that inference is cost prohibitive. We shall see.
Bob Matsuoka is CTO of Duetto and also writes about AI business at AI Power Ranking.
Related reading:
OpenAI’s ‘Apple Moment’: Building a Walled-Garden AI Stack — Where I first made the Apple analogy, three weeks and $10 billion into OpenAI’s acquisition run.
It’s the Harness, Stupid — The April piece where I already conceded model parity, under a section header that says so. Everything I got wrong after that is in this article.
I Switched to Claude.AI from ChatGPT As My Main AI Assistant — The May 2025 call I’m revisiting here, with the context-management and stability reasons that drove it.
AI Power Ranking — Tool comparisons and benchmarks for AI practitioners.
LinkedIn Newsletter — Strategic AI insights for CTOs and engineering leaders.



