Jul 28, 2026 · 12:19 AM
Subscribe
Home Ai

OpenAI quietly slashed GPT-5.6 Sol's reasoning power by 87% four days after launch

OpenAI slashed GPT-5.6 Sol's Max-tier reasoning budget by 87% four days after launch with no notice, then called it experimental parameter tuning when developers revolted. The incident reveals a structural risk in building on AI APIs that most startup contracts don't account for.

Julian Lim
· 5 min read · 547 reads
OpenAI quietly slashed GPT-5.6 Sol's reasoning power by 87% four days after launch

OpenAI did change GPT-5.6 Sol's reasoning settings after launch, then rolled the change back when users noticed. The specific 87% claim doesn't check out, but the trust problem for startups is real.

Sam Altman's launch-day pitch for GPT-5.6 Sol was simple: better work for fewer tokens. CNBC's July 9 transcript has the OpenAI CEO saying Sol was 54% more token efficient on agentic coding tasks, while being as good as or better than the strongest rival models. That part checks out. The mess came after launch, when Codex and ChatGPT Work users started comparing notes on context limits, usage burn and what some called a Sol nerf.

The original claim was sharper than the evidence supports. Searches turn up reporting and developer discussion around OpenAI changing reasoning effort, sometimes called juice values under the hood, but I couldn't verify the article's specific claim that the Max-tier reasoning budget fell from 960 to 128 overnight. The 87% figure depends on that number. It has to go.

What remains is still serious. According to a July 17 AI Deep Signal roundup of Tibo Sottiaux's X thread, OpenAI said it had deployed inference optimizations that should give GPT-5.6 Sol users roughly 10% more usage, rolled the Codex product context limit back from 372,000 to 272,000 after higher-than-intended usage charging, reverted experiments that changed reasoning efforts, and was fixing heavier-than-expected multi-agent behavior at high and xhigh settings. That is not the same as the clean story users thought they had bought on July 9.

Users felt it anyway.

Sottiaux, described in multiple reports as an OpenAI Codex engineering lead, also said the 8 million milestone covered active users across Codex and ChatGPT Work, not Codex alone. That distinction matters. If you run a startup on top of these tools, a combined product-family milestone tells you demand is surging, but it doesn't tell you how much capacity exists for your exact workflow at 3 p.m. on a weekday when everyone else is pushing the same model.

The process is the problem.

Even if OpenAI's explanation is true, the uncomfortable fact is that production users experienced changing behavior before they got a clear explanation. Enterprises don't usually expect infrastructure vendors to experiment with core performance settings in ways that change cost, context handling or perceived capability without advance notice. You wouldn't tolerate a database vendor quietly reducing your memory allocation to understand demand. You shouldn't treat AI infrastructure differently just because the control surface is fuzzier.

The invisible model setting

OpenAI's own model page lists GPT-5.6 Sol as the flagship tier, with Terra as the balanced option and Luna as the low-cost model. The pricing is clear: Sol costs $5 per million input tokens and $30 per million output tokens, Terra costs $2.50 and $15, and Luna costs $1 and $6. The same OpenAI documentation lists a 1.05 million-token context window, a 128,000-token max output and a February 16, 2026 knowledge cutoff.

Those numbers are useful. They aren't enough.

Benchmarks are snapshots of a model under a stated setup. OpenAI says Sol Ultra reached 91.9% on Terminal-Bench 2.1, while standard Sol scored lower on the same benchmark. Third-party benchmark trackers such as evals.report label that 91.9% result as an OpenAI-reported Sol Ultra score dated July 9. Fine. But if context limits, reasoning effort, multi-agent behavior and product metering can change around the model, the benchmark isn't the whole thing you're buying.

A startup choosing between Sol, Terra and Luna has to model more than token price. Sol may be the right call for long coding agents, multi-step debugging and work where getting the answer wrong costs more than the model run. Luna may be fine for summaries, classifications and disposable first drafts. Terra will probably carry a lot of ordinary production work because it sits in the middle. That is the easy decision tree.

The hard part is stability.

If you price your own product around Sol's launch behavior, then OpenAI changes the settings under load, your margin can move without your product changing at all. If you sell customers on answer quality and the upstream model starts taking shorter reasoning paths, your support team owns the complaint. Your users won't care that the root cause was a context-limit rollback or a multi-agent billing side effect.

What you should do now

You need your own evals. Not a slide with vendor benchmarks, but a small set of real tasks from your product that run every day against the models you rely on. Measure answer quality, latency, token use, failure rate and cost per completed task. Then keep the history. When the line moves, you'll know whether the model improved, degraded or merely got cheaper in a way that doesn't help your users.

Contracts need the same clarity. If you're selling an AI feature to customers, define what you control and what depends on OpenAI, Anthropic, Google or whichever provider sits underneath. Don't bury that in vague platform language. Name the dependency. Name the risk. If a provider changes model behavior, your customer shouldn't discover the boundary only after something breaks.

The Sol episode doesn't prove OpenAI acted maliciously. It proves something more practical: model capability is now an operational dependency, not a fixed product description. That's the real issue for startups. You can build on GPT-5.6 Sol, and in many cases you probably should. Just don't build as if the July 9 version is frozen in place.

Also read: Y Combinator fills an NBA arena with AI founders as Sam Altman and Jensen Huang tell 6,000 applicants the window is nowOpenAI brings real-time interruptible voice AI to enterprise workspaces and launches Presence for customer-facing agentsGlossGenius rebrands as Genius AI and closes a $44M Series D at a $1.15 billion valuation

TOPICS
Julian Lim is an entrepreneur, technology writer, and a researcher. He started JL Data Analysis after graduating from NUS in Intelligent Systems. Julian writes about technology innovations and entrepreneurship on Business Times, Asia Pacific Magazine and occasionally contributes to Startup Fortune.
Related Articles
More posts →
Loading next article…
You're all caught up