On August 6, 2026, OpenAI made a change to GPT-5.6 Sol in ChatGPT. The company said the update would produce "more direct responses, tighter formatting, and less unnecessary detail." It also introduced a slider allowing users to choose how much thought ChatGPT puts into a response.

Users noticed something else: the model felt dumber.

A Japanese market research team using Codex Sol MAX had been getting 12 out of 10 performance consistently. On the morning of the change, that dropped to 8 out of 10. The team's report noted that the model that once spent over ten minutes iterating, testing approaches, and refining logic had "completely vanished." Across Codex user communities, complaints were remarkably consistent: the model was faster, but it refused to dig deep anymore.

OpenAI denied any "nerfing," claiming it was "just running an experiment." But the experiment turned a hidden dial that users could not see. Community members discovered an internal OpenAI parameter never made public: the "juice value" — the model's thinking budget. Prior to the change, Sol's max setting corresponded to a juice value of 960. After the change, that same setting returned 128. A separate change cut usable context from 372k back to 272k.

OpenAI is not the only company facing these accusations. Anthropic users across X, GitHub, and Reddit have been swapping anecdotes about Claude feeling "nerfed." AMD AI Senior Director Stella Laurenzo wrote that "Claude has regressed to the point it cannot be trusted to perform complex engineering." According to community analysis, Claude's thinking median dropped from 2,200 characters to 600. On BridgeMind's BridgeBench test, Opus 4.6 accuracy fell from 83.3 percent to 68.3 percent, dropping from second to tenth place.

On April 23, Anthropic acknowledged that a system-prompt instruction added on April 16 to reduce verbosity, combined with other prompt changes, had harmed coding quality across multiple Claude models. The company later reverted the change. The incident showed that a model could appear measurably worse without any secret reduction in its underlying weights. A product-layer change intended to make responses more efficient was enough to degrade real-world performance.

The Cost Pressure

The economic pressure shaping modern AI systems is significant. Every additional unit of reasoning, context, and output costs money.

The numbers illustrate the pressure. OpenAI's GPT-5.6 Sol API pricing stands at $5 per million input tokens and $30 per million output tokens. The lower-cost Luna tier runs $0.20 input and $1.20 output. That spread — 25x on input, 25x on output — demonstrates that intelligence, latency, and cost are already treated as adjustable product variables.

For Anthropic, the pressure is similar. Claude Code subscription fees are $400 per month; API costs for the same usage can run $42,000 per month. One analysis found that reasoning models burn large volumes of internal "thinking" tokens that get billed as output, sometimes 100x what the final answer contains.

When a company serves billions of queries per month, even a 10 percent reduction in per-query compute translates to millions of dollars in savings. Some vendors have introduced techniques such as dynamic reasoning budgets, cache optimization, and routing downgrades, prioritizing overall system throughput and response speed. These hidden adjustments — reducing reasoning depth per query to control operating costs — are not announced to users.

This is not a conspiracy. It is arithmetic. When a model's inference cost is a recurring expense that grows with every user query, and when those costs are measured in billions of dollars annually, the pressure to optimize is structural. Once the initial hype fades, or when massive inference costs begin to hurt financial results, companies quietly adjust parameters behind the black box.

Meta's Avocado: A Different Kind of Reasoning Problem

While OpenAI and Anthropic face user backlash over quality cuts, Meta is dealing with the opposite problem: its reasoning model isn't good enough to ship.

In March 2026, Meta was forced to delay the release of its next-generation AI model, codenamed Avocado. The model, which the company had been developing for months, underperformed against rivals from Google, OpenAI, and Anthropic in internal tests for reasoning, coding, and writing.

Meta's internal testing showed that Avocado outperformed the company's previous AI models and did better than Google's Gemini 2.5 from March 2025, but still lagged behind Gemini 3.0 from November 2025. The delay was a significant embarrassment for a company that had been spending aggressively on AI — $72 billion in 2025 and up to $135 billion projected for 2026.

Meta's AI division even discussed the possibility of temporarily licensing Google's Gemini to support its AI products — an extraordinary admission that its own model wasn't ready. The Avocado delay exposed a structural gap at Meta: while the company had bet billions on AI infrastructure, its models were falling behind in reasoning capability, the one capability that matters most for next-generation AI.

As one analysis put it, "the industry's leading players are no longer distinguished by whether they have a model, but by reasoning capability, engineering efficiency, inference cost, and iteration speed." Meta, for all its spending, was failing on the first metric.

The China Difference

The cost-cutting dynamic affecting U.S. AI labs is not universal. Chinese AI companies face the same economic pressures but have taken a different path — one that does not require reducing user-facing quality.

DeepSeek's V4-Pro model, after a permanent price cut, charges $0.435 per million input tokens and $0.87 per million output tokens. Cache hits drop the effective input price to as low as $0.0037 per million tokens. Through OpenRouter, with an 86 percent cache hit rate, the weighted input price users actually pay is $0.064 per million tokens — less than one-fiftieth of GPT-5.6 Sol's input price.

Zhipu's GLM-5.3-Flash, released in August 2026, is priced at one-tenth of its predecessor GLM-5.3 and one-fortieth of Anthropic's Claude Opus 4.8. The model is built on a new architecture designed specifically for "extremely low cost," achieving three times the end-to-end service performance on the same hardware.

The difference in approach is structural. Chinese AI companies are competing on cost efficiency from the architecture up. They are not reducing reasoning depth on existing models; they are building new models that are cheaper to run. The price cuts are permanent, not experimental. The quality does not degrade because the cost savings come from better engineering, not from dialing down the thinking budget after users have already paid.

The AI landscape has shifted from a "brute-force" scaling race to an efficiency-driven race. Chinese labs have focused on that efficiency from the start — not because they are more virtuous, but because the constraints they face (export controls, chip shortages, domestic competition) forced them to. U.S. labs, by contrast, built their business models on scaling first and optimizing later. Now that the scaling has reached its limits, the optimization is happening in ways users can feel.

The User Impact

The consequences of these changes extend beyond benchmark scores or response times. AI systems are now embedded in professional work, research, coding, editing, design, and communication. When a tool that once handled complex instructions reliably begins losing context, disregarding constraints, or producing inconsistent results, users absorb the cost through repeated corrections, failed generations, duplicated work, and additional verification. What appears on a corporate dashboard as an efficiency gain can become hours of uncompensated labor for the people using the AI product.

The emotional burden can also be greater than with conventional software. A software program that crashes presents an obvious technical failure. A conversational AI system can instead acknowledge an instruction, appear to comply, and deliver something that is subtly wrong — requiring the user to detect the error, diagnose the failure, and compensate for it.

The message from OpenAI, Anthropic, and Meta is the same: reasoning is expensive. The question is whether the cost is being borne by the companies that build the models or by the users who rely on them.

Sources: Milwaukee Independent (September 6, 2026); 36Kr (June 30, 2026); 36Kr (July 15, 2026); The New York Times (March 12, 2026); Reuters (March 13, 2026); Bloomberg (March 13, 2026); ITHome (March 13, 2026); NetEase (March 16, 2026); OpenAI API pricing documentation; DeepSeek official pricing.

Disclaimer

The information provided in this article is for general informational and educational purposes only. It does not constitute legal, financial, or professional advice. The author and publisher are not responsible for any actions taken based on the content of this article. Readers should consult qualified professionals for advice specific to their situation. All trademarks and references to third-party products, services, or organizations are the property of their respective owners. The performance data and benchmarks discussed are based on specific research studies and may not generalize to all use cases or environments. As of the publication date, the AI landscape continues to evolve rapidly, and readers should verify current information independently.

Limitations

This analysis is based on reporting and public data available as of the article date; figures may be revised as sources update.

Forecasts from third-party analysts can change with market conditions.

Cost and pricing examples are point-in-time estimates; actual rates vary.

Country and company comparisons rely on public reporting, not operational data.

This sector moves fast; timelines and deal terms may be updated later.

Company deals and regulatory rulings may evolve; verify current status.

AI infrastructure is changing quickly; claims can become outdated soon.


Sources

  1. Milwaukee Independent (September 6, 2026)
  2. 36Kr (June 30, 2026)
  3. 36Kr (July 15, 2026)
  4. The New York Times (March 12, 2026)
  5. Reuters (March 13, 2026)
  6. Bloomberg (March 13, 2026)
  7. ITHome (March 13, 2026)
  8. NetEase (March 16, 2026)
  9. OpenAI API pricing documentation
  10. DeepSeek official pricing.

The information provided in this article is for general informational and educational purposes only. It does not constitute legal, financial, or professional advice. The author and publisher are not responsible for any actions taken based on the content of this article. Readers should consult qualified professionals for advice specific to their situation. All trademarks and references to third-party products, services, or organizations are the property of their respective owners. The performance data and benchmarks discussed are based on specific research studies and may not generalize to all use cases or environments. As of the publication date, the AI landscape continues to evolve rapidly, and readers should verify current information independently.

Limitations: This analysis is based on reporting and public data available as of the article date; figures may be revised as sources update.; Forecasts from third-party analysts can change with market conditions.; Cost and pricing examples are point-in-time estimates; actual rates vary.; Country and company comparisons rely on public reporting, not operational data.; This sector moves fast; timelines and deal terms may be updated later.; Company deals and regulatory rulings may evolve; verify current status.; AI infrastructure is changing quickly; claims can become outdated soon.