For years, the AI industry operated on a simple formula. Double the compute, double the data, double the parameters — and performance would scale predictably. The Kaplan scaling laws, formalized in 2020, promised that model performance improved smoothly with increased training compute. The Chinchilla scaling laws, refined in 2022, added that model size and training data must scale in equal proportions.

This paradigm drove the industry from GPT-3 to GPT-4. It worked for several years. By 2025, however, the cracks were becoming impossible to ignore. Training frontier models cost close to $1 billion. High-quality internet text was becoming scarce. OpenAI co-founder Ilya Sutskever publicly stated that the era of pure pre-training scaling was entering a plateau, and that intelligence growth would need to shift toward a new “research era.”

The scaling laws had not died. They had split into three.

The Pre-Training Plateau

The problem is not that scaling laws are incorrect. The problem is that they assumed an unbounded supply of fresh, high-quality pre-training data. That assumption no longer holds.

A July 2026 paper from researchers at the Beijing Academy of Artificial Intelligence (BAAI) formalized what many had suspected: pre-training is entering a regime where compute grows faster than the availability of high-quality data. The paper proposed Compute-Data (CD) scaling laws, introducing a “token-effectiveness function” that quantifies the diminishing value of derived tokens — generated through repetition or paraphrasing — relative to fresh tokens. The functional form of this function implies diminishing returns when substituting compute for data.

In plain English: you can keep throwing more compute at pre-training, but each additional unit of compute buys you less improvement than it used to.

This is not a niche academic finding. It is the industry consensus. As one 2026 analysis put it, “pre-training has consumed the internet‘s high-quality text.” The simple formula that worked from GPT-3 to GPT-4 — 10x the parameters, 10x the data, 10x the compute = 10x better AI — no longer delivers the same returns.

The industry has responded not by abandoning scaling, but by expanding it. As Jensen Huang put it in early 2026, there are now three scaling curves operating simultaneously.

The Three Curves

The traditional pre-training scaling curve is the first: more data, more parameters, more compute to train larger models. This curve is decelerating.

The second curve is post-training scaling. Through techniques like RLHF, DPO, and reinforcement learning, models are optimized after the initial training run. A 2026 ACL study on the Qwen2.5 series (0.5B to 72B parameters) found that reinforcement learning post-training follows its own predictive power-law — and that performance tends toward saturation as model scale increases.

The third curve is test-time scaling — also called inference scaling. Instead of training a larger model, you give the model more time to “think” during inference, allowing it to reason through problems via extended chains of thought.

The test-time scaling curve is the one that has transformed the industry in 2025-2026. Reasoning models like OpenAI‘s o1 and DeepSeek-R1 proved that increasing inference-time compute can unlock intelligence that pre-training alone cannot achieve.

The T² Revolution

The most consequential development in 2026 has been the integration of these curves into a unified framework.

In April 2026, researchers at the University of Wisconsin-Madison and Stanford University introduced Train-to-Test (T²) scaling laws — a framework that jointly optimizes model size, training tokens, and inference samples under fixed end-to-end budgets.

The findings were striking: when you account for inference costs, the optimal pre-training strategy shifts dramatically toward smaller models trained for much longer — far beyond the Chinchilla recommendation of roughly 20 tokens per parameter. T² scaling suggests ratios that are orders of magnitude higher.

In practice, this means that heavily overtrained smaller models can outperform larger, less-trained models when inference costs are factored in. As one industry observer put it, “when you account for inference compute, radical overtraining becomes compute optimal.”

This is not a theoretical curiosity. It is already reshaping how AI labs allocate their budgets. According to Deloitte‘s 2026 TMT Predictions report, the post-training phase overall now uses roughly 30 times the compute of the original base model training. Long-chain reasoning consumes more than 100 times the compute of a simple inference request.

The China Playbook

The shift from pre-training to test-time scaling has been particularly consequential for Chinese AI labs.

DeepSeek-R1, released in January 2025, was among the first to demonstrate test-time compute scaling at scale. The model exposed full reasoning chains in `<think>` tags and matched o1-level performance while remaining fully open-source. DeepSeek V3’s training cost was just $5.57 million — a fraction of GPT-4o‘s estimated $100 million. A 2026 Wedbush analysis noted that the AI landscape had shifted from a “brute-force” scaling race to an efficiency-driven “reasoning” race.

Chinese labs have also pioneered more efficient architectures. DeepSeek’s Multi-head Latent Attention (MLA) dramatically reduces inference costs — a structural advantage that allows them to deploy test-time scaling at scale without bankrupting themselves.

The result is a fundamental divergence in approach. U.S. labs continue to pour billions into pre-training larger models. Chinese labs are focusing on post-training and test-time scaling, achieving competitive performance at a fraction of the cost.

What This Means

The scaling laws have not died. They have fractured. Three distinct scaling curves now operate in parallel, and the optimal strategy depends on your constraints — compute budget, inference volume, and use case.

For most enterprises, the implications are clear: pre-training scale alone is no longer the differentiator. The real gains are in post-training and test-time scaling. The companies that optimize across all three curves will be the ones that capture the next wave of AI value.

For a deeper look at how this shift has unfolded — and why it changes everything about the AI race — see our earlier piece on the transformation of scaling laws.

Sources: BAAI Compute-Data scaling laws paper, arXiv:2607.25271 (July 2026); Ultralytics Scaling Laws Glossary (2026); (January 2026); ACL 2026 RL post-training scaling study; Train-to-Test (T²) scaling laws, arXiv:2604.01411 (April 2026); Wedbush “The Great Reasoning Shift” (January 2026); DeepSeek-V3 technical report; DeepSeek AGI Roadmap Issue #1365; Deloitte 2026 TMT Predictions report.

Disclaimer

The information provided in this article is for general informational and educational purposes only. It does not constitute legal, financial, or professional advice. The author and publisher are not responsible for any actions taken based on the content of this article. Readers should consult qualified professionals for advice specific to their situation. All trademarks and references to third-party products, services, or organizations are the property of their respective owners. The performance data and benchmarks discussed are based on specific research studies and may not generalize to all use cases or environments. As of the publication date, the AI landscape continues to evolve rapidly, and readers should verify current information independently.

Limitations

This analysis is based on reporting and public data available as of the article date; figures may be revised as sources update.

Forecasts from third-party analysts can change with market conditions.

Cost and pricing examples are point-in-time estimates; actual rates vary.

Country and company comparisons rely on public reporting, not operational data.

This sector moves fast; timelines and deal terms may be updated later.

Company deals and regulatory rulings may evolve; verify current status.

AI infrastructure is changing quickly; claims can become outdated soon.


Sources

  1. BAAI Compute-Data scaling laws paper, arXiv:2607.25271 (July 2026)
  2. Ultralytics Scaling Laws Glossary (2026)
  3. (January 2026)
  4. ACL 2026 RL post-training scaling study
  5. Train-to-Test (T²) scaling laws, arXiv:2604.01411 (April 2026)
  6. Wedbush “The Great Reasoning Shift” (January 2026)
  7. DeepSeek-V3 technical report
  8. DeepSeek AGI Roadmap Issue #1365
  9. Deloitte 2026 TMT Predictions report.

The information provided in this article is for general informational and educational purposes only. It does not constitute legal, financial, or professional advice. The author and publisher are not responsible for any actions taken based on the content of this article. Readers should consult qualified professionals for advice specific to their situation. All trademarks and references to third-party products, services, or organizations are the property of their respective owners. The performance data and benchmarks discussed are based on specific research studies and may not generalize to all use cases or environments. As of the publication date, the AI landscape continues to evolve rapidly, and readers should verify current information independently.

Limitations: This analysis is based on reporting and public data available as of the article date; figures may be revised as sources update.; Forecasts from third-party analysts can change with market conditions.; Cost and pricing examples are point-in-time estimates; actual rates vary.; Country and company comparisons rely on public reporting, not operational data.; This sector moves fast; timelines and deal terms may be updated later.; Company deals and regulatory rulings may evolve; verify current status.; AI infrastructure is changing quickly; claims can become outdated soon.