AI RACE— The AI Race
Research

The Cost of Matching Frontier AI Performance Is Collapsing by up to 13x a Year

Research from Epoch AI and MIT reveals that the price of achieving fixed benchmark scores is dropping at an unprecedented rate, driven by a combination of algorithmic progress, cheaper hardware, and intense market competition.

09/24/2026, 23:00
Chi phí hiệu năng AI đang lao dốc với tốc độ chưa từng có trong lịch sử công nghệ

AI Benchmark Costs Plunge at Historic Rates

The cost of achieving fixed performance milestones on key artificial intelligence benchmarks is falling faster than that of any previous transformative technology. According to findings from Epoch AI, a research organization monitoring AI trends, the price of reaching a set capability level on selected benchmarks has dropped by an average of 47 percent per quarter since 2023—amounting to roughly a 13-fold reduction per year.

A parallel study by researchers at MIT, led by Hans Gundlach, tracked comparable trends using data from AI evaluation platform Artificial Analysis between April 2024 and November 2025. The MIT team observed costs declining between 5x and 10x annually across a broader set of models. While these numbers reflect plummeting market prices to replicate past capabilities, researchers note that the cost of running peak state-of-the-art models on complex tasks can actually increase due to the heavy computational demands of newer reasoning architectures.

Breaking Down Algorithmic Progress, Hardware, and Pricing Wars

The decline in operational costs to match former frontier models is steep. Epoch highlighted OpenAI’s o3 model, which reached a 75 percent accuracy score on the PhD-level GPQA Diamond science benchmark in early 2025 at an estimated cost of 30 cents per question. Within 18 months, a model from the GPT-5.6 family achieved the identical score for just 0.04 cents per question—a 1/725 price reduction. To illustrate the scale, Epoch noted that if car manufacturing improved at the same rate, a €50,000 vehicle would drop to under €70. The release of OpenAI’s lower-cost GPT-6 Sol and Luna models indicates that the baseline cost continues to drop.

However, market prices do not reflect algorithmic improvements alone. When MIT researchers isolated market dynamics by examining open models and removing the impacts of cheaper hardware and aggressive pricing competition, they calculated the underlying rate of pure algorithmic efficiency gains at roughly 3x per year. Epoch's higher 13x figure captures the full market environment without disentangling hardware and competitive discounts.

Additionally, higher benchmark marks do not always equate to leaner systems. MIT found that modern reasoning models often boost accuracy by throwing significant amounts of extra test-time compute at problems, increasing the processing power and cost per query even as base token pricing declines.

Test-Time Compute, 'Benchmaxxing,' and Real-World Trade-Offs

The divergent metrics highlight the growing disconnect between raw benchmark scores and practical deployment costs. Because composite benchmark results merge improvements in data, architecture, training efficiency, and test-time compute, top scores can conceal higher execution expenses. Researchers also contend with "benchmaxxing," where developers optimize systems specifically around public tests. Epoch attempted to account for this by incorporating a private test, "Mystery Game Puzzles"; costs on that benchmark declined at the slowest rate in their sample, pointing either to less optimization on unreleased tests or differences in task formats.

For enterprise adopters, general price drops do not automatically dictate model choice. Operational requirements often outweigh cost-per-token metrics: highly discounted models with significant latency are unsuitable for real-time customer chatbots, while slow reasoning models can bottleneck automated pipelines. Conversely, deploying a more expensive frontier model can yield lower total operating costs if superior first-pass accuracy eliminates the need for repeated queries and error-correction loops.

◗ Sources

The Decoder09/24

Related stories