source avatarFelipe Demartini

Share

Epoch AI has just measured the cost of AI “thinking,” and the decline is the fastest ever seen for any transformative technology. Epoch is an independent research institute that tracks AI’s progress using its own data: model performance, training costs, chips, and data centers. In their September 22 report, Luke Emberson and David Roodman analyzed the cost of achieving the same level of performance across five benchmarks (mathematics, hard sciences, and skill-based games) over approximately three years. The central figure: since 2023, the cost of a given performance level has fallen by about 47% per quarter—equating to 13 times cheaper per year. Historically, this is four times faster than DNA sequencing, six times faster than computing, 18 times faster than lithium batteries, and 54 times faster than electricity from the start of the century until 1973. Evidence suggests this pace has been in effect since at least November 2021, when OpenAI fully opened the GPT-3 API. On January 31, 2025, OpenAI’s o3 scored 75% on the GPQA Diamond (a doctoral-level multiple-choice exam in physics, chemistry, and biology) at roughly 30 cents per question. Less than 18 months later, GPT-5.6 Luna achieved the same score for just four one-hundredths of a cent ($0.0004)—a 725-fold reduction. Epoch compares this to the price of a new car dropping from $50,000 to $69. The pace varies by domain: slower in game puzzles (39%–43% per quarter), faster in mathematics (50%–52%). Prices drop most sharply immediately after a model becomes the new state-of-the-art. Across the five main benchmarks, performance costs fell by an average of 66% per quarter (75 times cheaper per year) at the debut of each new state-of-the-art model; two years later, that rate halved to 32% per quarter (4.7 times cheaper per year). Their hypothesis: early on, labs charge a premium; then open and closed competition drives prices down. They acknowledge limitations: labs may be overtraining for benchmarks; acing tests isn’t necessarily useful work; real users don’t always choose the cheapest frontier model; the window is only three years; and averages can be calculated in multiple ways. For the AI stack, the takeaway is clear: capabilities that were once premium quickly become commodities—and what unlocks mass adoption is cost per task, not just the existence of a new model. For API providers, above-normal profits on any specific model appear fleeting. Epoch itself notes this dynamic is both cause and consequence of the race. In their chart, the fiercest competition near the top occurs primarily among closed models. For investors, inference margins are compressing. Volume and staying one quarter ahead on cost matter more than narratives of an “eternal” model. The premium for being the cheapest at the top vanishes rapidly. Chip and energy costs have risen. The cost of “thinking” has plummeted.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.