← All entries

The Receding Frontier

Every week someone sends me the same chart. Inference prices going down and to the right, the line bending toward zero, the caption some version of: intelligence is becoming free. The chart is true. I've checked the underlying numbers and they hold. What I've come to believe is that it answers a question almost nobody asked, and that the answer points the opposite way from how people read it.

Here's the number the optimists point at, and it's real. Epoch AI tracked the price of running a model at a fixed capability level over time, controlling for what a model can actually do rather than what it's called. To clear GPT-4-level performance on GPQA Diamond, a set of PhD-level science questions, the price fell about 40x per year. Across their basket of benchmarks the median was 50x per year, with a tail running from 9x to 900x. After January 2024 the median accelerated to 200x.

You can feel this in the catalog. In March 2023 the original GPT-4 cost $30 per million input tokens and $60 per million output tokens. Sixteen months later, in July 2024, GPT-4o mini launched at $0.15 and $0.60, scored 82% on MMLU, and beat the original GPT-4 on most things people bother to measure. On input tokens that's a 200x cut for a model that's better. If your workload is something a 2023 frontier model could already do, you're living in a deflation most industries never get to see.

Now the part the chart leaves out. Ask a different question. Forget what last year's capability costs today. What does it cost to get this year's hardest thing done, correctly, the first time? That curve looks nothing like the first one.


Two prices, moving apart

In December 2024 OpenAI's o3 was scored on ARC-AGI-1, a benchmark of visual reasoning puzzles built specifically to resist memorization. On the semi-private set, a low-compute configuration scored 75.7% at about $26 per task. A high-compute configuration of the same model scored 87.5%. That last 11.8 points of accuracy cost roughly $4,560 per task, about 172x more compute. Same model, same week. The gap between good and frontier was a factor of 172 in price.

So there are two prices for "intelligence" and they're moving in opposite directions. A capability we already had falls 40x a year. The capability we don't yet have, the last increment on a problem that resists us, runs into the thousands of dollars for a single answer. The first curve is the one on the chart. The second is where the work is.

The frontier moved to the meter

This wasn't always true in the same way. For years you bought capability once, at training time, and then served it cheaply. The model was a fixed asset and inference was the marginal cost of reading from it. Reasoning models broke that arrangement. They buy accuracy at the moment you ask, by thinking longer, which means spending tokens. o3's high-compute run wasn't a bigger model. It was the same model permitted to burn more tokens per question. Capability stopped being a thing you amortize and became a thing you meter.

Once you see it that way, the falling price stops looking like progress toward free intelligence and starts looking like something stranger. A capability gets cheap precisely when it stops being the frontier. The cheapness is the evidence. When o3's 87.5% drops from $4,560 a task to $5 a task, and it will, that won't mean the frontier got cheap. It'll mean the frontier left. By the time you can afford this year's hardest answer, it'll be a commodity, and the thing worth paying for will be some new task that's expensive again. Cheap tokens are a receipt for capability you no longer need.

Which side of the 172x

This reframes a lot of strategy that takes the chart as the whole picture. "Wait for it to get cheaper" is a fine plan for any capability you can name today and a guaranteed loser at the frontier, because the frontier is defined as the part that hasn't gotten cheap yet. The margin doesn't sit with whoever resells commodity tokens at the floor. That floor is a competitive bloodbath dropping 40x a year, and the price of last year's intelligence is heading toward the cost of the electricity to serve it. The margin sits with whoever owns a task hard enough that the high-compute answer is still worth $4,560. The whole game is which side of that 172x you're standing on.

What I'm actually claiming

Here's the belief, stated so it can be wrong. Through the end of 2027, the price of frontier-quality completion on the hardest contemporary reasoning benchmark won't fall at anything close to the 40x-per-year rate the commodity curve falls at, and it'll stay at least two orders of magnitude above the cost of a median commodity task. The technology won't stall. The redefinition is the mechanism: every time a hard task gets cheap, "the hardest contemporary benchmark" relocates to a new task that isn't. The deflation is real, and it'll keep chasing a target that moves by exactly enough to stay out of reach.

What would change my mind: a frontier reasoning capability, top-decile on a genuinely novel benchmark at release, falling into commodity-token range, say under a few dollars per task, within twelve months of release, with no new harder benchmark having displaced it as the thing the labs race on. If that happens, the curves are converging and the meter was a phase. I don't think it will. I think the meter is the new shape of the business.

The chart isn't lying. It's a receipt. It tells you, with great precision, the falling price of everything we've already figured out.