The cheapest token does not win — the cheapest useful state change does

AI Tokenomics: Why the LLM Price War Is Not a True Cost Competition

Contents

When major AI providers cut their prices, it is often read as clear evidence of technical progress: the model has become more efficient, so inference must be cheaper. That may be true. But it is only one possible explanation.

Today’s competition between large LLM providers blends at least three distinct things: technological efficiency, strategic pricing, and economic value for the user. A low customer price therefore does not prove a low production cost. Nor does a low price per token say anything conclusive about the cost of completing productive work.

Subscription prices are not a window into cost structure

A recent SemiAnalysis analysis compares subscriptions through their API-equivalent usage value. From that perspective, Anthropic subscriptions provide roughly five times the value of comparable OpenAI plans. The analysis also estimates that subscriptions may account for about 10% of revenue while consuming more than 40% of inference compute.

That is a meaningful observation, but it is not proof that a provider subsidises its actual production costs in the same proportion. API list prices are selling prices, not disclosed marginal costs. Utilisation, hardware contracts, model routing, caching, quotas, and the distribution of light versus very heavy users all matter too.

The robust conclusion is therefore narrower: there can be a large gap between revenue and resource consumption. Without internal cost accounting, it remains unclear whether that gap comes from a technical cost advantage, lower margins, cross-subsidisation, or a mixture of all three.

This changes the competitive question. It is not only, Who can run AI most efficiently? It is also, Who can finance especially generous usage for the longest?

Haiku 5.5: A sharp price move, not a cost disclosure

Anthropic introduced Claude Haiku 5.5 on 7 October 2026. According to the announcement, list prices for requests up to 100,000 tokens are 90% lower than Haiku 4.5, while Anthropic puts the average reduction in cost per task at around 75%.

This is a substantial price move. But its cause cannot be inferred unambiguously from the outside. It may reflect a more efficient model architecture, better infrastructure and utilisation, different routing, lower margins, or deliberately aggressive market pricing. Several factors may be operating at once.

For that reason, marketing and list prices should not be confused with underlying inference economics. They show what a provider charges, not what an additional request costs it to serve.

Mistral and the economics of “good enough”

The alternative is not necessarily a stronger model, but a better-matched one. Mixture-of-Experts architectures can combine large model capacity with far fewer active parameters per request. Their economic value then comes less from winning a benchmark than from covering a large share of productive workloads with fewer resources.

For Saganode, this was the practical question: in graph extraction, Mistral was approximately 80–85% cheaper than the GPT alternative used previously. This is not a general ranking of models. It describes a specific workflow with a clearly defined quality threshold.

That is the point: economically optimal is not maximum model intelligence, but sufficient intelligence at the lowest total cost — including review and rework.

The token is the wrong economic unit

A model can be ten times cheaper per token and still cost more if it fails more often, requires more iterations, or creates extra review work. Conversely, a smaller model can be economically superior when it completes the task reliably enough.

A more useful metric is therefore:

C_useful = (C_inference + C_tools + C_retries + C_validation) / N_valid transitions

It measures the cost per successful, validated state transition. The numerator includes not only inference costs, but also tool calls, retries, and validation. The denominator is not answers or tokens, but results that correctly change a desired state.

This perspective also reveals the limits of coding and chat benchmarks. Measured model performance is not automatically productive value.

Architecture can avoid inference

Saganode illustrates another lever: an LLM extracts and interprets information once, then converts it into persistent, structured state. Later requests do not need to process the entire original context again; they can build on stored knowledge and handle only the changes.

Token-centred modelState-centred model
Load context again and againExtract information once
Have an LLM process the same subject againReuse persistent state
Generate an answer and lose the contextInterpret and validate only deltas

State is not free, of course: persistence, consistency checks, versioning, and graph maintenance also consume resources. In repeated workflows, however, the avoided redundant inference can substantially outweigh those additional costs.

The operational rule is simple: Buy intelligence once. Persist the result. Reprocess only the delta. For the technical side of this approach, see associative memory and graph-based structures for language models .

Three competitions, not one

LevelDecisive question
Capital competitionWho can finance low prices and high usage for the longest?
Technology competitionWho can produce useful inference most efficiently?
Architecture competitionWho needs the least inference for the same result?

The third level is often underestimated in public discussion. A provider can operate the most efficient model and still lose to a system that eliminates 90% of repeated model calls. And a provider can have the lowest token price without revealing a structural cost advantage.

Conclusion

The AI market currently often confuses subsidised consumption, technological capability, and economic productivity. As long as token consumption stands in for usage and benchmarks stand in for value creation, the efficiency of the overall system remains partly invisible.

In the long run, the winner will not automatically be the provider that sells the most tokens as cheaply as possible. A more economically robust goal is to produce the greatest durably useful state gain with the least resource input.

Sources: