There is a claim that goes around: token prices are subsidised, and when the money runs out everyone will have to run their own hardware. It is a reasonable thing to worry about, and mostly wrong, but wrong in an interesting way. The subsidy is real. It does not sit where people think it sits, and the corrections it predicts have already started arriving, in forms that never touch a pricing page.

Four numbers, all called margin

When someone tells you inference is profitable, or that it is a bonfire of venture capital, ask which number they mean. There are at least four, they differ by more than an order of magnitude, and people swap between them mid-sentence without noticing.

Here is the same company, OpenAI, in the same year, at four different heights:

Four numbers, all called margin
Four measures of the same company's margin, in the same year. Rungs 01 and 02 count compute only; rung 04 carries research, salaries, sales and training amortisation.

The first is marginal token cost, the compute to serve one more token to a paying customer, and it sits well below price. The second is compute margin, paid revenue minus the servers running live traffic: The Information put it at roughly 70% in October 2025, up from 52% at the end of 2024 and about 35% in January 2024. The third is blended gross margin, the same line once the free tier and flat-rate plans are carried: 33% for 2025, down from 40% the year before and against an internal forecast of 46%. The fourth is all-in, after research, salaries, sales and training amortisation: $13.07bn of revenue against roughly $34bn of cost, a $20.9bn operating loss, or about minus 160%, per audited documents obtained by Ed Zitron and independently verified by the Financial Times. Research and development alone was $19.18bn.

Both figures are accurate. Neither refutes the other. A company can make good money on the request you just paid for and still lose twenty billion dollars over twelve months.

Two things sit between them. Between the second and third, roughly 37 points of it (the 70% is an October rate, the 33% a full-year figure, so treat the gap as approximate): only around 50 million of OpenAI's 900 million-plus weekly users paid for anything at the start of 2026, so roughly nineteen in twenty consumed compute for nothing. That changed in February, when ads went live on the Free and Go tiers. Free users now produce ad impressions, a different business but no longer zero. Between the third and fourth: the $19bn of research, which is most of the loss and has nothing to do with who pays.

Anthropic, which sells far more of its capacity to developers and enterprises, has a different shape. Its 2025 gross margin came in near 40%, revised down from about 50% after inference costs on Google and AWS ran roughly 23% over plan, and about 38% once non-paying Claude users are counted: a two-point gap where OpenAI's is 37. SemiAnalysis put its inference line in the mid-60s by May 2026, up from 38% the year before. Those are two different rungs, the 40% a proper gross margin and the mid-60s compute only. The same reporting that produced OpenAI's 70% noted where the two companies split: OpenAI holds better compute margins on paid accounts, while Anthropic runs more efficiently on server spend overall.

So when you read that inference is profitable: probably rung two, and probably true. When you read that it is a money pit: rung four, also true. The argument only becomes interesting once both people name their rung.

Prices fall. Bills rise. Those are not the same curve.

Price per token for a fixed capability collapses. Epoch AI measured declines between 9x and 900x per year depending on which capability you hold constant, landing around 40x annually for GPT-4-level performance on graduate-level reasoning. The academic Price of Progress analysis is more conservative, putting frontier knowledge, reasoning and coding benchmarks at 5x to 10x a year, of which about 3x comes from algorithmic efficiency rather than cheaper hardware.

Cost per finished task also falls, more slowly. In November 2025 Anthropic cut Opus pricing by two thirds, from $15/$75 per million tokens to $5/$25, and shipped a model that at medium effort matched Sonnet 4.5's best SWE-bench score on 76% fewer output tokens. Cheaper tokens, fewer of them per task. And prompt caching bills a cache hit at a tenth of the input rate, so an agent loop that re-reads its context forty times pays full price for the delta, not the whole window.

Spend went up anyway. One five-person team reported spending in six weeks of 2026 what they had spent across all of 2025. Not because the same refactor got dearer, but because a refactor that was not worth attempting at $75 per million output tokens is worth attempting at $25, and the next one after it, and the one that runs overnight. Jevons, in short. Reasoning models bill you for thinking, agents loop, and once a task is cheap enough to hand off you hand off more of them.

Single-shot vs Agentic loop
The same request, two years apart. Price per token down two thirds, tokens per job up 43x, the bill up 10x, and in 2026 the agent did the edits.

The practical version: a lower per-token rate is worth less than you think and a higher cache hit rate is worth more. Measure cost per completed unit of work and total spend separately, or you will congratulate yourself on a discount while your bill doubles, and misdiagnose why.

Where the subsidy sits

A positive compute margin does not settle it. Training is a cost of the token as surely as the GPU-hour is. The model is the goods sold. If OpenAI had to recover $19bn of research through the price of tokens, tokens would cost more. What holds prices at compute-plus-margin is competition, and the competition is paid for by investors. That is a subsidy to you, one step removed. It persists exactly as long as someone is willing to fund the next frontier model.

Three things sit on top of that.

The free tier

Paid API traffic carries a positive compute margin at every provider where numbers have leaked. Free consumer traffic did not, and it is most of the volume. The correction for retail users has arrived, and it was ads, not a smaller free tier or a dearer API. Whether ad revenue per free user covers compute per free user is a number nobody has published.

Depreciation schedules

Hyperscalers write GPUs down over five to six years. In November 2025 Michael Burry argued the real economic life is two to three, and that the industry will understate depreciation by $176bn between 2026 and 2028, with Oracle overstating earnings by 26.9% and Meta by 20.8% by the end of that window. Nvidia's rebuttal is that chips get demoted rather than retired: frontier training, then inference, then batch work, and A100s are still earning years on. Both positions hold together, which is the frustrating part. Amazon already trimmed server life from six years to five in early 2025, so the direction is not in dispute. The number that matters is three. Schedules at three years put roughly $50bn to $60bn a year of extra depreciation onto the industry, and the margins of whoever owns the chips compress with it. Labs renting compute pay for it in the rental rate, to the extent the lessor priced it correctly, and the lessors are the neoclouds whose schedules Burry was complaining about. If CoreWeave's rates assume six-year chips, the shortfall arrives at Anthropic's cost of goods on the next contract renewal, one step removed.

Vendor financing

Nvidia signed a letter of intent for up to $100bn into OpenAI in September 2025. No contract followed. By February 2026 the talks were dead and Nvidia took a roughly $30bn equity stake in the funding round instead. Oracle is in for around $300bn of compute, Microsoft $250bn, Amazon $38bn. Microsoft and Nvidia put roughly $15bn into Anthropic, which then committed something like $30bn back to Azure capacity and Nvidia silicon. The IMF flagged the pattern in July 2026. The fair reading is Gil Luria's: there is a healthy part and an unhealthy part. Nvidia's shipped revenue is real, $81.6bn in the quarter to April 2026 and up 85% year on year. Whether every announced dollar of demand is real is a separate question. The $100bn that became $30bn says announcements are not bookings.

The consequence of all three is the same. The risk is not that your price doubles. It is that a provider stops existing. The two risks meet in one place: consolidation. Competition is what holds prices at compute-plus-margin, and competition is what the investors are paying for. A field of two or three frontier labs is the one path from "a provider stops existing" to "your price doubles", and it is the scenario the rest of this piece does not cover.

And one thing that is the opposite of a subsidy

SemiAnalysis puts Nvidia's gross margin on the silicon at about 75%, roughly four times cost of goods. Call that a tax: every inference provider pays it and passes it on to you. GPUs are something like half of a datacentre's total cost, with power, networking and buildings making up the rest, so the tax is not four times your bill. It is still the single largest compressible layer in anyone's cost base, and it means the compute floor is a price rather than a law of physics.

You can see this in the data. When six providers were benchmarked serving the same model in February 2026, Google turned in the best latency and interactivity of the group at a competitive price, most plausibly because it runs on its own TPUs and skips the markup entirely. Anyone with their own silicon has a structural cost advantage over anyone renting Nvidia's.

The correction already happened. It never showed up on a pricing page.

2025 was the year flat-rate agentic subscriptions got repriced without a single headline number changing.

Anthropic announced weekly rate limits on 28 July 2025, effective a month later, stacked on top of the existing five-hour window, and said fewer than 5% of subscribers would notice. Cursor moved from 500 fast requests to twenty dollars of included usage on 16 June 2025; the CEO apologised in early July and refunded the surprise charges. GitHub added premium request quotas. Every agent product shipped a credit system within about nine months of the same date.

It is rationing, because the headline price held. It is a price increase, because the same twenty dollars now buys less. It is both, and anyone insisting on one of them is arguing about vocabulary.

2026 added a quieter version. Opus 4.7 shipped in April with a new tokenizer that counts up to 35% more tokens for the same text, at an unchanged $5/$25. Per-token price flat, per-request price up by as much as a third. And at the small end, Haiku 4.5 launched at $1/$5 against Haiku 3.5's $0.80/$4, which is a headline increase on an existing line however you dress it.

Then there is the increase that was announced and did not happen. Sonnet 5 launched on 30 June 2026 at $2/$10, billed as introductory until 31 August, with $3/$15 to follow. A 50% rise on a shipping model, dated and on the pricing page. On 10 August Anthropic made the introductory price permanent and the rise never happened. Read that alongside the tokenizer: the loud increase was withdrawn, the quiet one shipped. Whether the withdrawal was discipline or arithmetic is a fair question. GPT-5.6 Terra sits at $2.50/$15 and Gemini 3.6 Flash at $1.50/$7.50, and $3/$15 may simply have stopped being sellable. Either way the mechanism is the same: the sticker is the last thing to move, and competition is what pins it.

The frontier has its own version of this, and I should be consistent about it. Opus has held $5/$25 from 4.5 through Opus 5. Fable 5, the model above it, launched at $10/$50. Anthropic calls that a new tier. If your work needs the frontier, your frontier price doubled this year, and calling it a tier rather than an increase is the same vocabulary move I mocked on rate limits two paragraphs ago. Relabelling is how frontier increases happen, the way limits are how subscription increases happen.

The sequel matters more than the event. Anthropic doubled off-peak limits in March 2026 and made the five-hour doubling permanent in May. Limits tighten and loosen with capacity. They are a supply signal, not a policy, and treating a cap as evidence of structural business failure was a bad read in 2025 and is a worse one now.

Telling rationing from a bad day

On 17 September 2025, after weeks of complaints about degraded output, Anthropic published a postmortem naming three overlapping infrastructure bugs: a context-window routing error affecting up to 16% of Sonnet requests on 31 August, output corruption on TPU serving between 25 August and 2 September, and a compiler miscompilation in the token selection path. The statement was unambiguous: quality is never reduced for demand, time of day, or server load.

A user could not have distinguished that from deliberate throttling. Same symptoms, same weeks, same forum threads full of confident diagnosis, some of it wrong.

If output quality is load-bearing for you, instrument it. Log tokens per completed task, time to first token, and a small fixed eval, per model, with dates. It takes an afternoon and it converts an argument you cannot win into a chart. Vibes are not evidence, and the provider is the only party holding the counterfactual.

The cheap floor is real, and it is not a trap. It is borrowed.

Hosted open-weight models sit roughly between $0.08 and $0.50 per million tokens for commodity and mid-sized models. DeepSeek lists V4-Flash at $0.14 in and $0.28 out. A 120B open model runs around $0.08 on DeepInfra. Larger open coding models climb into the low single dollars.

Is that a land grab that vanishes once the market consolidates? The available evidence says no. SemiAnalysis ran the numbers on Crusoe in February 2026 and found gross margins up to 83% on input tokens and 45% on output, with depreciation counted in cost of goods. An independent estimate the same month, triangulating six providers against the InferenceMAX benchmark, put the range at 30% to 60% at around 50% utilisation. Sacra separately estimates Fireworks near 50% and Together near 45%, with Together's per-token API priced close to breakeven.

How that second number was produced matters, because the method is the only thing holding it up. You cannot measure a provider's concurrency from outside, but you can measure time to first token and output tokens per second per user. Plot those against the InferenceMAX throughput curves, back out tokens per GPU, then apply a GPU cost and a utilisation rate. It covers one model (DeepSeek R1-0528) at one workload shape, 8k in and 1k out, on a single day in February. The author is upfront that margins should fall as sequences get longer, and that 50% utilisation is generous for the smaller providers.

So the companies are equity-funded, generously (Together raised $800m in July 2026, Fireworks $1.5bn the same month), but the tokens are not being sold below the compute that serves them. Plan on the floor persisting.

What the floor rests on is upstream. Together and Fireworks paid nothing to train DeepSeek or Qwen. The cheap tier exists because DeepSeek, Alibaba and Meta keep releasing weights a step behind the frontier, for reasons of their own, and the day that stops the floor does not rise so much as go stale. That is the fragility worth watching, more than any one provider's margin. Two smaller caveats. These are compute-line estimates on chat-shaped workloads, a few thousand tokens in and a thousand out, and long-context agentic traffic runs thinner. And margin-positive is not the same as surviving; Groq raising at half its previous valuation in August 2026 is a reminder that the model can work while the company does not.

The numbers, ranked by how much they will survive scrutiny

Table
FigureValueSourceConfidence
Opus price cut, Nov 2025$15/$75 to $5/$25Vendor pricing pageBand 1, verifiable
Hosted open-weight pricing$0.08 to $0.50Provider price cardsBand 1, verifiable
OpenAI FY2025 revenue and operating loss$13.07bn / -$20.9bnLeaked audited docs, FT-verifiedBand 1, verifiable
Capability-adjusted price decline5x to 900x / yrEpoch AI, Price of ProgressBand 2, range, method-dependent
OpenAI compute margin~70%The Information, Oct 2025Band 2, reported, not audited
Anthropic inference margin, 2026mid-60sSemiAnalysis estimateBand 2, modelled
Hosted provider gross margin30 to 60%Benchmark-derived, Feb 2026Band 2, modelled
Public cloud share of production inference56% to 41%Broadcom survey, 1,800 IT leadersBand 2, vendor-commissioned, seller of private cloud
Self-host breakeven vs frontier API2 to 10M tok/daySitePoint, Lyceum, AlpackedBand 3, assumption-driven
DeepSeek cost-profit ratio545%DeepSeek, self-publishedBand 3, theoretical only
Reliability of commonly cited figures. Each cell names its band: 1 is the strongest and 3 the weakest.

That last row deserves a warning label. DeepSeek published it themselves in March 2025: $562,027 of hypothetical daily revenue against $87,072 of daily GPU rental. It assumes peak load all day, prices every token at the more expensive model's rate, and excludes training, research, salaries and amortisation. DeepSeek said plainly that actual revenue is substantially lower. It gets cited as proof that inference is wildly profitable. It is not evidence of that, and quoting it will cost you credibility with anyone who has read the original.

Should you buy hardware

On price alone, mostly no. This section is for one person or a small team weighing an API bill against a box under the desk. A rack-scale buyer has a different question, and it comes at the end.

Against frontier APIs, published breakeven analyses cluster between two and ten million tokens a day at high utilisation. Against hosted open weights the number becomes absurd, north of 190 million tokens a day or simply unreachable on a single card. If your case for a GPU is the monthly bill, check the arithmetic before the purchase order. A hosted open-weight endpoint beats you on unit cost and will keep doing so.

Buy hardware for the things price does not capture: data that cannot leave the building, latency you control, a bill that does not vary, air-gapped work, vendor risk. Broadcom's Private Cloud Outlook 2026 found public cloud's share of production inference falling from 56% to 41% while private cloud rose to 56%. Respondents cited cost and cost predictability first, then security, governance and control, with sovereignty rising but secondary. Two things to hold against that. Broadcom sells private cloud. And the respondents were enterprises with existing datacentres and heavy data gravity, for whom cost really is the reason. That does not generalise to one developer with a workstation.

If you do run locally, the binding constraint is prefill, and token price does not come close. The numbers below are mine: one card or one box, llama.cpp, cold prompts of 16K to 28K tokens.

Table
Strix Halo, 128 GB unifiedRadeon AI PRO R9700, 32 GB
Prefill, dense 8B at Q8, 20K prompt~235 tok/s~3,200 tok/s
Prefill, MoE with 3B to 4B active, 16K to 32K prompt~800 tok/s~1,500 tok/s
Prefill, sliding-window coding model, 28K prompt~1,400 tok/s~4,100 tok/s
Decode, 26B to 35B MoE48 to 64 tok/s77 tok/s, flat to 131K context
Cold 128K read, best caseabout 90 secondsabout 30 seconds
Prefill and decode on the two boxes I run, one instance per card, no clustering. Prefill is cold and the best each box has produced. A different llama.cpp build on the APU measured 4.5x lower on the same prompt.

Tuned public builds land in the same band. Soothill's August 2026 run, with hipBLASLt and rocWMMA flash attention, holds about 1,350 tok/s on a 30B-A3B MoE at 32K context and 340 to 520 tok/s on a 122B MoE. The Strix Halo wiki puts the same 30B MoE at 660 to 755 tok/s on a short prompt and about 50 tok/s once the context is 130K deep. The published 340 tok/s for a 120B model, against roughly 1,700 on a DGX Spark, is an untuned build. Treat the 128K row as a lower bound. Attention's quadratic term slows the tail of a long prompt, so a real cold read takes longer than the division suggests.

Cold is the operative word. An agent loop keeps its KV cache between turns, so the per-turn cost is the prefill of the new delta, not the whole window. The 90 seconds is what you pay on the first turn, on every cache miss, and every time the context is compacted or the prefix changes, which in a long coding session is more often than you would like. That cost decides whether the machine is usable at all, and it is why the two boxes have different jobs. The APU holds a model too big for any single card and serves one agent at a time. The discrete card prefills three to thirteen times faster on the same prompt, depending on the model. With llama.cpp's one-prefill-at-a-time serving, give each agent its own card and the loops never queue behind each other. A batched server on one card would change that arithmetic, at the cost of memory headroom. Cost per token is not in the top three considerations for either.

At rack scale the constraint moves. Datacentre hardware batches prefill across users, so prefill stops deciding anything and utilisation decides everything. That is the same number the hosted-margin estimate above assumed at 50%. At that tier, whether a provider makes money and whether you should self-host are one calculation.

Where prices go from here

Three forces, pulling in different directions.

Downward: hardware cost per token keeps improving, algorithmic efficiency contributes around 3x a year by itself, competition at the commodity tier is brutal, and Nvidia's 75% margin is a large compressible layer sitting inside everybody's cost base.

Upward: demand for peak capability grows faster than cost deflation, reasoning and agent loops inflate tokens per task, and depreciation reality has to land somewhere eventually.

Sideways: capital discipline. If funding tightens, the correction arrives the way it did in 2025 and 2026, through limits, tiers, tokenizers and routing, rather than a headline increase. Providers have learned that raising the sticker price is the one move that makes the news.

What I expect

  • Commodity and mid-tier per-token prices keep falling. High confidence.
  • Headline prices on a named model stay flat to down, and the frontier price rises by relabelling. Opus stays $5/$25 while the model above it launches at $10/$50, and the next one above that will too. The small tier is where the plain increases sneak in, as Haiku already showed. Medium-high confidence.
  • Your monthly spend rises regardless, because consumption grows faster than price falls. High confidence, and the one that catches people out.
  • Limits, tokenizers and routing, rather than prices, remain the primary adjustment mechanism, tightening and loosening with capacity. Medium-high confidence.

Signals worth watching

Four things that would change the read, as opposed to the many that would not.

  • A frontier lab raises the headline price on a shipping flagship. Not a new tier above it, not a tokenizer change, not a small-model bump. An actual increase on the model you are already using. Haiku 4.5 is the nearest miss so far, and Sonnet 5's scheduled rise was withdrawn before it landed. On the flagship line it has not happened, and the most likely way it does is a lab exiting and the survivors finding they no longer have to compete.
  • A hyperscaler cuts GPU depreciation to three years. Six to five is done. Three would confirm the accounting critique and compress reported margins across the sector in the same quarter.
  • Frontier-adjacent open-weight releases slow down. The cheap floor is downstream of DeepSeek, Alibaba and Meta shipping weights. A large hosted provider failing or repricing by more than 2x would be a symptom. The release cadence is the cause.
  • Local models close the agentic gap. Qwen3.6-27B reports 77.2% on SWE-bench Verified against 80.9% for Claude Opus 4.5 in the same table, and the frontier has moved since. Inside JetBrains' coding agent the same model resolved 38.9% of 522 tasks against 51.5% for GPT-5.5, at three percent of the cost per run, and was the weakest of the four models tested at finding the root cause. That gap, not price, is what keeps work on the frontier.

Until one of those moves, the sensible response is the boring one. Keep cheap hosted open weights for the base load, keep frontier capacity for the work that needs it, measure cost per completed task and total spend as two numbers, and own hardware only where residency, latency or air-gapping decide the question. Buying the GPU comes last on that list, and for most people it stays there.