AI Compute Costs: You Are Building on a Subsidy
Five companies will spend about 725 billion dollars on AI infrastructure in 2026. Amazon around 200 billion. Google 175 to 185 billion. Meta 115 to 135. Microsoft on a run rate near 150. That is more capital than the annual output of Switzerland, aimed at one category of machine.
Founders read that number as a weather report. Big spending, big future, carry on. It is not weather. It is the reason the number in your cost model looks the way it does. The price you pay for intelligence today is not the cost of producing it. It is a promotional rate, funded by the largest coordinated capital deployment in the history of software, and the people funding it are quite open about the fact that they are losing money to hold that rate.
I run two companies where AI writes most of the first draft of almost everything. I have built cost models on API prices twice. Both times I was modelling a price, and I thought I was modelling a cost. Those are different things, and the difference is the entire subject of this piece.
Here is the part nobody plans for. Everyone in my circle has a mental model for prices falling. Nobody has one for prices falling and their business getting worse because of it. Both directions kill companies, and they kill different ones.
Table of Contents
- The problem: you priced against somebody else’s loss
- The Subsidy Line
- The three price regimes
- What actually breaks when the price goes up
- What actually breaks when the price goes down
- The Volume Trap: why cheaper never felt cheaper
- The Price-Exposure Ladder
- The Two-Way Stress Test
- A worked example: repricing a document product
- The contrarian take
- What to do Monday morning
- FAQ
The problem: you priced against somebody else’s loss
Start with the arithmetic that everyone quotes approvingly. GPT-4 launched in March 2023 at 60 dollars per million output tokens. By 2026 a frontier-class model sits near 15. Input pricing on the cheap tiers has fallen past a dollar per million. Across the broader curve, per-token cost for a fixed capability level has come down roughly 99.7 percent. Anthropic cut again on June 30, 2026, putting Sonnet 5 at 2 dollars in and 10 out as an introductory rate against a 3 and 15 standard. OpenAI is reported to be weighing its own cuts in response.
Now put that next to a second set of numbers that rarely appears in the same paragraph.
OpenAI is burning roughly 8 billion dollars a year and is tracking toward something near 14 billion in losses for 2026. The Bank for International Settlements used its 2026 Annual Economic Report, published June 28, to name an AI capex bust as one of three interlocking risks to the global financial system, alongside circular financing arrangements and sovereign debt. Moody’s counted roughly 662 billion dollars of data center leases signed but not yet commenced, sitting off balance sheet. Morgan Stanley and JPMorgan both estimate the technology sector needs to issue around 1.5 trillion dollars of new debt over three years to fund the build. PIMCO expects capex to consume about 94 percent of hyperscaler operating cash flow through 2026 and 2027.
Those two sets of facts are the same fact. The price fell because an enormous amount of capital decided that market share now is worth more than margin now. That is a legitimate strategy. It is also a decision, made by somebody else, that can be unmade by somebody else.
Your cost model does not know this. Your cost model has a number in a cell.
The uncomfortable law underneath all of it:
A margin that depends on someone else’s losses is not a margin. It is a loan, and the terms have not been written down.
This is not the same problem as platform risk, which is about a vendor changing policy, cutting access, or competing with you directly. It is not the same as the commoditization clock, which is about the model absorbing your feature. Both of those are about the vendor doing something to you. This one is quieter. Nobody does anything to you at all. A price simply moves, in either direction, and your business turns out to have been a bet on that price staying still.
The Subsidy Line
Here is the model I use now. Two curves, one gap.
The first curve is what you pay. It is public, it is on a pricing page, it falls reliably, and it is the only curve most founders have ever looked at.
The second curve is what it costs to produce. That includes the depreciation on the hardware, the power, the cooling, the debt service, and the amortised cost of training the model in the first place. It is not public. It falls too, because chips and serving stacks genuinely do get better. It just does not fall as fast as the price does, because the price is being pushed down by competition and the cost is only being pushed down by engineering.
The vertical distance between them is the Subsidy Line. Today it is negative in your favour. You are paying less than production cost, and the gap is being covered by equity, debt, and a shared belief about 2030.
The Subsidy Line is not a moral problem. Nobody is being tricked. It is a planning problem, because a business built inside that gap has an implicit assumption baked into every projection: that the gap stays open, at roughly the current width, for as long as the projection runs.
Two things follow from the picture, and the second one is the one founders miss.
First, the gap can close from above. Capital gets more expensive, returns disappoint, financing pulls back, and the price walks up toward the cost line. That is the reset most people vaguely fear and almost nobody has modelled.
Second, the gap can stay open and still ruin you, because a price that keeps collapsing destroys any business whose value proposition was the price. If what you sold was cheap access to a capability, and access gets cheaper for everyone including your customer, you have not been helped. You have been erased. This is the failure mode that the thin wrapper has always had, described in cost terms rather than product terms.
So there is no safe direction. There is only exposure, and exposure is measurable.
The three price regimes
Prices for a commodity input do not drift randomly. They sit in regimes, and each regime rewards a different kind of company. Naming the three of them is most of the work, because once you can name the regime you are in, you can ask the only question that matters: what happens to me in the next one?
Regime one is Subsidy. Price sits below production cost. Capital is buying share. Usage is encouraged, quotas are generous, and every founder in the category feels like a genius because the unit economics of their demo are spectacular. The winning strategy here is speed: grab workflows, grab data, grab habit, because inputs will never be this cheap relative to their true cost again. The losing strategy is to mistake the environment for skill.
Regime two is Reset. Price converges toward production cost, and possibly overshoots while capacity is scarce. This is what happens if financing tightens, if a capex cycle disappoints, or simply if a market consolidates enough that the survivors stop paying to hurt each other. In Reset, the metered part of your product suddenly has a real cost attached, and every generous quota you shipped in Regime one becomes an invoice. I wrote about the specific version of this that hits acquisition in free tier economics. Reset is the general case.
Regime three is Deflation. Price keeps collapsing, faster than anyone expected, because efficiency compounds and competition is brutal. This looks like the best possible outcome and is the one that quietly destroys the most companies, because in Deflation the thing you were charging for becomes something your customer can do themselves for almost nothing. Your COGS improves and your pricing power evaporates at the same time, which is a very strange feeling to watch happen in a spreadsheet.
Most founders I talk to have implicitly bet on a fourth regime that does not exist: prices fall a comfortable amount, competition does not intensify, and their pricing holds. That is not a forecast. That is a wish with a chart attached.
| Regime | What the price is doing | Who wins | Who dies |
|---|---|---|---|
| Subsidy | Below production cost, held there by capital | Whoever converts cheap inputs into durable workflow position fastest | Nobody yet. That is the trap. |
| Reset | Converging up toward cost, possibly overshooting | Products priced on outcomes, with cost passed through or capped | Flat-rate products with uncapped consumption |
| Deflation | Collapsing faster than anyone modelled | Products whose value was never the compute | Products whose pitch was cheap access to a capability |
What actually breaks when the price goes up
Run the upward case properly and it stops being abstract fast.
The first thing that breaks is not your margin. It is your pricing architecture, because most AI products have a fixed price on the outside and a variable cost on the inside. That mismatch is survivable at a small gap and fatal at a large one. A flat monthly fee against metered consumption is a short position on your own input price, taken without meaning to.
The second thing that breaks is your heaviest users, who are almost always your best logos. Consumption in these products is not normally distributed. A small number of accounts drive most of the tokens, and those accounts are usually the ones in your case studies. A price move does not hurt you evenly. It hits precisely where your reference customers live.
The third thing that breaks is the agentic part of your product, and this is the one that surprises people. An agent loop hits the model ten to twenty times per completed task: plan, act, observe, adjust, retry. Every one of those hops is billed. A single-shot feature and an agent feature that produce identical output for the user can differ by more than an order of magnitude in cost, and the agent feature is the one you have been demoing. If you have not instrumented which loops are running away, you cannot see this coming, which is the cost-side version of the argument I made about agents failing silently.
The fourth thing that breaks is the story you tell investors. The gross margin gap is already visible without any reset at all. ICONIQ’s 2026 survey puts average AI product gross margin at 52 percent, improved from 41 percent in 2024, against 70 to 80 percent for classical software. Inference alone runs around 23 percent of revenue at scaling-stage AI B2B companies. Those are Subsidy-regime numbers. A meaningful reset does not nudge them. It takes them somewhere that changes what kind of company you are, which is the deeper point of the gross margin playbook.
What actually breaks when the price goes down
Now the direction everyone assumes is safe.
When the price of an input falls hard and fast, three things happen at once, and only the first is good for you.
Your cost of delivery improves. Good. Enjoy it, because the other two are arriving on the same truck.
Your competitors’ cost of delivery improves identically. Cost advantage that comes from a public price list is not an advantage, it is a coincidence you share with everyone. If your pitch included any version of “we do it cheaper,” a general price drop removes the only reason a customer chose you while making it trivially easy for three new entrants to appear.
Your customer’s cost of doing it themselves improves too, and this is the killer. The build-versus-subscribe line moves every time inputs get cheaper. Work that was clearly worth outsourcing at one price becomes obviously worth insourcing at one twentieth of it. You do not lose to a competitor. You lose to your customer’s weekend.
Here is the test I now apply to any feature: if the model got ten times cheaper and twice as good tomorrow, does my customer need me more or less? More means the value was never the compute. Less means you were selling access, and access is exactly what is being given away.
| Failure surface | Breaks if input price rises 5x | Breaks if input price falls to a twentieth |
|---|---|---|
| Flat pricing over metered cost | Yes. Margin inverts on heavy accounts first. | No. Margin improves. |
| Value proposition is “cheaper access” | Partly. You can raise price with the market. | Yes. The reason to buy disappears entirely. |
| Deep agent loops per task | Yes. Cost multiplies by loop depth. | No. This is where you get to spend more. |
| Generous free tier | Yes. Acquisition cost rises with no revenue attached. | No, but every rival can now match it. |
| Customer could plausibly build it | No. Rising prices push them back to you. | Yes. The insourcing line moves past you. |
| Value sits in data, accountability, workflow | No. Pass the cost through and keep going. | No. Cheaper inputs make you more useful. |
Look at the last row for a second. Exactly one profile survives both columns, and it is not the cleverest product. It is the one whose value was never denominated in compute in the first place.
The Volume Trap: why cheaper never felt cheaper
There is a reason the historic price collapse has not shown up as relief in anybody’s accounts. Per-token cost for a fixed capability fell roughly 99.7 percent. Enterprise AI cloud spending went up about three times in a single year. Total enterprise AI spend rose 320 percent in 2025 while unit prices kept falling.
William Stanley Jevons noticed this in 1865 with coal: efficiency gains in steam engines did not reduce coal consumption, they increased it, because cheaper coal made a thousand previously uneconomic uses viable. The same mechanism is running now, with three compounding parts.
New workloads cross the viability line. At 60 dollars per million output tokens you deployed AI only where the return was unambiguous. At a fraction of that, a long tail of marginal use cases becomes affordable, then default.
The workloads themselves got heavier. Single-shot completion gave way to reasoning traces, tool calls, retries, and multi-agent handoffs. Cheaper per token, vastly more tokens per unit of useful work.
And usage scales with success. When something works and costs little, teams use it constantly. Nobody ever responded to a working, cheap tool by using it the same amount as before.
That is the Volume Trap: unit price falls, unit consumption rises faster, and the bill goes up while every dashboard says costs are improving. It is also why “prices will keep falling” is not the reassurance founders think it is. Falling prices have not once, at any point in this cycle, produced a falling bill. They have produced a bigger appetite. I made the spending-side argument in the efficiency trap. The Volume Trap is its cause.
Note what this does to the two-way test. Deflation does not protect you from Reset. You can be crushed by a price increase applied to a consumption base that only exists because prices got cheap. Cheap inputs do not just lower your cost, they enlarge your exposure to the next move.
The Price-Exposure Ladder
Exposure is not a single number for a company. It is a property of each thing you sell. Two features in the same product can sit at opposite ends of the range, which is why company-level answers to this question are always useless.
Five rungs, from most exposed to least. Value migrates downward.
Rung one, pass-through resale. You buy tokens and sell tokens. Your margin is the spread, and the spread is set by a competitor’s marketing budget. Prices rising compresses you. Prices falling invites entrants who will run the spread to zero. There is no version of the future where this rung is comfortable, which is the whole reason the thin wrapper conversation never ends.
Rung two, metered feature at a flat price. The most common rung and the least examined. You charge 49 dollars a month, some accounts cost you four dollars to serve and some cost you sixty, and you find out which is which during a reset. Every uncapped flat plan is a bet that consumption stays where it is.
Rung three, fixed-work product. Compute per unit of work is bounded by design. One document, one call, one reconciliation, with a known ceiling. You still carry price risk, but it is arithmetic rather than surprise, and you can reprice cleanly because you know what a unit costs.
Rung four, outcome-priced. The customer buys a resolved ticket, a closed book, a filed return. How many tokens it took is your problem, which sounds worse and is much better, because it converts an exposed cost line into an ordinary engineering incentive. Get more efficient and you keep the gain instead of passing it to a customer who never asked for it. This is the direction I argued for in revenue models for AI products and again in the case against per-seat pricing.
Rung five, compute-independent value. The data you caused to exist, the workflow you sit inside, the accountability you carry, the integrations nobody wants to rebuild. Input price moves here are weather. They change your cost and not your position.
Most products span three rungs and have only ever thought about one. The audit is worth an afternoon.
The Two-Way Stress Test
Here is the actual test. It takes about ninety minutes and it uses numbers you already have.
Take your product one feature at a time. For each, run both arms.
The up arm. Multiply your current per-unit input cost by five. Not because five is a forecast, but because it is roughly the distance between a subsidised price and an unsubsidised one with a normal margin on top, and it is a number your finance model can actually digest. Now ask: does this feature still make money at current pricing? If not, can I raise the price without losing the account? If not, can I cap or meter it without breaking the promise I made? If all three answers are no, that feature only exists because of the subsidy.
The down arm. Divide your current per-unit input cost by twenty. Now ask: does my customer still need me? Does a competitor with no distribution and a weekend now match my core function? Was my pricing power ever attached to the difficulty of the work? If the honest answer is that cheap inputs make me less necessary, that feature is a Deflation casualty, and no amount of margin improvement saves it.
Four outcomes, and only one of them is a business.
Fails up, survives down: you have a subsidy-dependent feature that at least is not commoditised. Cap it, meter it, or reprice it, and you are fine.
Survives up, fails down: you have a cost-efficient feature nobody will need. This is the dangerous quadrant because the financial metrics look healthy right up until the demand disappears.
Fails both: shut it down or fold it into something else. It exists because of a moment, not because of a need.
Survives both: this is the thing. Fund it, deepen it, and organise the rest of the product around it.
I ran this on my own products in an afternoon, and the honest result was that one feature I was proud of failed the down arm badly. Its whole pitch was doing something expensive cheaply. Cheap was never mine to own.
Score it. One row per feature, honest numbers, no rounding in your own favour.
| Question | Exposed answer | Durable answer |
|---|---|---|
| What does one unit of work cost me? | I know my total API bill, not the per-unit number | Per feature, per completed unit, loop depth counted |
| What is my worst account costing me? | Have not looked, consumption is not tracked by account | Top decile identified, margin known, capped |
| What happens at five times the price? | Margin inverts and I have no repricing path | Cost passes through, or is bounded by design |
| What happens at a twentieth of the price? | My customer or a weekend project replaces me | Cheap inputs make my position stronger |
| Where does my value live? | In access to a capability, priced on a public page | In data I caused, workflow position, accountability |
| Which regime am I planning for? | The one that does not exist: gentle decline, stable pricing | All three, with a written answer for each |
Anything in the middle column is not a flaw in your product. It is an assumption you inherited from the regime you happened to start in.
A worked example: repricing a document product
Abstractions are easy to nod at. Here is the concrete version.
Imagine a product that reads inbound contracts and produces a risk summary. Pricing is 199 dollars a month per workspace, unlimited documents, because unlimited tested better than metered and inputs were cheap enough that it did not seem to matter.
Median workspace processes 40 documents a month. A document runs about 30,000 input tokens and produces 4,000 out, and because quality mattered the pipeline makes four passes: extract, classify, cross-reference, summarise. Call it 120,000 in and 16,000 out per document once the passes are counted.
At a cheap tier, roughly a dollar per million input and six per million output, that is about 0.12 in input and 0.10 in output. Around 22 cents a document, so 8.80 dollars a month for the median workspace against 199 in revenue. Gross margin looks like classical software. Everyone is happy.
Now the top decile. Those workspaces process 600 documents a month, not 40. That is 132 dollars of compute against the same 199 in revenue. Margin on your best-looking logos is already thin, and this is before anything moves.
Run the up arm at five times. The median workspace costs 44 dollars, still fine. The top decile costs 660 dollars against 199 of revenue. You are now paying 461 dollars a month for the privilege of serving your reference customer, and every renewal conversation with them is a conversation about a number you cannot support. Note what the reset did: it did not lower your margin evenly, it inverted it precisely where your case studies live.
Run the down arm at one twentieth. Compute is now about a penny a document. Wonderful. Except at a penny a document, the customer’s own engineer can build a serviceable version in a week, three competitors launch at 29 dollars a month, and your 199 dollar price has nothing holding it up except the thing you have not built yet.
So what does the fix look like? Not a price change. A shift down the ladder.
Move to a fixed-work unit so cost is bounded: an included allowance of documents, overage priced above your own worst-case cost. That handles the up arm on its own. Then, for the down arm, put the value somewhere compute cannot reach. The accumulated history of every contract this customer has ever reviewed and what they decided. The clause library that reflects their own risk tolerance, learned from their edits. The audit trail their counsel actually signs. None of that gets cheaper when tokens do. All of it gets better.
The moment you make that shift, cheap inputs stop being a threat and become a discount on something you were going to do anyway. That is the whole objective. Not to predict the price. To stop caring which way it moves.
The contrarian take
The standard advice on this topic is to optimise: route cheap queries to small models, cache aggressively, trim prompts, batch where you can. Red Hat has documented enterprise deployments cutting compute cost around 70 percent with tiered routing while holding quality. That advice is correct and I follow it.
It is also almost entirely beside the point, and the reason is uncomfortable.
Optimisation is a multiplier on your exposure, not a reduction in it. If you cut cost per task by 70 percent and your value proposition was still denominated in compute, a five times price move takes you from broken to slightly less broken, and a twenty times price collapse still removes your reason to exist. You have improved a ratio while leaving the denominator someone else’s decision.
The founders I watch getting this right are doing something that looks lazier and is much harder. They are not chasing efficiency. They are moving the thing they charge for until it no longer lives on the price curve. That is slow, it does not produce a satisfying metric, and it is the only move that pays off in both directions.
Now the honest counter, because this argument has a soft spot and I would rather name it myself.
The subsidy might not close in any way that matters. If efficiency keeps compounding, production cost may fall to meet the subsidised price rather than the price rising to meet cost, and the gap could simply close from below with nobody harmed. There are serious people who expect exactly that, and the historical record on compute prices is broadly on their side. If they are right, the up arm never fires and everyone who spent 2026 defensively repricing looks overcautious.
I would still run the test, for one reason. The down arm fires in that world too, and it fires harder. A world where compute costs collapse to nothing is a world where selling access to compute is worthless. The optimistic scenario does not spare you. It just picks a different arm.
And the pessimistic scenario is not fringe. When the BIS devotes a chapter of its annual report to circular financing and a capex bust, when hyperscalers are depreciating hardware over five to six years that many observers think lives two to three, when 662 billion dollars of lease commitments sit off balance sheet and capex is projected to eat 94 percent of operating cash flow, the probability of a financing pullback is not zero. You do not have to believe it will happen. You only have to build something that does not require it not to.
What to do Monday morning
Five things, in order, none of which require a strategy offsite.
One: get your real per-unit cost, by feature. Not your total API bill. Cost per completed unit of customer-visible work, separated by feature, with agent loop depth counted. Most teams cannot produce this number, which is itself the finding. An afternoon of instrumentation gets you there.
Two: find your top decile of consumption and check the margin there. Sort accounts by tokens consumed. Look at the top ten percent. That is where a reset lands first and it is usually where your logos live. If margin is already thin there in Subsidy conditions, you do not have a future problem, you have a current one you have not looked at.
Three: run both arms on your three biggest features. Five times up, twenty times down, ninety minutes. Write the four-quadrant result down where your team can see it. The conversation this starts is more valuable than the numbers.
Four: put a ceiling on anything uncapped. Not a price rise, a boundary. An included allowance with overage above your worst-case cost, or a hard cap with an upgrade path. You are not trying to make money on the cap. You are trying to stop holding an unbounded liability against a bounded price, which is the specific mistake that turns a reset into an emergency.
Five: pick one compute-independent asset and start compounding it this quarter. The decision history, the correction log, the customer-specific rules, the integration nobody wants to redo. It does not have to be impressive at the start. It has to start, because that class of asset only accumulates with calendar time and no amount of money buys you back the quarter you skipped. This is the same argument I made about what to build when building is free, arriving from the cost side.
The founders who come out of the next regime in good shape will not be the ones who guessed the direction. They will be the ones who arranged not to need a guess. If you want the wider view of where those positions sit, the AI-native founder playbook and the AI opportunity map both start from the same place.
FAQ
What are AI compute costs actually made of?
For most founders, the visible cost is API pricing per million tokens, split between input and output. Underneath that sit GPU depreciation, power, cooling, networking, debt service on the hardware, and the amortised cost of training the model. You pay the first number. The second set determines whether the first number is sustainable. In 2026 the price you pay is generally below the cost of production, with the difference covered by investor capital.
Will AI inference costs keep falling?
Per-token cost for a fixed capability level has fallen roughly 99.7 percent since early 2023, and there is real engineering behind that, so continued decline is a reasonable base case. Two things complicate it. First, prices are also being held down by competition funded by losses, and that part can reverse if financing conditions change. Second, falling unit prices have not produced falling bills at any point in this cycle, because consumption rises faster than price falls. Plan for a cheaper unit and a larger bill at the same time.
What is the Subsidy Line?
The Subsidy Line is the gap between what you pay for AI inference and what it costs to produce. Today that gap is in your favour, funded by equity and debt in a capital cycle you do not control. It matters because any business plan built inside the gap carries a hidden assumption that the gap stays open at roughly its current width, and nobody has written that assumption down.
How do I stress test my AI product against price changes?
Run both arms per feature. The up arm multiplies your per-unit input cost by five and asks whether the feature still makes money at your current price, and whether you could raise or cap it if not. The down arm divides that cost by twenty and asks whether your customer still needs you when the work becomes nearly free. A feature is durable only if it survives both. Passing one arm is common. Passing both is the whole test.
Why is a falling AI price dangerous for my startup?
Because a general price fall improves your cost, your competitors’ cost, and your customer’s cost of doing it themselves, all at once, and only the first is good for you. If any part of your pitch was that you do something expensive cheaply, cheap inputs remove the reason to buy from you while making it trivial for new entrants to appear. The build-versus-subscribe line moves every time inputs get cheaper, and it can move straight past you.
What is the Volume Trap?
The Volume Trap is what happens when unit price falls, unit consumption rises faster, and the total bill goes up while every efficiency dashboard shows improvement. It is the Jevons paradox applied to tokens. Cheaper access makes marginal use cases viable, agent workflows consume ten to twenty model calls per task instead of one, and teams use tools more when the tools work and cost little. Total enterprise AI spend rose 320 percent in 2025 while unit prices kept falling.
Should I switch to usage-based pricing to protect my margin?
Usage-based pricing solves the up arm and does nothing for the down arm, so it is half an answer. Passing costs through protects you if input prices rise, but if input prices collapse, usage-based pricing means your revenue collapses with them. The stronger move is to price on the outcome, where the token count becomes your internal engineering problem and efficiency gains stay with you, or to move value into assets that are not denominated in compute at all.
Is the AI capex buildout actually at risk?
Serious institutions think it warrants attention. The BIS named an AI capex bust and circular financing among the top risks to the financial system in its 2026 annual report. Analysts have flagged the gap between five to six year book depreciation on AI hardware and a two to three year economic life, roughly 662 billion dollars of off balance sheet lease commitments, and an estimated 1.5 trillion dollars of new tech-sector debt needed over three years. None of that is a prediction of collapse. It is a reason not to build a company that requires a collapse never to happen.
What does compute-independent value look like in practice?
Data you cause to exist rather than data you bought, such as the record of what your customers decided and corrected over time. Workflow position, where you sit inside a process that would be painful to rearrange. Accountability, where you carry a risk the customer does not want to hold. Integrations and system-of-record status. These share one property: they get more valuable when compute gets cheaper, because cheap compute makes them easier to act on and harder to replace.