Pricing Under Cheap Inference

· 26 min read

The price of a million tokens hit its lowest point of the year this month, and half my founder friends read it as good news. It is good news for your cost line. It is a quiet emergency for your price line, and almost nobody is treating it that way.

A cost win that quietly becomes a revenue loss

When the raw material of your product gets cheaper, the instinct is to celebrate. Lower cost, better margin, more room to compete on price. That instinct is correct for most of business history and wrong for this specific moment, because the raw material here is not getting a little cheaper. It is falling off a cliff, and it is dragging your pricing anchor with it.

A unit of frontier-class intelligence that cost around 30 dollars per million tokens in early 2023 sells for under a dollar now. Blended inference fell roughly 67 percent in a single year, from about 18.40 to 6.07 dollars per million tokens between the first quarter of 2025 and the first quarter of 2026. By one measure, the price of equivalent capability dropped 280-fold in two years. Anthropic cut prices, OpenAI followed, Google pushed a fast tier close to free, and Chinese labs kept undercutting the floor. The direction is not in dispute. The cost of the compute inside your product is heading toward zero.

Here is the part founders miss. The moment your cost approaches zero, any price you set as a markup on that cost approaches zero too. If you priced your product the way software has always been priced, as some multiple of what it costs you to deliver, then a 90 percent cost cut is not a margin gift. It is a countdown on your revenue.

You can already see the squeeze in the margin data. ICONIQ’s 2026 numbers put the average AI product gross margin at 52 percent, against the 80 percent that defined mature software. Bessemer pegs AI-first companies at 50 to 60 percent. Public software companies disclosing AI-driven margin pressure are reporting gross margins 10 to 17 points below their pre-AI baseline. The tempting read is that inference is too expensive. The truer read is that these companies priced against a cost line and the cost line moved.

I have watched founders respond to falling token prices by cutting their own prices to stay competitive, then watched a rival cut faster, then watched the whole category race toward a floor that keeps dropping because the thing it is anchored to keeps dropping. That is not a strategy. That is being tied to a falling object. The way out is to stop pricing the tokens and start pricing the thing the tokens produce.

The framework: the cost-price decoupling

The central move of this whole essay is one idea, and every tactic below is a way of acting on it. Cost and price used to travel together. In a deflating-cost world they have to be pried apart on purpose.

Think of two lines on the same chart over time. One line is the cost of doing the work, which is collapsing. The other line is the value of the work to the person who needs it done, which barely moves. A resolved support ticket is worth roughly the same to a company whether the model burned a thousand tokens or a million. A correctly reviewed contract, a closed sale, a monitored network, a cleaned dataset, these are worth what they are worth because of the outcome, not because of the compute. The two lines used to sit close together. Now they are splitting apart, and the space between them is widening every quarter.

The Cost-Price DecouplingThe cost of the work falls. The value of the outcome does not. Anchor to the wrong line and you fall with it.dollars per unit of worktime and model generations20232026cost and price started here, togethervalue of the outcome (flat)cost of the work (falling)the deflation dividendthis gap is yours to keep or give awayprice anchored to valueprice anchored to cost

The law is simple to state and hard to live by. When the cost of the work falls toward zero, price stops being a markup on cost and becomes a claim on value. Anchor your price to a number that is falling and your revenue falls with it. Anchor it to the outcome and the deflation lands in your margin instead of your customer’s savings.

Notice what this reframes. Cheap inference is not primarily a cost event. It is a pricing event. The founders who treat it as a cost event go hunting for ways to spend less per call, which is worth doing, but it answers the small question. The founders who treat it as a pricing event ask a bigger one: now that the work is nearly free to perform, what is it actually worth to have done, and how do I charge for that instead? This is the difference between running a thin passthrough on somebody else’s model and running a business. It is the same instinct behind the wrapper trap, where the product is only a light coat of paint over an API, except here the trap is not thinness of product, it is thinness of pricing logic.

The three pricing bases, and which one survives deflation

Every price you have ever set was anchored to one of three things. I call them the three pricing bases, and the whole game right now is migrating down the list before your category forces you to.

The first base is cost. You add up what it costs to deliver and mark it up. Cost-plus is the oldest pricing logic there is and it is the one that dies fastest in a deflation, because the number you are marking up is the number falling off the cliff. The second base is usage. You charge per seat, per call, per token, per some proxy for how much of the thing the customer consumed. Usage pricing feels modern and it powers a lot of good businesses, but per-token usage is quietly cost pricing wearing a nicer outfit, because the token is exactly the unit whose price is collapsing. The third base is outcome, sometimes called value. You charge for the job getting done: a resolved ticket, a booked meeting, a cleared claim, a closed deal. Outcome pricing is anchored to the one line on the chart that is not falling.

Table 1. The three pricing bases under falling costs
Pricing base What you anchor to Does it have a floor? What deflation does to you
Cost-plus Your compute bill, marked up No. Falls with the cost. Revenue tracks your cost to zero. Every price cut invites a deeper one.
Usage (per token or per call) Units of consumption Weak. The unit itself is deflating. Price per unit sinks while usage climbs. You can lose on both axes at once.
Outcome (per job done) Buyer-recognized value Yes. Set by what the result is worth. Cost falls, price holds, the gap becomes your margin. Deflation works for you.

The market is already moving down this list, just slower than the cost curve is moving. Usage-based pricing went from roughly 30 percent of software companies in 2019 to about 85 percent by 2024. Hybrid models, a base fee plus a variable unit, sit near 43 percent today and are projected to reach 61 percent by the end of 2026. Pure outcome pricing is still young, but the companies using outcome components report 31 percent higher retention and 21 percent higher satisfaction. The catch worth respecting: 78 percent of the firms pulling outcome pricing off had five or more years in market, which tells you it is a capability you grow into, not a switch you flip on day one.

The practical answer for most founders is not to leap straight to pure outcome pricing, which is hard to meter and hard to sell early. It is to stop anchoring on cost, move usage up to something closer to a business unit than a compute unit, and add an outcome component where you can measure the result cleanly. That is a different discussion from the menu of revenue models you pick from at launch. That menu tells you the shapes available. This tells you which way to slide along it as your cost line drops out from under you.

The Jevons trap: cheaper tokens, bigger bills

There is a second reason per-token pricing is dangerous, and it is the one that surprises people most. Cheaper tokens do not lower the bill. They raise it.

This is the Jevons paradox, the old observation that when a resource gets cheaper, total consumption rises faster than the price falls, because the low price unlocks uses that were not worth it before. It is playing out in inference right now with almost comic force. A simple chatbot answer used to be one round trip. An agentic workflow, where the model plans, calls tools, checks its work, and iterates, hits the model 10 to 20 times for a single task. Goldman projects that the shift to agentic patterns could raise total token demand as much as 24-fold. So even as the price per token fell 280-fold, total enterprise AI spend rose about 320 percent over the same stretch. Enterprise AI infrastructure spending went from 11.5 billion dollars in 2024 to 18.3 billion in 2025 to 37.5 billion in 2026. In one survey, 73 percent of enterprises reported that their AI spend blew past what they had budgeted.

The Jevons TrapPrice per token falls. Usage per task rises faster. The bill goes up anyway.Price per tokenfalls 90 percent or moreTokens per taskrises up to 24 times (agentic loops)The billUPfor you and your customerPrice the tokens and you charge less per unit while the units explode. You lose on both.

Sit with what that means for your price. If you built a product that charges a markup on tokens, the deflation cut your revenue per unit and the Jevons effect multiplied the units your customer has to buy. Your price per call went down. Your customer’s total bill went up. You made your own revenue smaller and your customer’s invoice scarier at the same time. It is the worst of both worlds, and it is the default outcome of naive usage pricing. This is the pricing-side cousin of the efficiency trap, where a metered cost you thought you controlled quietly runs away from you.

Cutting the bill through smarter engineering, the kind of routing and caching I wrote about in multi-model routing, is real and worth doing. It protects your cost line. But it is a cost move, and this is a price problem. You cannot engineer your way out of a pricing mistake. You can only re-anchor.

Who keeps the deflation dividend

Every time inference gets cheaper, a pool of savings appears. I call it the deflation dividend. The work costs less to do, so there is money that used to go to compute and now goes to somebody. The question that decides your business is which somebody.

Here is the thing founders do not fully register: who keeps the dividend is not decided by the market. It is decided by your pricing model. If you price cost-plus or per-token, you hand the dividend straight to the customer, because their bill drops as your cost drops. That can be a deliberate acquisition strategy, cheap AI as a wedge, and sometimes it is the right call. But most founders are giving the dividend away without deciding to, simply because their pricing was built in a world where cost and value moved together. If you price on the outcome, the dividend stays with you. The customer pays for the resolved ticket, the resolved ticket is worth what it is worth, and the falling cost of producing it becomes expanding margin on your side of the table.

Table 2. Passthrough versus capture: who keeps the deflation dividend
Pricing model Where the dividend goes When it is the right call
Cost-plus / per-token passthrough To the customer, automatically Rarely on purpose. Only as a deliberate, temporary land-grab where price is the wedge.
Hybrid (platform fee plus outcome unit) Split, and you set the split Most companies, most of the time. Predictable for the buyer, capture for you.
Outcome / value To you, as widening margin When the result is clean to measure and clearly attributable to your product.

Bessemer, in a study of AI founders who solved their pricing, landed on the hybrid as the practical default: a platform or retainer fee that covers capacity and onboarding, plus a variable unit tied to work the buyer already recognizes as valuable, like resolved tickets or reviewed claims. That structure is powerful because it lets you keep a real share of the dividend while giving the customer the predictability they need. The base fee is your floor. The outcome unit is your upside. And because the outcome unit is priced on value, the falling cost of inference under it turns into margin, not a discount you were forced to pass along. This is the same move that shows up in the gross-margins playbook, seen from the revenue side rather than the cost side.

The value floor: why outcome prices do not fall

The reason outcome pricing survives a deflation is that it has a floor and cost pricing does not. A cost-anchored price can fall forever, because the cost it tracks can fall forever. An outcome-anchored price stops falling at the value of the outcome, because below that number the deal still makes obvious sense for the buyer. A resolved support ticket that saves a company 8 dollars in human handling time is worth some fraction of 8 dollars no matter how cheap the tokens get. That fraction is the floor. It does not care about the price of compute.

The clearest proof is in the companies that already price this way. Intercom’s Fin charges 0.99 dollars per resolution. You pay when it works, and the price is set against the value of a solved conversation, not the tokens spent solving it. Sierra charges around 1.50 dollars per resolved interaction, has no public pricing page at all, serves 40 percent of the Fortune 50, and reached a 15.8 billion dollar valuation on the back of that model. Decagon starts near 95,000 dollars a year and commonly charges per conversation the agent handles. Notice that none of these prices is a markup on tokens. Notice too that Sierra can decline to publish a price at all, which is a thing you can only do when your number is anchored to value the buyer feels rather than a cost the buyer can look up.

Table 3. What the outcome-priced players actually anchor to
Product The unit they charge What that unit is anchored to
Intercom Fin $0.99 per resolution A solved conversation. Pay only when it works.
Sierra About $1.50 per resolved interaction The value of a resolution. No public price at all.
Decagon From ~$95k per year, often per conversation Handled volume, floored by a platform commitment.

There is an honest counterweight here, and skipping it would make the argument weaker. Outcome pricing puts the risk of the model’s misses on you. If you charge only for resolutions and your agent resolves 60 percent of tickets, you eat the cost of the 40 percent it touched and failed. That is why the pure per-resolution players still tend to sit on a platform fee or a minimum, and why per-conversation pricing, where you get paid whether or not the issue is solved, exists as a hedge. Moving toward value pricing is not free. It transfers delivery risk to you, which is exactly why it also transfers the margin to you. You are being paid for carrying the outcome, and carrying the outcome is the business.

Why outcome pricing is a capability, not a switch

It would be dishonest to make this sound easy. You do not flip a switch on Monday and start charging for outcomes on Tuesday. The data says so plainly: 78 percent of the companies successfully running outcome pricing had five or more years in market before they got there. That is not because outcome pricing is a secret only veterans know. It is because charging for a result requires three capabilities you have to build, and each one takes time.

The first is measurement. You cannot charge for a resolved ticket until you can prove, cleanly and to the buyer’s satisfaction, that a ticket was resolved. That means instrumentation, logging, and a definition of the outcome that survives an adversarial customer who would rather it did not count. The fights over outcome pricing are almost never about the price. They are about the definition. What is a resolution? Does a deflected question count, or only a fully closed case? Whose judgment settles a dispute? Build the meter before you touch the price, because the meter is the actual product of a pricing change.

The second is delivery reliability. Outcome pricing only makes sense once your success rate is high enough that eating the misses still leaves a profit. Early on, when your agent resolves 30 percent of cases, charging per resolution would either bankrupt you on the failures or force a price so high the buyer balks. You earn the right to outcome pricing by getting good at the outcome. Until then, a hybrid with a healthy base fee is not a compromise, it is the correct structure for your maturity.

The third is trust. A buyer will only accept a value-anchored price from a vendor they believe will define the outcome fairly. That is why the companies that get there tend to have tenure. So the path is staged, not sudden. Start by moving off cost-plus to a usage unit that is a business unit rather than a compute unit. Add a small outcome component beside a base fee. Grow the outcome share as your measurement and reliability improve. Each step moves you rightward on the map without betting the company on a meter you have not yet built. The mistake is not pricing on cost forever. The mistake is thinking you can skip the staircase and land on pure outcome pricing in one jump.

Copilot versus agent: attribution is what defends a price

Not every product can charge for outcomes, and the reason is worth understanding because it tells you what to build. Outcome pricing needs an outcome you can point at and say, that happened, and my product is why. The more attributable the result, the more defensible the price. The softer the result, the faster the price gets negotiated down toward cost.

This is the real line between a copilot and an agent, and it is a pricing line before it is a product line. A copilot sits next to a human and makes them faster. Its value is real but soft: the human still did the work, the time saved is fuzzy, and different users get wildly different benefit. That softness compresses willingness to pay, because the buyer cannot cleanly attribute a dollar figure to your product, so they anchor back to a per-seat number and squeeze it. An agent completes the workflow end to end and hands back a finished result. Its value is hard: the ticket is resolved, the claim is processed, the meeting is booked, and the attribution is unambiguous. That hardness defends the price, because the buyer is negotiating against a number they can see.

So the pricing question and the product question turn out to be the same question. If you want to escape cost-anchored pricing, build toward attributable outcomes, not just faster assistance. The value has to be legible to the buyer or it collapses back to a proxy for cost. This connects to a shift I have written about in agents as your customer and the broader move where the buyer’s own agent starts comparing your attributable results against a rival’s. In that world, a legible outcome is not just easier to price. It is the only thing that gets selected.

The decoupling map: where to build

Put the two forces together and you get a simple map. One axis is what your price is anchored to, from cost and tokens on the left to value and outcome on the right. The other axis is how visible your value is to the buyer, from soft copilot help at the bottom to attributable agent results at the top. Where you land decides whether the deflation is your friend or your funeral.

The Decoupling MapAnchor your price to value, and make the value legible. Then deflation feeds your margin.Squeezedclear value, but yougive the dividend awayDurable Capturevalue anchor plusattributable resultbuild hereRace to Zerocost anchor, soft value,nothing to defendFragile Premiumhigh price the buyercannot see the reason forprice anchored to cost / tokensprice anchored to value / outcomeattributable (agent)soft (copilot)

Bottom left is the Race to Zero. You anchor to cost, your value is soft, and you have nothing to defend when a rival cuts price. This is where most thin AI products sit, and it is where the deflation eats you. Top left is Squeezed. You actually deliver clear, attributable value, but you still price on cost or per seat, so you are handing the dividend to your customers while doing the hard work. This is the most common unforced error I see: a genuinely good product leaving money on the table because its pricing was set before the founder understood the decoupling. Bottom right is Fragile Premium. You charge a value-shaped price but the buyer cannot see why, so the number feels arbitrary and gets negotiated down at every renewal. Top right is Durable Capture. Your price is anchored to value, the value is legible, and the falling cost of inference under it flows into your margin. That is the only quadrant where cheap inference is unambiguously good news for you, and it is the one to build toward.

The arrow runs from bottom left to top right, and it runs through two different kinds of work. Moving right is a pricing and packaging change. Moving up is a product change. Most founders can start moving right this quarter and should, while they do the slower work of moving up.

Two founders and one price cut

Let me make the decoupling concrete with a scene, because it is easy to nod at a chart and still price the old way out of habit. Two founders build AI contract review tools. Same category, same underlying models, same quality. On the same morning, their model provider announces a 50 percent price cut on inference. One external event, hitting both of them identically.

Founder A prices per page reviewed, a rate she set by marking up her compute cost per page. When the cut lands, her cost per page halves, and she does what feels honest and competitive: she halves her price to pass the savings along. Her revenue per document drops by half overnight. A week later a rival, pricing the same way, cuts deeper to win a deal. Now she cuts again. She is in a race whose finish line keeps receding, because the number everyone is anchored to keeps falling. Her product got no worse and her business got much weaker, and she never made a single bad decision in isolation. Each cut felt reasonable. The anchor was the problem.

Founder B prices per reviewed contract, at a number set against what a clean contract review is worth to the customer, roughly what they would have paid a junior lawyer or a review service to do it. When the same 50 percent cut lands, his cost to produce a review halves too. But the value of a reviewed contract to his customer did not change, so he holds his price. The customer does not complain, because the price was never framed as a markup on tokens, and a reviewed contract is still worth exactly what it was worth yesterday. The entire cost saving falls into his margin. He now has more gross profit per contract than he had last week, from the identical event that cut Founder A’s revenue in half. He reinvests some of that margin into higher accuracy and better reporting, which makes his outcome more attributable, which defends his price even harder at renewal.

Same market, same models, same price cut, opposite results. The only variable that differed was what each founder anchored the price to. That is the whole essay in one scene. The deflation is not doing something to these founders. Their pricing choice is deciding what the deflation does.

The contrarian take: cheap inference is a pricing-power test

Here is the thing almost everyone gets backwards. Falling inference costs are not a gift to whoever builds on AI. They are a test, and the test is about pricing power, not compute.

The comfortable story is that cheap intelligence lifts all boats, that as the cost of the smart part drops, anyone with an idea can build a profitable product on top. The uncomfortable truth is that cheap intelligence is a solvent. It dissolves any pricing power that was secretly resting on the cost of compute. If the only reason your product could charge money was that running a model was expensive and you did it for the customer, then as running the model gets free, so does your product, and your price goes with it. The deflation does not reward everyone. It rewards the founders who had a claim on value that never depended on compute being expensive in the first place, and it exposes everyone who was quietly selling access to a cost.

This is why I think the biggest pricing mistake of this era will not be charging too much. It will be reflexively passing the savings through. Founders are trained to think that lowering price when your costs drop is honest, competitive, customer-friendly. In a normal market it is all three. In a deflating-cost market it is often just the slow surrender of your entire margin to a customer who never asked for it and would have paid the same for the outcome. The models are commoditizing, which I have called the commoditization clock, and the switching cost between them is nearly zero once you have any kind of router in place, a fragility I dug into in vendor lock-in and switching costs. When the input is a commodity you can swap in one config change, the only durable price you can charge is a price on something that is not the input. That something is the outcome, and the willingness to hold a price on it is the actual test.

Put bluntly: the question is not whether you can build now that inference is cheap. Almost everyone can. The question is whether you can charge, and charging requires a claim on value that the falling cost of compute cannot touch. That claim is your real product. Everything else is a passthrough with a login screen.

What to do Monday morning: the decoupling test

This is the part you can act on without a strategy offsite. Run every price you charge through one question, then re-anchor the ones that fail.

The decoupling test is a single line. For each SKU, ask: is the number I charge anchored to a number that is falling? If your price is a markup on compute, a per-token rate, or anything that shrinks automatically as inference gets cheaper, it fails the test and it is quietly bleeding. If your price is anchored to a resolved job, a delivered outcome, or a business result the customer already values, it passes.

For everything that fails, here is the order of operations I would run.

First, find the outcome. For each product, name the single unit of value the buyer would recognize and pay for on its own: a resolved ticket, a qualified lead, a reconciled invoice, a cleared document. If you cannot name it, that is your real problem, and it is a product problem wearing a pricing costume. Fix the product until there is a nameable outcome.

Second, put a floor under it. Move to a hybrid before you move to pure outcome pricing. Set a platform or minimum fee that covers your capacity and your delivery risk, then add a variable unit tied to the outcome. The base protects you from the 40 percent the model misses. The unit gives you the upside and the dividend.

Third, make the value legible. If the buyer cannot see and attribute the result, your value-shaped price will erode at renewal. Instrument the outcome. Show the resolved count, the hours returned, the dollars saved, in the product itself, so the number you charge always sits next to the value it produced. Legibility is what turns a fragile premium into a durable one.

Fourth, decide the dividend on purpose. Choose, explicitly, whether the next inference price cut goes to your customer as lower price or to you as higher margin. Sometimes passthrough is the right, deliberate wedge. Just never let it happen by default because your pricing model made the choice for you. Founders who think about their own incompressible core, the part of the business only they can do, already know that the durable money sits on the things compute cannot commoditize. Price those. Rent the rest.

Do this once and you stop being tied to a falling object. The cheaper inference gets, the better your business does, because you finally arranged your prices so the deflation lands on your side of the ledger. That is the whole trick. Cost is falling toward zero either way. The only decision left to you is who gets to keep the difference. For the wider founder playbook this sits inside, see the AI-native founder playbook and the broader AI opportunity map.

FAQ

Does cheap inference actually hurt my business, or is it good news?
It depends entirely on how you price. For your cost line it is good news. For your revenue line it is dangerous if your price is anchored to cost, because a price set as a markup on compute falls as compute falls. The way to make it purely good news is to anchor your price to the outcome your product produces, so the falling cost turns into widening margin instead of shrinking revenue.

What is wrong with usage-based or per-token pricing?
Per-token pricing is cost pricing in disguise, because the token is exactly the unit whose price is collapsing. Worse, the Jevons effect means agentic workflows consume far more tokens per task, so your price per unit falls while the units explode. You can lose on both axes at once. Usage pricing can still work if the unit is a business unit, like a processed document, rather than a raw compute unit.

Should I lower my prices when my inference costs drop?
Not automatically. Passing the savings through is the reflex, and in a deflating-cost market it often just hands your entire margin to a customer who would have paid the same for the outcome. Choose the passthrough deliberately when cheap price is your acquisition wedge. Otherwise, hold the price on the outcome and keep the deflation as margin.

What is outcome-based pricing, in one sentence?
You charge for the job getting done rather than for access or consumption, like Intercom’s Fin at 0.99 dollars per resolution or Sierra at around 1.50 dollars per resolved interaction, so the price is anchored to what the result is worth instead of what it costs to produce.

Isn’t outcome pricing risky since I eat the cost of the misses?
Yes, and that risk is the reason it carries the margin. If your agent resolves 60 percent of tickets, you absorb the cost of the 40 percent it failed. That is why most outcome players sit on a platform fee or minimum, and why a hybrid, a base fee plus an outcome unit, is the practical default. You are being paid to carry the outcome, and carrying it is the business.

Why can Sierra charge a premium while a thin wrapper cannot?
Because Sierra’s price is anchored to an attributable result, a resolved interaction, that the buyer can see and value on its own, while a thin wrapper’s value is a proxy for the cost of the model call it makes. When the input is a swappable commodity, the only defensible price is one anchored to something the input cannot commoditize. Attribution is what defends the number.

How is this different from just improving my gross margins?
Improving margins is a cost-side move: spend less per call through routing, caching, or smaller models. This is a revenue-side move: change what your price is anchored to so the falling cost flows into your margin instead of your customer’s savings. You need both, but you cannot engineer your way out of a pricing mistake. Re-anchoring is the fix that cost work cannot substitute for.

I am a solo founder without a pricing team. Where do I start?
Run the decoupling test on every SKU: is the number you charge anchored to a number that is falling? For anything that fails, name the outcome the buyer would pay for, put a base fee under it to cover your risk, instrument the value so it is visible, and decide on purpose whether the next price cut goes to the customer or to you. That is a one-afternoon audit, and it is the highest-return pricing hour you will spend this year.