Service as Software: The Accountability Trap

· 26 min read

Service as software is the phrase every deck is built around right now. The pitch is clean: stop selling tools that help people do work, and start selling the finished work itself. The prize behind it is enormous. Foundation Capital put a number on it, a $4.6 trillion pool of salaries and human services that software can now go after, next to a classic software market worth a few hundred billion. Early in 2026 the public markets started pricing the shift, and hundreds of billions of dollars came off software valuations in the stretch people started calling the SaaSpocalypse. The direction is real. What almost nobody hands you is the bill that comes with it.

Here is the durable version, the part that will still be true after the current headlines age out. When you move from selling a tool to selling work, you are not just changing your price tag. You are changing who is on the hook when the work is wrong. That single move, who owns the outcome, quietly resets your margins, your defensibility, your hiring plan, and the multiple the market will pay for you. Most founders climb toward the biggest number without noticing they crossed the line where they stopped being a software company at all.

I have watched teams celebrate a switch to outcome pricing as a win, then spend the next year discovering they had signed up to run a services firm with venture expectations bolted on top. This post is the map I wish they had first. It is a ladder of what you can sell, a line that tells you when the economics flip, and a floor that tells you how high you can safely climb.

What this covers

The seat is dying, and the replacement is not free

For twenty years the software business ran on one unit: the seat. You sold access to a tool, per user, per month. The customer supplied the labor and kept the result. That model built the most profitable category in business history because the marginal cost of the next seat was close to zero. You wrote the software once and sold it a million times.

Agents break the seat. When one person with a good agent does the work of five, paying per person stops making sense to the buyer, and the vendor who insists on it leaves money on the table. So pricing is moving to the thing the buyer actually wants, the outcome. In customer support that shift is already public and specific. Fin charges about $0.99 per resolution with no platform fee. HubSpot moved its customer agent to $0.50 per resolved conversation, down from a dollar. Salesforce launched its help agent at $2.00 per conversation. Sierra runs a platform and setup cost north of $200,000 a year before per outcome charges even start. Same category, and the published per resolution rates already span from fifty cents to two dollars, a four times gap on what is supposedly the same unit of value.

This looks like a pricing change. It is not. It is a change in what you are selling, and that change reaches all the way down to the shape of your company. The seat sold access. The resolution sells work. When you sell work, the customer is no longer the one doing it, which means the customer is no longer the one responsible for it. You are. That is the whole story of service as software in one sentence, and it is the part the TAM math skips.

The stakes show up first in the one number founders defend hardest, gross margin. Classic software ran at 80 to 90 percent. AI application companies are landing at 50 to 60. ICONIQ pegged the average AI product gross margin at 52 percent early in 2026, up from 41 in 2024, which tells you the direction of travel but not the destination. Bessemer found fast scaling AI companies sitting near 25 percent gross margin in their early innings. The reason is simple and permanent: every unit of work you sell burns real compute, and that spend lands in cost of goods sold, not in fixed research and development. The more work you sell, the more it costs you to sell the next unit. That is not a software cost curve. I dug into why the old benchmark is gone in the piece on AI gross margins, and it is the ground this entire post stands on.

The Accountability Ladder: five things you can actually sell

Every AI product sits on one of five rungs. The rungs are not defined by how smart the model is or how much of the work is automated. They are defined by one question: when the work is wrong, whose problem is it? I call this the Accountability Ladder, because as you climb it, accountability for the outcome slides from the customer to you.

Services economics: each unit needs fresh delivery and riskSoftware economics: build once, sell many, marginal cost near zeroThe Accountability LadderThe Reuse Line1. Toolcustomer owns outcome2. Copilotcustomer still executes3. Supervised Agentshared accountability4. Managed Outcomeyou own delivery5. Guaranteed Outcomeyou own the downsideaccountability shifts to you

Rung one is the tool. You sell access to software and the customer does the work and keeps the result. This is the seat, the classic model, and it is not going away, it is just no longer the only option. Rung two is the copilot. You sell suggestions inside the flow of work. The human still drives and still owns whatever ships. A coding copilot that proposes a function is on this rung, because you never claimed the function was correct, you claimed it was a suggestion.

Rung three is the supervised agent. Now the agent does the work and a human approves it before it counts. Accountability is genuinely shared here, which is why this rung is both the most common landing spot for serious AI products and the most confusing one to price. Rung four is the managed outcome. You deliver finished work and you stand behind the delivery. A support agent that resolves the ticket end to end and bills per resolution lives here. Rung five is the guaranteed outcome. You get paid on the result, and if the result does not happen, you eat it. Contingency deals, revenue share, and risk bearing agents sit on this top rung.

The ladder is the map. The three ideas that make it useful are what the rungs hide: a line where the economics flip, a floor that caps how high you can safely go, and a set of budgets that get bigger and scarier as you climb. Take them one at a time.

Rung by rung: what changes as you climb

The mistake I see most often is treating the climb as a pure upgrade. Higher rung, bigger price, bigger market, done. But every single thing that made software a great business changes as you go up, and it changes in the wrong direction. Here is the full picture in one place.

Rung What you sell Who owns the outcome Pricing unit Budget you eat Margin reality
1. Tool Access to software Customer Per seat, per month Software budget 80 to 90 percent
2. Copilot Suggestions in the flow Customer Per seat plus usage Software budget 60 to 75 percent
3. Supervised Agent Work done, human approves Shared Per task or action Edge of software and labor 50 to 65 percent
4. Managed Outcome Finished work you deliver You (delivery) Per outcome, per resolution Salary and services budget 25 to 55 percent
5. Guaranteed Outcome The result, or no pay You (result and downside) Percent of value, contingency Risk budget Variable, can go negative

Read the margin column top to bottom. That descent is not a failure of execution. It is the price of the climb, baked in. The buyer pays more per unit at the top, which feels like progress, but you spend more to deliver each unit and you carry risk you never carried before. The revenue line looks better and the business gets harder. Two things move against you at the same time, and only one of them shows up in the headline number.

Notice also the pricing unit column. It is the tell. The moment your unit stops being “access” and becomes “a thing that happened,” you have crossed from selling potential to selling performance. A seat is a promise the customer will find value. A resolution is a claim that value was delivered. You can be wrong about the second one in a way you can never be wrong about the first, and being wrong now costs you money instead of costing the customer time.

The Reuse Line: where software economics end

Somewhere on this ladder you stop being a software company. Not in spirit, in economics. I call the boundary the Reuse Line, because the thing that flips is reuse. Below the line, you build an artifact once and sell it many times, and the next unit costs you almost nothing. Above the line, every unit of value you sell requires fresh delivery, fresh compute, fresh review, and sometimes a fresh human. The artifact stops being the product. The delivery becomes the product.

The Reuse Line sits between rung three and rung four for most categories. At rung three you are still shipping a product that a customer operates, even if a person on your side reviews edge cases. At rung four you are delivering finished work, which means someone or something on your side touches enough of each unit that your cost scales with volume. Cross that line and your business obeys services math, no matter how much of the work is done by a model. The compute in your cost of goods sold is the visible half. The invisible half is the quality assurance, the exception handling, and the account management that a delivered outcome demands and a sold tool does not.

This is why the gross margin data is not a temporary dip that better models will fix. Better models lower the compute cost per unit, which helps, but they do not move you back below the Reuse Line, because the review and the liability do not disappear when the model improves, they just get rarer and more expensive per incident. The founders who think a cheaper model will restore software margins on a rung four product have misread which cost is killing them.

Signal Below the Reuse Line (software business) Above the Reuse Line (services business in software clothing)
Cost of the next unit Near zero, you reuse the artifact Rises with volume, compute plus review per unit
What scales The product The team and the compute bill
Gross margin 70 percent and up 25 to 60 percent
Where the moat lives Product and data Delivery quality and trust
Who you hire to grow Engineers Delivery, review, and operations
How the market values you Software multiple Services multiple

That last row is where the money is made or lost. The public markets already know the difference, which is part of what the 2026 revaluation was about. A dollar of revenue from a reusable product is worth far more than a dollar of revenue from delivered work, because one compounds and the other has to be re earned every time. If you price and pitch like a software company but operate above the Reuse Line, you will eventually get repriced as what you are. The switching dynamics that protect a real product do not protect delivered work the same way, which is a theme I traced in the commoditization clock.

The Verifiability Floor: how high you can safely sell

If climbing is dangerous, why does anyone go above rung three? Because the price at the top is set against a much bigger budget, and for some outcomes the climb is worth it. The question is not whether to climb but how high you can climb without buying a liability you cannot control. The answer is a floor, and it is set by verification, not by model capability.

Here is the rule. You can safely sell an outcome only as high as you can verify that outcome cheaper than you are paid for it. If checking whether the work is correct costs you more than the margin on the sale, you have not sold a product, you have sold a promise you cannot afford to keep. I call this the Verifiability Floor. It is the highest rung where your cost to confirm the result stays below your price for the result.

The Verifiability Flooreasy to verifyhard to verifyhow high you sell (Tool to Guaranteed Outcome)Safe zoneverify cheaper than you are paidLiability zonehigh rung, cannot check the result cheaplysupport resolutioncheckable, safe highlegal or financial callhard to check, risky highThe floor is the diagonal: it rises as your outcome gets harder to verify.

Customer support is the category everyone points to because the outcome is unusually checkable. Did the customer get their answer and stop asking? You can measure that at scale, cheaply, and even the vendors know it. Zendesk shipped a change in 2026 where every billed resolution gets verified by the agent and by a separate evaluation model, with spam and routine chatter excluded. That is a company building its business exactly at its Verifiability Floor and no higher. The reason support agents can sell a rung four managed outcome and survive is that the outcome is cheap to confirm.

Now watch what happens when the outcome is not cheap to confirm. Some vendors count a resolution when the customer simply stops responding after a timeout, which quietly books an abandoned conversation as a success. Buyers have complained about abandoned chats billed as resolutions with no way to dispute them. That is what a business looks like when it sells above its Verifiability Floor: it is forced to fudge the definition of the outcome, because it cannot actually prove the outcome happened. The fudge works until a big customer audits it, and then the whole revenue line is suspect. If you cannot verify it, you will end up faking it, and faked outcomes are a lawsuit or a churn event waiting for a date.

This is also the cleanest reason two companies on the same rung can have wildly different fates. It is not that one has a better model. It is that one picked an outcome that is cheap to verify and the other picked one that is not. Verifiability, not intelligence, sets your ceiling. I made the technical version of this argument in the reliability paradox, where the same idea decides how much autonomy an agent can hold. Here it decides how much of the outcome you can sell.

The two budgets: software line, salary line, risk line

The reason the climb tempts everyone is the budget. Each rung reaches into a different pocket, and the pockets get dramatically bigger as you go up. This is the real content of the $4.6 trillion number. It is not a software number. It is a salary number.

The budget gets bigger as you climbSoftware linea few hundred billionpriced vs other toolsSalary lineabout $4.6 trillionpriced vs a person’s costRisk lineopen endedpriced vs cost of wrongbuyer: IT or opsbuyer: the P and L ownerbuyer: whoever owns downside

At rungs one and two you eat the software line. It is small and capped and it gets compared to every other tool the buyer already pays for. Your ceiling is whatever the category norm is, and the buyer is an IT or operations line manager who thinks in terms of tool budgets. This is a fine place to live, but it is a small pond, and everyone else in it is racing the same commoditization clock.

At rung four you eat the salary line. Now you are not priced against a tool, you are priced against the cost of the person whose work you replaced. That is a jump of ten times or more in what a single account can pay, and it is the entire reason service as software is exciting. But the buyer changes too. You are now selling to the person who owns that headcount and that budget, who will hold you to the standard they held the person to. Their expectations do not scale down because a model is doing the work. If anything they scale up, because they expect the machine to be more consistent than the human.

At rung five you eat the risk line, and the risk line has no ceiling, in either direction. You are priced against the cost of being wrong, which is wonderful when you are right and catastrophic when you are not. A contingency deal on a large outcome can pay more than any subscription ever would. It can also cost you more than you made, because you signed up to own the downside. Most founders who chase rung five have modeled the upside and never modeled the tail. I wrote about who actually funds this kind of ambition in the AI capital stack, and the short version is that risk line businesses need a very different balance sheet than software ones.

Owning the downside: the top rung is a risk book

The top of the ladder is not a bigger software business. It is a risk book with a software front end. When you sell a guaranteed outcome, you are underwriting, whether or not you use that word, and underwriting has its own brutal discipline that most software founders have never had to learn.

The liability is not hypothetical, and it does not wait for you to be ready. A Canadian tribunal decided a case in 2024 where Air Canada was held responsible for wrong information its chatbot gave a grieving customer about bereavement fares. The airline argued, in writing, that the chatbot was a separate entity responsible for its own statements. The tribunal rejected that flatly and said the company owns what its bot says, the same as any other page on its site. The dollar amount was small, a few hundred dollars, but the principle is enormous and it is now precedent: you own your agent’s output. When you climb to a rung where the agent acts on the customer’s behalf and you sell that action as an outcome, you have accepted every version of that liability at scale.

This changes what you have to build. A tool needs to work. A guaranteed outcome needs to work, and needs to be provably not your fault when it does not, and needs a reserve for when it is your fault anyway. That means contracts with real definitions, an appeals process, logging that would survive a dispute, and capital set aside against the tail. None of that is software. All of it is the cost of standing behind work. The founders who get burned are the ones who priced the outcome like a feature and discovered the liability like an accident. If you are going to run agents that act, read how they get pulled when they misbehave in the piece on decommissioning agents, because on rung five a misbehaving agent is not a bug ticket, it is an open position on your risk book.

There is a buyer side to all of this too. The customer who buys an outcome is making the mirror image of your decision, handing you accountability they used to hold. That is a bigger act of trust than buying a tool, and it is why the sale is slower and the proof burden is higher. I looked at the customer’s version of this calculation in hire versus automate, and the same logic that makes them cautious about firing a person makes them cautious about trusting your outcome.

Where the moat lives when you sell work

When you sold a tool, your moat lived in the product and the data around it. Features were hard to copy, integrations were sticky, and the data you gathered made the next result better. That is still true below the Reuse Line. Above it, the moat moves, and founders who keep defending the old one get surprised by how fast new competition shows up. How you set the price itself, as inference keeps getting cheaper, is a related but separate question that I covered in pricing under cheap inference. This section is about what protects you once you have picked a rung.

When you sell work, the customer is not comparing your feature list to a rival’s. They are comparing your outcome to the one they used to get from a person, and to the outcome a competitor’s agent promises. The model doing the work is largely the same model your competitor can call. So the scarce thing is not the intelligence. It is the trust that your outcome is correct, the proof that it was correct last time, and the consistency that lets a buyer stop checking your work. Trust, proof, and consistency are the moat on the work rungs, and none of them come from the model.

This is why a track record matters more here than it ever did with tools. A tool earns trust in a demo. An outcome earns trust over hundreds of delivered units without a bad one. That history is genuinely hard to copy, because a new entrant with the same model still has zero delivered units and zero proof. The moat is the ledger of outcomes you can stand behind, plus the verification machinery that lets you stand behind them cheaply. I made the broader version of this case in the data moat test: raw data is not an advantage, a working loop that turns data into a better and provable outcome is.

The flip side is that this moat decays the moment you ship a bad outcome at the top of the ladder, because trust is asymmetric. It takes hundreds of clean resolutions to build and one wrong guaranteed outcome to break. That fragility is exactly why the commoditization pressure I traced in the piece on where software’s moat moved hits work businesses differently. Your defense is not a feature others lack, it is a reputation others have not earned yet and could take from you with one strong quarter of their own delivery.

The wrapper: keeping software economics on a services rung

The best service as software companies do something that sounds contradictory. They price at the salary line but they fight to keep as much of the work as possible below the Reuse Line. The tool for that fight is the wrapper, the software you build around the work so that the reusable part stays reusable and only the genuinely variable part costs you per unit.

Start by splitting the work into two piles. One pile is the part that is the same across every customer and every unit: the pipeline, the standard steps, the checks, the formatting, the common cases. Build that once, reuse it, and it stays software. The other pile is the part that is truly different each time: the odd exception, the judgment call, the messy input. That pile is where your cost per unit lives, so the whole game is shrinking it. Every exception you turn into a standard step moves work from the expensive pile to the cheap pile, and drags your margin back toward software territory.

The second move is to make the outcome standard enough that verifying it is cheap. A narrow, well defined outcome is both easier to deliver reliably and easier to check, which means it sits higher under your Verifiability Floor. This is the practical reason winners sell narrow outcomes rather than broad ones. A narrow outcome keeps the reusable pile large and the verification cost small at the same time. The instinct to widen the promise to grow the market usually does the opposite of what founders hope, because it inflates the variable pile and pushes the outcome above the floor.

The third move is to keep humans only where the work genuinely cannot be standardized, and to treat every hour of that human time as a signal of where to build next. The human touch is not the product, it is a to do list for the wrapper. I wrote about protecting that irreducible core in the AI native founder playbook, and the same logic applies here. The parts you cannot compress are where your real work is, and the parts you can compress belong in software. Do this well and you can charge against the salary line while running much closer to software margins than your rung would predict, which is the entire prize worth chasing in this shift.

The contrarian take: winners stop one rung early

Here is what most of the market has backwards. The prevailing move is to climb as high as you can, because the budget is bigger up there and outcome pricing earns a richer valuation story. The race is upward. I think the race is a trap, and the companies that win the next decade of this will deliberately stop one rung below where they could go.

The rung that maximizes the value of your company is not the highest rung you can technically reach. It is the highest rung that stays below your Verifiability Floor and keeps you under the Reuse Line for as much of your revenue as possible. Climb past your Verifiability Floor and you do not become an AI company with a bigger market, you become a staffing firm with venture burn and a lawsuit calendar. Climb past the Reuse Line without meaning to and you keep software expectations while running services economics, which is how you end up with a great revenue chart and a valuation that keeps getting cut.

The strongest businesses I see in this shift are doing something specific. They sell the outcome where it is cheap to verify, and they refuse to sell the outcome where it is not, even when a customer offers to pay for it. They build the software wrapper so that the reusable part of the work stays reusable and only the truly variable part touches human hands. They price against the salary line but they operate with as much of the product below the Reuse Line as they can engineer. In other words, they treat the ladder as a menu with a ceiling, not a staircase to sprint up.

This is the same discipline that separates a real data advantage from a pile of logs, which I broke down in the data moat test. The pattern repeats: the winners are not the ones who claim the most, they are the ones who claim exactly what they can defend and not one inch more. Ambition in this game is not how high you climb. It is how precisely you choose your rung.

What to do Monday morning

This is not a theory you file away. It is a set of five checks you can run on your own business right away, and each one produces a decision, not a feeling.

First, locate your current rung. Ignore your marketing and look at your invoice. What is the unit you actually bill? Access is rung one or two. A task or action is rung three. A resolution or a delivered outcome is rung four. A share of the result is rung five. Be honest, because everything downstream depends on knowing where you really stand, not where the pitch says you stand.

Second, compute your Verifiability Floor. For the outcome you sell or want to sell, ask what it costs you to confirm it happened, correctly, without the customer’s help. If that cost is a fraction of your price, you have room to climb. If it approaches your price, you are at your floor. If it exceeds your price, you are already above it and should either fix verification or step down a rung before it catches up with you.

Third, find your Reuse Line. Look at the cost of delivering your next hundred units. If it is basically flat, you are below the line and you have software economics, protect them. If it rises with volume, you are above the line and you are a services business, so plan hiring, margins, and fundraising as one. Do not let the pitch deck decide this. Let the cost curve decide it.

Fourth, price the rung, not the tool. If you have climbed to the salary line, stop benchmarking against other software and start benchmarking against the cost of the work you replaced. Underpricing a rung four outcome against rung two tools is one of the most common ways AI companies leave the majority of their value on the table. The budget is bigger up there, so charge into it, as long as you can verify what you are charging for.

Fifth, choose your target rung on purpose and build for it. Decide the highest rung that stays inside your Verifiability Floor, and architect the product so the reusable parts stay reusable and only the genuinely variable work costs you per unit. If you decide to go above the Reuse Line, do it with open eyes, staff the delivery function, and hold reserve against the tail. The worst outcome is climbing by accident. The best is climbing exactly as far as your ability to verify allows, and stopping there with intent.

Service as software is a real shift and the budget behind it is real. But the phrase hides the cost, and the cost is accountability. You do not get paid for the software, you get paid for the work your software is trusted to own. Sell an outcome you cannot verify, and you have traded a software business for a liability. Pick your rung like the whole company depends on it, because it does.

FAQ

What is service as software?

Service as software is a model where a company uses AI to perform work internally and sells the finished outcome, instead of selling a tool that helps the customer perform the work. Foundation Capital coined the term and sized the opportunity at about $4.6 trillion, the pool of salaries and human services that software can now go after, next to a classic software market worth a few hundred billion. The core change is what you sell: work, not access.

How is service as software different from SaaS?

SaaS sells access to a tool per seat, and the customer does the work and owns the result. Service as software sells the result itself, priced per outcome. The deeper difference is accountability. In SaaS the customer is responsible when the work is wrong. In service as software you are, which resets your margins, your hiring, your defensibility, and how the market values you.

Why are AI company gross margins lower than SaaS margins?

Because every unit of work you sell burns real compute, and that spend lands in cost of goods sold rather than fixed research and development. Classic SaaS ran at 80 to 90 percent gross margin. AI application companies land closer to 50 to 60 percent, and fast scaling ones have been seen near 25 percent early on. The more work you sell, the more the next unit costs, which is services math, not software math.

What is the Reuse Line?

The Reuse Line is the point on the accountability ladder where software economics end and services economics begin. Below it, you build an artifact once and sell it many times at near zero marginal cost. Above it, each unit of value needs fresh delivery, compute, and review, so your cost scales with volume. It usually sits between the supervised agent rung and the managed outcome rung.

What is the Verifiability Floor?

The Verifiability Floor is the highest rung you can safely sell, set by verification rather than model capability. You can sell an outcome only as high as you can confirm that outcome cheaper than you are paid for it. If checking the result costs more than your margin, you have sold a promise you cannot afford to keep, and you will be tempted to fudge the definition of the outcome.

Is outcome-based pricing better than per-seat pricing?

It can be, because it ties your revenue to the metric the buyer cares about and reaches a much larger budget. It is also riskier, because it makes you accountable for the result and only works when the outcome is cheap to verify. Per resolution rates in customer support already range from about fifty cents to two dollars, so even within one category the right number depends on how checkable the outcome is.

Who is liable when an AI agent makes a mistake?

The company running the agent, in the cases decided so far. A Canadian tribunal held Air Canada responsible in 2024 for wrong information its chatbot gave a customer, and rejected the argument that the chatbot was a separate entity. The higher you climb the accountability ladder, the more of that liability you own, which is why guaranteed outcomes function as a risk book, not just a product.

Should a startup climb to the highest rung of the ladder?

Usually not. The rung that maximizes company value is the highest one that stays below your Verifiability Floor and keeps most of your revenue under the Reuse Line. Climbing past your ability to verify turns you into a staffing firm with venture burn, and climbing past the Reuse Line by accident gives you software expectations on services economics. Winners tend to stop one rung early on purpose.