The Data Moat Test: When Data Is an Asset
A dead airline just sold its email inbox for ten million dollars.
Spirit Airlines stopped flying in May 2026. When the estate went to auction, Google paid ten million dollars for roughly a hundred million internal emails, five hundred million Teams chats, and the OneDrive and SharePoint files behind them. Google outbid Mercor, which offered seven and a half. The one thing Google did not want was the customer list. Passenger profiles and frequent flyer records were carved out of the sale. Google wanted the internal exhaust, the messy human back-and-forth of people running a company, not the thing Spirit spent two decades trying to own.
Sit with that for a second. The part of Spirit worth bidding on was not the brand, the routes, the loyalty program, or the fifty million customer relationships. It was the inbox. And it was only worth money once the company was dead.
That is the whole lesson, if you know how to read it. The data gold rush is real. Labs are running out of public text to train on and they are paying serious money for private corpora that were never on the open web. But most founders are reading the wrong signal. They see a price tag on somebody else’s data and conclude that their own pile of logs and records is a moat. It usually is not. A thing that only fetches a bid at your bankruptcy auction was never protecting you while you were alive.
So here is the durable version, stripped of the headline. Data has a price. Price is not a moat. And the gap between the two is where founders quietly waste years defending a hoard that would not survive a single motivated competitor. This is the data moat test, the four conditions that decide whether the data you have is an asset you can defend or just a cost you have not noticed yet.
Table of contents
- The goldmine that wasn’t
- The data moat test
- Lock one: can anyone else get it
- Lock two: does it rot
- Lock three: a pile or a loop
- Lock four: is it welded to a workflow
- The gold rush is a liquidation sale, not a moat market
- Two companies, same data, one moat
- What most founders get wrong about data
- What to do Monday morning
- Frequently asked questions
The goldmine that wasn’t
Every deck I have seen in the last three years has a version of the same line. We are sitting on a goldmine of proprietary data. It shows up on the moat slide, right after the team slide, as if the volume of rows in a database were self-evidently a wall around the business.
It is almost never true, and the reason is simple. Having data and defending with data are two different things. A goldmine is defensible because the gold is in the ground under land you own and nobody else can dig there. A hard drive full of records is not like that. If a competitor can scrape the same thing, buy the same thing, or generate a good-enough version of the same thing, then your pile is not a wall. It is a warehouse. And a warehouse costs money to keep.
This is the part founders miss. Data is not free to hold. Every record you store carries a running bill: storage, pipelines, engineers to keep it clean, and a fatter target for the day you get breached. Most company data is what I call exhaust, the byproduct of operating that piles up whether you want it or not. Server logs. Support tickets nobody reads twice. Half-finished CRM fields. Telemetry you collect because collecting it was the default. Exhaust is not an asset. It is a liability wearing an asset’s clothes, and the disguise holds right up until a regulator or an attacker sends you the invoice.
The stakes are not academic. When a founder believes the data pile is a moat, three bad things follow. Strategy gets lazy, because why sharpen distribution or product when the data supposedly protects you. Security gets loose, because a hoard you think is valuable is one you keep instead of delete, which grows the blast radius of every breach. And fundraising gets built on a claim that does not survive diligence, so the story cracks at the worst possible moment. I have watched teams pour a year into a data platform meant to widen a moat that a mid-size rival replicated in a quarter with a scraper and a checkbook.
The training-data boom makes this worse, not better, because it dangles proof. If Google will pay ten million for a dead airline’s emails, surely my live company’s data is worth a fortune. Maybe. But worth a fortune to a buyer at auction and worth a moat to you as an operator are not the same measurement, and confusing them is the single most expensive data mistake a founder can make. Related to what I argued in the commoditization clock: the thing you assume is protected is usually the thing about to get copied.
The data moat test
So how do you tell the difference between data that is an asset and data that is just expensive to keep? Run it through four locks. A pile of data is a moat only if it clears all four. Clear none, and what you have is exhaust. Clear some but not all, and you have something real but partial, an asset with a known expiry, which is worth knowing so you can price it and defend it honestly.
Here are the four locks, in the order they actually bite.
Read left to right, this is a filter, not a checklist you get to grade on a curve. Each lock removes data that cannot do the job. Proprietary removes anything a rival can also get. Fresh removes anything that rots faster than you renew it. Looped removes dead stock that sits there without improving the product. Wedged removes data that a departing customer or engineer can simply carry out the door. What comes out the far end is small, and that is the point. Real moats are narrow. The pile is wide because most of it is exhaust.
The four locks map onto a plain question each: can anyone else get it, does it rot, is it a pile or a loop, and is it welded to something. Let me take them one at a time, because the failure mode at each is specific and the fix is different.
| Dimension | Data exhaust | Data asset |
|---|---|---|
| Who else can get it | Anyone with a scraper or a budget | Only you, or a tiny set of holders |
| What time does to it | Rots, and you rarely renew it | Rots, but you renew it faster than rivals |
| How usage changes it | Sits there, a dead stock | Feeds a loop that improves the product |
| What it is attached to | Nothing, it is portable | A workflow the customer lives inside |
| On the balance sheet | A cost and a risk you carry | A compounding thing rivals cannot match |
Lock one: can anyone else get it
The first lock is proprietary, and it is where most data piles die. The test is not do you have the data. The test is whether you are one of the few who has it, and how long that stays true. I think of it as the replication clock, the amount of time it takes a motivated competitor to stand up their own copy of what you consider precious.
There are three ways a rival gets your data, and each runs a different clock. They can scrape it, if it leaks into public surfaces or lives on the open web. They can buy it, from a broker, a partner, or the same source you used. Or they can generate it, with synthetic data that is close enough for the job. If any of those clocks is short, your data is not proprietary, it is just currently uncopied, which is a very different and much weaker thing.
Scraping kills more data moats than founders admit. If your special corpus is product reviews, job listings, restaurant menus, or anything that shows up on the public internet, then your pile is a snapshot of something a crawler can take too. The web already ran out of easy training text, which is why labs are hunting private corpora at all, but public-facing data is still the cheapest thing in the world to duplicate. If a competitor can rebuild it with a weekend and a crawler, you do not own it. You are renting attention to it.
Buying is the clock that shortens fastest right now. There is a whole market of firms whose job is to sell the thing you thought only you had. Data brokers, labeling shops, and marketplaces have turned into serious businesses. Mercor, an expert-data marketplace, was in talks in mid-2026 to raise at around a twenty billion dollar valuation, up from ten billion less than a year earlier, on more than two billion in annual revenue and a network of over thirty thousand vetted professionals paid around ninety-five dollars an hour. When there is a liquid market for a category of data, no single holder in that category owns a moat. They own inventory. The a16z partners Martin Casado and Peter Lauten made this point years ago in an essay called The Empty Promise of Data Moats: for most companies, what looks like a data moat is really just data scale, and scale is buyable.
Generating is the newest clock and the one that scares hoarders the most. Synthetic data will not replace every kind of real signal, and the sharpest reasoning and rarest edge cases still need humans, which is exactly why frontier labs each spend on the order of a billion dollars a year on human data. But for a large middle band of use cases, a competitor can now manufacture a training set that is good enough to close the gap. If your data’s only advantage was volume, synthetic generation is the flood that levels it. This is the same defensibility problem I laid out in the AI wrapper trap: if the thing under you can be reproduced cheaply, sitting on top of it is not a position.
The honest version of lock one is a stopwatch, not a yes or no. Ask how many months of runway your data advantage really has before a rival with money and intent has their own. If the answer is measured in weeks, stop calling it a moat and start calling it a head start, and treat it accordingly.
Lock two: does it rot
The second lock is freshness, and it is the one that quietly turns yesterday’s asset into today’s cost. Almost all data decays. Prices change, people move, preferences drift, the world updates, and the corpus you were so proud of last year describes a reality that no longer exists. The investor Abraham Thomas put it plainly in his writing on data and defensibility: keeping a corpus fresh is a constant expense that grows with scale, not a one-time win you get to bank.
This is why the Spirit inbox is such a clean teaching case. A hundred million emails from a company that no longer operates is a perfect snapshot of a frozen moment. That is genuinely useful to a lab that wants raw human text to train on, because for training, a fixed pile of authentic conversation has value the day it is bought. But notice what it is not. It is not renewing. Nobody at Spirit is writing new emails. The corpus is a fossil, valuable to a paleontologist, useless as a living animal. A fossil can be sold. A fossil cannot compete.
The founder version of lock two is the renewal loop. Do you have a live source that refreshes this data faster than anyone else can refresh theirs. A maps company whose users report a closed road in real time has a renewal loop. A fraud system that sees new attack patterns the day they emerge has a renewal loop. A pile of last year’s transactions has a fossil. The difference is not the data, it is whether there is a pump behind it that keeps the water moving while your competitor’s puddle evaporates.
Freshness also flips the cost question. For a fossil, storage is pure expense, you pay to keep something that gets less true every day. For a renewal loop, the same spend buys you a moving target that rivals have to chase. The exact same line item is a cost in one case and an investment in the other, and the only thing that separates them is whether new data flows in faster than old data rots. When founders tell me their data is a moat, my first question is not how much do you have. It is how fast does it go stale, and what is your pump. If there is no pump, you are guarding a melting asset, which connects to the deeper point I made in pricing under cheap inference: things that feel permanent in an AI market are usually decaying under you.
Lock three: a pile or a loop
The third lock is the one everybody gets wrong when they say the words data network effect. The claim is that more data makes the product better, which attracts more users, who generate more data, which makes the product better still, a flywheel that a latecomer can never catch. It is a beautiful story and it is mostly false, because most of what founders call a data network effect is really just a data pile getting bigger.
Here is the distinction that matters. A stock is a pile that sits there. A loop is a pile where usage measurably improves the product on a dimension the customer actually feels, in a way a competitor starting cold cannot match. Adding rows to a stock gives you diminishing returns, the millionth record teaches the model almost nothing the nine hundred thousandth did not. A real loop is different, because the value is not in the volume, it is in the closed circuit between use and improvement. Casado and Lauten’s whole point was that scale gets mistaken for a network effect, and the economics run backward: the cost of adding genuinely new, useful data rises while the value of each extra record falls.
The cold-start test is how you tell them apart in one move. Imagine a well-funded competitor launches tomorrow with zero of your data. On the dimension your customer cares about, speed, accuracy, relevance, coverage, how far behind are they, and does the gap widen or close as you both run. If the honest answer is that they catch up in a quarter because they can buy or generate their way to parity, you have a stock. If the answer is that every day of your operation pushes the gap wider on something the customer can feel, and they cannot shortcut it, you have a loop. The loop is the only version of a data advantage that behaves like a moat, because it is the only one where standing still is not the same as falling behind.
Loops are rarer than the pitch decks suggest, and they tend to live in narrow places. Vertical products often have them, because the data is specific enough that no general pile substitutes for it, which is one more reason the winners in AI keep going narrow, a pattern I broke down in why vertical AI winners go narrow. If you cannot point to the loop, the closed circuit where use improves the product on a felt dimension, then you do not have a data network effect. You have a big table, and big tables are for sale.
Lock four: is it welded to a workflow
The fourth lock is the one founders forget, and it is often the one doing the real work. Data on its own is portable. It can be exported, copied, subpoenaed, or walked out the door by a customer who decides to leave. A moat has to survive that, and the way data survives it is by being welded to a workflow the customer lives inside every day.
Think about the difference between two companies that hold the exact same records. One stores them in a database and offers an API. The other has built the records into the daily motion of the customer’s job, the screens they open every morning, the approvals that route through it, the reports their own boss expects in that format, the muscle memory of a team that has done it this way for two years. The first company’s data can leave in a CSV. The second company’s data is load-bearing. Pulling it out means rebuilding how work happens, and that cost, not the data itself, is the wall.
This is the same insight behind switching costs generally, and it is why data plus workflow beats data alone every time. I wrote about the mechanics of this in the switching cost trap: the lock-in that matters is not the data you store, it is the work that would have to be redone to leave. A customer will tolerate a mediocre product to avoid re-teaching their whole team. They will not tolerate anything to move a spreadsheet.
The workflow wedge also explains why distribution keeps beating data on the moat slide. If you own the surface where the work happens, you own the place the data is created and consumed, and that position renews every one of the other three locks for you. Fresh data flows in because the work flows through you. The loop closes because usage and improvement happen in the same place. And proprietary holds longer because the data is entangled with a process nobody wants to rebuild. This is why I keep pointing founders back to distribution as the real front door, and to the argument in who your customer really is. The workflow is where a data advantage stops being portable and starts being a position.
Run the four locks together and you get a strict definition. A data moat is proprietary enough that rivals cannot cheaply copy it, fresh enough that it does not rot faster than you renew it, looped so that use improves the product on a felt dimension, and wedged into a workflow that makes leaving expensive. Miss any one and you have a weaker thing with a shorter life, which is fine as long as you know which one you have and stop paying moat prices for exhaust.
The gold rush is a liquidation sale, not a moat market
Now back to the ten million dollars, because the training-data boom is exactly the thing confusing founders, and it deserves a clear-eyed look. The demand is real and it is structural. Epoch AI, a research group that studies this, projects with high confidence that the stock of publicly available human text will be fully used up somewhere between 2026 and 2032. Models eat data faster than the internet makes it. So labs have turned to private corpora, and a market has formed around it. The AI training dataset market was worth roughly three and a half billion dollars in 2025, the data-labeling market is on a path from under four billion in 2024 toward seventeen billion by 2030, and content owners are cashing in. News Corp signed with OpenAI for more than two hundred fifty million dollars over five years. Reddit licenses to Google for a reported sixty million a year. Shutterstock’s AI licensing revenue reached around a hundred thirty-eight million in 2024.
Those numbers are real and they prove one thing precisely: data has a price. They do not prove your data is a moat, because a moat and a sale are answers to different questions. A sale asks what a buyer will pay for a static corpus once. A moat asks what protects your business day after day while you operate. The Spirit auction is a liquidation. The estate is selling off the last thing of value from a company that already lost. That is the opposite of a moat. A moat is what stops you from ending up on that auction block in the first place.
| Question | The sale (what labs pay for) | The moat (what protects you) |
|---|---|---|
| What the buyer wants | A big static pile of authentic human data | Nothing, a moat is not for sale |
| When it pays off | Once, at the moment of transfer | Every day you keep operating |
| Does it need to be fresh | No, a fossil is fine, even ideal | Yes, it must renew faster than rivals |
| What it protects | Nothing, it is being handed away | Your position against a live competitor |
| Spirit Airlines | Ten million at auction, after death | Zero, the data never saved the airline |
The trap in the gold rush is that it teaches founders the wrong lesson at the exact moment they are eager to hear one. The right lesson is not my data is worth money so I have a moat. It is my data might be worth money to a buyer, and that is a completely separate fact from whether it defends me. The founders who confuse the two start managing their company like an estate sale, hoarding every byte in case a buyer materializes, when they should be asking which slice of their data actually clears the four locks and building the rest of the business as if the data will not save them, because for most companies it will not. The funding side of this same mania, where the money chases the story faster than the fundamentals, is something I traced in the AI capital stack.
Two companies, same data, one moat
Make it concrete. Two startups sell scheduling software to medical clinics. Both hold years of appointment records, no-show histories, and provider availability. Same category, roughly the same volume of rows. One has a moat and one has exhaust, and the four locks tell you which is which.
Company A stores the records and reports on them. A clinic can export its history any time, the data describes last year’s patterns, and nothing about using the product makes the underlying data smarter. If a competitor launched tomorrow with a clean database and a good sales team, they would sign clinics on features and price, and the data gap would close in a quarter because the new clinics generate their own records fast. Company A has a stock. It is a fine business, maybe, but the data is not the wall. The wall, if there is one, is elsewhere.
Company B built the records into the clinic’s daily flow. The no-show model retrains on every appointment and gets measurably better at predicting which slots will empty, so the front desk overbooks smarter and revenue climbs, a dimension the clinic feels in dollars every week. The predictions live inside the booking screen the staff use all day, the reminders route through it, and two years of tuned behavior would have to be rebuilt to leave. New attack of a competitor does not close the gap, it watches it widen, because Company B’s loop compounds on something the customer cares about and cannot be bought in bulk. Company B has a moat, and notice it is not because it has more data. It has the same data, wired differently.
The map is worth keeping in your head because it separates two things founders mash together. The bottom-right quadrant, the bankrupt-airline pile, is data that is hard to replicate but static. It has real sale value, which is exactly why Google bid on Spirit, and zero moat value, which is exactly why Spirit went bankrupt holding it. The top-right quadrant is the only real moat, and the only way to get there is to add a live loop and a workflow to data that is genuinely hard to copy. Most companies think they are top-right. The four locks usually put them bottom-left, in exhaust, which is the one quadrant where the right move is to store less, not more.
What most founders get wrong about data
Here is the thing almost everyone gets backward right now. The data gold rush is being read as proof that data is the ultimate moat, when it is actually proof of the opposite. Labs are paying for data precisely because data by itself is a commodity that can be bought, and they are buying it the way you buy fuel, not the way you buy a fortress. Fuel has a price. A fortress does not, because it is not for sale and its whole value is that you keep it.
The a16z crew were right in 2019 and they are more right now: for most companies, the data moat is empty, and the real defenses are distribution, switching costs, and workflow. What the last few years added is not a refutation of that, it is a footnote. There is a narrow, real data asset, and it lives in the top-right corner of the map, proprietary data caught in a compounding loop and welded to a workflow. That asset is rare, it is small, and it is almost never the giant undifferentiated pile the moat slide is bragging about. The founders who win with data are not the ones with the most rows. They are the ones who found the one loop worth building and ignored the exhaust.
So the sharpest way to say it is this. If your data is only worth something at your bankruptcy auction, it was never a moat. It was inventory you had not sold yet. A moat is what keeps you from ever reaching the auction. And the test for which one you have is not how many terabytes you are sitting on, it is whether the data clears all four locks while you are alive and operating, not just whether a buyer would take it off your hands when you are gone.
Let me give the other side its due, because there is a real counterweight. A big static pile is not worthless, and pretending it is would be its own mistake. If you genuinely hold a rare corpus, the bankrupt-airline quadrant is a real, if one-time, source of value, and licensing it can fund the business or subsidize the loop you actually want to build. The error is not valuing the pile. The error is mistaking its sale price for a moat and running your strategy as if the hoard protects you. Sell the fossil if someone wants it. Just do not confuse selling a fossil with owning a living moat, and do not let the check convince you to stop building the thing that actually keeps competitors out.
What to do Monday morning
Enough theory. Here is how to run the data moat test on your own company this week, and come out with a strategy instead of a slogan.
First, list what you actually have. Not the aspirational data platform, the real tables, logs, and records sitting in your systems right now. Write them down in plain language: appointment histories, support transcripts, usage telemetry, whatever it is. You cannot audit a hoard you have never inventoried.
Second, run each item through the four locks and be brutal. Can anyone else get it, by scraping, buying, or generating, and how many months until they do. Does it rot, and do you have a live pump that renews it faster than rivals. Is it a dead stock or a loop where use improves the product on a dimension the customer feels. Is it welded to a workflow, or could a leaving customer export it in a CSV. Mark each item with how many locks it clears. Most will clear zero or one. That is normal, and it is the first honest picture of your data most teams ever draw.
Third, act on the quadrants. The exhaust, the stuff that clears zero locks, is a liability, so store less of it, delete what you can, and shrink the blast radius and the bill. Do not defend it, prune it. The bankrupt-airline piles, rare but static, are for selling or licensing, not for building strategy on, so price them and move on. And the one or two things that could reach the top-right corner, that is where your data investment goes, all of it, into closing the loop and wedging it deeper into the workflow.
Fourth, fix your own narrative. If your pitch, your strategy doc, or your own head says the data is the moat, replace it with the specific true version: here is the one data loop we own, here is why a rival cannot cheaply copy it, and here is the workflow it is welded to. Everything else is exhaust, and we treat it like exhaust. That single edit will make your strategy sharper and your diligence survivable, and it will stop you from pouring a year into defending a pile that a competitor would replicate in a quarter. This is the same discipline I argued for in finding the incompressible core: know exactly which small thing is load-bearing, and stop pretending the rest is.
The founders who get this right end up with less data and a stronger position, because they stopped confusing volume with defense. For the wider map of where the real defensible positions sit in an AI market, the whole argument connects back to the AI-native founder playbook, and to the way even hard assets are shifting in the atoms premium. Data can be part of a moat. It is almost never the moat by itself, and knowing the difference is worth more than another terabyte.
Frequently asked questions
What is a data moat?
A data moat is a data advantage that actually protects a business from competitors over time, not just a large volume of data. To count as a moat, the data has to be proprietary enough that rivals cannot cheaply copy it, fresh enough that it does not rot faster than you renew it, part of a loop where use improves the product on a dimension customers feel, and welded to a workflow that makes leaving expensive. Most company data fails at least one of those tests and is therefore not a moat.
Is proprietary data really a competitive moat?
Sometimes, but far less often than founders assume. Proprietary only means uncopied right now, and the value depends on how long that stays true. If a competitor can scrape the same public data, buy it from a broker or marketplace, or generate a good-enough synthetic version, then your data is not defensible even if you are currently the only one holding it. Proprietary is the first lock, not the whole moat.
Why did Google pay ten million dollars for a bankrupt airline’s data?
Google bought bankrupt Spirit Airlines’ internal emails, chats, and files, around a hundred million emails and five hundred million Teams messages, to train AI models, because the public web is running low on fresh human text and a large authentic corpus has training value. Notice that Google skipped the customer list and wanted the internal exhaust. It is a one-time purchase of a static pile, which is a sale, not a moat. The data was worth money at auction and never protected Spirit while it operated.
Do data network effects actually exist?
They exist, but they are rarer than the pitch decks claim. Most so-called data network effects are really data scale effects, which means more rows with diminishing returns rather than a compounding loop. A true data network effect requires that usage measurably improves the product on a dimension the customer cares about, in a way a competitor starting from zero cannot quickly match. If adding data mostly just makes the pile bigger, that is scale, not a network effect, and scale is buyable.
What is the difference between a data stock and a data loop?
A stock is a pile of data that sits there, where the millionth record teaches you almost nothing new and value flattens then decays. A loop is a closed circuit where using the product generates data that improves the product on something the customer feels, which drives more use and more data. A stock is for sale and gives diminishing returns. A loop compounds and is hard to catch. Only the loop behaves like a moat.
How do I know if my company’s data is valuable?
Run the four locks. Ask whether a rival can cheaply get the same data, whether it rots faster than you renew it, whether use improves your product in a way a cold-start competitor cannot match, and whether it is bound to a workflow that makes leaving costly. Count the locks each dataset clears. Data that clears all four is a moat, data that clears one or two is a partial asset with a known expiry, and data that clears none is exhaust you should store less of.
Can AI companies just use synthetic data instead of buying real data?
For many use cases, yes, and that is exactly why volume-only data advantages are weak. Synthetic data can close the gap for a large middle band of tasks. It struggles with the rarest edge cases and the sharpest human reasoning, which is why frontier labs still spend on the order of a billion dollars a year each on real human data. If your data’s only edge was quantity, synthetic generation is the flood that levels it. If your edge is a proprietary live loop, synthetic data does not substitute for it.
If labs are paying for data, doesn’t that mean my data is a moat?
No. It means data has a price, which is a different fact from whether your data defends your business. A price is what a buyer pays for a static corpus once. A moat is what protects you every day you operate. Selling data is a liquidation event, most visible when a company is already dead, as with Spirit Airlines. If your data is only worth something at your bankruptcy auction, it was inventory, not a moat.