AI Decision Fatigue: The Hidden Cost of AI Oversight
I watched a founder run six AI tools at once last month. One for code, one for copy, one for research, one for email, two for whatever the first four missed. She shipped more than any solo founder I know. She also looked exhausted in a way that had nothing to do with hours worked. By 4pm she could not decide what to have for lunch, let alone which of the six outputs in front of her was actually good. She had automated the making and kept all of the deciding, and the deciding was eating her alive.
There is a name for this now. A Harvard Business Review study from researchers at BCG and UC Riverside surveyed almost 1,500 workers and found that 14 percent of people using AI at work report what they call “AI brain fry,” the mental fatigue that comes from supervising machine output past the point your brain can hold it. Reporters at CNN, CBS, and Fortune picked it up because the finding is counterintuitive. The people most drained by AI are not the ones ignoring it. They are the ones babysitting it.
That is the durable story under the trend, and it is the one worth keeping. The tools will change. The names will change. What will not change is that AI moves the expensive part of your job from doing the work to checking the work, and checking is a different kind of tired. This post is about that shift, why it hits founders hardest, and how to run your own judgment like the scarce resource it has become. If you have felt busier and more productive and somehow more fried at the same time, you are not soft. You are paying a tax nobody put on the invoice.
Table of contents
- The real cost was never execution
- The Oversight Tax, and the Judgment Budget it drains
- The three loads that make oversight heavy
- Why checking is more expensive than making
- The Tool-Count Curve: why your fourth tool slows you down
- The Attention P&L: AI is a cost shift, not a cost cut
- The founder trap: you cannot delegate the oversight of your delegation
- The contrarian take
- What to do Monday morning
- FAQ
The real cost was never execution
For most of the last two decades, the bottleneck in building anything was execution. You had an idea, and then you spent weeks turning it into working software, a finished draft, a real campaign. Execution was slow, so execution was where the cost lived. We hired for it, trained for it, and paid a premium for people who could translate intent into output.
AI collapsed that cost. A solo founder can now go from idea to working prototype in a day. The making got cheap, and when a cost falls that fast, everyone assumes the whole job got easier. It did not. The work did not disappear. It moved.
What is left when the machine does the making is the part the machine cannot do: deciding whether the output is any good, whether it is right, whether it fits the specific thing you were actually trying to build. That is judgment, and judgment did not get cheaper. If anything it got more expensive, because now there is far more output demanding it. It is the skill I called the real moat in the taste moat, and AI raised its price rather than lowering it. I wrote about the flip side of this in what to build when building is free: when building costs nothing, choosing what to build becomes the whole game. Oversight is the same principle pointed at your own attention. When output costs nothing, judging output becomes the whole job.
Here is the trap. Making is generative. It has momentum, a sense of progress, the small hit of a thing coming into existence. Checking is the opposite. It is evaluative, adversarial, and thankless. You read something you did not write, hunt for what is wrong with it, and the reward for doing it perfectly is that nothing bad happens. Your brain treats those two activities very differently. One feels like building. The other feels like grading a stack of papers that never ends. AI did not save you the papers. It printed more of them and handed you the red pen.
The Oversight Tax, and the Judgment Budget it drains
Let me give you the two ideas the rest of this rests on. The first is the Oversight Tax. The second is the Judgment Budget it draws down. Together they explain why the founder shipping the most can also be the founder thinking the worst by dinner.
The Oversight Tax is the cognitive cost of supervising AI. Every output an AI hands you has to be read, judged, and either accepted, corrected, or thrown out. That act is not free. It costs attention and it costs decision energy, and the bill grows with three things: how much you have to verify, how many tools you are switching between, and how much raw output is coming at you. It shrinks with one thing: how clearly you specified what “correct” means before the machine started. Hold that shape in your head, because it is the whole lever.
The Judgment Budget is the other half. You do not have unlimited good decisions in a day. Psychologists have studied this for years under the label decision fatigue, and the pattern is consistent. The quality of your choices degrades as you make more of them. Your first few calls of the day are sharp. Your fiftieth is a coin flip you dressed up as a decision. You have a finite daily supply of high-quality judgment, and it drains whether you spend it on something that matters or on approving an AI-drafted Slack message.
Now put the two together. AI multiplies the number of decisions demanded of you without multiplying the supply of judgment you have to meet them. That is the entire problem in one sentence. The machine can generate a thousand things to decide about. You can still only decide well a couple of dozen times before the quality falls off. The Oversight Tax is what you pay. The Judgment Budget is the account it is drawn from. And most founders are running that account into overdraft every single day and calling the overdraft “productivity.”
The HBR researchers found the sharpest version of this in the data. Workflows where people mainly oversee AI, meaning they review, correct, and interpret model output, carried 14 percent more mental effort, 12 percent more fatigue, and 19 percent more information overload than workflows where AI simply replaces a task and hands back a finished result. The most draining thing you can do with AI is not use it. It is supervise it. That distinction is the whole ballgame, and almost nobody designs their day around it.
The three loads that make oversight heavy
“Oversight is tiring” is true but useless. To do anything about it you have to break the tax into its parts, because each part has a different fix. There are three, and once you can name which one is hitting you, you can actually cut it.
| The load | What it is | What drives it up | The tell |
|---|---|---|---|
| Verification load | The effort to decide whether one output is correct. | Vague specs, high stakes, output that looks confident but may be wrong. | You reread the same paragraph three times and still are not sure. |
| Switching load | The cost of moving your attention between tools and contexts. | More open tools, more tabs, more half-finished threads. | You forget what you were checking the moment you switch back. |
| Decision-volume load | The drain from the sheer number of small approve or reject calls. | More output, faster generation, no batching. | Trivial choices start to feel weirdly hard by afternoon. |
Verification load is the one people notice, because it feels like work. You are staring at a block of AI-written code or a market analysis and trying to decide if it is right. The catch is that AI output is engineered to look correct. It is fluent, formatted, and confident whether or not it is true. Fluent-but-possibly-wrong is the single most expensive thing a human can be asked to evaluate, because you cannot skim it. You have to actually reconstruct the reasoning to catch the one number that is off. That reconstruction is the synthesis work AI still cannot do for you, and it is exactly where your oversight time goes. I made this exact argument for machine reliability in AI agent observability: the failures that cost you are the silent ones, the outputs that look fine and are not. Every silent failure mode is a verification load your brain has to carry.
Switching load is the one people underrate. Every time you jump from your coding tool to your writing tool to your research tool, you pay a reset cost. Your working memory dumps and reloads. The BCG research puts a number on the aggregate effect, which I will get to, but the mechanism is old and well understood: context switching is expensive, and supervising many AI tools is context switching wearing a productivity costume. Each tool has its own prompt style, its own quirks, its own way of being wrong. You are not just checking outputs. You are re-learning the personality of a different machine several times an hour.
Decision-volume load is the quiet killer. It is not any single decision. It is the count. AI is a decision-generating machine. Accept this draft or regenerate. Keep this function or refactor. Use this framing or that one. None of these is hard on its own. Two hundred of them is a different animal. This is where the classic decision-fatigue research bites, and it is why the founder who felt sharp at 9am cannot pick a lunch spot at 1pm. The lunch decision is not the problem. It is the two hundred micro-approvals that came before it, each one a small withdrawal from the same account.
Why checking is more expensive than making
Here is the part that surprises people, and it is the part that turns this from a productivity gripe into something you can build a strategy around. The claim is not just that checking is annoying. It is that checking is structurally more expensive than making, and we have known this since 1983.
A cognitive psychologist named Lisanne Bainbridge wrote a short paper that year called “Ironies of Automation.” Her subject was industrial control rooms, not AI, but the finding transfers exactly. Her central point: the more you automate a system, the more demanding the human’s remaining job becomes, not less. When you take the routine work away from a person and leave them only with monitoring and the rare critical intervention, you have handed them the hardest possible task. You have to stay alert to something that is usually fine, ready to catch the moment it is not, using skills you no longer get to practice because the machine does the practicing.
Two of her observations matter for anyone running AI today. First, human vigilance collapses fast under passive monitoring. The research she drew on shows attention degrading significantly after about thirty minutes of watching a system you are not actively operating. You cannot supervise well for hours. Your brain is not built for it. Second, monitoring uses a completely different and more taxing mode than doing. When you make something, you are inside the task, and the task carries you. When you monitor something, you are outside it, holding a model of what “correct” looks like and comparing reality against it, continuously, with no momentum to ride. That is why an hour of reviewing AI output can leave you more wrung out than three hours of writing the same thing yourself.
The code world has already run this experiment at scale, and the numbers are brutal. One 2026 study found senior engineers spend an average of 4.3 minutes reviewing an AI-generated suggestion versus 1.2 minutes for human-written code. Same length, roughly four times the review time, because you cannot assume good intent from a machine that will confidently hand you a plausible bug. Teams using AI ship 21 percent more tasks and merge 98 percent more pull requests, but pull-request review time jumps 91 percent. The work did not vanish. It piled up at the review gate. And the METR study from 2025 is the one that should stop every founder cold: developers using AI believed they were 20 percent faster, and were actually measured 19 percent slower on real tasks. They felt the speed of generation. They did not feel the drag of verification, because the drag is invisible from the inside. That gap between felt productivity and real productivity is the Oversight Tax hiding in plain sight.
This is also why raising your specification quality is the highest-return move you can make, and it is the one lever in the Oversight Tax that points down. If you define “correct” sharply before the machine runs, verification gets cheap, because now you are checking against a clear standard instead of a vague feeling. If you define it loosely, every output becomes a fresh negotiation with yourself about whether it is good enough. I made the full case for this in spec-driven development. The spec is the cheapest place to catch an error, and it is also the cheapest place to lower your own oversight burden. A tight spec does not just improve the output. It protects your Judgment Budget.
The Tool-Count Curve: why your fourth tool slows you down
Now the number that should change how you set up your desk. BCG researchers looked at how productivity moves as people add AI tools, and the shape is not a line going up. It is a hill. One tool helps. Two tools help more. Three tools is the peak. Add a fourth, and self-reported productivity starts falling. Push to four or more and it drops off a cliff.
Read the curve as two forces fighting. Each tool you add gives you more raw output, which pulls the line up. Each tool you add also gives you one more thing to supervise, one more context to switch into, one more personality to track, which pulls the line down. For the first few tools, the output gain wins. Around the third, the two forces balance. After that, the supervision cost wins, and every tool you add past the peak is net negative even though it feels like more firepower. You are not getting more done. You are getting more to check.
The BCG and HBR work found the human cost that sits under the curve too. Workers who reported brain fry had 33 percent more decision fatigue and 39 percent more major errors than those who did not. Sit with that second number. The point of oversight is to catch mistakes. Past a certain tool count, the oversight itself becomes so draining that it manufactures more mistakes than it catches. You cross from supervising your tools to being supervised by them, and the reliability you were buying goes into reverse.
This is the practical heart of the whole post. The instinct when you feel behind is to add a tool. The curve says that past three, adding a tool is the thing making you feel behind. The move is not addition. It is subtraction. And subtraction is hard, because every tool has a champion in your feed telling you it is the one that changes everything. It might be. That does not mean your brain has room to supervise it.
The Attention P&L: AI is a cost shift, not a cost cut
The reason smart founders keep overloading is that they are running the wrong mental accounting. They see AI as a cost cut. Task used to take two hours, now it takes twenty minutes, book the savings. But that is only one line of the ledger. AI cuts your execution cost and raises your oversight cost at the same time. Whether you actually come out ahead depends on the net, and the net is not automatic.
| Line item | The cost-cut story you were sold | What actually lands on the ledger |
|---|---|---|
| Execution time | Drops hard. Hours become minutes. | True. This line really does fall, and it is the line everyone measures. |
| Oversight time | Roughly zero. The tool handles it. | Rises, often more than execution fell. This line is invisible on most dashboards. |
| Judgment reserves | Untouched. You saved energy. | Drawn down by every micro-approval, so the decisions that matter get your worst self. |
| Net effect | Pure gain. Do more with less. | A gain only if you cut decisions, not just add output. |
The trap is that the execution line is easy to see and the oversight line is nearly impossible to see. You can feel the two hours you saved. You cannot feel the extra forty minutes of checking spread across the day in twenty-second slices, or the fact that your 3pm strategic call was worse because your Judgment Budget was already spent approving drafts. So the ledger looks like a pure win, and you add another tool to win harder, and the invisible line quietly swallows the visible one.
This connects to something I argued in the AI efficiency trap. Cheaper per-unit almost never means cheaper in total, because falling cost raises volume, and volume brings its own bill. The attention version is exact. Cheaper output per task raises the number of tasks, and every task drags an oversight cost behind it. You optimized the cheap thing and multiplied the expensive one. The founders who win the AI era are not the ones with the lowest execution cost. They are the ones who keep their oversight line and their Judgment Budget on the same dashboard as their output, and manage the whole P&L instead of one flattering line of it.
The founder trap: you cannot delegate the oversight of your delegation
Everything so far applies to any knowledge worker. This section is why it is worse for founders specifically, and why “just hire someone to check the AI” does not save you the way it sounds like it should.
When you delegate work to a person, you can also delegate the judgment. You tell a senior engineer to own the checkout flow, and they own the deciding too. Their taste, their standards, their calls. Your oversight of them is light and periodic, because you are supervising a person who has their own judgment to spend. That is the whole point of hiring well. You buy back your own attention. I laid out when this trade makes sense in hire versus automate.
AI breaks that trade. When you delegate to a model, you keep all the judgment. The model executes but does not own the outcome, does not carry your standards, does not feel the stakes, will not push back when the brief is wrong. So the deciding does not transfer. It concentrates. Every AI you deploy routes its judgment calls back to exactly one place, which is you. Ten agents does not mean ten judgments distributed. It means one judgment, yours, now responsible for ten streams of output. You did not build a team. You built ten funnels all pointed at your own attention.
And you cannot fully escape by hiring a human to review the AI, because now you have to oversee the overseer. Did they actually check it or did they rubber-stamp it? Their incentive under volume is to skim, because skimming feels like keeping up. So you are back to supervising, one level removed, and often the removal just adds a layer of false comfort. This is the founder-shaped version of Bainbridge’s irony. The more of your company you automate, the more the few remaining human judgments matter, and the more of them collapse onto the founder who cannot practice every skill they are now the last check on. This is also why the apprenticeship gap is dangerous. If the machine does the reps, the humans never build the judgment to catch it when it is wrong, and the buck slides further up to you.
The way out is not to oversee everything better. You cannot. It is to be ruthless about what deserves your judgment at all, and to let a lot of low-stakes output ship with almost no oversight on purpose. That is a strategic choice, not a lazy one, and most founders never make it explicitly. They oversee everything a little, which is the worst setting, because it spends the Judgment Budget evenly across things that do not matter and things that decide the company.
The contrarian take
Here is what most of the discourse gets wrong. AI is pitched as a multiplier on your attention. Point it at your inbox, your codebase, your research, and free yourself to think about the big things. The pitch is that it hands your attention back to you.
It does the opposite. Every tool you add is a new thing to supervise, and supervision is the exact resource the pitch promised to free. AI does not give you attention back. It borrows against it. It is a loan, and like any loan it feels like money in hand right up until the bill arrives, which for attention arrives at about 4pm as the inability to make one more good call. The scarce resource was never execution. It was judgment. And judgment is the one thing you cannot buy more of by adding tools, because every tool spends it rather than supplies it.
So the founders who are actually winning this era are not the ones running the most agents. They are the ones supervising the fewest things, deliberately. They picked their three tools and refused the fourth. They wrote specs tight enough that verification got cheap. They decided in advance which decisions get their real judgment and let the rest ship at speed with a shrug. It is part of why individuals keep outrunning big companies here, the pattern I traced in the AI adoption paradox: fewer things to supervise means faster, better judgment. They treat their own attention as the constraint the whole company is built around, not an unlimited input to be spent chasing every new capability. The person drowning in six tools thinks the person with three is behind. The person with three has just noticed that the game was never how much you can generate. It was how well you can decide, and how long you can keep deciding well.
The hardest part to accept is that better models do not fix this. A smarter model produces better output, which lowers your error rate, which is good. But it does not lower the number of decisions demanded of you, and by making the output more trustworthy it can quietly make you check less carefully, which raises your exposure at exactly the moment it feels safest. The better the AI gets, the more the bottleneck becomes you. That is not a temporary state on the way to full automation. For anything that carries real stakes, it is the destination. I made the trust version of this case in the trust gap, and it holds here: knowing when to spend your judgment is the skill, and no model upgrade hands it to you.
What to do Monday morning
Enough diagnosis. Here is what I would actually do, and what I have watched work for founders who climbed out of the six-tool hole.
Run a tool audit and cap it at three. List every AI tool you touched last week. Rank them by how much they actually moved your week, not how impressive they are. Keep the top three for active daily use. Everything else goes to occasional or gone. The Tool-Count Curve is not a suggestion. Your fourth active tool is very likely net negative, and killing it will feel like falling behind for about two days and like relief after that.
Batch your verification. The switching load is self-inflicted when you check each output the instant it appears. Instead, let outputs queue and review them in dedicated blocks. Generate in the morning, verify in a focused afternoon window, so your brain holds one “checking” context instead of whipping between making and checking two hundred times. And keep those blocks under thirty minutes with real breaks, because that is roughly where vigilance falls off a cliff. You cannot review well for two hours straight no matter how disciplined you are.
Set a Judgment Budget and defend it. Decide how many genuinely high-stakes decisions you will make in a day. For most people the honest number is small, maybe three to five. Protect those slots. Make them early, before the micro-approvals have drained the account. Everything below that line either gets automated, delegated with the judgment attached, or shipped with deliberately light oversight. The goal is to stop spending your best decisions on your least important ones.
Raise your specs, because it is the only lever that points down. Before you hand a task to AI, write what “correct” looks like in enough detail that checking becomes comparison instead of negotiation. This is the single highest-return habit in the whole system, because it attacks verification load at the source. A sharp spec turns a draining judgment call into a quick match against a standard. A vague one turns every output into a fresh argument with yourself.
Use an oversight dial, not an oversight switch. Not everything deserves the same scrutiny. Match your oversight to the stakes, and let the low-stakes stuff run nearly unwatched on purpose.
| Task class | Example | Oversight setting | Judgment-budget cost |
|---|---|---|---|
| Reversible, low stakes | Internal notes, first-draft copy, throwaway scripts. | Ship with a glance. Fix later if it matters. | Near zero. Refuse to spend judgment here. |
| Repeatable, medium stakes | Customer emails, routine code, analysis you will act on. | Spot-check against a spec or checklist. | Moderate. Systematize so it stays cheap. |
| Irreversible, high stakes | Pricing, legal, security, anything customer-facing and hard to undo. | Full human judgment, rested, unrushed. | High, and worth it. This is what the budget is for. |
Do these five and the tax does not disappear, because it cannot. Oversight is the real job now. But you stop paying it on things that do not matter, you stop paying it in the most expensive way, and you keep enough of your Judgment Budget in reserve for the decisions that actually decide whether the company works. That is the whole objective: not to check less because you care less, but to spend your judgment where it changes the outcome. If you want the wider system this sits inside, it is part of the founder operating system I keep coming back to, and it pairs with what to learn in the AI era, because the meta-skill under all of it is knowing where your attention is worth the most.
FAQ
What is AI decision fatigue?
AI decision fatigue is the drop in decision quality that comes from making too many judgment calls about AI output in a day. Every accept, reject, or correct is a small withdrawal from a finite daily supply of good decisions. Because AI generates far more things to decide about than it removes, it pushes many people into overdraft, where their later choices get sloppy. It overlaps with what researchers now call AI brain fry, the broader mental fatigue from supervising machine output past your cognitive capacity.
How many AI tools should a founder actually use?
Research from BCG suggests productivity rises through the first two tools, peaks around three, and declines at four or more, because each additional tool adds output but also adds a supervision and context-switching cost. A practical rule is to keep three AI tools in active daily use and be ruthless about everything past that. The fourth tool usually feels like more capability while quietly making you slower.
Why does reviewing AI output feel harder than doing the work myself?
Because it is a different and more taxing mode of thinking. Making something has momentum and carries you through the task. Checking something means holding a standard in your head and comparing external output against it continuously, with no momentum to ride. Cognitive research going back to Lisanne Bainbridge’s 1983 work on automation shows monitoring is structurally more draining than doing, and that human vigilance degrades sharply after about thirty minutes of passive watching. AI output makes it worse because it looks confident whether or not it is correct, so you cannot safely skim it.
Does AI actually make me slower?
Sometimes, and you often cannot feel it. A 2025 METR study found developers using AI believed they were about 20 percent faster while measuring roughly 19 percent slower on real tasks. The generation speed is visible and the verification drag is invisible from the inside, so felt productivity runs ahead of real productivity. Whether AI nets out faster for you depends on your oversight cost, which most people never measure.
What is the Oversight Tax?
The Oversight Tax is the cognitive cost of supervising AI. It rises with three things, how hard each output is to verify, how many tools you are switching between, and how much raw output you are producing, and it falls with one thing, how clearly you specified what correct means before the machine ran. It is a tax because it is real, it is paid in attention and decision energy, and it never shows up on the invoice that made AI look free.
What is a Judgment Budget?
A Judgment Budget is the idea that you have a finite daily supply of high-quality decisions, and it drains whether you spend it on something important or trivial. The move is to decide in advance how many high-stakes calls you will make, protect those slots by making them early and rested, and push everything below that line into automation, real delegation, or deliberately light oversight, so your best judgment is not spent approving AI-drafted Slack messages.
Can I just hire someone to check the AI for me?
Partly, but less than you would hope. When you delegate to a person you can hand over the judgment along with the task. When you delegate to AI the judgment stays with you, and if you hire a human to review the AI you now have to oversee the overseer, whose incentive under volume is to skim. It helps for repeatable medium-stakes work with a clear checklist. It does not remove the founder as the last real judgment on the decisions that matter most.
Is this the same as regular burnout?
It is related but more specific. Regular burnout usually tracks total hours and emotional load. AI decision fatigue tracks the volume and mode of oversight, and it can hit even in a week where you worked fewer hours, because the exhausting part is the density of judgment calls, not the clock. The tell is feeling productive and drained at once, and finding small decisions strangely hard by afternoon. The fix is structural, cutting the number of things you supervise and how you supervise them, not just resting more.