The Reps Problem: Why AI Delegation Erodes Judgment
Entry-level jobs are collapsing, and every headline blames the robots. The real story is quieter and it applies to you too, even if you run a company of one.
I watched a founder ship a payments integration in a weekend using an agent. Clean code, tests passing, live by Sunday night. Six weeks later a currency rounding bug ate a few thousand dollars in reconciliation errors, and he could not find it. Not because he is not smart. Because he had never once written a payments integration by hand, so he had no feel for where they rot. The agent gave him the output. It could not give him the thing that would have let him catch the output when it was wrong.
That gap has a name now, and it is showing up in the labor data before it shows up in anyone’s self-assessment. Entry-level job postings in the US are down about 35 percent since early 2023. Junior software and data roles have fallen by as much as 67 percent. A Stanford analysis of payroll data found that workers aged 22 to 25 in the most AI-exposed jobs saw employment drop 13 percent since late 2022, and young software developers specifically fell around 20 percent. In companies that adopted generative AI, junior headcount dropped 9 to 10 percent while senior headcount held steady.
Everyone reads that as a hiring story. It is actually a learning story wearing a hiring story’s clothes. The work that is vanishing is the exact work that used to turn a beginner into someone with judgment. And once you see the mechanism, you notice it is not only happening to 23-year-olds at big firms. It is happening to you, every time you hand a task to a model instead of doing the rep yourself.
The missing rung nobody is pricing in
PwC looked at what happened to entry-level roles that survived and found something stranger than deletion. The roles did not disappear so much as mutate. Entry-level positions in highly AI-exposed occupations are now about 7 times more likely to demand skills that used to appear much later in a career: strategic decision-making, stakeholder management, judgment. Job descriptions for people with zero to two years of experience dropped by 29 percentage points in their share of postings. The rung did not just get taller. It got sanded off, and the next rung up now asks for things you could only have learned by standing on the rung that is gone.
Think about what a first job actually taught, back when it existed. First-pass document review. Reconciling a messy spreadsheet until it tied out. Writing a memo, getting it back bleeding red, rewriting it. Chasing a bug for three hours and finally understanding why it happened. None of that grunt work was valuable for its output. A senior could have done all of it in a fraction of the time. The point was never the memo. The point was the thousand small collisions with consequences that slowly compounded into the thing we call judgment.
AI is very good at exactly that grunt work. That is the trap. The tasks a model absorbs first, the summarizing, the formatting, the routine drafting, the initial research, are the same tasks that fed the accidental apprenticeship. So we hand a beginner a tool, tell her to be efficient, and she is. She ships more. She also stops doing the slow reps that would have built her instinct for when something is off.
The legal world named this early because its apprenticeship was so explicit. First-pass document review, initial case research, issue spotting, chronology building, diligence lists. Every one of those was a task AI reached for first, and every one was also the exact work that used to turn a first-year into someone a partner could trust. When a firm hands a junior a model and tells her to be efficient, she becomes more productive on Monday and less on track to develop judgment by year three. The firm sees only the Monday number. The year-three cost lands on a different page of the ledger, long after anyone is still watching the trade that caused it.
I wrote earlier about the apprenticeship gap as a team problem, the question of how you build an organization when the training rung is gone. This is the other half of it, the personal half. Forget your team for a second. What happens to your own competence when you can delegate anything? That is the question almost nobody is asking, because the short-term feedback is so good it drowns out the signal.
The framework: the judgment ladder
Judgment is not a thing you download. It is the top of a ladder, and every rung below it is made of reps you actually did. Here is the pipeline that turns effort into instinct, and the exact point where delegation cuts it.
Read it from the bottom. Reps are the raw material: doing the hard part yourself and being wrong in private a few hundred times. Enough reps and you start seeing patterns, this bug is that kind of bug, this customer is that kind of customer. Enough patterns and the recognition goes below conscious thought and becomes intuition, the quiet sense that something is off before you can articulate what. Intuition plus a few high-stakes moments where you were forced to decide with incomplete information becomes judgment, the ability to act well when the playbook stops applying.
Every rung is built only from the one beneath it. There is no shortcut from the ground floor to judgment. You cannot read your way there, and you certainly cannot prompt your way there. When you delegate the reps to a model, you are not skipping the boring part and keeping the good part. You are removing the only material the whole structure is made of. The output arrives. The ladder does not get built. And the cruel part is that from the outside, and even from the inside, the shipped work looks identical either way.
The four kinds of reps, and which one AI eats first
Not all reps do the same job, which is why “just do some things manually” is not specific enough advice. Reps come in four flavors, and AI is hungriest for the one that feeds all the others.
| Rep type | What it builds | What AI does to it | Safe to skip? |
|---|---|---|---|
| Production reps | Fluency. Doing the core task fast and correctly without thinking about mechanics. | Absorbs them almost completely. This is what agents are best at. | No, these feed the rest |
| Judgment reps | Taste for what good looks like, and a nose for what is subtly wrong. | Quietly starves them, because you never see the wrong versions you did not produce. | No |
| Recovery reps | Knowing how to debug, unstick, and dig out when things break. | Removes the small breakages that used to train you, so only large ones remain. | No |
| Taste reps | A point of view on quality that is yours, not the model’s average. | Nudges you toward the median output of everyone who came before. | Rarely |
Production reps are the entry point, and they are the ones AI takes first and most fully. That would be fine if production reps were self-contained. They are not. Doing the task by hand a hundred times is how you accumulate the raw examples that later become pattern recognition, taste, and recovery instinct. Take away the production reps and you have not just outsourced typing. You have cut off the supply line to every higher skill.
Judgment reps are the sneakiest loss because the deprivation is invisible. When you write ten bad drafts and one good one, the nine bad drafts are not waste, they are the contrast that teaches you what good means. When the model hands you a competent draft on the first try, you never generate the bad versions, so you never build the internal comparison. You get output without the education that used to come attached to it. Anthropic’s own 2026 research found that learners who used AI to delegate coding scored about 17 percent lower on comprehension than those who used it to ask conceptual questions. Same tool. The difference was whether they did the reps or handed them off.
This is why “AI makes everyone a senior” is backwards. It makes everyone’s output look senior while quietly capping how senior the person underneath can become. I have written before about taste as a moat and synthesis as the skill that compounds. Both of those are built from taste reps and judgment reps. If you delegate the inputs, the moat never forms.
The delegation trap: why it hides from you
The reason this is dangerous rather than merely sad is that the mechanism is self-reinforcing and it disguises itself as success the entire way down. A 2026 Yale study modeled human skill and AI reliance as a coupled system, and the finding is worth sitting with: AI assistance can strictly improve your short-run performance while causing persistent long-run performance loss compared to never using it at all. The driver is a feedback loop between delegation and practice. You delegate more precisely because the assisted work looks better than your own, which means you practice less, which means your own work gets worse, which makes delegating look even smarter.
What makes the loop nearly impossible to feel from the inside is that the Yale authors traced the problem to stability, not to any bad incentive or misalignment. You are not being reckless. Each individual decision to delegate is locally correct. The assisted work really is better this week. The loop is not made of mistakes, it is made of a hundred reasonable choices that add up to a slope. That is exactly why it evades detection. There is no moment where you feel yourself getting worse, because in the short run you are not getting worse, your output is getting better. The decline is in a capability you are not currently using, so you do not notice its absence until the day you need it and reach for it and it is not there.
I called a nearby version of this the velocity illusion, the gap between how fast you feel and how much you actually ship. The reps problem is the deeper layer under it. Velocity is about this week’s output. Reps are about next year’s competence. You can be winning the first while quietly losing the second, and the first will happily lie to you about the second.
The competence curve: fast now, fragile later
If you plot skill over time, the delegation path and the no-delegation path do not just diverge in size. They cross. Early on, the person leaning on AI is ahead, because the tool is carrying them past their actual ability. Later, the person who did the reps pulls ahead and keeps climbing, while the delegator plateaus and then slides, because the underlying skill was never built and has started to decay.
The green line is the person who kept doing the hard parts. Slower at first, because reps are slower than delegation, then compounding and never really stopping. The red dashed line is the delegator: a fast early rise powered by the tool, a crossover point where the reps person catches up, and then a long slide below the baseline as the un-practiced skill decays. The shaded region on the right is the part nobody prices in at the start. Call it the fragility gap, the distance between how capable you look and how capable you actually are when the tool is removed or the situation goes off-script.
Notice this is not the same shape as a cost curve where one line is flat and another falls. Both competence lines start at the same point and both move. The delegation line is a sugar high: it buys you a real, visible boost that you pay back with interest later. And here is the uncomfortable implication for anyone building fast right now. If you are shipping impressive things well beyond your actual skill, you are not necessarily a fraud. You are on the red line, in the early part where it looks great. The question that decides your next two years is whether you are also, on the side, doing enough reps to build the green line underneath.
Why seniors are safe and juniors are not
Here is the asymmetry that explains both the labor data and your own risk. Delegation is safe in proportion to the reps you have already banked. A senior engineer who spent ten years writing systems by hand can hand the boilerplate to a model all day, because she already has the patterns, the intuition, and the judgment. For her, the tool is pure amplification on top of a foundation that is already poured. She reads the AI’s output and instantly senses when it is subtly wrong, because she has produced ten thousand versions of it herself.
| Who | Reps already banked | What delegation does to them | Exposure |
|---|---|---|---|
| Senior who did the reps | Full ladder built before AI | Pure amplification on a poured foundation. Can supervise output well. | Low |
| Junior who never did them | None, and the training tasks are gone | Output looks senior, skill stays at zero. Cannot catch bad output. | High |
| Founder who delegates everything | Deep in a few areas, zero in most | A junior to his own tools in every area he never practiced. | Hidden and uneven |
The junior is exposed because he has nothing banked and the tasks that would let him bank anything have been automated out from under him. His output looks senior. His skill sits at zero. He cannot catch the model’s mistakes because he has never made those mistakes himself and paid for them. This is not a knock on juniors. It is a description of a rung that got removed while they were mid-climb.
The founder is the interesting case, and probably the one reading this. A solo founder is deep in maybe two or three areas and a total beginner in everything else, which is most of the company. The moment you can delegate legal, design, marketing copy, data analysis, and half your engineering to models, you get a company that ships in all of those areas while you personally remain a junior in all of them. You become a junior to your own tools. And unlike the actual junior at a firm, nobody is going to redline your work and teach you. You will just have output you cannot fully evaluate, in domains where you cannot tell great from plausible, which is precisely the trust gap that bites hardest when it matters most.
This is also why the incompressible core of a one-person company is not a nice-to-have. The areas where you keep doing the reps are the only areas where you can actually run the company rather than nominally own its output. Everything else, you are trusting on faith.
The absence compounds, not just the skill
Reps compound, and that is the whole reason the ladder is worth climbing. Each rep you do makes the next one more valuable, because you bring more pattern to it, so you extract more from it. Judgment is not a fixed prize you unlock once. It is a rate. People with judgment get more out of every new experience than people without it, which is why the gap between the practiced and the unpracticed widens over a career instead of closing.
What almost nobody says is that the absence compounds the same way, just pointed downhill. When you skip reps, you do not only fail to gain skill. You lose the ability to tell which reps you should not have skipped. Judgment is also the thing that tells you where your judgment is thin. Take it away and you lose the meter, so you start delegating the load-bearing tasks with the same easy confidence as the throwaway ones, because you can no longer feel the difference between them. The person most likely to over-delegate is the person who has already delegated enough to stop noticing.
This is what turns a personal habit into a company-shaped risk. A founder who has quietly become a beginner across most functions does not experience it as danger. He experiences it as a smooth-running operation, because everything ships and nothing is obviously on fire. The erosion is silent by construction. You would need the exact judgment you traded away in order to notice that you traded it away. So the failure does not announce itself as decline. It shows up much later as a decision that goes badly in a domain you were sure you had handled, and even then you may not connect it back to the hundred small delegations that hollowed the domain out.
The contrarian take: delegation is a loan
The standard framing says AI delegation is a multiplier, a strict amplifier on what you can do. I think that framing is quietly wrong in the one way that matters. Delegation is not a multiplier. It is a loan against your future competence, and like most loans it feels free at signing because you do not read the repayment terms until later.
A multiplier acts on a capability you already have. A loan gives you spending power you have not earned yet, on the promise of paying it back. When a senior delegates, that is amplification, because the capability exists and the tool multiplies it. When a beginner delegates, that is a loan, because the capability does not exist yet and the tool is fronting it. The output is the same either way, which is exactly what makes the loan invisible. You are spending competence you have not built, and the bill arrives as a decision you cannot make well, on a day you cannot predict, in a domain you thought you had covered.
The sharpest way I can put it: you can delegate the task, but you cannot delegate the learning the task would have given you. That learning was never a side effect of the work. For the first several years of any skill, the learning was the entire point of the work, and the output was the byproduct. AI inverts that. It gives you the byproduct and keeps the point. So the friction you are so happy to remove, the struggle, the rewriting, the debugging, the being wrong in private, was not the cost of getting good. It was the tuition. Remove the tuition and you do not get the education for free. You just do not get the education.
There is a second-order version of this that should worry anyone betting on AI supervision as a career. The whole optimistic story is that humans move up the stack: the model does the work, you supervise it. But supervision is a judgment skill, and judgment is the top rung of a ladder built from reps. You cannot supervise well in a domain where you never did the reps. So the plan of “let AI do it and I will judge the output” only works for people who already climbed the ladder before AI kicked it away. For everyone climbing now, the plan quietly assumes the exact skill the plan destroys.
Where the bill comes due
Abstract erosion is easy to ignore, so here is what the loan actually looks like when it gets called. Three shapes, all of them ordinary.
The first is the one I opened with. Something you shipped breaks in a way that requires the skill you delegated to build. The payments bug that eats reconciliation for six weeks, because nobody on the project ever developed a feel for where payments code rots. The output was fine the day it shipped. The bill arrived later, denominated in the exact judgment you did not bank, and no amount of prompting closes the gap in the moment, because you do not even know what you are looking for.
The second is quieter and usually more expensive: a decision you cannot evaluate. A vendor sends a proposal. A candidate walks you through their approach. A contractor hands you an architecture. In a domain where you did the reps, you would smell the problem in ten seconds. In a domain where you delegated your way to fluent-looking output, you nod, because it all sounds plausible, and plausible is the ceiling of what you can assess. You do not make an obvious mistake. You make an invisible one, and you make it with confidence, which is worse than making it with doubt.
The third is the slowest. You stop being able to tell your own good work from your own mediocre work. Taste is a rep. When you have generated enough bad versions of a thing by hand, you can feel the quality of a new one instantly. Delegate the generation long enough and that internal gauge drifts toward the model’s average, so your standards quietly track down to competent-and-forgettable without you ever choosing it. The bill here is not a bug or a bad hire. It is a slow slide into producing work that is fine, forever, and never knowing it could have been more.
The pattern across all three is timing. Delegation moves the cost from now to later and from visible to invisible. You feel the benefit today, in output you can point at, and you pay the cost on an unknown future date, in a currency, judgment, that shows up on no dashboard. That mismatch is the whole reason smart people walk into this with their eyes open and still get caught. The feedback that would warn you runs on a delay measured in months. The feedback that rewards you is instant.
When delegation is actually fine
None of this means do everything by hand like it is 2010. That would be its own kind of malpractice, and I delegate constantly. The point is not to refuse the tool. It is to know which reps are load-bearing and which are genuinely safe to hand off. There are three honest cases where delegation costs you nothing.
First, reps you have already banked. If you have written five hundred cold emails by hand and you have the instinct for what converts, letting a model draft the next hundred is pure amplification. You will catch the bad ones instantly because the ladder is already built. Delegate freely in your areas of real depth. That is the reward for having done the reps.
Second, work with no learning left in it for you. Some tasks stopped teaching you anything years ago. Reformatting a CSV, converting a file, writing the same boilerplate config for the hundredth time. If a task has no remaining rep value, if you would learn nothing from doing it again, the only thing it costs you is time, and delegation is the right call. The test is not “is this hard.” The test is “is there a rep left in this for me.” If the answer is no, hand it off without guilt.
Third, genuinely throwaway output where being wrong is cheap and reversible. Exploratory drafts you will heavily rewrite, first-pass research you will verify anyway, scaffolding you will replace. Here the model is a fast start, not a substitute for your skill, as long as you actually do the verifying and rewriting yourself, because that verification is itself a judgment rep. This connects to knowing your fallback competence, the floor you can operate from when the tool is wrong or gone. Delegation is safe exactly when your fallback in that area is solid. It is dangerous exactly when the tool is the only thing standing between you and a task you could not do without it.
What to do Monday morning: the rep audit
Concrete moves, not vibes. Here is what I actually do to keep the ladder building while still using AI for most things.
Run a rep audit on last week. List the ten most important things you shipped. For each one, mark whether you did the core rep yourself or delegated it. Then mark, honestly, whether that was an area where your ladder is already built or one where it is not. Delegating in a built area is amplification. Delegating in an unbuilt area is a loan. Count your loans. If most of your important work is loans, you have a competence problem hiding behind a productivity number.
Pick one or two skills to keep hand-built. You cannot do everything by hand and you should not try. Choose the one or two capabilities most central to your edge, the ones in your incompressible core, and do those reps yourself on purpose, even when the model could do them faster. Take the slower path in the places where being genuinely good is the whole game. Delegate aggressively everywhere else.
Do the rep before you delegate it, at least sometimes. For a skill you want to actually own, do the task yourself first, then ask the model and compare. The comparison is where the learning lives. You see what you missed, you see what it missed, and you build the judgment rep that pure delegation would have skipped. You do not have to do this every time. You have to do it enough that the ladder keeps growing.
Grade the AI against the source, not against your gut. When you review model output, do not just skim it for plausibility, because plausibility is exactly what the model is optimized to produce. Check it against the actual source, the real data, the real code path, the real document. That act of verification is a recovery rep and a judgment rep at once, and it is the single highest-value habit for staying able to catch bad output. I wrote about this inversion in reading as the new bottleneck. Reading critically is now a rep you have to protect.
Set a struggle tripwire. Notice when you have not been genuinely stuck on anything in a while. Being stuck is not a failure state, it is the feeling of a rep in progress. If weeks go by and nothing has made you struggle, you are almost certainly delegating all your reps away. Deliberately take on one thing that is hard for you, in an area you care about, and do it without the model. That discomfort is the ladder getting built. For a fuller version of this, see what is worth learning in the AI era and the founder operating system it fits inside.
FAQ
Does AI delegation really make you worse, or just faster?
Both, on different timelines. A 2026 Yale study modeling skill and AI reliance as a coupled system found that AI assistance can improve short-run performance while causing persistent long-run skill loss compared to not using it, driven by a feedback loop between delegating more and practicing less. You get faster now and, in the un-practiced skill, worse later. The two effects are not in tension. They are the same mechanism seen at two time horizons.
Is this just the old “calculators ruined math” panic?
No, and the difference is scope. Calculators removed arithmetic reps but left the reps that build mathematical judgment, setting up problems, choosing methods, sanity-checking answers, mostly intact. Today’s tools absorb the entire chain from problem framing to finished output, which means they remove the judgment reps too, not just the mechanical ones. When a tool can do the thinking and not just the calculating, the reps it removes go much higher up the ladder.
How do I know if I am on the safe side of delegation?
Ask whether your ladder in that specific area is already built. If you have done hundreds of reps by hand and can instantly sense when the output is wrong, delegation is amplification and you are safe. If you could not do the task without the tool and cannot reliably catch its mistakes, delegation is a loan and you are exposed. The test is per skill, not global. You can be safe delegating in your two areas of depth and exposed everywhere else at the same time.
I am a solo founder. Which reps should I keep?
Keep the reps in your incompressible core, the one or two capabilities that are the actual source of your edge, plus verification reps everywhere, because being able to check output is what keeps you from flying blind. Delegate aggressively outside the core. The mistake is not delegating. It is delegating in every area at once so that you become a beginner across your whole company while the output looks professional.
What about juniors, is their situation hopeless?
Not hopeless, but the path changed. The old accidental apprenticeship, where grunt work quietly taught judgment, is mostly gone, so the reps now have to be deliberate rather than incidental. The most useful moves are to verify AI output against source material before anyone senior sees it, which practices the judgment rep directly, and to seek real-time mentorship and high-stakes exposure rather than waiting to absorb judgment from redlined documents. The reps still exist. They just no longer arrive automatically inside a first job.
Can AI ever build judgment instead of eroding it?
Yes, when you use it to interrogate rather than to offload. Anthropic’s 2026 research found learners who used AI for conceptual inquiry scored well above those who used it to delegate the coding, a gap of more than 45 points on comprehension in that study. Same tool, opposite outcomes. If you use a model to ask why, to challenge your reasoning, to explain the thing you are stuck on, it can accelerate rep-building. If you use it to skip the reps, it starves them. The verb matters more than the tool.
How is this different from the apprenticeship gap or skill atrophy?
Skill atrophy is about losing a capability you already had. The apprenticeship gap, which I covered separately, is the organizational problem of how to build a team when the training rung is gone. The reps problem is the personal mechanism underneath both: reps are the unit that converts practice into judgment, and delegation cuts the supply of reps. Atrophy is the symptom in people who had the skill. The reps problem is why new skill never forms in the first place, whether you are a junior or a founder standing in for one.
Isn’t some competence loss an acceptable trade for the speed?
Sometimes, and that is the point of the rep audit rather than a blanket rule. Trading competence for speed is fine in areas where you will never need to operate without the tool and where being wrong is cheap. It is a bad trade in your core, in anything high-stakes, and in anything where you need to catch the model when it is subtly wrong. The error most people make is not choosing the trade. It is making the trade everywhere by default, without noticing they made it, because each individual delegation felt locally free.
This is part of an ongoing series on building well as a solo founder in the AI era. If it was useful, the founder operating system ties these pieces together.