Execution Is Cheap Now — Evals Are the Job
Claude is my command center now. Gmail, Slack, research, Jira tickets — all of it routed through one place instead of five tabs and the context switch between them. The speedup is real, and honestly it's the least interesting thing about the change.
What I keep circling back to is this: if the thing I used to be good at was execution, and execution just got close to free, what exactly is my job now?
This post is me working that out in public. Some of it I'm confident about. Some of it I think is probably wrong, and I've marked those parts.
TL;DR
- Routing a whole workday through one assistant didn't make me faster in the normal sense — it changed the shape of the work. The bottleneck moved.
- As AI absorbs execution, the human job moves up a level: from doing the thing to deciding what's worth doing and judging whether it worked.
- Which makes evals the real skill. Execute → review → learn → execute. Which is just Lean Startup running at a much higher clock speed.
- I have four candidate answers for what stays human — learning speed, emotion, self-knowledge, parallel agency — and I only actually trust one and a half of them.
The part I'd overstate if I weren't careful
My gut says the throughput gain on certain tasks is somewhere between 100x and 1000x. I want to be honest that this is a feeling, not a measurement — I haven't instrumented my own workday, and anyone quoting a number like that (including me) should be treated with suspicion.
But the directional claim I'll defend: the tasks where the gain is enormous are the ones that used to require me to manually carry context between five tools. Read the thread, find the doc, summarize it, write the ticket, ping the person. None of those steps were hard. Stitching them together was the whole cost. That cost is now near zero, and near-zero costs don't make you 20% faster — they delete a category of work.
The job moves up a level
Here's the concrete version. I'm not going to hand-build a website again. I'm not going to hand-build a Facebook clone. I'll ask for it, and I'll get something running.
Which means "can I build this?" has stopped being an interesting question. The interesting question — the one that was always the actual hard question, just hidden behind months of implementation — is:
Why would this win? Against an entrenched competitor, in a market that's now moving faster than it was, with distribution I don't have yet?
That question doesn't get cheaper when code gets cheaper. If anything it gets harder, because everyone else's build cost dropped at the same time as mine. The moat was never the building.
So the scarce skill shifts from execution ability to strategy, judgment, and taste. I think that part is close to obviously true, and it's also the part people find least comforting, because those are exactly the skills nobody knows how to train deliberately.
The new bottleneck is evals
If AI does the execution, my job is evaluation. Did this actually work? Is this actually good? Why did it fail? What do I change?
Put like that it's a loop: execute → review → learn → execute.
And yes — I noticed halfway through thinking about this that I had reinvented build–measure–learn. It's Lean Startup. The framework didn't change. What changed is the clock speed of the execute step, which means the measure and learn steps are now the entire bottleneck. When a build took a quarter, sloppy evaluation was survivable — you had a quarter to notice. When a build takes an afternoon, bad judgment compounds faster than good execution can rescue it.
The practical consequence: I should be spending my skill-building time on getting better at evaluation, not on getting better at the thing being evaluated. Writing good rubrics. Knowing what "good" looks like before I see the output. Being able to say why something failed rather than just that it did.
What I think doesn't get automated (and where I'm shaky)
This is the part I'm least sure about, so I'll make the arguments and then attack them.
1. Learning speed on genuinely new domains. My claim is that humans still generalize into unfamiliar territory faster than current systems do. My evidence is thin and I know it: I have friends working in robotics who think the field is meaningfully further out than the hype suggests, which is a real data point that intelligence and physically competent, adaptive learning are not the same thing.
But notice how narrow that actually is. "Embodied AI is harder than people think" is a claim I believe. "Human learning speed generally beats AI learning speed" is a much bigger claim, and the robotics anecdote does not get me there. I've been letting one piece of evidence carry more weight than it can hold.
There's also a tension I haven't resolved. I think AI's accumulated knowledge will vastly exceed any individual human's. I also think humans pick up new things faster. Both can be true — one is a stock, the other is a rate — but I should decide whether that distinction is actually my point, because as stated it reads like a contradiction.
2. Emotion as a driver. Anger, competitiveness, the specific fury of watching someone do a thing badly that you care about — these generate drive. AI doesn't have them.
Weakest argument in the post, and I'm keeping it in because I think there's something there I haven't found yet. "Anger drives us" is true. It does not follow that anger produces better strategic judgment than a well-evaluated process would. Motivation and judgment are different variables and I've been quietly treating them as one.
3. Knowing why you exist. This one I'd separate sharply from "having goals." Plenty of systems have goals. The people I'd call top tier seem to have something else — an actual answer to why they're here, and they're driven by that answer rather than by incentives someone else set. I don't know how to test this claim, which is a problem for it being a thesis. It's a conviction.
4. Parallel agency. One human directing many AI processes at once is doing something structurally different from any single instance. Not smarter — higher in the stack. This one I'd defend, though I suspect it's temporary in its current form: it's an org design advantage, and org design advantages get productized.
Honest scoring of my own list: (4) I'd defend, (1) is half-right and over-argued, (2) and (3) are convictions dressed as arguments. I'd rather publish that assessment than pretend the list is four for four.
The uncomfortable practical version
Underneath all of this is something much less philosophical.
If execution is cheap, you have to prove your value concretely and fast. My working number is 30 to 60 days — within that window you should be able to sit down with someone and show them why they're paying you. Not describe your responsibilities. Show the judgment call you made and what it produced.
Because if execution is the cheap part, and you can't demonstrate judgment or results, there's nothing left to stand on. That's not a threat, it's just what the sentence means.
I think that's actually the whole thing, and the robotics and consciousness material above is partly genuine belief and partly me building a case for why humans still matter — which is a thing I need to believe in order to know where to invest my own skill-building time. Worth naming that motive out loud rather than pretending I arrived here neutrally.
If you've found a way to actually get better at evaluation — rubrics, review habits, anything that made your judgment sharper rather than just faster — I'd genuinely like to hear it. That's the skill I'm most trying to build right now, and I don't have a method yet.