01The 2,700-line spec challenge
Let’s say you had this spec here. This spec is 2,700 lines. Imagine you gave it to a general-purpose engineer who has some experience in your codebase but is, you know, kind of like a new hire. How long do you think it would take them to build?
First of all, how long would it take them just to read it? A couple of days to wrap their head around it? Needless to say, this would take a human engineer, I don’t know, three months to build, if it’s even possible. I don’t think a single person could handle this spec, because first you’d have to read it, then break it down into projects and parts.
Now the question is: could an AI pull off the implementation of this spec?
The last time I had an agent try to implement a spec was back in March of this year, and it failed catastrophically. It failed so hard that I dedicated my life to tearing apart the scam of AI coding. I was like:
“Dude, what is this s***? They promise you that agents can do all this.”
Back in March, you’d give an agent a spec that wasn’t even this long or detailed, maybe 500 or 600 lines, and the agent would say, “Say less, bro, I got this.” Then it would go drown in its own vomit, leave your codebase an absolute mess, give up and declare, “Done!”
YouDid you finish?
AgentOh, I finished, all right.
Finished what? The turn? “I finished processing tokens.”
Did it finish the feature? Not even close. Did it leave the codebase functional? Not even close. Back then, smart engineers ran into those hurdles and thought, “I need to design systems to scaffold these AIs so they can actually get through a spec.”
02The shift to evolutionary coding
As you may have seen in my recent videos, Opus 5.5 has turned me into somewhat of a believer in what agentic engineering can actually finish. I’ve seen it do things it simply couldn’t do back in March.
Right now I’m building an app called Enjoy, and I’m doing it fully agentically. You may call it vibe coding. I call it evolutionary coding:
- You let the code evolve to satisfy environmental pressures: the test suite.
- If it passes those pressures, you let it proliferate.
- Biological genomes are an absolute mess, but they work.
Enjoy gives non-developers access to the terminal-based AI coding agents running on their own computers. As I talked with people, several told me, “I have a non-technical co-founder I’d love to share my project with so we can collaborate.”
Long story short, I needed to completely redesign how the teams feature worked. I was dreading it, because it was massive. We had built the app one way, and now we had to re-architect everything for multi-user collaboration. In 2020 this would have been a three- to four-month project: I’d have to stop everything, hide in a cave, not shower for weeks and just wallow in technical misery.
Instead, I started with a massive voice memo, thinking out loud through the requirements. I worked through it with Claude, and it generated this 2,700-line spec. My first reaction was:
“Oh, you built a spec, did you? What the f*** do you want me to do with a spec? You’re telling me you can build out an entire feature from a 2,700-line document?”
I figured the project was doomed and that I’d have to break the document into 10 or 20 micro-tasks by hand. But I decided to try anyway.
03Running Claude Ultracode
I switched Claude Code to Ultracode. Ultracode is Claude’s nuclear option: it spins up dozens of sub-agents and runs for hours. I’m on the $200 plan, and I gave it a single instruction: “Build the whole thing end to end.”
Here’s how the run went:
-
Claude starts, saying it will implement the full computers-and-devices model across the relay, billing, desktop, web and OS targets.
-
No code commits yet. Claude is still drafting and finalizing the spec, already 2,600+ lines.
-
Claude spawns four critics to review the spec.
-
The first implementation commits land.
-
Over 60 sub-agents are running concurrently.
-
45 of 53 active verification agents have finished.
-
The full run completes.
Ultracode spawned adversarial reviewers on its own, without being asked. Eight reviewer agents produced 34 findings. Before anything was fixed, each finding went to two or three independent skeptic agents whose only job was to disprove it: three for high-severity findings, two for the rest. That added roughly 80 short-lived, read-only verification agents on top of the original reviewers.
In roughly seven hours, it implemented the full spec.
Did it work? It did. The end result was flawless. It cleanly handled a flow that spans every surface: browser pairing, desktop client authentication, relay coordination and end-to-end encryption. And it didn’t even use up the 5-hour usage limit on my plan.
04Code quality vs. measurable verification
Skeptics will naturally ask: “What about code quality? Isn’t it just slop? How can you trust end-to-end encryption written by an agent?”
With modern frontier models, hallucination is rarely the problem when you give them enough reasoning effort. When asked to audit their own work, they are intensely skeptical of themselves and double-check edge cases again and again. They can hold the thread across a 10- to 20-hour run without devolving into incoherence.
I don’t care whether we call it intelligence, pseudo-intelligence or next-token prediction. What matters is that it built the system end to end, with high fidelity to the spec and an enormous suite of passing tests.
Line-by-line code review is irrelevant
DHH recently posted:
It takes a while to build the confidence, but there’s no future where you’re manually reviewing every line of agent code. Not much acceleration in that. You need adversarial agent reviews, you need automated testing, and maybe you spot check. That’s it. From prompt to production!
I agree. If you’re manually reading every line of agent-generated code, your velocity drops right back to 1x. The gains from agentic engineering are too big to bottleneck behind a human inspecting diffs.
2025Vibe coding
2026Evolutionary coding
Major engineering teams are already making this shift. Shopify and Coinbase both recently announced moves from React Native back to fully native mobile apps. How do they justify the overhead? Engineers no longer have to write and maintain two codebases by hand. An agent can build the iOS app and automatically keep the native Android app in sync with it.
05The real future: CI over code inspection
Many engineers argue: “There’s no way you can operate a product if you can’t reason about every line of your codebase.”
Does a CEO need to understand every line of code their employees write? If you treat the agent as an engineer, the job shifts from inspecting syntax to validating outcomes.
Human-driven builds
Often suffer from tunnel vision, brittle assumptions and light test suites.
~500 tests
Agent-driven builds
Come with comprehensive test suites built around environmental pressures, and CI stays green.
~10,000 tests
It’s like self-driving cars: you don’t judge one by the elegance of its driving algorithm. You judge it by the crash rate.
The bottleneck in software engineering is no longer writing or reading code. The bottleneck is the CI environment.
If your automated test suite puts strict, rigorous pressure on the agent’s output, clean architecture emerges functionally rather than cosmetically.
If you stay dogmatic about manual, line-by-line pull request reviews, you risk anchoring your career to an obsolete paradigm. The future belongs to verified, agentic delivery.