Skip to content

i stopped reading the diff first

An agent wrote the code, then wrote the proof the code works, out of the same misunderstanding. One artifact in that pull request survives that problem.

There’s a pull request open right now with 63 files changed and about 11,000 lines touched. Somebody on your team opened it. They wrote maybe 400 of those lines by hand. An agent wrote the rest, and it took an afternoon.

Now you’re the reviewer. Go.

If you read every line at the pace you’d read a normal PR, that’s a full day, probably more, and at the end of it you’d have caught the naming inconsistencies and missed the thing that actually breaks. Meanwhile two more PRs landed in your queue. Your product manager has a demo Thursday. The business wants the feature because features are the only reason anybody is paying for any of this.

The standard advice is „make the PRs smaller.“ That’s a fossil from a world where a ticket was small because typing was slow.

So I don’t read the diff first anymore. I read the end-to-end tests. There are twelve of them, sitting next to 340 unit tests the agent also wrote, and the ratio isn’t what matters. A unit test can only be checked against the code, and reading the code is the thing you can’t afford. An end-to-end test is written in the language of the person who asked for the feature, which means you can check it without opening a single implementation file.

That’s the argument. The rest of this is why it holds, and where it stops working.

small PRs were never the point

The advice was always „keep pull requests small and focused.“ Good advice. I gave it for years. But look at why it worked, because the reason matters more than the rule.

Small PRs worked because tasks were small. The ticket said change the button from blue to green. An engineer opened the codebase, found the button, changed the hex value. One file, one line. Later we got slightly more ambitious and a normal ticket became three files and 200 lines, still trivially reviewable, because a human typing at human speed over two days produces roughly that much code.

The PR was small because the unit of work was small. Small PRs were a downstream effect of human typing speed, not an independent virtue we chose.

That constraint is gone. When an agent can implement a whole feature vertically, from migration to endpoint to UI to tests, in an afternoon, the ticket grows to fill the new capacity. And it should. Nobody is going to artificially cap a story at 200 lines when the tooling can deliver the entire thing, and any manager who tries will lose to the manager who doesn’t.

So yes, you can still split the PR. Sometimes you should. But splitting an 11,000-line change into eleven 1,000-line changes doesn’t reduce the reading. It just gives you eleven opportunities to approve something you didn’t read, which is worse, because now you’ve laundered the lack of review through a process that looks rigorous.

That’s the trap. The teams complaining loudest about giant PRs mostly aren’t reviewing the small ones either.

the first pass isn’t yours anymore

Some of the old review job genuinely can be handed off, and this part I’m happy about.

Think about what you actually did on a first pass. You looked for obvious bugs. Null handling, off-by-ones, a missing await, an error swallowed in a catch block. You checked whether they reinvented a helper that already exists three directories over. You flagged style drift, because six people with six different habits in one codebase turns into a mess nobody wants to touch.

That’s real work, and it’s mostly pattern matching against a known set of mistakes, which is exactly the shape of work these tools are good at. CodeRabbit, Copilot’s reviewer, whatever your org bought, point it at a diff and it’ll surface the obvious defects and the reinvented wheels before a human opens the tab. Feed it your conventions too, the things you say in review for the hundredth time, because every rule you write down is a comment you never type again.

But be honest about what you got. You automated the first pass. The first pass was never the pass that caught the expensive bug. It was the cheap pass, the one a linter was already half doing. You didn’t buy confidence. You bought back the hours you were spending on things a machine should have been doing since 2015, and now you get to spend those hours on the hard part.

The hard part is still sitting there. Sixty-three files of it.

the agent grading its own homework

Not the unit tests. There are 340 of them and the agent wrote every one. Reading 340 unit tests is worse than reading the implementation, because a unit test tells you the function does what the function does. The agent wrote the code and then wrote the proof that the code does what it does, and those two things came out of the same understanding, including the same misunderstanding. If it built the wrong thing correctly, every unit test passes. Green means nothing about whether the feature is right.

Integration tests are better and still too low. You’re now checking that two components the agent wrote talk to each other the way the agent thought they would.

So go to the end-to-end tests. And be suspicious of me while you do, because I just spent three paragraphs arguing that agent-written tests prove nothing, and now I’m sending you to tests the same agent wrote.

twelve tests in the language of the ticket

The difference isn’t who wrote them. It’s what language they’re in. A unit test asserts against the code, so judging it means reading the code. An E2E test asserts against the ticket. A user does this, then this, and this should happen. The agent still wrote it, but it wrote it in a language where you’re competent to catch it lying.

That’s why twelve of them beat 340. Twelve you can read. Twelve you can hold in your head. Say that 63-file PR was a CSV import. Here’s the whole E2E list, which is the actual thing I open first:

uploads a valid CSV, sees the rows in the table
uploads an empty file, sees "no rows found"
uploads a .pdf, sees a file type error
uploads a 60MB file, sees a size limit error

You don’t need the implementation to review that. You need to know what your users do. You’re not asking „is this code correct,“ you’re asking the question that actually matters. Did this thing do what we asked, and is the list of things it handles the right list?

absence is invisible to a test suite

Look at that list again. Four scenarios, four passing tests, nothing wrong with any of them.

Nobody wrote the one where the CSV parses fine but carries a duplicate ID in row 4,000. That’s the one your actual users hit every Monday, when the overnight export double-writes a batch, and it’s the one that corrupts the table. The agent never wrote that test because the agent never knew Monday was a thing.

That’s the half no tool does for you. A missing test raises no flag, fails no build, appears in no diff. Every test passes. The scenario just isn’t there.

Finding what’s missing is domain knowledge and memory of how this system has broken before. The habit that works: read the E2E list against the last three incidents this system had, and against whatever the support queue asks about most. If a scenario burned you in March and isn’t in that list, you’ve found the review comment worth writing.

make it show you

That’s the same trick as reading the E2E list, pointed at a different target. Get the system to speak a language you’re competent to judge. The tests state the intent. Now make the software prove it delivered.

For anything with a UI, I want Playwright screenshots or video of the flow. Not a description of the flow. The actual pixels. It takes about a minute to wire up a run that walks the happy path and drops images, and I trust thirty seconds of watching the thing work more than an hour of reading the code that supposedly makes it work.

Plenty survives a passing test suite. Text overflowing its container. A modal you can’t dismiss on mobile. A loading state that flashes in a way no test asserts against because nobody writes an assertion for „this feels broken.“ All of that is invisible in a diff and obvious in a screenshot.

Backend, same principle, different artifact. You don’t get pictures, so you get output. Run it, capture the responses, look at the actual JSON, the actual status codes, the actual log lines when you feed it something malformed. Make the system emit evidence about itself rather than asking the code whether it intends to work.

I’d rather have one video of the feature working than a hundred green checkmarks.

the honest limits

Two things I’d flag before you take this as a complete method, because it isn’t one. And the frame for both is that reading every line wouldn’t have caught them either. Let’s not pretend the old way was working.

E2E tests are slow and they rot. Push everything into them and you’ll have a 40-minute pipeline that people start skipping, which lands you somewhere worse than where you started. Keep the E2E suite small enough that it’s the thing you actually read. When it’s too big to read in a sitting, it’s stopped doing this job.

And this doesn’t catch everything. It catches „the feature is wrong,“ which is the expensive category. It won’t catch the race condition under load or the security hole three layers down. Some of that is load testing, some is your agent hammering the API with garbage and handing you the report, some of it you find in production at 2am like you always have. You’re accepting a defect rate that’s nonzero, the way it always was.

what you’re actually being paid for

I want to name the uncomfortable part directly, because I think it’s why people resist this.

Reading code carefully felt like the job. It was the visible craft. You’d leave a sharp comment about a subtle edge case and feel like a senior engineer, because that was being a senior engineer for twenty years. Letting go of it feels like letting standards slip.

It isn’t. The standard didn’t drop, the leverage moved. Careful line-reading is now an instrument you aim at the 5% that scares you, the auth path, the money path, the migration that can’t be rolled back. It stopped being how you evaluate a change and became something you pick up on purpose.

Aiming it takes what no agent has. Knowing what your users do on a Monday, what the business actually needs, and where this codebase has hurt you before. That’s what you were being paid for the whole time, and you were spending it on brace placement.

The 11,000-line PR isn’t going away. It’s going to be 30,000 next year. The job was never reading all of it. The job is telling real evidence from a green checkmark wearing a costume.

Peace, nerds.

DSGVO Cookie Consent mit Real Cookie Banner