A replica that serves nobody but my test runner, a bake window that overrules me at 3am, and the expand-then-contract migrations that price it all in.
The single best thing I’ve ever added to a deploy pipeline is a replica that serves zero percent of traffic.
Nobody reaches it. Not one user. It sits in the cluster running the new version, with the real config and the real dependencies, and the only thing that can talk to it is my test runner, because there’s a header, and I’m the one who sets it.
That sounds useless. It’s the reason I don’t have a feeling in my stomach when the pipeline goes green.
nobody upstream of the pipeline is nervous anymore
My team is Claude Code and friends, building apps for me and for paying customers. No standups, no sprint planning, nobody arguing whether the ticket is a 3 or a 5.
An agent will happily produce a refactor at 2am that touches forty files, passes the tests it wrote for itself, and quietly breaks the one code path that only fires for customers on the annual plan. It is not malicious. It is not lazy. It’s just fast, and fast without a net means you find out from a customer.
Here’s the part that took me a while to name. When a human writes the change, someone upstream of the pipeline has a rough sense of „this feels risky, let me check the payment path.“ That intuition is doing more work in most teams than anyone accounts for. Take the human out and the intuition goes with it. Everything you were quietly relying on a nervous senior engineer to catch now gets caught by something mechanical, or it isn’t caught at all.
Some of that is replaceable directly. CodeRabbit reviews every PR, which means I use an AI to review the AI, and it has caught things I would have merged without blinking. An error swallowed here, a nil check there. It’s not a senior engineer. It’s a very patient reader who never gets bored on file thirty of a big diff, which is exactly where I stop reading carefully.
But code review only catches what’s visible in the diff. The rest has to be caught by the deploy itself, which means the gate had better not depend on me being awake.
header at the ingress, and the part that breaks it
Every environment gets a canary, and the first step is that replica serving zero percent of traffic.
That’s where the end-to-end suite runs. Against the actual new version, sitting in the actual cluster, with the actual config, serving exactly one client: my test runner.
The routing is the part people ask about, so: the decision lives at the ingress. The gateway looks for a header on the way in, and if it’s there, the request goes to the canary pods instead of the stable ones. Everyone else gets routed the way they always were and has no idea anything is happening. The catch is that the header has to survive the whole trip. If your service calls three others, they need to forward it, or your „canary test“ quietly falls back to stable services halfway through and passes for the wrong reason. Anything that leaves the request path entirely, a queue consumer, a cron job, doesn’t get covered by this at all. I test those separately and I don’t pretend otherwise.
If it fails there, nobody knows. There’s no incident, no rollback, no Slack message. The deploy just doesn’t happen.
no, this isn’t staging with extra steps
This is the objection I’d raise too, so let me kill it properly.
Staging is a different environment pretending to be production. Different data, different secrets, different scale, different neighbors, and a config file that drifted from the real one six months ago and nobody noticed because nothing broke loudly. Every staging bug you’ve ever chased down and discovered was „just a staging thing“ is that drift talking.
The canary isn’t pretending. It’s the same cluster, the same config, the same dependency instances, the same secrets, the same everything. The only difference between it and the version your customers are using is a header, and the header decides who gets routed there. There is nothing to drift, because there’s no second environment to drift from.
Which is also why rolling updates are a lie you tell yourself. You replace pods one at a time, and yes, technically there’s no downtime, but customers are hitting the new version from the moment pod one comes up. You found out it’s broken because a user found out it’s broken. Your rollback is a race against the support inbox. The difference isn’t the sophistication of the tooling. It’s who is holding the bug when it surfaces.
two versions live at once, and your schema has to agree
I’d rather say what this costs than have you find out on your own repo.
You run two versions at once, so you pay for both during every rollout, and more importantly your code has to tolerate both being live. Schema migrations are where this bites. Anything that isn’t backward compatible breaks the entire premise, so migrations get split into expand-then-contract steps across separate deploys, and that’s genuinely more work than writing the one migration you wanted to write.
Stateful services are harder still. If two versions can’t safely share the same data or the same connection, the zero-percent trick doesn’t save you, it just relocates the problem.
And the wall clock gets longer. Real bake windows across three environments mean a change can take hours to reach production. For me that’s fine, because nothing about this promises a fast deploy. It promises a boring one. If you’re in an incident and need a hotfix out in four minutes, you’re going to be tempted to bypass the whole thing, and you should have a deliberate answer for that ready before the night it happens.
For a small internal tool with two users, all of this is more than the job needs. I’d skip it and say so.
bake time is the part you can’t cut
Suite passes, so now real traffic. A trickle first. Small percentage, and I’m watching two numbers: latency and error budget.
Not watching in the „I have Grafana open in a tab“ sense. Watching in the sense that the rollout controller is watching, and if either metric goes outside its bounds, the rollout stops and reverses itself without asking me. The system has an opinion and it’s a better one than mine at 3am.
If the numbers hold through the bake window, more traffic. Hold again, more. Eventually the new version has everything, the old deployment dies, and the pipeline moves on to the next environment and does the identical dance. Dev, QA, prod. Same mechanism, same guardrails, same header, same suite.
The waiting is the step everyone wants to shorten, and it’s the one that has to stay. Some bugs need volume before they show up. Memory leaks need hours. That connection pool exhaustion you’re going to discover eventually needs a specific pattern of traffic that only exists on real users. Five minutes at ten percent finds none of that. Thirty minutes finds a lot of it. Cutting the bake window feels like speed, but it just moves the discovery to production and calls the pipeline fast because the part you watch finished sooner.
the bugs that ship are the ones nobody wanted to write a test for
All of this rests on the suite being worth running, and suites get thin for one reason.
The code that breaks is rarely code that was untested on purpose. It’s code where writing the test would have been miserable, so nobody did. Hard to test and untested are the same thing six months later.
That’s why I’m annoying about architecture even on projects where I’m the only human. Ports and adapters, domain logic that doesn’t import the framework, boundaries you can fake. Not because it’s elegant. Because it means the agent can write a test in twelve seconds instead of building a mock pyramid, and a test that’s cheap to write actually gets written.
A canary is only as good as the suite you point at it. Thin suite, and the zero-percent replica is an expensive way to confirm the service starts.
take the replica, not the repo
Setting all this up takes about two weeks per project before you write a line of feature code, and I did it enough times that I got sick of it and packaged it as Baukit. It is aggressively opinionated and very alpha, and if you adopt it today you’re adopting my taste along with my bugs.
But don’t make Baukit the takeaway. Make it the zero-percent replica. Deploy a version that serves nobody, run your tests against it, and only then let a human near it. That’s a weekend of work and it’ll do more for your uptime than any amount of shipping slower.
What it bought me isn’t a metric I can put on a dashboard. No deploy window. No Friday freeze, which was always a superstition dressed up as a policy anyway. I push, I go make coffee, and the thing that decides whether my users see this change is a set of metrics, not my nerve.
There’s a thread in r/engineeringmanagers I keep thinking about, where a new EM inherits a team that ships constantly into a permanently burning production, and the comments tell him to go fix the culture. I think they had the wrong noun. That team wasn’t shipping too fast. They were shipping fast into a system where every mistake went straight to a customer’s face. Slow them down and you get slow and broken, which is the worst possible outcome and also, funnily enough, what most companies pick.
Speed and safety aren’t a tradeoff. They’re both downstream of the same thing: how quickly you find out you were wrong, and how few people find out with you.
Peace, nerds.