Skip to content

2.9 GB to 30 MB: The Rewrite Worked Because the Tests Weren’t Mine

Two agents ported a Python service to Rust in two and a half days and cut idle memory by 100x. Verifying that number was the hard part, and the tests that proved it belonged to someone else.

The service did nothing. I want to be clear about that up front, because it’s the whole reason this story is annoying.

Users push data in. We store it. They query it back out, and we hand them some light analytics on top. Create, read, update, delete, plus a bit of grouping. It’s the kind of thing you’d expect to sit in a container quietly using a couple hundred megs and never appear in a meeting.

We measured it at 2.9 GB.

The latest release spiked to 4.9 GB, settled back to 2.5, and I stopped trusting the graph entirely. For a REST API in front of a database. Not a model server, not a video pipeline. A CRUD wrapper, holding more RAM than my first three laptops combined.

The obvious reaction is to go find whoever wrote it. Except we didn’t write it. It’s an open source component we deploy, which turns out to be the detail that makes everything else here work.

It idles at 30 MB now, in Rust, ported by two agents over two and a half days. I’ll get to how. But the rewrite is the boring half. The reason I believe the result is that the test suite I checked it against was written by people who have never heard of me.

where 2.9 GB actually goes

We could read the source, deploy it however we wanted, and profile it, so we did all three.

The application server was the first chunk. Gunicorn, and Uvicorn in the newer release, running the standard configuration we’d inherited: 4 workers, roughly 128 threads each. That accounted for something like 300 MB on its own. Not the whole problem, but a lot of memory to spend before a single request arrives.

Then you look at what a worker actually is. Every worker is a process, and every process carries its own Python interpreter, its own imports, its own copy of everything the app touched on startup. Four workers means four of those. The async processes running alongside them in the container had the same story. You’re not paying for your code. You’re paying rent on however many interpreters you decided to run, and the standard config decided for you.

Here’s the part that stings. None of that memory is doing work. It’s not cache, it’s not buffered data, it’s not something the service will use later under load. It’s the fixed cost of the runtime being present, multiplied by process count, sitting there whether you serve ten requests a day or ten thousand.

So we tuned it. You can always tune it. Drop the workers, cut the threads, trade throughput for footprint, and get to a number that’s still embarrassing. That’s the ceiling, and the ceiling was too low to matter.

The normal next move is to fork it and fix it properly. I don’t think that’s a move. The moment we fork, we own a Python codebase we didn’t write, and we drift a little further from upstream with every release they ship and we don’t. That’s a permanent tax on a team with a limited budget for maintaining things, spent on a component that is one small piece of what we actually sell, and whose best case is „the memory number is now merely bad.“

That was the real dead end. Not technical, financial. Any path ending in „and now we maintain Python we didn’t write, forever, for a modest improvement“ was already dead. Which left one interesting question: was there a path with a return big enough to pay for its own maintenance?

rewrites got cheap while nobody was looking

Two things happened in the last year that changed the math on this, and if you haven’t updated your priors since, you’re working from an old price list.

In May, Jarred Sumner merged Bun’s rewrite from Zig to Rust: over a million lines and 6,755 commits, ported in about 11 days by a fleet of Claude agents running in parallel. Zig’s creator called it unreviewed slop, which is a fair thing to worry about and also didn’t stop it from shipping.

The closer analogue is LiteLLM, who started moving their gateway from Python to Rust for exactly the reasons I’d been staring at. Their published numbers: 359 MB down to 32 MB, throughput from 453 to 6,782 requests per second, per-request overhead from 7.5ms to 0.05ms.

I read those the week I was watching our own graph, and the resemblance was hard to ignore. Same shape of service, same runtime, same problem, roughly an order of magnitude of memory sitting on the table.

the actual trick is that the tests already existed

So: same idea, our codebase. I set up Fable 5 and GPT-5.6 Codex in an agent loop and pointed them at a full rewrite in Rust.

That setup is the part everyone asks about, and it’s the least interesting thing here. An AI rewrite of a real service is not a prompting problem. It’s a verification problem. Two agents will happily generate forty thousand lines of very confident Rust, and the question that decides whether any of it is usable is how you find out what’s wrong. If your answer is „read it,“ you don’t have an answer. Nobody is code-reviewing a from-scratch reimplementation into correctness.

We got lucky in a specific, structural way. The thing is a REST API. The contract is HTTP in, HTTP out, and the implementation behind it is an internal detail. Swap the entire language underneath and the interface is still supposed to behave identically.

And the open source project has a big test suite.

So the loop wasn’t „write Rust and hope.“ It was: run their tests against the Python implementation, run the same tests against the Rust one, and treat every difference as a defect. The tests became an executable specification that neither model wrote, couldn’t rationalize, and couldn’t quietly weaken to make a failure disappear. That last property is the one that matters. A test the agent wrote is a test the agent will happily rewrite when it goes red.

Two and a half days of continuous agent work later, the rewrite was running. I tested it myself across at least ten rounds. Most of the functionality was there.

Note „most.“ I’m not going to tell you it emitted a flawless port on the first pass. It didn’t, and any post that tells you otherwise is selling something. But „most of a service, verified against a borrowed test suite, in two and a half days“ was a good enough result to justify measuring it properly.

the number that ends the argument

Idle memory on the Rust version: about 30 MB.

Against 2.9 GB, that’s roughly 100x. Depending on what work the service is doing, it lands somewhere in the 100 to 150x range. Big enough that when I first saw it I assumed I’d measured the wrong container.

But a single replica isn’t the number that matters, and this is where the business case actually comes from.

In production you don’t run one. You run at least three, spread across availability zones in a region, because that’s what high availability costs. And every customer gets their own instance. So the per-replica saving gets multiplied twice before it reaches a finance conversation: three replicas, then however many customers you have.

Per instance, that’s 10 to 15 GB saved. Take an arbitrary 100 instances and you’re at roughly 1,000 GB of RAM that no longer needs to exist. Price that against what you pay for memory and it’s comfortably a six-figure annual number. It gets better with every instance we sell, which is the direction you want your savings to scale.

Latency came along for the ride, up to 60x faster. That was a real complaint from real users, so it isn’t a vanity metric. And the worker configuration problem just evaporated. There’s nothing to tune, because Rust uses the whole CPU by default. The tuning dial that used to trade throughput against memory doesn’t exist anymore.

maintenance is cheaper than the spreadsheet says

The obvious objection is the one I raised against forking. Now you maintain your own implementation, forever. Fair, except a port isn’t a fork: there are no merge conflicts to reconcile every release, and I’m not designing anything. Upstream decides which features exist, and I decide which of them we implement and when.

That matters because porting a feature that already exists in another language is the single thing current agents are genuinely good at. Not designing systems. Not making judgment calls. Translating a known, working, tested implementation from one language into another, with the original sitting right there as the reference. That’s the shape of the work, and it’s the shape that keeps working.

Meanwhile the savings could pay for a couple of engineers outright. When the maintenance burden is smaller than that, the argument is over.

what I’d actually tell you to do

Don’t rewrite your software because a language is faster. You know Rust is faster than Python. That fact has never once been the deciding factor, and „we rewrote it because the benchmarks looked nice“ is how teams burn a year and ship nothing.

Rewrite when you can write the business case on one line first. Ours was: the waste is fixed per process, we run three replicas per customer, and we have a lot of customers. The multiplication is what made it worth doing. The same 2.9 GB on one internal service nobody replicates is a shrug and a Jira ticket that never gets picked up.

Then check whether you can verify the result without trusting the thing that produced it. A stable interface and a test suite you didn’t author is the cheat code. If you have those, an AI rewrite is a measurable experiment. If you don’t, it’s a pile of plausible code and a long unpleasant quarter.

And the price of finding out has genuinely dropped. Tokens aren’t free, and anyone telling you AI made this free is not paying the bill. But spending one to four thousand euros on a proof of concept to test a six-figure annual saving is not a hard call. That’s not a technology decision. That’s arithmetic with a good ratio.

The thing I keep coming back to is that the Python implementation was never bad code. It’s a reasonable project doing a reasonable job. Nobody made a mistake.

What made the bill invisible is that the two numbers that multiply into it are owned by different people. Infrastructure sees a couple of gigabytes per replica and correctly shrugs, because a couple of gigabytes is nothing. Sales sees the customer count go up and correctly celebrates. Neither dashboard is wrong, and nobody’s job description includes multiplying one by the other. That’s where the money was. Not in the code, in the gap between two teams who were each doing their job properly.

So go multiply your own two numbers. Then go find out who owns the tests, because that’s the one that decides whether you could ever prove the fix worked.

DSGVO Cookie Consent mit Real Cookie Banner