Zum Inhalt springen

My Rust rewrite cut memory 100x. It still should not have come first.

Dieser Artikel ist auf Englisch.

Vampire Survivors shipped in a browser game framework and moved to Unity two years later. Slow code under real load is evidence you have customers. Here is the crossing point where a rewrite stops being a hobby.

Luca Galante was unemployed in 2020 when he started building Vampire Survivors in Phaser 3, a free web game framework. JavaScript. In a browser engine. For a game whose entire design is „put a truly stupid number of things on the screen at once.“

Minute one, you have a handful of enemies and one weapon throwing one projectile. Minute twenty-five, the screen is a solid wall of skeletons, and every one of them needs a position update, a collision check, damage math, a death check, an XP drop. The workload doesn’t grow gently over the run. It explodes, on purpose, because the explosion is the fun.

Phaser was never going to hold that forever. It held it long enough.

The game hit Steam Early Access in December 2021 and blew up. Then, in August 2023, not quite two years after the explosion, it moved to Unity. Not because Phaser was embarrassing. Because Phaser had become the thing standing between a proven hit and the platforms it hadn’t reached yet. It ran fine on PC. The problems showed up on mobile and console, where the entity counts the game demanded were more than the original setup could carry.

Read that order again, because there’s a version of it that never happens. In that version, 2020 Luca sits down and thinks: this design is going to need thousands of entities on screen, so I should learn a real engine first, get comfortable with sprite batching, maybe look into an ECS since that’s how you do this properly. And that guy is unemployed for another eight months, and he ships nothing, and none of us ever hear the name.

The rewrite was not the correction of a mistake. The rewrite was a prize he won by shipping something people wanted.

you have not earned your rewrite yet

That is a sentence I have to keep repeating to myself, because every instinct I have as an engineer pulls the other way. You know how to make it fast. You can see the query that will fall over at ten thousand rows. You know which language would handle this properly, and you know it isn’t the one you’re about to reach for. And none of that matters yet, because the thing you actually don’t know is whether anybody wants this at all.

So build the slow version, in whatever language you already dream in. The speed that matters at zero customers is the speed from your head to a working thing in someone else’s hands, and every hour you spend on an architecture that survives scale is an hour spent on a problem you have not earned.

You only get to have engine problems if enough people are playing your game to cause them.

the rewrite that was arithmetic

I ran the same trade at my own desk, from the other end, and the numbers were not subtle.

I rewrote an open-source controller from Python to Rust. Not for the fun of it, though I will admit Rust is fun. Latency was too high and memory use was bad, and both of those showed up on a bill. The rewrite cut memory by 100 to 150 times and improved latency by somewhere between 10 and 60 times depending on which route you hit.

Those numbers look like a slam dunk, and I want to be careful about how they read. They were not a slam dunk at small scale. At small scale, the Python version was fine. A bit heavy, a bit slow, and invisible against everything else we were paying for. The Python version was not wrong. It was correct for its size.

What changed is that the cost curve got steep. Every unit of growth started charging us more than the last one, and at some point the line crossed the cost of my time and the ongoing cost of somebody having to own Rust code. That crossing point is the entire decision. Before it, the rewrite is a hobby with a business justification stapled on. After it, the rewrite is just arithmetic.

Which means you have to be able to see the curve. A single slow endpoint is a fix. A cost curve bending upward is a forecast, and the only way to have one is to have been running long enough under real load to draw it. The design doc has a guess about this. The running system has an answer, and it is usually a surprising one: the endpoint people actually hammer is rarely the one you hardened, and the bottleneck you spent a week worrying about frequently never shows up at all.

when optimizing first is the correct call

I don’t want to turn this into „always ship the sloppy version,“ because that’s how you get the other kind of disaster.

Google and Amazon put you through design review before you write meaningful code, and people love to mock that as bureaucracy. It isn’t, for them. When you’re inside a company like that and you’re building something adjacent to an existing business, the demand question is largely settled before you start. You know traffic is coming, because it’s already there, sitting in a neighboring product that’s going to spill into yours. Under those conditions „we’ll find out if anyone wants it“ is not a real risk. The expensive risk moves somewhere else. What happens when this thing is load-bearing for a hundred million people and you got the data model wrong three years ago.

Where it goes wrong is cargo culting. A four-person startup that reads the Amazon engineering blog and decides they need the same review gates has copied the ritual and skipped the reason. Those gates exist to manage the cost of being wrong at enormous scale. If your scale is forty users, being wrong costs you a weekend, and the gate costs you more than the mistake would have.

So the question isn’t „should I design carefully.“ It’s „do I already know the customers are coming.“ If yes, design hard. If no, that uncertainty is your actual biggest problem, and no amount of architecture makes it smaller.

you can swap a language. you cannot swap a data model.

This is the honest limit of everything I’ve said, and it deserves more than a footnote.

A language is contained. The controller I rewrote kept the same HTTP routes, the same request shapes, the same responses. Callers never knew. That’s why the rewrite was schedulable: I could put six weeks against it, and the blast radius stayed inside my own repo.

A data model leaks. Say you shipped fast and made user identity a single email column, because on day one every user is one person with one email. Two years later you have companies with shared accounts, people who changed employers, an audit log keyed on that column, three integrations customers built against the API that exposes it, and a support team with runbooks that assume it. None of that is your code. You cannot rewrite other people’s workflows in six weeks, and no AI is going to do that port for you either, because the spec isn’t in your old implementation. It’s in a thousand customers‘ heads.

So the asymmetry is worth internalizing before you take my advice. Ship the sloppy implementation. Be slower about the shapes that other systems will grow into: your schema, your public API, your identifiers. Those are the decisions where a design review at four people is still cheap and still correct.

AI made rewrites cheap. it did not make owning code cheap.

Rewriting is one of the things AI is best at. The behavior already exists, the edge cases are already encoded, the spec is sitting right there in the old implementation. A port that used to be a quarter of work is now a fraction of that, and you’re mostly paying in tokens.

Which is real. It’s also the smaller half of the cost.

The bill you keep paying is that you now own that code. Somebody has to read it at 3am when it’s misbehaving. Somebody has to extend it next quarter. If your team writes Python and the new thing is Rust, you haven’t just changed a language, you’ve changed who on the team can respond to an incident, and the answer might currently be one person. There’s a real onboarding cost and it doesn’t show up anywhere in the token spend. It shows up later, as the thing that slows everything down, and it never appears on the invoice that made the rewrite look cheap.

„python doesn’t scale“ loses that meeting, correctly

Eventually somebody asks why you’re spending six weeks on a rewrite instead of the feature request sitting next to it. „The p99 is 400ms and it should be 40ms“ is an answer. „Python doesn’t scale“ is a vibe, and it loses, correctly.

It is also the test that gets you an exception to a company’s approved-language list. Saving a large amount of money wins that room. „I like this language better“ does not, and everyone in it can tell the difference.

Which gives you something to do with the embarrassment you feel about your own codebase. Ask what it’s evidence of. If it’s evidence that real load broke real code and you have the graph to prove it, the rewrite is a prize and you should go collect it. If it’s evidence that you read a blog post about a language last weekend, it’s a hobby, and the feature request should win.

Galante spent the better part of two years shipping his game in the wrong engine. That is the only reason anyone knows his name.

DSGVO Cookie Consent mit Real Cookie Banner