I stopped writing most of my code, so slow builds became a cost I don’t pay and type errors became my review pass. Four apps, one Rust core, and the fifteen-minute CI bill that came with it.
I picked Rust for a shared backend that four apps will sit on, and the requirement that decided it is one I would not have written down two years ago.
Not performance. Not the type system, at least not for the reason you’d guess. It’s that I don’t write most of the code in this repo anymore. I describe, an agent produces, I review. Once that’s true, the compiler stops being a tax on typing and becomes the first thing in the pipeline that has actually read the whole change.
That inverts the scorecard. Every complaint anyone has about Rust is a writing-side cost. Slow builds, verbose error handling, a compiler that argues with you. I’m barely on the writing side. Everything I bought sits on the reviewing side, where I now spend my days.
So I paid for it. Fifteen-minute CI builds, and a target/ directory that eats 100 gigabytes of my laptop. Those are real and I’ll show you the invoice rather than wave it away. But when an agent hands you plausible code, plausible is the dangerous word, and something has to read it before you do.
the first cut is arithmetic, not taste
I’m building four apps. A fitness tracker, a meditation app, a journaling app, and one that blocks other apps across your phone, your desktop, and your browser. Same reason behind all four: the people around me are heavier and more welded to a screen than they were ten years ago, and I would like to do something about that other than complain.
Four apps means four backends, unless you’re not an idiot, in which case it means one shared core and four thin things on top. That core is going to outlive every one of these products. There will be a fifth app on it.
So the first constraint is that the thing has to run in a container so small that hosting it rounds to free. This is not a benchmark contest. It’s a solo builder with four products and no revenue looking at a hosting invoice. If each service idles at a gigabyte, four apps in three environments is a number I have to justify to myself every month. If each service idles at forty megabytes, I stop thinking about it entirely, and thinking about it is the expensive part.
That single line kills Python and Java before we’ve discussed a single interesting property of either. Both of them can sit at a gigabyte of RAM without doing anything remarkable. The JVM in particular is a fantastic piece of engineering that assumes you have a machine, and I want to pack a lot of small things onto one cheap box. Same story for anything that ships a runtime and a garbage collector tuned for throughput on hardware I’m not renting.
I like Python. I’ve shipped plenty of it. It’s just not what you build a shared core on when the shared core has to be cheap to keep alive in five copies.
There was a second constraint on the list, the ecosystem, and it’s the one that never eliminated anything. I went in expecting to hit a wall and didn’t. REST APIs go through Axum, OpenAPI specs come out of utoipa, which was the piece I worried about most, because „we have a web framework“ and „we can generate a spec my client tooling actually consumes“ are very different maturity levels. Publishing my own crates was uneventful, which is the highest compliment you can pay a package manager.
So the pile is Rust and Go. Which is where most stack posts end, with a shrug about how both are great and it depends on your team. But I don’t have a team. I have agents.
the compiler is the first thing that actually reads the code
An agent produces a lot of that code in a hurry, and it reads correctly. It uses real function names. It has the shape of something that works.
In a dynamic language, the way you find out it doesn’t work is by running it, in the specific circumstance that trips it, which might be a code path nobody exercises until a user in production does. The gap between „the agent wrote it“ and „I know it’s broken“ is measured in days and sometimes in incidents.
In Rust, a whole category of that gap closes before the code runs at all. Wrong type, forgotten error case, a value used after it moved, a lifetime that doesn’t hold, an enum match missing a variant. Those aren’t test failures. They’re not failures at all, they’re refusals. The code never becomes a binary.
This is the old „if it compiles, it probably works“ line, and I used to think it was a bit of a meme. It reads differently when the code isn’t yours. When you wrote it yourself, you carry a model of what it’s supposed to do, and the compiler is mostly confirming what you already believed. When an agent wrote it, you have no such model, and the compiler is the first entity in the pipeline that has actually read the whole thing carefully.
So the strong type system stopped being an aesthetic preference and turned into infrastructure. I build an abstraction, I encode the rules in types, and the compiler checks every piece of code that touches it against those rules, whether I wrote it at 2am or an agent generated it in a loop I wasn’t watching. The abstraction is enforced instead of documented. Everything in the core carries doc comments for the same reason, so that reviewing a generated call site doesn’t mean opening three files to find out what the thing was supposed to do. That’s what broke the tie in Rust’s favor rather than leaving it level with Go. Go’s type system is fine. It is not doing this job.
everyone measures the corpus. the corpus isn’t the variable
The obvious objection is that all of this is downstream of training data, and on training data the old languages win by tonnage. Python since 1991, Java since 1995, an ocean of both in every model, against a Go that showed up in 2012 and a Rust that hit 1.0 in 2015. But volume and consistency are different quantities. Thirty years of Python includes thirty years of tutorials, abandoned Stack Overflow answers, three incompatible eras of packaging, and a decade of Python 2 that is now just wrong. A model trained on all of that has seen the good and the bad and doesn’t get told which was which. Rust’s corpus is smaller and mostly from a world where the modern idioms already existed. Less material, less contradiction.
I want to be honest about my confidence here, because this is the weakest link in the argument and I’d rather name it than let you find it. I can’t hand you a benchmark. What I can tell you is what I observe in my own repo, which is that agents produce Rust I have to argue with less than the Python I used to get. That’s my experience, not a study. Which is exactly why I don’t rest the decision on it. If the corpus argument turns out to be wrong, the compiler argument still stands, because that one doesn’t depend on the agent being good. It depends on the mistakes being catchable before they run.
now the bill
Start with the disk. A single project’s build artifacts run somewhere between 50 and 100 gigabytes. That’s not a typo and it’s not one weird dependency tree. Every crate you pull in gets compiled locally, in multiple configurations, and it all lives in target/. If you’ve got a few projects going, you will run out of disk on a laptop that felt roomy last year. cargo clean becomes muscle memory. It’s annoying and there’s no clever fix, just a habit.
Then compile time, which everyone knows about already. What people undersell is that it hurts in two different places, and only one of them is the one they complain about.
Locally, honestly, I don’t care. I hit build, I go read the diff the agent produced, the build finishes. It fits in the gap I was going to spend on review anyway. Slow local builds are a papercut.
CI is where it stops being a papercut, because CI is metered. You pay for pipeline minutes. Every push, every PR, every deploy. A build that takes fifteen minutes instead of three isn’t an inconvenience, it’s a line item that grows with every commit you make for the rest of the project’s life. That’s a real cost that shows up on a real invoice, and it’s the one thing about this decision I’d call a genuine loss rather than a tradeoff.
So compare those minutes against what they replace, which is a test suite you would otherwise write and maintain to catch the same class of mistake. A test suite is code I have to author, keep current, and trust. The compile is a check I can’t skip and never have to update. Priced against the tests I’m not writing, fifteen minutes of CI stops looking like pure loss.
You can also fight the number directly. Cache aggressively, split the workspace so a change doesn’t rebuild the world, don’t rebuild dependencies that didn’t move. It gets better. It does not get to „fast,“ and anyone telling you otherwise is comparing against a cold Docker build.
Still a real bill. I just think I’m getting something for it.
where the reviewer fights back
The honest version of this argument has to admit where it goes bad, and it goes bad in a specific place.
A compiler that refuses wrong code also refuses code that is merely unfamiliar. Agents thrash against the borrow checker. You’ve seen it if you’ve worked this way: a model that clones its way out of a lifetime error, or wraps something in an Arc<Mutex<_>> because that made the message go away, or circles the same three failed fixes until you step in and do it yourself. That’s a refusal costing you review time without catching a real bug. The strict reviewer generates false positives, and false positives are not free.
I think the trade still comes out ahead, because a loop I can see is cheaper than a bug I can’t. Thrashing happens in front of me, in a build I’m already watching. The dynamic-language alternative fails silently and shows up later, in production, on a path nobody tested. But if your agent is spending more of its time fighting the language than writing the feature, that ratio can flip, and pretending otherwise would be dishonest.
what I’d tell you to copy, and what I wouldn’t
Don’t copy Rust. Copy the elimination.
Write your constraints down in the order they actually bind, then let them delete candidates. Cheap to run removed two languages before I’d thought about anything fun. Compile-time enforcement broke the tie between the two that survived. The ecosystem check never fired.
Your list will differ, and it should. If you have four engineers who all write Go and ship weekly, that’s a constraint that outweighs everything I’ve said, and Go idles small too. If you’re building something where time-to-first-endpoint is the whole game because you don’t yet know if anyone wants it, pick the language your hands already know and worry about the container size when the container size costs you money.
But run the agent requirement, because it’s not a question about the language in isolation. It’s a question about how quickly you find out the generated code is wrong. Languages that answer that at compile time are worth more than they were when you were the one writing every line.
the invoice lands on the side I left
For most of my career this swap would have been a terrible one. I was the writer, the reviewer was me on a good day, and paying in build minutes to buy a second reader made no sense when the first reader already knew what the code was for. What changed isn’t Rust and it isn’t the compiler. It’s that I stopped being the first reader.
So the question isn’t which language you like. It’s what fraction of last week’s code you actually typed. Work that number out honestly, because it decides which costs on your invoice are taxes and which ones are insurance. Mine turned over last year, and I don’t think it’s turning back.