Ports and adapters, three test layers, Prometheus by name. I got the parts that look like code.
I got tired of paying a subscription to log that I ate rice.
That’s genuinely how this started. Every fitness app on the market wants nine euros a month to be a calorie counter and a stopwatch. I’ve been building software for a living for years. I looked at that and thought what everyone reading this has thought at least once: I could do this myself. In a weekend.
So I did it properly. I wrote the feature set into a Markdown file, handed it to Fable to review and extend into a real implementation plan, and it caught gaps I’d left. GPT-5.6 Codex implemented, Fable orchestrated. Two frontier models, a written plan, and someone who has shipped production systems steering them.
Four hours later I opened it and it worked. Not „compiled.“ Worked. Screens, navigation, data going into a database and coming back out. That’s the video everyone makes. Prompt, cut, working app, look how easy this is now.
Nobody posts week three.
The weekend ended six weeks ago and the app still isn’t deployed. This isn’t an „AI can’t code“ post, and I’d close the tab on that too. The agents wrote good code. That was never the bottleneck.
The bottleneck was everything that isn’t writing code, which turns out to be roughly the entire job. None of it is the part I could describe in a Markdown file. Most of it is the part I only know from having been on call.
you cannot prompt for what you forgot
The requirements weren’t exotic. Native mobile on Android and iOS, browser access and probably a PWA, and one backend fully decoupled from every client. That last one is the only genuinely opinionated decision in the project, and I’ll come back to why, because it’s also the reason I’m still going.
Then I started clicking around as a user instead of an author, and the gap opened. Not bugs yet. Absences. Things that were so obvious to me I never thought to write them down, and so invisible to a model that it never thought to ask.
Authentication, for one. I wanted Keycloak, Clerk, and WorkOS as interchangeable identity providers, because I refuse to hardcode my own vendor lock-in on day one. Obvious to me. Nowhere in the plan, because it lives in the part of my head labeled „how software is built,“ not the part labeled „features of a fitness app.“
So I re-prompted, and the agents integrated it. Competently. That’s the trap, actually. Every one of these gaps is cheap to fix once you notice it. The cost isn’t the fix. The cost is that noticing requires a person who already knows what a finished system looks like, using the thing, with taste, over and over. There is no prompt for that. You cannot ask a model to tell you what you forgot, because the entire problem is that neither of you knows it’s missing.
Some of what’s missing isn’t knowledge, though. It’s judgment.
agents have no eyes
The design wasn’t broken. It was worse than broken, it was fine. Default spacing, default palette, every element aligned to the grid and dead on arrival. It looked like a settings page wearing the costume of an app.
This is the failure mode nobody warns you about, and it’s structural. An agent can verify that a button works. It cannot look at the screen and feel that something is off, because looking at the screen and feeling something is the one input it doesn’t have. Correctness it can check. Taste is not a test that passes.
What worked was routing around it. I took the current app into Claude Design, gave it the feature set, and had it regenerate the design as an actual design artifact. The output was better than what the coding agents produced on their own, which was a low bar it cleared easily. Exported it, handed it back to Claude Code, had Fable implement it against the real screens.
The screens that came back looked like someone had made choices. The reason it worked is narrow. I stopped asking a coding agent to have an opinion about visual design and gave it a design to implement instead. Translating a spec into components is what these things are excellent at. Deciding what the spec should be is what they’re worst at. Every hour I’ve saved with agents came from getting that boundary right, and every hour I’ve wasted came from getting it wrong.
the bugs do not stop
Then I used my own app for a week, and found bug after bug after bug.
Not crashes. Crashes would be a gift, crashes are loud. These were the quiet ones. A number that’s right until you cross midnight. State that survives a refresh in one tab and not the other. An edge case in unit conversion that only shows up at exactly the weight I lift, which is the kind of bug you only find by being the user.
So I did what any reasonable person does when they’ve automated one thing badly. I automated the thing on top of it. I wrote an agent loop to do QA. It runs, it exercises the app, it writes me a report.
It works. It finds real bugs. And it never finishes. Every pass surfaces a fresh handful of small wrongnesses, and I’ve come to think that’s not a defect in my loop, it’s just what the thing is. The QA agent is a bug generator with extra steps, because the same process producing the fixes is producing the next batch. Somewhere in there is a fixed point where it converges. I have not found it and I’m no longer sure it exists.
I asked for the whole stack. I got the parts that look like code.
I didn’t only ask for features. I asked for the way I build things. Clean structure, ports and adapters, domain logic that doesn’t know what a HTTP request is. Unit tests, integration tests, end to end tests. I named all three. And my stack, by name: Postgres, Valkey, Prometheus for metrics, Grafana for dashboards, Tempo for traces, Loki for logs.
Structure? Got it. It’s good, honestly. Tests? Written, and they run.
Observability? It wired up the libraries. Metrics get emitted. Traces exist, technically. And then nothing. No dashboard that tells me anything. No alert that fires before I notice. No span where I’d actually want one, which is the boundary between my code and the thing that’s slow. It produced the artifacts that look like observability from the outside.
Which makes sense when you think about what „done“ means to each of us. To an agent, done is a green test and a satisfied instruction. Prometheus configured, done. To me, done means I get paged before a user notices, and I can answer why in four minutes at 3am. That definition doesn’t live in the code. It lives in having been woken up.
nothing was written as if it would fail
Resilience went the same way. Every network call assumes the network. Every dependency assumes the dependency. It’s not that the agents write fragile code. They write for the happy path unless you list every unhappy one, and knowing which unhappy paths are worth paying for is most of what senior engineering is.
The fair objection here is that this is a fitness app with one user. No SLA, no pager, nobody to page. An engineer who quietly deprioritized Grafana dashboards and circuit breakers for a personal project would be making a defensible call, and you could argue the agent made the same one.
Except it didn’t make a call. Making that call would require weighing my context against the cost, and nothing in the output suggests that happened. It installed Prometheus because I wrote Prometheus. It emitted metrics because that’s what the library does when you wire it up. What came back is the shape that pattern-matches the words I used, which is a different thing from a judgment that happens to agree with it.
So I got a codebase that passes review at a glance and has none of the properties that make software survivable. And I asked for them. In writing. By name.
and it still isn’t deployed
The app runs on my machine. It is not in anybody’s hands, including mine on my phone, because getting it there is its own project and it’s the least automatable work in the whole thing.
CI, containers, environments, secrets, migrations, rollback are fine. Tedious, well-trodden, agents help. That’s the easy half.
Then there’s the App Store. Signing certificates. Screenshots at specific dimensions. A review run by humans who reject you for reasons you get to guess at, on their schedule. Play Store has its own parallel maze. None of this is hard, exactly. All of it is a browser, a form, an account, and a wait, and there’s no API to point an agent at because Apple would very much prefer there wasn’t.
You could point a computer-use agent at it and watch it click through screenshots, but today I’d spend longer supervising that than doing it myself, and a silently misconfigured signing identity means debugging a machine’s mistake in a domain whose error messages were already hostile to humans.
So the last mile of shipping software, the one that decides whether a thing exists or not, is a UI written by a company with no incentive to let you automate it. That’s not a model capability problem. No amount of scaling fixes it.
the fast part was never the job
The claim isn’t fake. The initial build genuinely was hours, and four years ago it was weeks. That’s real and I’m not giving it back.
But agents didn’t make building software faster. They deleted the visible part of it and left everything else standing at full size. Typing used to sit on top of the judgment, the taste, the argument about which failure modes deserve engineering, and the afternoon lost to a provisioning profile. It hid how much of the job those were. Take typing to zero and none of them shrink. They just become the whole surface of the project, which is why six weeks feels so much slower than it used to even though nothing actually got slower.
The people selling you the weekend aren’t lying about the first four hours. They’re selling you a demo and calling it a product, and the difference between those two things is the entire profession.
the valuable part was never the app
So why am I still spending evenings on a calorie counter. Because of that one opinionated decision, the backend decoupled from every client. I want to point my own agent at my own training data and ask why my squat stalled, and have it write next month’s block, and dump a meal into it in plain text and let it do the macro math. Every app is bolting a chatbot into a sidebar right now, and I think that’s backwards and temporary. What survives is an assistant that already knows your calendar, your health data, your inbox, and can reach into any backend that will have it.
Which means the valuable part of my app was never the app. It’s an API that isn’t hostile to a client I didn’t write, and nobody sells that, because no subscription business survives it. That’s worth six weeks of evenings. It’s the one thing the agents couldn’t shorten and the one thing I couldn’t buy.
Just not this weekend. Peace, nerds.