Skip to content

My Tests Passed. My App Was Broken in 26 Places.

Every test I wrote passed honestly. The bugs that stopped the release were the ones I never thought to write down.

I clicked around my own fitness app for a week and came away pretty pleased with myself.

Test suite green. Unit, integration, end to end, all passing, all written against requirements I wrote myself. Every evening I opened it, logged in, added a workout, logged a meal, closed it. Everything I touched worked. I was about a week from putting it in front of actual humans.

Then I pointed an agent at it with a real QA prompt. Twenty-six findings came back, two of them P1, along with a recommendation not to ship.

The tests weren’t wrong, which is the uncomfortable part. Every one passed honestly against exactly what I told it to check. A test is a written-down expectation, and the bugs that kill you are the ones nobody wrote down. Manual clicking is supposed to cover that gap, and it doesn’t, because you walk the happy path you built. Your hands know where the app is solid and they route around the parts that aren’t, without ever telling your brain they did it. That week of everything-works was a tour of my own muscle memory.

1,490 calories and 242 grams of protein

The report opens politely. Sound local-first core, broad automated coverage, a mostly coherent design system. Clean historical sync loaded 104 sessions, 1,040 sets, 416 meals, and 364 health records with no data loss and totals that matched Postgres.

Then it tells me not to ship, and gives two reasons.

The first is that onboarding can talk itself into nonsense. Answer the questionnaire a particular way and it hands you six high-volume strength days plus a cardio day, warns you repeatedly that you’re above thirty weekly sets per muscle, and sets a target of 1,490 calories with 242g of protein and 23g of carbs. One questionnaire, guidance that contradicts itself, and a create button sitting there enabled.

The second is that shopping schedule rows collapse into one-character columns at 412px, an ordinary phone width, and again at exactly 1024px, the desktop rail breakpoint.

Below those: progress mixing date scopes and rendering bodyweight dates as 0... Hydration drawing eight 250ml blocks against a 2,500ml goal. Ramped workout targets summarized as four identical sets. Return-after-break logic cutting the weight while copying ten old sets forward. Fast reloads on a fresh device tripping an identity rate limit, showing no server data, and calling an HTTP 429 a server outage.

Both P1s were in screens I had opened dozens of times. I never typed a wrong value into my own app because some part of me knew better than to try, and I never resized the window to 412px because I develop at 1440. A fitness app is nothing but numeric input and I had validated almost none of it beyond what my own fingers naturally reached for.

why Playwright never replaced the human week

Every app I use is a little bit broken. Spotify does something wrong to my queue on a regular basis. The Deutsche Bahn app has a talent for showing me a connection it can’t actually sell me, which is a delightful thing to discover on a platform with four minutes to go. These are not garage projects. These are companies with QA departments, test suites, staging environments, budget.

And the bugs ship anyway, and nobody treats it as an emergency, and we’ve all quietly agreed that software is like this now.

More test cases isn’t the answer, because those companies have plenty. Written-down expectations catch what someone thought of. The bugs that survive to production are, by definition, the ones nobody thought of. You cannot write your way out of a blind spot.

Playwright doesn’t fix that either, and it’s worth being precise about why. A Playwright test says click this selector, expect this text. Somebody has to already know what to click and what to expect. It’s the transcript of an exploration a human did once. It replays curiosity. It doesn’t have any.

So the only thing that ever found the unwritten bugs was a person with time and no investment in the thing working. That person costs a week per release, which is why exploratory QA got budgeted, then trimmed, then skipped.

Which brings me to the part I got wrong for a while.

„do some QA“ is not a prompt

My first attempt was useless. I typed something close to „do a QA test on the app and figure out if there are problems,“ read what came back, and felt nothing. It had clicked around. It reported some observations. Nothing in it made me want to open my editor.

The agent did what I asked. „Is this app okay“ has no natural answer, and a model asked to find problems in an unfamiliar app will find something, and something is worthless.

The fix took six rounds and turned out to be four things: a map of the app, a stack of personas, a definition of what counts as a bug, and a report format I would actually read.

the app has to describe itself first

The first one isn’t a prompt at all. I wrote a features.md that lists every feature in the app and how a person uses it.

Not a spec. Not user stories. A tour. Log in with a provider. Land on a dashboard showing your current status, whether today’s workout is still undone, whether the day’s nutrition is missing. Then the workout half: build templates, build training plans, run a session off a plan, or start a free workout with no plan at all. And on down the list.

That file is the difference between an agent wandering and an agent working. Without it, the agent explores whatever the UI makes visible from the home screen, and anything three clicks deep behind a state it doesn’t know how to reach is invisible. With it, coverage is something I control instead of something I hope for.

It’s also just useful on its own. Writing every feature down in one place made me notice two I could no longer explain the purpose of. That was free.

give it people to be, not tasks to run

The part of the prompt that changed the most was the opening, where I stopped describing a task and started describing a person.

You are acting as a Principal QA Engineer, Senior Product Designer, Senior UI/UX Engineer, and adversarial end user. Your goal is not merely to verify that the application works. Your goal is to actively try to break the product, uncover confusing experiences, discover visual inconsistencies, find unexpected user journeys, identify missing states, and expose anything that would make a real user hesitate, misunderstand the interface, lose trust, or abandon the product. Assume that anything you do not test could contain a bug.

Then I give it people to be. A first-time user. An impatient one. A confused one. A power user. A user who makes mistakes, or changes their mind halfway through a flow. A user with incomplete data, and a user with years of it. Someone returning after a long absence.

That list does more work than any instruction about what to click. „Test the login flow“ gets you the login flow. „You are a user who changes their mind halfway through“ gets you a half-filled workout, a back navigation, and a question about what happened to the sets already logged. The archetype generates the test cases, which is the whole point, because generating test cases is the exact thing I’m bad at for my own app.

What an agent has that I don’t is that it doesn’t know better. It has no investment in the app working. It types nonsense into the field, follows the dead link, hits back four times in a row for no reason. It’s the world’s most patient bad user, and a bad user is exactly who’s about to download your app.

what counts as a bug, and what a report has to look like

Correctness against the feature list is the boring part, and it catches the boring bugs, which are still bugs. Then hostile input, everywhere a field exists: negative reps, a body weight of zero, nine hundred characters in a name field, text where a number goes, numbers with commas in a locale that wants dots. Then weird behavior, which is the fuzziest instruction and the one that pays best. Navigation that loops back on itself, buttons that do nothing, a back button leaving you somewhere impossible, state that survives a refresh in one place and vanishes in another. I described the shape of that rather than enumerating cases, because enumerating cases is exactly how my test suite failed and I wasn’t going to repeat it in a prompt.

The output contract matters as much as the hunting instructions. Screenshots on every finding, because a bug report without a picture is a claim I have to go verify by hand and I will not do it. Every finding graded P1 through P4, so I’m reading a queue instead of a wall. A report organized by feature, so I can see that plans are solid and nutrition is a swamp. And a rule I had to learn the hard way: if everything is marked P1, nothing is.

finding and fixing want opposite postures

One report is a to-do list, and a to-do list you run once is worth about one afternoon. The loop is where it compounds.

Split it in two: a skill that finds and a skill that fixes. Run finding. Feed the report to fixing. Run finding again on the result. Repeat.

They want opposite postures, which is the whole reason to keep them apart. Finding should be adversarial and paranoid, which is what the persona stack buys. Fixing should be surgical and narrow. Mash them together and you get an agent that grades its own work, which reliably discovers that its work is fine.

Don’t try to write the finding prompt perfectly up front, either. I tried, and that’s how I got the useless report. Write a small one, run it, and pay attention to the part where you’re disappointed, because that’s the spec. Every time I thought „that’s not what I meant,“ the fix was a sentence I’d left out. Then hand the prompt back to the agent and ask what’s ambiguous. That works better than it has any right to.

it does not converge

Every fixing pass is new code, and new code has new bugs, so the next finding pass has something to say. The findings do get smaller. P1s stop appearing and you end up arguing with an agent about P4 spacing, which is a good place to be but is not zero.

Two other costs worth naming. Twenty P2s is a lot of report to read, and the ratio gets worse as the real bugs run out, so eventually you’re paying an attention tax on cosmetic complaints. And agents do file confident findings for behavior that’s actually correct, which burns a fix pass and, worse, can talk you into changing something that was fine. So I read every P1 myself before it goes to the fixer. That part stays manual, and I don’t see it becoming automatic.

It’s a way to push the floor up, not a proof of correctness. You stop when the top of the report is boring. You don’t stop when the report is empty, because it may never be, and waiting for empty is how you never ship.

the habit outlived the expense

Cutting exploratory QA was the right call for twenty years. A week of a person’s attention is real money, and the bugs it catches are the ones you can’t point at in advance, which makes it the easiest line in any budget to defend cutting. Nobody was being stupid.

Then the audit above ran overnight against every feature in the app for the price of some tokens, and the math that justified the cut stopped holding. Not gradually. It just stopped, recently enough that most teams are still running the old calculation and don’t know it.

Which is the part I’d actually think about, past the QA angle. That cut was rational once and became a habit, and the habit didn’t notice when the reason expired. It’s worth asking what else in your process is a sensible response to a price that no longer exists.

I put this off longest because manually clicking around felt like it was working.

It wasn’t working. It was just quiet. Peace, nerds.

DSGVO Cookie Consent mit Real Cookie Banner