Subscribe
Digital Strategy Dispatches · Build notes

Your Users Are Too Busy to Test Your App. Build a Test Harness.

"Test harness" doesn't roll off the tongue like "bacon," but it's just as tasty. It does the grunt work of testing so your app ends up a lot less rickety than you fear it is.

Key takeaways

  • Write a brief with hard rules first: never test against production, no real member or voter data.
  • Tell the model to use best practices, then make it explain the plan. The plan review is where most of the value shows up.
  • Keep checkpoints where you, the person who knows the work, make the calls. Fixes, merges and going live stay yours.

What do you do when your users are too busy, or too leery, to help small-group test an app?

Build a test harness. I know “test harness” doesn’t roll off the tongue like “bacon,” but it’s just as tasty. It does the grunt work of testing so your app ends up a lot less rickety than you fear it is.

Why I needed one

I built an election protection app for my org. Volunteers at polling places use it to file incident reports, hit a panic button if something goes wrong and send check-ins through the day. Staff watch it all on a live dashboard, and the app texts the right people when something serious comes in.

Election Day is four weeks out. The people who would normally test it are organizers and volunteers who are already stretched thin, and some of them are understandably nervous about poking at a tool they’ll depend on. Asking them to click through a test script was not going to happen at the scale we needed.

So I handed the job to Claude Code and asked it to build a test harness: a set of automated tests, fake users and fake data that exercise the app the way real people would.

Ask for the best and smartest

Tell your model (Claude Code, Codex, Antigravity, whatever you use) to use the best and smartest practices to do it.

Why that modifier? Focus. Instead of looking for a solution, it’s now looking for best practices, and the most efficient of those. At least that’s my AI hope.

Make the model review the plan with you, so you can see and question what makes it “best” and “smartest.” That review is where a lot of the value showed up for me. Before writing a single test, the model read my brief and pushed back on it:

  • My plan had no dates. Five checkpoints, Election Day 29 days out and no deadline on any of them. It suggested a code freeze in the final week so the last stretch is checking, not building.
  • I had the riskiest tests last. The tests for who can read and write what in the database were near the end of my plan. Those were the biggest known risk and needed no setup, so it moved them up.

None of that is magic. It’s what a good senior engineer would say in a plan review. The difference is I got it in a minute, could ask “why” on each point and then decide.

What the harness actually did

It found real problems before any tests ran. The first phase was just recon: read the code and map every screen, every user role, every form and every place the app reaches out to the world. That alone surfaced three serious security holes in the live app, including a way for an ordinary account to give itself top-level admin rights and a bug that re-sent old alert texts to staff every time the server restarted. No volunteer would have found those by clicking around. We wrote and reviewed the fixes right away.

It never touched the live app. Everything runs on a local copy of the database on my machine, set up so it physically can’t reach the real one. Every text and email the app tries to send lands in a test inbox instead of someone’s phone. All the users and reports are made up, and deliberately messy: missing fields, duplicate names, weird addresses, very long text.

It defined “correct” before testing for it. Before any test was written, the model drafted the nine tasks that cannot fail on Election Day, such as filing a report with no cell signal or hitting the panic button. For each one it wrote who does it, what should happen, how fast and what must never happen. It marked every expected result it guessed from the code as “inferred,” so I knew which ones needed my judgment. Where our field manual and the code disagreed, it flagged it as my call, not its call.

It found the bugs real users would hit. Once the tests ran, the harness logged 32 findings. A few of my favorites, because they’re exactly what a stressed volunteer would run into:

  • The panic button texted some staff twice for one emergency.
  • A volunteer who ignored the phone’s “share your location?” prompt couldn’t submit a report at all. The form just sat there.
  • An automatic check-in feature was quietly failing and retrying every minute.
  • Leaving a half-written report to check another screen threw the draft away.

It stopped and asked. The plan has a checkpoint at the end of each phase where I review what it found and decide whether to keep going. It was not allowed to change the app’s code without proposing the exact change first. Any fix it wrote went up for my review, and merging and going live stayed my decision. When it got something wrong, like which of two copies of the code on my laptop was the real one, it said so plainly and reversed itself.

It told me what it couldn’t cover. The final report lists the gaps: tests it wrote but couldn’t run yet, and areas it only reviewed by reading. That list is as useful as the findings, because it tells me where I still need a human.

It’s reusable. The harness lives outside the app and points at it through a config file. The next app gets the same treatment with a new config, not a new project.

What I’d tell another nonprofit developer

You don’t need a QA team to take testing seriously. You need:

  1. A written brief with hard rules. Mine started with “never run tests against production” and “no real member or voter data.”
  2. A model told to use best practices, and a habit of making it explain its plan.
  3. Checkpoints where you, the person who knows the work, make the calls.

Your volunteers’ time is better spent on the real thing. Let the harness do the grunt work.

What’s next

I set this up in Claude to test Claude-built projects, but I’m likely to run the same harness in Codex next. Different take, different token budget too.

#notAISlop

More dispatches

All dispatches
Get the dispatch

What I built, where I stand, what I saw.

One issue a week. No filler.