Building apps with AI AI writes fast. I make sure it's right.
By Jop Middelkamp ยท Updated 11 October 2026
I build apps with AI coding agents. They write code faster than anyone can type and happily churn out the boilerplate and tests nobody likes writing. They also hand you code that's almost right, with total confidence, and shipping that is how AI projects go wrong. So I treat everything an agent writes like a pull request from a new colleague: often good, and always checked.
Where AI goes wrong
The research on this is pretty clear:
- It feels faster than it is. In METR's 2025 trial, experienced developers expected AI to speed them up, still felt faster afterwards, and were measured 19% slower. METR has since called that result out of date and sees a small speedup in newer data, but the lesson stands: you can't judge AI speed by feel.
- In Stack Overflow's 2025 survey, 66% of developers named AI answers that are almost right, but not quite, as their biggest frustration.
- Veracode, which sells security scanning, found that AI-generated code passed its security tests only 56% of the time in 2026, virtually unchanged from a year earlier, while the models themselves kept getting smarter.
- Code models also make up dependencies. In a USENIX Security 2025 study, almost one in five package suggestions from code models named packages that don't exist. Attackers can register those names and wait.
- Agents with too much access break things. In July 2025, Replit's agent deleted a production database during a code freeze. Incidents like that share one cause: broad permissions and nothing to stop a destructive command.
How I keep it honest
Each of those problems has a counter-move:
- Every feature starts as a short written plan, so the agent knows what done looks like before it writes a line.
- An agent stops when the work looks done, as Anthropic's own Claude Code guide admits, so my tests and CI decide whether it really is.
- Agents work on small, scoped tasks, and I read each change before it's merged.
- Agents get test credentials, never production ones. Publishing and anything that deletes data go through me.
- Login, secrets and personal data get a line-by-line review, as the UK's National Cyber Security Centre advises for AI-assisted code.
- A new dependency only goes in after I've checked that it exists and is the real thing.
AI inside your app
Putting an AI feature in an app brings its own problems. The same question can get a different answer every time. The model you build on gets retired: Anthropic retired Claude Sonnet 4 and Opus 4 about 13 months after launch. And according to OWASP, it's unclear whether prompt injection can ever be fully prevented.
So the model call lives on a server, never in the app. That keeps API keys out of the app, which Google's Firebase docs insist on, and I can switch models without waiting for an app update. I test AI features with evals: a fixed set of example inputs with known good answers, rerun after every prompt or model change.
The rules count too. Since November 2025, Apple requires explicit permission before an app shares personal data with a third-party AI. Since 2 August 2026, the EU AI Act requires that people know when they're talking to an AI and that AI-generated content is marked. That part wasn't postponed. I build the consent screen and the AI notice in from the start.
What I won't promise
I won't promise you "10x faster". The best controlled studies show modest gains, and Google's DORA research found that AI amplifies whatever a team already does, good habits and bad ones. You get the speed AI really does give, with checks that keep its mistakes out of your app.
See it for yourself
My own app Ergates, a mobile app for Hermes Agent AI assistants, is built this way: design documents first, a test suite, and CI that runs on every change. The code is open on GitHub.
Let's build something
Want to build an app with AI, or add AI to the one you already have? Send me a few lines about it.