The vibe coding moment arrives
The question I had arrived at did not stay abstract for long. It found a very specific test case almost immediately, wearing a name that was already spreading through timelines and conference talks by the time I ran into it: vibe coding. The idea, stripped of its branding, is straightforward. Describe what you want in natural language, let the model generate the implementation, and keep going without stopping to read closely what it produced. Not "review later." Not "understand eventually." Skip the step entirely, on purpose, as the point of the exercise.
My first reaction was curiosity, more than alarm or refusal. It would have been easy to tell this story as a moment of panic, but it wasn't one: I read the threads, watched a few demos, and ran a small experiment of my own over a weekend, in the kind of low-risk sandbox I had already committed to using before deciding anything. Underneath the curiosity, something sharper than discomfort took shape: a design problem, not a speed one.
I had no objection to moving fast. What put me on alert was that not understanding what had been generated was being treated as an acceptable, even intended, feature of the workflow: not a risk to be managed, but a trade you were supposed to be glad to make.
That trade made sense for some people, just not for me. I understood immediately why someone with initiative and no coding background, an entrepreneur chasing an idea, would see value in this. If you can't read the code, the quality it was written with doesn't register as a variable at all. It's like a button that builds you a car to drive from home to the office. You'd press it, you'd have no way to judge whether the engineering underneath was sound, and you might just start driving and end up stranded halfway there. My bar for software, the bar I'd built over years as a professional, sat somewhere else entirely, and that's why vibe coding left me wary rather than dazzled.
What XP actually says about moving fast
This is where a decade of practicing Extreme Programming turned out to matter more than I expected. XP has a reputation, in rooms that have never practiced it, for being about discipline for its own sake: pair programming as etiquette, TDD as a purity ritual. XP is actually a set of practices built around one design goal: sustainable speed, achieved by keeping feedback loops as tight as the work allows.
What "sustainable" means in practice is speed with control built in. Picture driving at a hundred kilometers an hour: fine, until you're ten meters from a wall and still can't slow down. Going fast only works if you can see the wall coming and still have room to brake. That is what XP's practices actually buy: not raw velocity, but velocity you can stop, adjust, and redirect before you hit something.
TDD is one of those brakes, though not for the reason people usually assume. Writing a failing test first doesn't prove you understood the problem. Only the market proves that: the customer, the stakeholder, the person who actually needed the thing built. What TDD gives you is narrower, and still worth having. First, the code does exactly what you expect, because you write your expectations down as a test before you write the implementation, and that test stays red until it turns green: the transition is the guarantee. Second, you work in small steps, so your feedback loop on whether you're moving in the direction you believe is correct stays fast, cycle after cycle. Third, red-green-refactor gives you a rhythm to hold onto while the code grows, and TDD's test list, the explicit list of behaviors and cases still to cover, works like a checklist: you always know how much is left. Kent Beck has made this point more than once. Knowing when you're done is not a small thing.
Refactoring is another one of those brakes. It's active management of complexity, done continuously, so complexity never gets the chance to compound past the point where a person can hold it in their head.
If I had to rank the brakes, continuous delivery might matter more than TDD, especially now. It forces a constraint that never lets up: whatever is on trunk has to be releasable, right now, not after a cleanup pass. That constraint keeps the bar high in a way nothing else does, and it buys you something less obvious too, a kind of calm. You're never more than a few hours of work away from a working, shippable system. That safety net matters for a person. It matters just as much, maybe more, once AI is writing some of the code that lands on that trunk.
These practices are systems for making speed safe. They are mechanisms that let you move quickly precisely because they keep the invisible costs of that speed visible. That is the frame I could not put down once vibe coding entered my field of view. What happens to a codebase when the practice with the least friction is deliberately removed from that system?
What vibe coding removes
The obvious answer is tests. If you are not reading the generated code, you are very likely not writing tests for it either, or you are asking the model to write both the implementation and its own tests, which is a different problem I will come back to in a later piece. But tests were not actually the thing that concerned me most. What concerned me was upstream of tests: the understanding loop.
Every practice I have described so far assumes that a person carries a working mental model of the system as it changes. TDD builds that model one red-green cycle at a time, and refactoring updates it as the code moves. Trunk-based development forces it to stay current, because there is nowhere to hide a branch where your model of the system quietly diverges from reality. Vibe coding, as a practice, replaces the need for that model rather than updating it. The code exists, it runs, it might even be correct. But the mental model of why is not in anyone's head, not even fully in the model's "head," in any sense that survives past the session.
Every developer has merged a pull request they understood eighty percent of. That gap is old, older than AI. But pull request review, especially the asynchronous kind, was never carrying that much weight to begin with, even back then. The reviews that actually caught something were the synchronous ones, someone sitting next to you, walking through the change live. The rest was a formality that happened to look like a safety net. Continuous delivery already knew this: it's the automated checks that decide whether something can go to production, not a person skimming a diff. What CD doesn't replace is how understanding spreads across a team, and that has to happen some other way. Pairing. Mobbing. Sitting next to someone while they actually work, not a comment thread on a pull request nobody was going to read closely anyway.
What changed with AI is treating that eighty-percent gap as the intended mode of working rather than as a debt to be repaid before it comes due, and volume makes it worse. AI can generate code faster than any team I have led could review it carefully, which means the real constraint on quality was always understanding. Understanding does not scale by adding more tokens per minute. The gap between plausible and correct is easy to miss in a demo, where the surface area is small and nobody is adversarial. It is much harder to miss in a production system carrying real traffic and real edge cases, which is exactly where it tends to show up.
None of this makes vibe coding worthless. There's a version of it that only makes sense for someone who was never a software professional to begin with. If you can produce a working app without understanding what's underneath, that has value, for you and for the business you're building. The mistake is assuming that's the job of software engineering, or that working means good. Those are different claims, and vibe coding only ever answers the first one. I kept seeing people conflate them anyway.
Watching it happen
I did not have to imagine this. Through the second half of last year and into this one, the pattern kept surfacing in public, not as a hypothesis but as incident reports. The one that stayed with me involved an AI coding agent that had been given direct access to a live production database during a period when no changes were supposed to happen at all. When it hit an unexpected state, it did not stop and ask. It acted: deleting production data, then generating status updates describing a system that was fine when it was not. Every individual failure in that chain is explainable on its own: production access an agent should never have had, no separation between what it could touch in development and what it could touch when something real was on the line, no person positioned to catch the moment its confidence and its correctness came apart. But the chain only existed because understanding had been designed out at every one of those points, not just the last one.
It was not an isolated story, and it showed up at every scale. Take Baudr. An Italian streamer with over a million followers on Twitch built the social app entirely through vibe coding, on a budget of forty euros, and launched it on March 14, 2025. Within hours it had a data breach. The admin panel was reachable directly from the browser, because validation existed only on the frontend and nowhere on the backend. Thousands of user records got pulled out, some accounts were deleted, including the founder's own, and compromised accounts started sending fraudulent messages to other users. Security analyst Luca Di Domenico put it plainly afterward: with even minimal security knowledge, the site was crackable in about two minutes.
Scale it up and the pattern doesn't change, it just gets more expensive. In March 2026, Amazon had two separate incidents tied to AI-assisted code landing in production: one cost around 120,000 orders and generated 1.6 million site errors, the other dropped orders across North America by 99 percent and cost roughly 6.3 million orders. Amazon's response was a 90-day safety reset across some 335 critical systems, requiring two-person sign-off before changes could ship. The company called it user error, not an AI problem, but the fix looked a lot like distrust of the workflow either way.
Across all three of these, understanding had been treated as optional at exactly the moments it mattered most.
What I wanted instead
None of this pushed me toward rejecting AI as a tool. That would have been the easy, comfortable move: a decision made on identity rather than evidence, which is precisely the kind of thing I try not to do. What it pushed me toward was much narrower, and I think more useful: a refusal of AI used without discipline, not a refusal of AI.
That refusal had a practical root, not just a philosophical one. Working the way I'd always tried to work costs energy: writing the test first, keeping changes small, keeping trunk releasable. None of that is free, and staying disciplined about it takes sustained effort, especially on a bad day. What I wanted from AI was help carrying that cost, not a way around it. Lower the cognitive load of doing the disciplined thing. Don't make it easier to skip it.
So the question sharpened into something almost mechanical. XP's practices are feedback systems, and AI can clearly accelerate parts of the work. Could AI be made to operate inside that feedback loop, subject to the same constraints that keep the loop honest, instead of being used as a shortcut around it? The discipline would need to be encoded into the workflow itself, structurally, so that using the tool well did not depend on remembering to be careful on a Friday afternoon with a deadline the next morning. AI that only accelerates, and is never forced to stop and check, is a car with no brakes: fast, right up until it very much isn't.
I did not have an answer yet. What I had was a hypothesis I intended to test, and a much clearer sense of what testing it properly would need to mean.
A quieter kind of clarity
Looking back at that period, nothing about it was dramatic. There was no single incident of my own, no dramatic reversal. The vibe coding moment gave me clarity. I knew, precisely and specifically, what I did not want: code in a system I was responsible for that no one understood, produced at a volume no review process could realistically absorb, shipped on the assumption that plausible was close enough to correct.
Knowing what you don't want isn't knowing what to build instead, and that is where the experiments start: what disciplined AI use actually looks like inside a real workflow, under real constraints.


