AI Can Write the Code. Someone Still Has to Own the Decision.

Also published on LinkedIn, where the discussion happens.

I had an interesting conversation recently with someone working in Big Tech, about whether code review and architectural understanding are becoming irrelevant. Their take: the future is “focus on requirements, ship behind a toggle, watch the metrics.” If the metrics look bad, roll back. If they look good, keep going. Clean code and architecture, in their view, are a human indulgence the industry can no longer afford, and arguably a waste of effort AI will route around anyway.

I’ve spent over a dozen years on one codebase - a promotions engine with tens of thousands of commits and 500+ contributors across ten teams - and I think they’re half right in a way that matters a lot, and wrong in a way that will cost someone a very expensive year of their life.

Where they’re right

Feature toggles plus observability is a genuinely better default than it used to be. It decouples deploy from release, it makes mistakes reversible, and it removes a huge amount of the ceremony that used to substitute for actually knowing whether a change worked. If your system is young, loosely coupled, and small enough that any one engineer can hold it in their head, “ship it behind a flag and watch the dashboard” beats a week of pre-emptive design review nearly every time. And that side keeps getting stronger on its own terms, independent of anything AI does - contract tests, chaos engineering, differential testing against real traffic, all shrinking the set of failures that need a human’s scar tissue to catch at all.

Two futures, not one

Before going further, a bifurcation worth naming, because it changes what kind of claim this essay makes. Software that’s cheap to rewrite and cheap to get wrong can keep moving toward toggles and metrics - and should, since occasionally shipping something quietly wrong costs less than reviewing everything deeply. Software that’s expensive to rewrite and expensive to get wrong - regulated, financially load-bearing, decades-old, too entangled to start over - doesn’t get to make that trade; payments, identity, and core trading systems carry the same shape, which is why every example from here on comes from that side, not from some microservice nobody would miss. The two futures sort systems by one question: can this be rewritten if it turns out to be wrong? They run in parallel rather than one winning. The one I work on can’t. A lot of software can, and where it can, I’d expect AI-first development to win outright - a real share of what’s currently called software engineering there will disappear. That’s not a hedge against this essay. It’s the boundary of where it applies.

There’s a way the two futures get confused that has nothing to do with which one is correct: prevention is invisible, shipping is visible, so an org underpays whoever caught a problem early and overpays whoever shipped around it, until something expensive breaks. That’s mispricing, not proof the judgment was unnecessary. The fix isn’t a better argument that something might break - “might” is unfalsifiable - it’s pricing review the way an options desk prices a hedge, by the exposure it’s written against, not by whether the tail event happens. That cost is never exact, but treating it as zero is the real absurdity.

Where it breaks

On the expensive-to-be-wrong side of that split, the toggle-and-observe model still assumes the thing you’re watching will tell you when something is wrong, in time, in a way you can interpret. On a system old enough and tangled enough, that assumption fails quietly, for years, and the bill arrives all at once.

Take the cache stampede bug in a memoizer class. InMemoryMemoizer.retrieve() had a non-atomic get-then-remove-then-get sequence that could cause a thundering herd of recomputation under load, the kind of latent race no dashboard was ever likely to flag as more than occasional noise. I found it by asking an assistant, during an unrelated review, what the weaknesses of that class were, and it surfaced the non-atomicity unprompted. Neither half found it alone: the question came from someone curious enough to ask it adversarially, the answer from an AI actually reading the implementation. Understanding why let the fix be one line - a Map#compute() call replacing a 47-line abstraction - instead of a toggle papering over a failure mode nobody understood.

Or a restored monitoring signal that fired on a distributed lock failing to acquire, rather than on the operation that lock protected actually failing. The dashboard was working perfectly. It was just measuring the wrong thing.

The part that should worry the “AI understands better than you” crowd

Here’s the detail I didn’t expect to be useful: I built my own RAG assistant, grounded directly in our own 12-year-old codebase’s source, to answer engineering questions about it. If AI genuinely has a cleaner model of a codebase’s architecture than the humans who built it, this is the tool that should prove it. Instead, in my own five-week log of its answers, its fabrication rate swung from roughly 6% on a good day to more than 40% on a bad one. Ask it the same trivia question about the codebase four times and it invented a different commit-count statistic each time - not retrieved from source, invented, confidently, in the same tone as the correct answers. The single most common failure mode on record: citing sources without actually checking them.

To be clear, I like the tool, and it keeps getting better. But it’s direct, first-party evidence that fluency is not the same thing as grounded understanding, even when the model is pointed straight at your source code. A reviewer who trusts an AI’s architectural read because it sounds authoritative is making exactly the mistake that model has no defense against: nothing in “ship it and watch the metrics” catches a confidently wrong explanation, because the explanation isn’t the thing being measured.

The obvious rebuttal: the rate will keep falling. Models get better, grounding gets better, maybe my assistant is sitting at 1% in a year. Grant all of that - it still doesn’t touch the two things actually carrying this argument. A model that never fabricates is still not the party who answers for a wrong call; accountability doesn’t improve with capability, it’s a structural fact about who’s on the hook, and it has never been the model. And a better model might even reconstruct why that dashboard was watching the wrong condition from incident history instead of needing it written down - a real possibility, not a wall it can’t climb. What doesn’t improve along with it is who decides the reconstruction is trustworthy enough to act on. A falling fabrication rate, or a rising ability to infer context nobody recorded, makes the first-pass read more trustworthy. It doesn’t make the model accountable for being wrong about either one.

The question decides the answer

The cleanest example I have is from this week, and to be fair to that model, it’s a good instance of it.

A colleague opened a PR to shrink the stored state of a tiered rewards feature. The feature keeps a copy of each tier’s progress, so every user action gets recorded once per tier, and the suspicion was that these documents had grown large enough to hurt performance. Nobody had measured that yet. The change dropped three context identifiers tied to the triggering action from those copies. The commit was co-authored with an AI assistant, and it was careful work: the safety reasoning was correct, and the tests proved that progress and payouts came out the same either way. It sat behind a runtime toggle, off by default, with a stated plan to revert it “if we do not observe a material benefit.” Requirements, toggle, metrics. If that model works anywhere, it should work here.

The AI did well at the tasks it was given. The assistant that co-wrote the change correctly answered “is it safe to stop storing these?” Four rounds of an AI code-review bot answered “does this change do what it says?” and caught real problems: the toggle reached beyond this one feature into every rule type sharing the same stored-state mechanism, and switching it off wouldn’t restore the entries already stored lean. What none of them did was change the question. Every step that made the change simpler came from a human doing exactly that.

The first new question was mine: is this field used at all? One of the three identifiers turned out to be written when the record is built, copied when two records are merged, and never read by anything. That isn’t a hard find. Once you ask the question, it’s a straightforward search. Afterwards I gave the PR to Claude for a full review, and when I asked whether it would have spotted the dead field unprompted, its own guess was probably not. It had been checking the PR’s claim that these ids are never read back, so it searched where the PR pointed. I wouldn’t call that proof of a ceiling, since it’s a guess about a run that never happened. But it matches what I saw: the framing of the change became the boundary of the review. And the answer to the wider question was simpler than the toggle. You don’t put a flag on a field nobody reads. You delete it.

The second came from other reviewers: should this be a global switch at all? A runtime flag leaves documents partly lean and partly full, with nothing recording which is which, waiting for the day someone actually needs that data. The fix they settled on was to let each rule type decide what it stores, so the data stays aligned with the rule that owns it. That’s a call about where responsibility sits in the design, not a fact anyone can look up.

The third question sat underneath both: is this worth doing? The benefit was a hypothesis, while the costs were concrete: a new runtime flag, mixed documents, and a wider blast radius than intended. Weighing those, and asking for numbers before building rather than after, takes knowing how the system actually behaves in production. That knowledge isn’t necessarily encoded in the code itself, so it won’t appear in a prompt unless someone already knows to supply it.

That’s the limit of “ship behind a toggle and watch the metrics.” A toggle can tell you whether a change is safe to undo. It can’t tell you whether the change was worth making, or whether it was the right change at all. AI may well get better at generating, challenging, even proposing better questions than the one it’s handed - that’s not the limit, and I wouldn’t bet against it. The limit is who authorizes which questions the system gets to answer on the organization’s behalf, and who answers for it afterward. That’s still someone’s job - and on consequential technical decisions, it’s often the engineer’s.

The analogy that stood in for a check

A different example, from a code review this week. A field initializer built a metrics wrapper and handed it a Spring-injected dependency through a Supplier - () -> monitor - instead of the value directly. A reviewer flagged it: move this out of the field initializer, since the injected value isn’t set yet when field initializers run. Reasonable rule of thumb, and usually right.

I asked an assistant to weigh in. Its first answer agreed with the reviewer; pushed on the Supplier specifically, it corrected itself: a lambda closing over an instance field re-reads it live, so what actually matters is whether the wrapper’s builder invokes the Supplier eagerly or only later. Good reasoning, but it never opened the one file that would have settled that: the builder’s own code. Instead it reasoned by analogy to similar classes elsewhere in the codebase - plausible, probably right, still unverified, delivered with a confidence that buried the one clause that mattered, “as long as,” inside a closed-sounding sentence. Asked why its first pass missed this, it gave an honest account of applying a rule of thumb without checking it.

A second assistant, asked to judge the exchange, gave a fair critique (good reasoning, unverified premise) but didn’t open that builder file either. It evaluated the gap one level removed from the gap itself: pattern-match, reason well once pushed, then settle for analogy instead of the one check that would close it. Nothing about a newer model changes that shape, only how convincing the stopping point sounds. Nobody had opened the file, on either pass.

What actually carried the exchange wasn’t a better model - it was a better question. Nothing moved until someone who knew how Java closures and Spring’s bean lifecycle behave doubted the first confident answer enough to ask about lazy evaluation specifically. That’s the engineer’s part: formulating the one challenge the exchange turned on, then staying skeptical of the better-sounding answer that followed. Take away that depth, and a plausible guess plus a confident agreement don’t look like two unverified guesses. They look like consensus, and nothing in the toggle, the metric, or the tone gives that away. Auditing a verification path nobody else is auditing isn’t a nice-to-have; here, it was the only thing between a wrong comment and a merged PR.

So what’s the engineer actually for, now

Not typing the code - that part genuinely is becoming commoditized, and good riddance, typing was never the most valuable part. The job that’s left, and that becomes more valuable as code gets cheaper to produce, is:

  • Owning the consequence. When an architecture decision record I reviewed had good instincts but under-evidenced justification that kept shifting mid-thread, rejecting it wasn’t a style nitpick - it was refusing to let a decision ship without a traceable reason, because someone has to be accountable for it not working, and that someone is never the model.
  • Carrying tacit context that isn’t written down anywhere. Why a restored log line ended up checking the wrong condition. Why an earlier generic abstraction rotted. Some of that lives in no diff, no document, and no ticket. It lives in scar tissue. Writing more of it down - an ADR, a documented invariant, a decision log - helps, and reduces how much rides on one person’s memory. It doesn’t remove the need for someone accountable for whatever’s still unwritten.
  • Verifying instead of inferring. A separate cache-invalidation design, for a distributed database cluster, went through three rounds, each one correcting an assumption about its topology against the actual source rather than trusting a plausible-sounding guess. That discipline (check before you conclude) is exactly what a bad, 40%-plus fabrication day looks like when nobody applies it.
  • Authorizing the question. AI keeps getting better at generating and challenging questions; that was never where the limit sat. In that PR, each reframing pointed to something simpler than the toggle. Elsewhere, a redesign made a whole list of the AI’s own findings disappear at once instead of fixing each in place - it had been listing symptoms of a design nobody had questioned yet.

Those four create a fifth, less glamorous job as a side effect: someone has to curate the AI’s own output. Findings get marked resolved before they’re actually fixed, or lost entirely when a conversation moves to a replacement PR, unless someone keeps track of which ones still need an answer.

AI is a genuinely excellent first-pass reader of a codebase - often dramatically faster than a human at surfacing “here’s what this probably does.” What it doesn’t yet do is own being wrong, and it doesn’t carry forward the decade-plus of institutional memory that make a codebase this old navigable at all. Review and architecture aren’t a human indulgence we keep out of nostalgia. They’re the mechanism by which someone stays accountable for a system that will outlive the sprint, the quarter, and possibly the team that built it. That model optimizes for a world where everything can be rewritten. The one I work on has to survive in a world that can’t.

What happens with nobody in the loop at all

Push the thought experiment further: a codebase with no human review at all, every feature vibecoded start to finish, requirements going straight from a stakeholder’s prompt to shipped code. Does that scale? Not for the reason people usually reach for first, and not because any single change is wrong. Each one can be locally correct, tests passing, the feature doing what was asked. The failure is in the sum: nobody asks “should this exist” across a thousand changes, because that question only gets asked consistently by something with a stake in the system still making sense next year. A dead field doesn’t get deleted; it gets company.

Vision has the same problem, one level up. An org without someone willing to say no to a plausible feature request doesn’t fail to scale from bad code - it fails from sprawl. Saying no costs something now for a payoff nobody can point to yet, and that cost only gets paid by whoever is accountable for the system still making sense later. Nothing today makes AI accountable for that call, so there’s nothing stopping it from saying yes to everything plausible, right up until the system is unrecognizable.

Which is really the same question as whether AI can be the authority guiding the people asking for the features, not just the hands building them. I don’t think it can - not for a capability reason, since a well-grounded model could probably tell you what was decided and why most of the time, but for a structural one: authority requires an accountability structure, someone or something that answers for the call when it’s wrong, and nothing today gives AI one. That could change - governance, insurance, and liability frameworks are exactly the kind of structure that might eventually assign AI that role, the way product liability or corporate personhood evolved for other non-human actors. The claim here is about today’s absence, not a permanent ceiling. A company has no feelings either, but it has an accountability structure: a board, a budget, people who can be removed. AI itself sits outside that structure; someone else is always on the hook for a decision it’s plugged into. What it can do, usefully, is remember: surface the decision from two years ago that this request quietly contradicts, so a person with something to lose can decide whether to overrule it. That’s augmentation, not authority, and the authority still has to be held by someone who stays around long enough to find out whether they were right.

Is it still worth learning to code

A fair question falls out of all this: if typing was never the most valuable part, is it still worth a person learning to write code, instead of starting straight from directing AI? Yes - but for a different reason than it used to be. What broke the tie in the field-initializer story wasn’t taste, it was knowing, cold, that a lambda closing over an instance field re-reads it live, and that Spring’s field injection runs after the constructor finishes. That kind of knowledge gets built by writing code, watching it fail in exactly that way, and fixing it yourself - not by reading about it, and not by learning to prompt well. The point isn’t suffering through programming the old way for its own sake; it’s acquiring the causal understanding that lets you recognize when an abstraction is lying to you, and writing code was historically just the cheapest way to build that. AI is lowering the cost of producing code faster than it’s eliminating the need to understand why code behaves the way it does - and that gap is exactly what this section is about. What’s worth learning shifts: less syntax and boilerplate, more debugging and the runtime mechanics that only show up once something breaks in a way the abstraction promised it wouldn’t. The understanding still has to get built somehow, because it’s the same thing that lets you authorize a question instead of just accepting the one you were handed.

Not all legacy is the same

One objection: if deep understanding is what matters, doesn’t that just mean I end up the only person who can touch the system, stuck maintaining something nobody else wants to learn, the way a mainframe COBOL expert is valuable right up until it curdles into a trap? Maybe - it depends on which kind of legacy system is in question, and AI changes the answer differently for each.

Split legacy systems along two axes: still evolving or abandoned, and low-stakes or still load-bearing if something goes wrong. Abandoned and low-stakes - frozen, nobody deciding anything new, the occasional patch its only contact with a human - is exactly where AI dissolves the sole-expert problem: the tacit knowledge is already baked into decisions that aren’t changing, the task is bounded, and nobody is betting the architecture’s future on a CVE patch.

A system that’s still evolving and low-stakes - a hackathon prototype, a campaign tracker nobody will touch past this quarter - doesn’t need that treatment either. What doesn’t get the relief is evolving and load-bearing, regardless of age: active development keeps producing new decisions that need the scar tissue AI doesn’t have, and reading the code well doesn’t resolve who’s accountable for extending it correctly.

The dangerous quadrant is the one that looks safest from the outside: abandoned, but still load-bearing. Frozen not because it’s safe to ignore, but because everyone’s too scared to touch it, while it still runs something that matters. AI may navigate that system confidently and propose plausible changes without knowing which undocumented landmines it has failed to reconstruct, and if an org’s conclusion from that is “AI can handle the legacy stuff now,” that’s precisely the moment the last person who remembers why something is shaped that way gets let go - right before the one time it would have mattered.

So “stuck being the only one who can touch it” isn’t a fixed fate attached to deep expertise. It’s a risk that lives in one specific quadrant: still-critical, and either actively evolving or frozen out of fear rather than safety. Over a dozen years on this one puts me in exactly that quadrant, for what it’s worth - which is either validation of the argument or the textbook definition of being too close to it to see clearly. Probably some of both.

comments powered by Disqus