agentic coding · process · dialectics

To Vibe or Not to Vibe, That Is the Question

Every fork in software eventually dissolves when somebody moves up a level. Whether that works for agentic coding, and where the move is only laziness.

5 September 2026Sirsh AmarteifioDraft

Tabs and spaces. Vim and emacs. ORM and raw SQL. REST and GraphQL. Monorepo and polyrepo. Each one arrives as a genuine question, hardens into a tribe, produces a decade of conference talks, and then quietly stops mattering while everyone pretends they won.

Consider microservices. We fled the monolith because it was a big ball of mud, sprinted as far in the other direction as we could manage, and arrived somewhere with forty repositories, distributed transactions and a tracing bill. Then Prime Video's team published a post about moving a service back to a monolith46 and cutting infrastructure cost by around ninety percent, and half of engineering Twitter behaved as though a law of physics had been repealed. Nothing had been repealed. We had simply run to the far wall and discovered it was also a wall.

Now we have a new one. Karpathy named vibe coding in early 2025, it evolved fairly quickly into something people run as an actual agentic development process, and the tribes formed in about a week.

We want to argue about it, but not in the usual shape, where we explain why one side is stupid and you either nod or unfollow us. Both sides are describing something real, which is slower to read and gets about a tenth of the engagement.

The trick, stated up front

Almost every argument that gets genuinely resolved, as opposed to abandoned when everyone gets tired, gets resolved the same way. Somebody moves up a level of abstraction and the fork stops being a fork. The two things that could not both be true turn out to be two views of one thing seen from too close.

That is it. That is the whole move. It shows up in philosophy, in negotiation, in devops research, and in what we should be building for agentic coding.

It also fails in specific ways, because a trick that always works is not a trick, it is a religion.

Complementarity and synthesis

On this, it is often Hegel who first comes to mind, though not in the version most people carry around. That version goes roughly: two opposing ideas are each partly right and partly wrong, and they are destined to fuse into a third thing, which is how progress happens.

Thesis, antithesis, synthesis is not what Hegel called his own method.39

The terms are borrowed. Thesis and antithesis are Kant's, from the antinomies. The synthetic resolution is Fichte's, who answers the contradiction between self and not-self by positing a third concept, arrived at by finding the respect in which opposites are alike. Chalybäus packaged the triad as a summary of Hegel in 1837 and it stuck. Hegel does use the words, in his lectures, describing Kant. He also went after the schema directly, dismissing formalistic uses of triplicity as a dead scheme stamped onto a subject from outside the way you would apply a formula.18 His own names for the three moments are understanding, dialectical, and speculative. The correction gets overstated in turn. Being, Nothing and Becoming fit the triad neatly, plenty of serious readers still use it, and the real complaint is against applying it dogmatically, since Hegel has triads whose middle term is no strict opposite, and at least one section with only two terms.

What he actually gives you is two ideas.

The first is determinate negation. A position does not lose because a better one shows up and beats it. It falls apart from contradictions it was already carrying, and the way it falls apart constrains what can come next. Negation has content. You do not end up with nothing, you end up with a specific something.

Which is roughly how civilisations go. They are rarely finished off by whoever happens to be standing there at the end. The overextension, the fiscal rot, the thing the founding arrangement could not accommodate, all of that was in the building already, and the shape of the collapse decides what gets built on the site.

The second is Aufhebung, which Hegel points out in the Science of Logic means both to cancel and to preserve.19 The beaten position is demoted, not deleted. It stops claiming to be the whole story and becomes a true thing about a smaller region.

His line for it in the Phenomenology is that "the true is the whole".18 Every partial position is wrong as a totality and right as a moment.

We like this because it is a mechanism rather than a mood. It explains why the losing side's insight keeps crawling back out of the grave instead of staying in it.

Two people got to the same place without needing any of the machinery. Mill argued in On Liberty that opposing doctrines "share the truth between them"37, and that the heresy matters precisely because it is the half being sat on. Bohr designed a coat of arms in 1947 with the motto contraria sunt complementa and a yin-yang on it, which is either deeply profound or the physics equivalent of a tattoo you get at twenty two. His son reports him saying the opposite of a profound truth can be another profound truth,6 which is the kind of thing that sounds wonderful until you try to use it on a Tuesday.

We recently read Heisenberg's Physics and Philosophy,20 and what stayed with us was not the physics. It was watching a generation of physicists collide with a genuinely new philosophical idea and struggle with it in public. Wave-particle duality was not a puzzle you could solve by being cleverer. It was two descriptions that both worked and could not be held at the same time, and complementarity was Bohr's answer to that rather than a slogan for coats of arms. Heisenberg's account of how he eventually settled it in his own head is instructive, and we will not elaborate here.

The antithesis

Hegel does not only say oppositions are productive. He says the process gets somewhere. History has a direction, Spirit arrives, the contradictions work themselves out.

However Isaiah Berlin would have it that "Some among the Great Goods cannot live together."5,4 Liberty and equality. Mercy and justice. Spontaneity and planning. He calls it a conceptual truth, adds that we are doomed to choose, and that every choice may cost something that does not come back. What makes him awkward is that he concedes the first half. He accepts that both sides are tracking something real. He just denies that any frame satisfies both. The conflict is in the values themselves, not in our confusion about them, the choice costs you something real, and pretending otherwise hides the bill from whoever ends up paying it.

We find it too boring to simply agree with Berlin. The tension itself is evidence. Where two camps keep failing to resolve, on the questions we actually care about, that persistence is telling us something about the thing rather than about us. Heisenberg's physicists were not confused. They had found something that genuinely required two incompatible descriptions, and the trouble was that no single sentence would hold both.20 Iain McGilchrist works this ground at length in The Matter with Things33: some unities can only be approached from two sides, and the demand that they collapse into one is made by language rather than by the world. Think of a gestalt image. It is not the duck and it is not the rabbit. It is both, and you can only see one of them at any given moment. There are ideas shaped like that, we do not have the vocabulary for them, and so they keep getting mistaken for disagreements.

The rest of the opposition, briefly. Kierkegaard was first and hardest: Either/Or is a direct attack on mediation,25 on the grounds that the system dissolves the moment of actual choice and the person doing the choosing will not be absorbed into anybody's totality. Adorno keeps the dialectic and bins the reconciliation, inverting the slogan to "the whole is the untrue",1,2 because he thinks premature synthesis is how domination launders itself as reason. Weber leaves the value spheres in permanent war, old gods climbing out of graves.58 Schmitt makes antagonism constitutive; Mouffe accepts that and proposes agonistic pluralism instead,48,38 which is roughly "keep your enemy legitimate and stop pretending you agree". And Popper objected on technical grounds: "if contradictory statements are admitted, any statement whatever must be admitted".45 Treat contradiction as fertile and you have licensed everything. Popper would have hated this article.

The burden therefore sits with whoever claims the reconciliation. Show the frame doing the work in a case where the conflict was real and is now measurably smaller. Most candidates fail that, and they fail the same way, by relabelling a conflict as a balance and then treating the relabelling as the solution.

From the perspective of management science

Philosophy is all well and good, and it has the luxury of leaving the question open indefinitely. Management science does not. It is forced into a more decisive stance, because somebody has to centralise or decentralise by the end of the quarter, and the answer of it depends is not something you can put in a reorg memo. That constraint produced a more usable literature than the philosophers managed, mostly by accident.

Barry Johnson's Polarity Management separates the problems you solve from the polarities you manage.22 Problems have answers. Polarities like centralise against decentralise, or stability against change, have no answer, only better and worse oscillation. And what goes wrong is always the same. You over-correct straight into the downside of whichever pole you just fled.

Which is the microservices story exactly. We did not discover that microservices were wrong. We discovered the far wall.

James March put data behind the same structure in 1991 with exploration and exploitation.31 Tushman and O'Reilly answered with the ambidextrous organisation, structurally split so it can run both modes at once.57 Collins and Porras sold it to executives as the Genius of the AND,9 which is thin but got the vocabulary into boardrooms. Roger Martin relocated it in individual cognition as integrative thinking.32 Smith and Lewis work the same ground under paradox theory.49,50

Fisher and Ury turned it into something you can use in a meeting.14 Focus on interests, not positions. Positions are dichotomous by construction, because each one is a specific proposed answer. Interests usually are not. That is the level shift, wearing a suit.

Why nobody does this out loud

Integrative complexity, which is Tetlock's term for the capacity to hold competing considerations at once, predicts better judgment.55,56 It also collapses the moment people expect to be evaluated by an audience whose views they already know. Nuance reads as disloyalty, so we flatten ourselves for the room. His foxes beat his hedgehogs at forecasting and get a fraction of the airtime, because "it depends, here are four conditions" has never once gone viral.

The punishment has a vocabulary, and everyone reading this has been on the receiving end of it. Say both things are true at once and you are sitting on the fence, hedging, being evasive, refusing to commit. Say it in writing and you are both-sidesing. Update after new information and you are flip-flopping, which is the same accusation aimed at your past self, and the fact that updating is the entire point of thinking does not help you in the room. There is no comparable word for being confidently wrong at speed. The vocabulary is asymmetric because the social function is asymmetric. A tribe needs to know quickly whether you are with it, and a person weighing considerations is indistinguishable, from the outside, from a person about to defect.

So people learn to have the position first and the reasoning after, because the position is the thing being checked.

Mercier and Sperber argue that reasoning is argumentative in function.34,35 It evolved to make and break cases in front of other people, not to be quietly correct on your own. One-sidedness is close to the spec. Haidt puts intuition first and has reasoning conscripted after it.17 And on identity-loaded questions, Kahan found that more numerate people polarise more,23 because competence gets recruited to the side instead of to the answer. Your maths does not save you. It just makes you a better lawyer for whatever you already were.

The dichotomous performance is front-stage, in Goffman's sense, and not a report on what the actor thinks backstage.16 People are routinely more complementarity-tolerant inside their own heads than in anything they post. The divergence is structural rather than hypocritical, which is either comforting or bleak depending on the hour.

And every pressure outside the head points the same way. A recommender system optimises engagement, engagement responds to arousal, and being told you were right about something is more arousing than being told the question was badly posed. So the ranking learns to promote the flattest version of every disagreement, and we learn to write it. Headlines follow. Nobody has ever clicked on "it is more complicated than that", and the person who tried it once and watched the numbers has quietly stopped.

Notice what quietly is doing in that sentence. It hands you a secret and tells you that you are the one being let in, which is why every second headline these days has someone quietly abandoning something. An ingredient of clickbait, used here tongue in cheek.

The same incentives and pressures appear across most institutions. Journalism runs on a two-sided format that predates the internet. Conference talks need a position, not a range. Consultants sell frameworks with a recommended pole, because "manage the oscillation" does not invoice well. Analysts are graded on being right rather than being calibrated. Political parties are literally an architecture for suppressing the complementarity inside your own coalition. Even a standup is a bad venue for "we think we are both partly wrong", because the meeting is fifteen minutes and somebody has to pick a ticket.

None of these actors are behaving badly. Each one is responding to a local incentive that happens to punish the same thing. Which is why the private-public gap is so wide. The idea survives fine in your head, where nothing is measuring it.

Ripley calls them conflict entrepreneurs in High Conflict,47 people who make a living keeping a fight in its unproductive register. The quote-tweet dunk is a business model. Tannen documented the format problem in 1998,54 before any of the current machinery existed, which suggests the platforms optimised a tendency rather than inventing one.

When the trick is just laziness

On settled empirical questions, saying both sides have a point is itself the lazy move. The false balance literature in science communication is a catalogue of exactly this, and every one of those cases involved somebody feeling very even-handed while being wrong.

Which kind of disagreement it is decides everything. There are three.

Factual ones, about what is the case, have an answer. Go and find it.

Weighting ones, how much X against how much Y, are where the trick works. Go hunting for the frame, and expect something to be lost on whichever side you drop anyway.

Framing ones, where the argument is really about what the question is, are the easiest. The fork is usually something somebody invented, and it goes away once the thing is described properly.

Even then, a good reframing shrinks the tension rather than deleting it. Claiming deletion without demonstrating it is the premature synthesis Adorno was warning about, and he would be insufferable about it.

Right. Vibe coding.

Engineering is about efficiency and reliability. Those two words pull apart in your hands. We intentionally used at least two words here to describe engineering, and a pair that yield some tension. Efficiency sounds like lean, fast, minimal, nothing wasted. Reliability in practice often means slow and bubble-wrapped, with retries and circuit breakers and three separate places that check the same invariant in slightly different ways. What actually exists is a spectrum of tolerances, and where your team sits on it is a value weighting that everybody pretends is a fact.

Two serious results bracket this and they point opposite ways, which is the same pattern turning up in the engineering literature.

Erik Hollnagel's ETTO principle, the efficiency-thoroughness trade-off, says every operational act trades one against the other,21 that we make the trade continuously and mostly without noticing, and that it cannot be eliminated. You do not solve ETTO. You pick where to sit.

That is Berlin in a hi-vis vest. Same claim, arrived at from the shop floor rather than the seminar room, and none the weaker for it.

Then there is DORA. Its central finding is that speed and stability are not traded off against each other at all. High performing organisations ship faster and break less, and the two move together. The dichotomy everybody assumed turned out to be an artifact of large-batch gated process, and it dissolved under small batches, automated testing and fast feedback.15 That research was conducted before any of this existed, which turns out to matter a great deal.

Both findings are right, and they are not talking about the same thing. Hollnagel is describing a single act: a person deciding, right now, whether to run the extra check or ship. At that grain the trade is real and permanent, because the time spent checking is time not spent shipping and no amount of cleverness makes it otherwise.

DORA is describing a system: how work is batched, tested and released over months. At that grain the trade dissolves, because small batches and fast feedback make each individual act cheaper to verify, so nobody has to buy speed by skipping the check in the first place.

The tension is irreducible in the act and tractable in the process. Which is the level shift again, and it is the whole reason process design matters more than individual discipline.

Two camps on agentic development practice

One camp says delegate the implementation, work at the level of intent, ship an order of magnitude faster. The objection from inside their own tent is that if you have to review every generated line you have converted a writing cost into a reading cost, and reading unfamiliar code is not obviously cheaper than writing familiar code. The gain evaporates.

The other camp says you must understand what you ship. Responsibility attaches to the artifact, these generators produce plausible-looking code with subtle defects, and an engineer who cannot explain their own system cannot maintain or secure or debug it.

The second camp, if it actually exists as an absolute ideal, would be on the wrong side of history. Most of the people arguing for it do not hold it that strictly, which is worth saying before taking a swing at it.

The version of their argument we take seriously

It is not "read every line". It is Peter Naur, in 1985, in Programming as Theory Building40, arguing that the asset is not the program text. It is the theory in the programmers' heads: why the system is shaped this way, which invariants hold, what the design would do under a change it has not seen yet. Text is a residue of theory. His own formulation is that a program dies when "the programmer team possessing its theory is dissolved"40, and the corpse keeps running and producing useful results. Death only becomes visible when someone asks for a change and nobody can answer intelligently.

We already believe this, we just call it the bus factor, the number of people who would have to be hit by a bus before nobody left understands how the thing works. And anyone who has inherited a system knows the specific horror, which is not that the code is bad. It is that the code is fine and you have no idea why any of it is like that.

Dijkstra put it more sharply about natural language programming.11 Formal notation is not a barrier to be abstracted away, it is the instrument that makes precise reasoning possible. And Brooks drew a boundary that still holds:7 tooling attacks accidental complexity, while the essential complexity of deciding what the system should do is not reducible by any tool.

These are real and we are not going to wave at them.

The empty cell

Code review has always been partly theatre, which is the Goffman point again and neither camp says it out loud. LGTM on a nine hundred line diff. The approval that arrives four minutes after the pull request. Every one of us has rubber-stamped a migration we skimmed, and every one of us has also been the person whose careful review caught something real. Both are true, which is annoying.

So when the second camp says you must understand what you ship, part of what they are defending is a front-stage performance rather than a practice. Not all of it, and we are not accusing anyone of laziness. But the guarantee was never as solid as the ritual implied, and pretending otherwise makes it harder to build the thing that would actually be a guarantee.

Compilers, in the right order

Yang and colleagues built Csmith, pointed it at production C compilers in 2011, and reported more than 325 previously unknown bugs.59 Every compiler they tested could be made to crash and to silently emit wrong code from valid input. GCC. LLVM. The compilers under everything you have ever shipped. We were already trusting a demonstrably defective abstraction, cheerfully, for decades.

What made that rational was not the compiler being correct. It was differential testing, enormous shared test suites, decades of collective exposure, and eventually formal verification, with Csmith finding nothing in the verified core of CompCert, a C compiler whose optimiser carries a mathematical proof of correctness.27 Donald MacKenzie traces this in Mechanizing Proof.30 What counts as verification is a socially constructed standard that moves as the surrounding infrastructure matures.

Nobody hand-inspects assembly any more. The naive form of that comparison has a weakness. A compiler is deterministic, has a stable contract, and produces an artifact you throw away. Assembly is output. You keep the C. With a model, the generated code is the source you keep and extend and debug, so the abstraction never fully seals.

But the claim was never that compilers are trustworthy and models will be too. The claim is that trust in an abstraction has never come from the abstraction being perfect. It came from the control regime built around it.

Non-determinism proves less than people want it to

Models are non-deterministic, which is the technical objection people reach for, and it does not survive contact with what we already do. We already run on non-deterministic components. Human engineers are non-deterministic. The same person writes different code on different days and sometimes it is wrong, sometimes badly. Our response was never to individually verify each engineer. It was code review, CI, unit and property tests, staging, canary deploys, feature flags, blameless postmortems and rollback. We built an entire discipline around the assumption that the thing producing the code is unreliable, because it always was.

Distributed systems are non-deterministic too, and there the answer was formal methods at the specification layer, chaos engineering, SLOs and error budgets. Nancy Leveson's STAMP, a model of safety for whole systems, treats it as a control problem28 imposed on the system rather than a property of individually verified parts.

Non-determinism does not tell you to refuse the abstraction. It tells you which control regime the abstraction needs. That regime is statistical and adversarial rather than deductive, which is a real change, and it is the work nobody has finished.

So what is the actual question

Comprehension was never the guarantee, even before any of this existed. No CTO understands every line their organisation ships. Nobody has read their dependency tree. Your node_modules contains code written by strangers at three in the morning and you shipped it to production this week.

What we have always actually relied on is process. Tests that encode intent. Types and interfaces that constrain. Review that catches classes of error. Observability that surfaces failure fast. Rollback that bounds the damage.

So the open question is control and transparency. What does review look like when generation is cheap and reading is the bottleneck?

Every answer worth anything moves review effort off line-by-line inspection. Make the specification the reviewed artifact, so you review properties and invariants and contracts and treat the implementation as output that has to satisfy them. Keep machine-readable provenance of what was generated from what context at what revision, because transparency is a logging problem before it is an epistemic one. Verify by execution rather than by reading, with property-based testing and fuzzing and differential testing against a reference. Engineer the blast radius with aggressive modularity and reversibility, so an undetected defect costs a bounded amount instead of having its probability wished to zero.

That list is the DORA move again. Attack the tension at the level of process design instead of picking a pole.

What is in the head of the principal engineer

Start from the right diagnosis. People say AI produces security bugs, and empirically they are right. Pearce and colleagues built 89 scenarios around MITRE's Top 25 weaknesses, generated 1,689 programs with Copilot, and found roughly 40% of them vulnerable.43 About 39% of the top-ranked suggestions were, which matters more, because the top suggestion is the one a novice accepts.

Hold that next to a human baseline before drawing the obvious conclusion. Synopsys audited 2,049 commercial codebases the same year and found 81% contained at least one known vulnerability and 49% at least one high-risk.53 We were never operating from a clean floor.

And ask any of those models what is wrong with the code they just wrote and they will tell you, roughly as well as anyone would.

So the deficit is not in what the model knows. It is in what it attended to at the moment of writing. And harder, what it attends to as the code changes around it over months. That second one is the difficult part, and it is difficult for humans too.

This reframes the problem from representation to retrieval at the right moment, which is a far more tractable target. Bigger context windows do not fix it. Putting an invariant somewhere inside fifty thousand tokens does not make it salient, in the same way that your company wiki does not make anyone informed.

Think about what the staff or principal engineer holds. Not the one who started three months ago. The one who has lived in the codebase for years. They know that changing this here breaks that there, and they hold it alongside a working sense of which best practices apply and which ones they are allowed to break and why. Today that structure is oral, undocumented, and it walks out of the building when they do.

What already exists

Spec-driven development is the name for most of this, and it has a tool ecosystem. GitHub shipped Spec Kit, AWS shipped Kiro, and there is OpenSpec, BMAD, Tessl and cc-sdd, plus some flavour of it bolted onto every major coding tool. The practice sorts onto a spectrum. Spec-first means you write the spec and still maintain the code. Spec-anchored means specs drive generation and coexist with hand-written code. Spec-as-source means the spec is the only artifact humans edit and the code is a regenerated byproduct. Piskala's January 2026 preprint puts the distinction plainly:44 traditional specs are read by humans, these ones are executed by agents. There is even a constitutional variant, where a versioned machine-readable constitution encodes constraints pulled from the public catalogues of known software weaknesses.

Karpathy named vibe coding in a post on X in February 2025, describing a mode where you give in to the vibes and forget that the code even exists.24 Nine months later Collins made it word of the year for 2025, classified as a noun.10 By 2026 he was calling it passé.

The graph half exists too, and it is crowded. CodeGraph, GitNexus, Graphify, Understand-Anything, Serena, Augment's context engine spun out as a standalone server. Mostly local-first, mostly served to agents over MCP, the protocol they use to call outside tools, mostly pre-computing structure on device. On the research side there is RepoGraph, CodexGraph, LocAgent, CGM, GraphCodeAgent and Seddik's programming knowledge graph work. Several open-source projects launched with much the same architecture within weeks of each other in 2026, which is what convergence looks like from the inside.

What all of these hold is structure. Files, functions, classes, who calls whom, what imports what, which types flow where, and the transitive closure of that, so you can ask which parts of the system a given change could possibly reach. It is built by parsing rather than by prompting, which makes it exact and cheap to query. An agent can consult it before it touches anything, and the measured win is real, mostly in tokens and precision rather than in correctness. The structural layer is commoditised.

What the graphs do not hold

Structure is not the expensive part. Every one of those graphs can tell you that this function calls that one. None of them can tell you why the call has to happen in that order, who decided it, what broke the last time somebody reordered it, and which of the surrounding conventions are allowed to be violated when the deadline is real.

So put the three together. There is a body of knowledge that currently lives in one person's head and leaves when they do. There is a movement trying to make specifications the reviewed artifact rather than the code. And there are graphs that already model a repository well enough for an agent to query before it acts. Each has the shape of the other two's missing half.

What we think that converges on is a claims dependency graph. Take the spec-driven ambition of making intent something that can fail rather than something that quietly drifts, and put it on the substrate the graph people have already built. Nodes are claims about the system rather than symbols in it. Edges are dependency and blast radius. It is versioned next to the code, queryable by an agent, and it sits on top of the structural graph rather than replacing it.

This is the sense in which specs might compile. Not that a document is turned into a binary, but that each claim in it is a thing which can be checked, can fail, and can propagate its failure to the other claims that depended on it. That is what a compiler does with types, and it is roughly what the long-tenured engineer does in their head when somebody proposes a change.

The value is in the diff trigger, not in the graph. You are not trying to hold the whole picture at once. You want a change to land, the graph to name the small set of claims inside its blast radius, and only those to be re-asserted against the new state. You are not asking a model to comprehend the system. You are asking it to attend to a handful of specific things at exactly the right instant. Which is a close description of what the experienced engineer does, and it is the part that survives them leaving.

Three constraints decide whether this works or joins the graveyard.

Rot will kill it if anything does. Every previous attempt died the same death. UML round-tripping. Model-driven architecture. Architecture documentation. Decision records. Not because the ambition was wrong, but because a hand-maintained artifact drifts from the code and then starts actively lying, at which point it is worse than nothing. This is Lehman's law of continuing change and Parnas on software aging,26,42 operating on the documentation instead of the code. The only defence available is that a claim should be falsifiable by machine wherever it possibly can be, so a stale claim fails loudly instead of sitting there being confidently wrong. Prose claims are the rot vector. Keep them, rank them below executable ones, expire them without sentiment.

A claim has to carry its exceptions or you have built a linter. A node with nothing but an assertion in it recreates static analysis and earns the same fate, which is a rule added to the ignore file in week two. What makes it different is everything else in the node: the rationale, the scope, and the known exceptions with their justifications. The rules you are allowed to break. This is Chesterton's fence with a note nailed to it explaining who put it there and what happens if you move it, and no linter has ever held that.

The claims worth having are the ones you cannot derive. Static extraction gives you the cheap layer. This function assumes non-null. This handler is not idempotent. Fine, and largely already solved. The expensive knowledge is cross-cutting and implicit, of the form that this cache is only safe because writes are funnelled through one path in an unrelated module. No analyser sees that.

But it does surface as text, in specific places. Incident postmortems. Revert commits. The review comment that says careful, this breaks X. Every bug fix is a fossil of a claim that was violated. git blame is already archaeology, we just do it manually and only when something is on fire. Mining fix history is a high yield extraction source and almost nobody does it.

CodeQL is the closest existing thing,3 since it already treats a program as a relational database with a query language over it. What it lacks is the rationale and exception layer, and re-evaluation triggered by change. That gap is the proposal.

The numbers suggest we have a long way to go

DORA's classic result came from research conducted before any of this existed. The AI-era numbers point the other way.

In 2024 a 25% rise in AI adoption came with an estimated 1.5% drop in throughput and a 7.2% drop in delivery stability.12 Both directions bad. By 2025 the throughput sign had flipped positive, teams genuinely were shipping faster, and stability had not recovered. That is the shape of the thing. Speed came back. Stability did not. DORA's own framing is that AI is an amplifier, magnifying whatever the organisation already was, and they have given the reading cost a name. The verification tax. Time saved writing, re-spent auditing.

Which is the objection from inside camp A's own tent, measured, in a Google research report.

The rest of the telemetry agrees. Faros AI found review time up 91% and bug rate up 9% per developer alongside 98% more PRs merged, with almost no correlation at company level between AI use and output.13 LinearB's 2026 benchmark covered 8.1 million pull requests from 4,800 organisations across 42 countries. AI-assisted PRs run about 154% larger than human ones. They wait roughly five times longer for a reviewer to pick them up. And 67.3% fail on first review against 15.6% for human-written code.29 CircleCI analysed nearly 28 million CI workflows from 22,000 organisations. Average throughput grew 59% year over year, the largest jump in the seven years they have run the report. The median team got 4%. The top 5% got 97%. And for that median team, feature branch activity rose 15% while main branch throughput fell 7%, with main branch success rates at a five-year low of 70.8%.8 Feature branches are where you build. Main is where you ship. Worth noting that their throughput metric counts pipeline runs rather than deployments, so it measures activity rather than delivery, which if anything makes the main-branch decline more telling. METR ran a randomised controlled trial on experienced open source developers and found them 19% slower with AI while believing they were 20% faster. Their February 2026 follow-up cuts the other way:36 a newer cohort came in at 4% slower with a confidence interval running from 15% slower to 9% faster, and the researchers think developers are probably more sped up now than they were. Which is what a closing interval looks like. Sonar's 2026 survey of more than 1,100 developers puts AI at 42% of committed code, with 96% saying they do not fully trust it to be functionally correct and only 48% saying they always verify it before committing.51 Those last two numbers should not be able to appear in the same survey.

And people are quietly not reviewing. Faros AI's 2026 report has median PR review time up fivefold and 31% more PRs merging with no review at all,13 which is the theatre point with a number attached to it.

Naur is having a revival. Comprehension debt is the term, popularised by Arvid Kahl and Addy Osmani, with an ACM Queue piece extending it to cognitive and intent debt at team level.52 Osmani puts it as "AI writes faster, people understand less, the gap grows".41 Human review was a bottleneck, but a productive one, because reading a colleague's PR forced comprehension and then distributed it. Volume breaks that loop. The codebase looks healthy while the understanding hollows out underneath it, and nothing in your dashboard is measuring the thing that is disappearing.

None of this proves the position wrong. All of it describes the interval.

The interval argument is unfalsifiable unless something could settle it, so here is the something. If the spec and graph infrastructure matures over the next few years and delivery instability stays elevated anyway, then the abstraction is not sealing, and the other camp was describing something structural rather than something temporary. That is checkable, which is more than most positions in this fight can say.

Where this leaves us

Berlin gets a word here, and so does Hollnagel, which is convenient since they are the same person in different clothes.

There are domains where none of this delivers. Avionics. Medical devices. Cryptographic primitives. Anywhere an undetected defect costs an unbounded amount and no statistical control regime is enough. There the tolerance has to be different, and the answer is not a better process. ETTO does not get solved by continuous delivery. It gets moved.

Everywhere else we think this is inevitable. Generation got cheap, the control regime has not caught up yet, and control regimes have always caught up, because that is the only thing the industry has ever reliably done when handed an abstraction it could not fully inspect. Compilers took decades and a fuzzer. This will take its own version of both.

One camp is describing now. The other is describing later. It only looks like a contradiction because both of them insist on the present tense.

So the warnings are accurate and the conclusion drawn from them is wrong, which is a distinction worth holding onto while the numbers stay ugly. It is going to be bumpy. We are fairly sure that is the fun part.

If you got this far, thank you. You are evidently someone who does not need their content bite-sized and one-sided, which the incentives above suggest is a shrinking demographic.

We should confess we wrote this for ourselves. It was the writing of it that let us work out what complementarity actually looks like when applied to this thing we have started calling vibe coding, and we did not know where it landed until we got here.


References

  1. Adorno, T. W. (1951/2005). Minima Moralia: Reflections on a Damaged Life. Verso.

  2. Adorno, T. W. (1966/1973). Negative Dialectics. Continuum.

  3. Avgustinov, P., de Moor, O., Jones, M. P. and Schäfer, M. (2016). "QL: Object-oriented Queries on Relational Data." ECOOP 2016.

  4. Berlin, I. (1958). "Two Concepts of Liberty." In Four Essays on Liberty. Oxford University Press.

  5. Berlin, I. (1990). "The Pursuit of the Ideal." In The Crooked Timber of Humanity. John Murray.

  6. Bohr, H. (1967). "My Father." In S. Rozental (ed.), Niels Bohr: His Life and Work as Seen by His Friends and Colleagues. North-Holland.

  7. Brooks, F. P. (1986). "No Silver Bullet: Essence and Accidents of Software Engineering." Proceedings of the IFIP Tenth World Computing Conference.

  8. CircleCI (2026). The 2026 State of Software Delivery Report. https://circleci.com/blog/five-takeaways-2026-software-delivery-report

  9. Collins, J. and Porras, J. (1994). Built to Last. HarperBusiness.

  10. Collins Dictionary (2025). Word of the Year 2025: vibe coding. https://www.collinsdictionary.com/us/woty

  11. Dijkstra, E. W. (1978). "On the foolishness of 'natural language programming'." EWD667.

  12. DORA (2024, 2025, 2026). Accelerate State of DevOps and State of AI-Assisted Software Development reports. Google Cloud. https://cloud.google.com/devops/state-of-devops

  13. Faros AI (2025, 2026). AI Engineering Report and The Acceleration Whiplash.

  14. Fisher, R. and Ury, W. (1981). Getting to Yes. Houghton Mifflin.

  15. Forsgren, N., Humble, J. and Kim, G. (2018). Accelerate: The Science of Lean Software and DevOps. IT Revolution.

  16. Goffman, E. (1959). The Presentation of Self in Everyday Life. Doubleday.

  17. Haidt, J. (2012). The Righteous Mind. Pantheon.

  18. Hegel, G. W. F. (1807/1977). Phenomenology of Spirit, trans. A. V. Miller. Oxford University Press.

  19. Hegel, G. W. F. (1812/2010). The Science of Logic, trans. G. di Giovanni. Cambridge University Press.

  20. Heisenberg, W. (1958). Physics and Philosophy: The Revolution in Modern Science. Harper.

  21. Hollnagel, E. (2009). The ETTO Principle: Efficiency-Thoroughness Trade-Off. Ashgate.

  22. Johnson, B. (1992). Polarity Management: Identifying and Managing Unsolvable Problems. HRD Press.

  23. Kahan, D. M., Peters, E., Dawson, E. C. and Slovic, P. (2017). "Motivated Numeracy and Enlightened Self-Government." Behavioural Public Policy, 1(1).

  24. Karpathy, A. (2025). Post on X, February 2025, coining "vibe coding".

  25. Kierkegaard, S. (1843/1987). Either/Or, trans. H. and E. Hong. Princeton University Press.

  26. Lehman, M. M. (1980). "Programs, Life Cycles, and Laws of Software Evolution." Proceedings of the IEEE, 68(9).

  27. Leroy, X. (2009). "Formal Verification of a Realistic Compiler." Communications of the ACM, 52(7).

  28. Leveson, N. (2011). Engineering a Safer World. MIT Press.

  29. LinearB (2026). Software Engineering Benchmarks Report. https://linearb.io/blog/8-million-prs-engineering-productivity

  30. MacKenzie, D. (2001). Mechanizing Proof: Computing, Risk, and Trust. MIT Press.

  31. March, J. G. (1991). "Exploration and Exploitation in Organizational Learning." Organization Science, 2(1): 71–87.

  32. Martin, R. (2007). The Opposable Mind. Harvard Business Review Press.

  33. McGilchrist, I. (2021). The Matter with Things: Our Brains, Our Delusions, and the Unmaking of the World. Perspectiva Press.

  34. Mercier, H. and Sperber, D. (2011). "Why do humans reason? Arguments for an argumentative theory." Behavioral and Brain Sciences, 34(2).

  35. Mercier, H. and Sperber, D. (2017). The Enigma of Reason. Harvard University Press.

  36. METR (2025, 2026). "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity" and February 2026 follow-up. arXiv:2507.09089. https://arxiv.org/abs/2507.09089

  37. Mill, J. S. (1859). On Liberty. John W. Parker and Son.

  38. Mouffe, C. (2000). The Democratic Paradox. Verso.

  39. Mueller, G. E. (1958). "The Hegel Legend of 'Synthesis-Antithesis-Thesis'." Journal of the History of Ideas, 19(3): 411–414.

  40. Naur, P. (1985). "Programming as Theory Building." Microprocessing and Microprogramming, 15(5): 253–261.

  41. Osmani, A. (2026). "Comprehension Debt: The Hidden Cost of AI-Generated Code." O'Reilly Radar. https://www.oreilly.com/radar/comprehension-debt-the-hidden-cost-of-ai-generated-code/

  42. Parnas, D. L. (1994). "Software Aging." ICSE '94.

  43. Pearce, H., Ahmad, B., Tan, B., Dolan-Gavitt, B. and Karri, R. (2022). "Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions." IEEE Symposium on Security and Privacy. arXiv:2108.09293.

  44. Piskala, D. B. (2026). "Spec-Driven Development: From Code to Contract in the Age of AI Coding Assistants." arXiv preprint.

  45. Popper, K. (1940). "What Is Dialectic?" Mind, 49(196).

  46. Prime Video Tech (2023). "Scaling up the Prime Video audio/video monitoring service and reducing costs by 90%." https://www.primevideotech.com/video-streaming/scaling-up-the-prime-video-audio-video-monitoring-service-and-reducing-costs-by-90

  47. Ripley, A. (2021). High Conflict. Simon & Schuster.

  48. Schmitt, C. (1932/2007). The Concept of the Political. University of Chicago Press.

  49. Smith, W. K. and Lewis, M. W. (2011). "Toward a Theory of Paradox." Academy of Management Review, 36(2).

  50. Smith, W. K. and Lewis, M. W. (2022). Both/And Thinking. Harvard Business Review Press.

  51. Sonar (2026). State of Code Developer Survey.

  52. Storey, M.-A. (2026). On technical, cognitive and intent debt. ACM Queue.

  53. Synopsys (2022). Open Source Security and Risk Analysis Report.

  54. Tannen, D. (1998). The Argument Culture. Random House.

  55. Tetlock, P. E. (1983). "Accountability and Complexity of Thought." Journal of Personality and Social Psychology, 45(1).

  56. Tetlock, P. E. (2005). Expert Political Judgment. Princeton University Press.

  57. Tushman, M. L. and O'Reilly, C. A. (1996). "Ambidextrous Organizations: Managing Evolutionary and Revolutionary Change." California Management Review, 38(4).

  58. Weber, M. (1917/1946). "Science as a Vocation." In H. Gerth and C. W. Mills (eds.), From Max Weber: Essays in Sociology. Oxford University Press.

  59. Yang, X., Chen, Y., Eide, E. and Regehr, J. (2011). "Finding and Understanding Bugs in C Compilers." PLDI '11. https://www-old.cs.utah.edu/~regehr/papers/pldi11-preprint.pdf

← All writing