The Maintenance Tax

In 1986 Fred Brooks made a bet against his own industry. He wrote that there is “no single development, in either technology or management technique, which by itself promises even one order of magnitude improvement within a decade in productivity, in reliability, in simplicity” (No Silver Bullet).

He split the difficulty of software in two. Accidental complexity is the friction we invented and can therefore remove: assembly, batch compiles, manual memory, boilerplate. Essential complexity is the problem itself, and no tool touches it. His argument was arithmetic. Even if you drove accidental effort to zero, you would not get your tenfold, because accidental effort was never most of the job.

Forty years later, a coding agent delivers the tenfold on the accidental half. Boilerplate is free. Migrations are free. A test suite nobody wanted to write is free. Brooks was wrong about whether it would arrive and right about what it would buy us.

Because the other half did not move. It got heavier.

What software actually costs

Writing software has never been the expensive part. Maintenance is somewhere between 50 and 80 percent of total lifecycle cost, a range that has held across four decades of measurement and every attempt to measure it differently (O’Reilly’s 60/60 rule, and a survey of the estimates running from 67 to over 90 percent). Over a system’s life, keeping it alive costs two to four times what building it did.

Lehman named the mechanism in 1974 and it has never been repealed.

> As an E-type system evolves, its complexity increases unless explicit work is done to maintain or reduce it.

Not “may increase.” Increases (Lehman’s laws of software evolution). The default direction of a working system is toward nobody understanding it.

That is the tax. You do not pay it when you ship. You pay it every time someone has to reconstruct why the system is the way it is before they are allowed to change it.

For thirty years the tax had a hidden subsidy: the people who made the decisions were still in the room. The reconstruction was a conversation. Ask Priya why the retry budget is 3 and not 5, and Priya remembers, or half-remembers, and that was enough to move.

The tax has a new payer

An agent has no Priya to ask.

It has the repository, and the repository is the wrong artifact. Code is the output of a decision, not the record of one. It tells you the retry budget is 3. It cannot tell you that 3 was chosen because the upstream provider rate-limits on a sliding window and 5 tripped it in production on a Tuesday in March. That fact lived in a person, then in a Slack thread, then nowhere.

So the agent does what it must: it infers. It reads the code, forms a plausible theory of intent, and acts on the theory. Most of the time the theory is right. When it is wrong, it is wrong confidently, and it writes the wrong theory into more code, and now the wrong theory has evidence supporting it.

This is the actual bottleneck, and it is not generation. Every session starts from zero, reconstructs intent from artifacts that don’t carry it, and pays the reconstruction cost again. Ten sessions, ten reconstructions, ten slightly different theories of your system. The cheaper generation gets, the more often you pay.

Which produces the outcome everyone has now felt and few have named. Output goes up. Confidence goes down. You ship faster and trust the result less, so you review more, so the review queue becomes the constraint. The tax did not go away. It moved from the keyboard to your attention, which is the one input in this system that did not get ten times cheaper.

Nobody’s roadmap fixes this, because everybody is still optimising the half that Brooks already conceded.

Code is not the record

The instinct is to write it down. Teams have been writing it down for thirty years, and the docs directory is where the tax compounds fastest.

Prose documentation fails in one specific way: it is unfalsifiable. A paragraph saying “we retry three times with exponential backoff” is not connected to the code that retries. Change the code and the paragraph does not complain. It just becomes a lie that reads like the truth, and a lie that reads like the truth is worse than no document at all — because now the agent has a confident source, and the source is wrong.

Documentation rots silently. That is the whole problem, and it is a structural property, not a discipline problem. No amount of “we should keep the docs updated” changes a system where drift produces no signal.

So stop writing documents. Start keeping records.

What a record has to be

A record earns its keep only if it has four properties, and dropping any one collapses it back into documentation.

Atomic. One fact, one unit. Not a page about the auth system, but a single assertion small enough that a human can agree or disagree with it in ten seconds. Facts you can argue with individually are facts you can maintain individually.

Reviewed by a human, on the record. Not approved in a meeting. Approved in the artifact, with a name and a reason attached. Approval that leaves no trace is indistinguishable from no approval.

Locked. Once a human has stood behind a fact, nothing changes it silently. Not the agent, not a refactor, not a dependency bump. Changing it requires unlocking it, on the record, with a reason. This is the load-bearing property and the unpopular one, because it introduces friction exactly where everyone wants speed. That is the point. Friction on approved truth is the only thing that makes approval mean anything.

Drift-detecting. When code contradicts a locked fact, something must fail loudly — in CI, before the merge, not in an incident review six weeks later. A record that cannot notice it has become false is a document wearing better clothes.

Get those four right and the division of labour inverts. The agent operates: it drafts records, links them to code, restructures freely, runs the checks. Draft records are its workshop and should stay unfrictioned, because gating the workshop defeats the purpose. The human does one thing, which is the only thing that was ever scarce: says yes, this is true, and I’ll stand behind it.

You stop reviewing diffs. You review claims about your system. There are vastly fewer of them, they change vastly less often, and unlike a diff, a claim you approved once keeps paying out on every future session.

The world where the tax is small

Assume it works. What actually changes.

An agent joining your codebase reads what is true instead of guessing what was intended, and the guessing was the entire source of the confident-wrong failure mode.

Onboarding stops being archaeology. The reasons are in the record, next to the code they govern.

A decision made eighteen months ago is still legible, still attributed, still attached to the constraint that produced it. Institutional memory stops being a property of who happens to still work there.

And the thing that scares people most about agents — that they will erode a system faster than anyone can review it — becomes structurally difficult rather than culturally discouraged. The agent can move as fast as it likes across everything not yet locked. What you have approved stays approved until you say otherwise.

Long-lived software becomes cheap to keep alive. That is not a productivity story. It changes what a small team is allowed to attempt, because the reason small teams stop shipping is never that writing code got hard. It is that the system stopped fitting in anyone’s head, and now every change costs more than the last one.

Build the record

The industry is spending its whole budget making generation faster, which is optimising the half Brooks already told us was cheap. The unsolved half is the same one it was in 1986: knowing what your system is, why it is that way, and who decided.

We have never had a good primitive for that. We had tribal memory, and it worked as long as the tribe stayed. It doesn’t scale to a collaborator that arrives with no memory, works at machine speed, and forgets everything when the session ends.

I built DossierX because I needed this for my own projects and nothing existed. It is a directory of YAML claims, linted and cross-checked, rendered into something a human can read and argue with, where a locked claim cannot change without an approval on the record. Twenty-two commands the agent uses. One command I run.

It is one implementation and probably not the one that wins. The idea is not the tool. The idea is that the record has to become a first-class artifact of your system — versioned with the code, enforced by CI, owned by a human — or the maintenance tax will eat every gain the agents just handed us, and we will have built the fastest way yet to produce software nobody understands.

Ten times more code was never the goal. Ten times more code you can still stand behind is.