The specification is the product
When AI writes the code, the real work moves to what you fix up front. From a platform we are building fully spec-driven.
We are currently building a cloud platform of which, so far, not a single line of code has been written by hand. We do it that way because it is faster and better, not to make a point. What we write are specifications and standards. The code follows from those, and gets checked before it is allowed to stay.
To anyone commissioning software this sounds odd. Aren’t you paying for code? This article explains why that has become the wrong question, and what takes its place. It starts simple and turns technical, because the second half is for the people who want to set it up themselves.
Code is no longer the scarce part
For years code was the expensive part. Someone had to type it, and that someone was expensive and slow. Everything around it was subordinate to that one bottleneck: deciding what you want, writing down how it should work, checking that it is right.
That bottleneck is gone. An AI agent writes in an afternoon what used to take a week. That shifts the question that matters. It is no longer who types this, but how you know that what comes out is what you meant. An agent does exactly what you ask, including the things you left implicit. Whatever is not fixed, it fills in with what seems plausible, and plausible is not the same as correct.
So what we deliver has shifted. We still deliver working software, and the client sees no difference there. But the artifact we maintain, the one our time goes into and the one the quality comes from, is no longer the code. It has become the specification.
What a specification is for us
A specification, for us, is a bounded document that describes one thing: which problem we are solving, for whom, and how you can tell it is right. So it is not a ticket and not a loose wishlist.
That last part carries the most weight. Every requirement is written as a concrete case. Given this situation, when this happens, then that should be the result. So not “the input must be validated”, but something testable: given a file with several problems, when it is checked, then every problem is reported at once. Report only the first, and an agent needs as many attempts as it has mistakes.
Each of those cases becomes at least one test. That is how the bridge from text to code runs. The specification says what must be true, the test checks whether it is, and the code exists to turn that test green. A specification with open questions we do not approve. An agent that starts without answers will guess, and a plausible guess costs you more later than asking the question up front.
For the client that has a pleasant consequence. What we build sits in plain language on paper before it exists. That is not sales talk, it is the source the software comes out of. What is not in it, we do not build, and that is exactly where you want to keep your grip.
Why this is a better deal than hours of code
Under the old model you bought effort. So many hours, so much code. The quality was hidden in the craft of whoever typed it, and you saw it only once something went wrong.
Under this model you buy a described, testable outcome. You read the specification and form an opinion about it before a cent has been spent on code. If the wish changes, you change the specification and the code follows. And because every requirement hangs on a test, you know not only that it once worked, but that it still works today.
Up to here this is the business story. The rest is about how you make an agent hold to it, because without that the above stays a nice intention.
Steering is data, not text
The first mistake people make is to cram every rule into one big instruction file for the AI. That works until it gets long. After that it contradicts itself and falls behind, and the agent picks the rule nearest to its context. That is often not the right one.
Our main instruction file therefore opens by stating that it contains no rules itself. It is a signpost that says where every rule lives. The testing standard sits in one place, the architecture rules in another. Every fact exists in exactly one file. If the same fact appears in two places, that is a defect we clean up.
The test is simple. If this changed tomorrow, how many files would I have to edit? The only acceptable answer is one.
The neat effect shows up in the last step of every change. The review step that checks the work against the standards reads that standards folder the moment it runs. Changing a rule therefore changes what is enforced, automatically. You do not update a second list and no checklist lags behind. The rules are data, and enforcement fetches them live. That looks like a detail, but it is the difference between rules that are true and rules that were once true.
The fences are code, not text
A standard that lives only in a document is a request. Agents mostly hold to it, until it is inconvenient under pressure, just like people. So what really must not happen goes into the machinery instead of into text.
The quality gate is one command that runs four things in a row: style checking, static analysis at the strictest level with no exceptions, a check that enforces the architecture layers mechanically, and the full test suite. None of the four can be skipped. A block also refuses any attempt to bypass the checks when committing. The only way past a red test is to fix the test.
What makes this trustworthy goes one step further, and it is the part an experienced engineer looks up at: the fences themselves are tested. The rule that enforces the architecture layers is checked by a test that verifies every new namespace actually falls under a rule. Without that test, a layer added later would be guarded by nobody while the gate stayed cheerfully green. A check that quietly stops checking is more dangerous than no check, because you rely on it. So we test the guards the way we test everything else.
Making it impossible works better than guarding it
The strongest rule in the project is one you will not find written down as a rule anywhere.
There are fields a user must never set themselves, such as permissions and ownership. The obvious approach is a check: if this field comes in, reject it. But a check can be forgotten, and an agent adding a feature in a hurry forgets it just as easily as a human.
So that field is not in the check. It is not in the processing either. It simply does not exist there, because there is no branch in the code that could accept it. A guard can have a bad day, a missing branch cannot. What does not exist cannot go wrong.
That is the thinking behind the whole setup. Wherever we can, we make the wrong thing impossible rather than forbidding it. A prohibition leans on judgment, and judgment falls away first under pressure, whether it belongs to a human or an agent.
Where the human sits
It might sound as if the human has become redundant. In reality their work has moved to where it weighs most.
The human decides what gets built and what does not. Writing and approving the specification is the real work, because that is where the expensive mistakes live: a wrong assumption in a spec propagates into everything that follows. The human also sees what is missing, the gaps an agent would rather fill plausibly than report, and decides when an open question is a gamble you may not take yet.
The agent is excellent at executing. Turning an approved specification into code, test-first, case by case. That is not a small role, but it is execution, and a project rarely runs aground there. Projects run aground on undescribed assumptions. You either write those down or you let the agent invent them for you.
What this costs, and why it is still the cheapest path
To be honest about the price: this is more work up front. You write specifications and standards, make a test plan per requirement, and build fences that you then also test. For a two-week job it would be nonsense.
Most of this, though, is what good projects should always have had. The difference is that you can no longer skip it, because an agent working without it only makes the mess faster. The payoff is that the software does not hang on one head. The specification is the truth and the code follows from it, so in two years the next one, human or agent, starts with exactly the same bounded picture we have today.
The code was never the valuable part. The valuable part was always the clarity about what you actually wanted. We used to leave that unwritten because the code was the expensive thing. Now that code is cheap, that clarity turns out to be the product.
Hexxore builds software spec-driven, with AI agents and strict standards, and advises on how to set that up responsibly. Is that question live in a project of yours? Tell us.