A discipline
Verification Design
AI can be incredibly powerful. But it's not particularly reliable. That's a problem when relying on factual and accurate answers, but it's fatal when trying to work with agents. Agents break big problems down into a series of steps, then break those steps into discrete actions. Get any of those actions wrong and the whole task fumbles and falls over.
The solution has roots that trace back more than 5000 years. Writing began as a way of 'taking stock': inventories and ledgers and receipts. As numbers became more central to civilisation, techniques improved: Egyptian scribes used red ink to mark corrections in their accounts; Romans separated "income and expenses", but all of this could still run afoul of simple human errors.
Then - more than 700 years ago - the Florentines invented double-entry bookkeeping. If the books don't agree, you know you have an error. Even better, you can trace back to where the books did agree, and you've simultaneously located your error.
Humans are no more perfect than AI, so we learned how to build a 'harness' around humans to make their work trustworthy and verifiable.
We can take that principle and apply it to artificial intelligence, in a discipline known as Verification Design.
Verification Design uses imperfect cognition as raw material: let agents do their worst - but within harnesses that catch and exclude their mistakes. That simple process change transforms error-prone AI into something that only gets better over time, ratcheting up and never back. Verification Design enables antifragile AI - the harder you test it, the better it gets.
Verification Design is the flywheel that makes AI usable at scale. Without the burden of having a human in every loop, 'botsitting' agents, complex tasks fall within the scope of automation.
This has profound implications for business. As Verification Design becomes an embedded practice, the core of the firm vanishes into automation, while the edge retains and amplifies necessary human qualities such as judgement - what shall we do? Taste - how shall we do it? And connection - with whom shall we do it?
An introduction to Verification Design
“Get it in Writing”
Consider a trip to the auditor.
You’ve received a notice from the ATO. Now you have to present everything to justify the last seven years of your tax filings.
In order to prepare for that audit, you will need to gather up all of your records. Hopefully you still have them.
If you have one, you’re going to want to bring your accountant along with you, because they’re well versed at answering the kinds of questions that come from the auditor.
The auditor is the judge of the matter. They decide whether your answers are sufficient. They decide whether your records are correct. They decide whether you get a fine or you get the all clear.
Any time you front an auditor you have to hope for the best. Because the auditor is not your friend.
It’s unnerving. These adversarial situations are designed this way by intent. Because adversarial frameworks are the most reliable way we have to get to the truth...
Read the rest of the essay here...
The theory behind Verification Design
The papers
Near-zero-cost cognition turns tokens into a quasi-currency: infrastructure mints them, and harnesses spend them to find value. Durable advantage shifts to physical assets, relationships, regulation, taste and trust, while businesses built on expensive human cognition are repriced.
AI becomes trustworthy when its outputs are checked against standards it cannot game, with machine-checked proof the strongest judge. Human work moves to setting and maintaining the specifications that govern those loops.
Post-Watershed firms replace transactions with self-healing loops that generate, check and repair work under rules governing the checks. Advantage moves to what resists copying - non-mintable assets and a harness record - while the firm contracts toward a human, a loop, finance and insurance.
Recursive AI self-improvement can remain governed only if delays slow it enough for human institutions to act. As those delays erode, pan-insurance becomes a fast, economy-wide regulator, alongside a boundary where the state must intervene.
Seven principles for loop governance separate the arrangements that survive from the human choices that must be protected. Together they make autonomous machine work governable by shrinking the ungovernable decisions to a manageable scale.