The harness thesis
Reliability is the frontier, and it does not live in the model. It lives in the harness around it.
Capability is not reliability.
In late 2025 a six-person startup called Poetiq reached the top of the ARC-AGI-2 leaderboard, a hard public test of machine reasoning, at roughly half the cost of Google’s own entry. They trained no models. They built a better harness around someone else’s.
AI capability scales along two axes. One is model scale, where hundreds of billions of dollars have gone. The other is harness sophistication: verification, iteration, adaptation. The returns have inverted. Improving the harness now returns more per euro than the next training run.
Meanwhile roughly 88% of enterprises use AI in some form and 6% to 7% get value at scale. That gap is not a capability gap.
Two ways reliability fails
- Formally. 95% accuracy per step compounds to 36% success across twenty sequential steps. And a model cannot reliably check its own work, because the verifier inherits the generator’s blind spots.
- On preference. Models produce the stochastic averages of what they were trained on. They cannot learn the specific judgment that makes an individual company competitive: the senior counsel’s risk tolerance, the creative director’s taste, the thing your best operator does without being able to explain it.
The harness has three parts.
Make yourself independent of any one provider’s pricing policy or release cadence, and put the intelligence in the infrastructure around the model.
- The orchestratorIteration and budget
- Governs the loops, allocates test-time compute where it pays, and stops once the evidence has converged. Done well it matches models fourteen times larger, which turns inference budget into a reliability dial you can set.
- The verification layerHeterogeneous checking
- Formal provers where truth is binary. Parameterized simulation where it is not, producing calibrated confidence under assumptions you can read. A shared blind spot is not a check.
- Preference learningCapturing tacit judgment
- Reads the implicit signals, the choices and edits and rejections, and learns what your people consider good without touching model weights.
Robust verification of open-world claims eventually needs the same capability as generating those claims reliably. There is no cheap verifier for hard problems: the verification gap and the AGI gap are one gap.
Where the value lands, and when.
-
Foundation loops
Formal verification reaches production reliability for code and compliance, using external orchestration.
-
Domain loops
Semi-automated verification deploys for legal and medical analysis, with human oversight in the loop.
-
Differentiation loops
Preference learning externalizes the capture of expertise. Generic AI becomes personal.
-
World model integration
Causal simulation migrates inward, approaching the threshold where verification and generation stop being distinguishable.
What to do about it
- If you run a company: the bottleneck is verification infrastructure, not model capability. Codify your constraints and start capturing preference signals now. The judgment data you accumulate this year is the moat next year.
- If you invest: value shifts off the foundation models, which are capex-heavy with binary outcomes, and onto the harness layer, which is opex-dominant, iterative, and modular.
More content from yours truly.
- Modelling a 1-Person AI Company The conductor-and-orchestra model for an AI-native firm, and the experiment of testing it on this one. Essay
- The AI Value Gap Why 95% of enterprise AI fails. The failure is a deployment playbook with two holes in it: no learning loop, and a default to build rather than buy. Essay
- AI Isn’t Magic Three tiers of mastery, and why most people stop at the first one. Essay
- AI Boom or Bubble Echoes of dot-com. Which parts actually rhyme, and which parts do not. Essay
- The Harness Thesis The full paper, with the citations behind every number on this page. January 2026, DOI 10.5281/zenodo.18416449. PDF
Focus on what matters most.
My position is that too many people get lost in the hype around the latest model. What will decide the future of your business is a foundation solid enough to deploy any future model on, and capturing today what makes you special.
Plenty of companies are losing exactly that right now. They hand the work to a model that returns the average, and the thing that made them them quietly leaves the building.