← Back to list

How Big a Mistake Will We Allow AI to Make? A response to Demis Hassabis’s article.

A debate running about “AI regulation” is about who should hold the gate (control and approval), and how big the gate is. My question is…

Tony Fish · 2026-07-17 12:47 · 10 claps · 20.1 min read
#ai-governance #governance #ai #demis-hassabis #ai-regulation
Open on Medium ↗
Wiki topics: AI · AI · General 🏃 · Running & Endurance

How Big a Mistake Will We Allow AI to Make? A response to Demis Hassabis’s article.

A debate running about “AI regulation” is about who should hold the gate (control and approval), and how big the gate is. My question is whether a gate is the thing we should be building at all.

copied from Demis’s article, linked below

copied from Demis’s article, linked below

In truth, my head is exploding from reading Demis Hassabis’s article and the volume of responses that have piled in around it, including comments. Gary Marcus, Mark Daley, a dozen or more industry analysts asking whether FINRA is the right template, and, running underneath all of it, a White House AI adviser insisting “there will not be an FDA for AI”. Lest we forget a16z Marc Andreessen’s views. There are many strong and credible voices.

The opinions are in distinct camps, and it takes time to read all this, digest it, and then look for clarity (sense-making). My honest difficulty is that everyone is defending their own stance, model, and lens, which makes it genuinely hard to see a broader picture from narrow views. My first instinct was that they were protecting a narrative, and I have come to think the truth is more interesting and more unsettling: I don’t think they are avoiding or hiding from the question I want to ask; I do believe that they simply cannot see it from where they are standing. Nor, most days, can I see my own as I get barraged by other strong stances.

The question I keep exploring is “**How big a mistake will we allow these systems to make, and how would we even know?”**

Of note: I cannot find this question in any of the original posts, comments or responses. They are far cleverer than me, but it is because the frame/ lens they are all working on makes my question invisible/disappear, or, for the fake news lovers, wrong.

One thing before we start, because the ground is crowded, and I do not want to be mistaken for being in certain camps. This is not an argument for fewer rules. The “no FDA for AI” position is already the loudest sound in Washington, and I want no part of that pulpit. I am arguing for something FAR MORE demanding than a rule, not less; it is a thing that regulation, by its nature, cannot build.

Everyone is arguing calibration. Almost nobody is arguing category.

Reading the original, comments and responses side by side, and for me, a pattern emerges.

Starting with what Demis actually proposes, because it is thoughtful, carefully worded and deserves to be restated faithfully. He wants a new “Frontier AI Standards Body,” a federally overseen public-private partnership modelled on FINRA, funded largely by industry, with a board that includes independent technical experts and open-source voices. It would set benchmarks that define which models count as “Frontier-class”, and test them in the domains that frighten us most: cybersecurity, biological threats, deception, and the loss of control over agentic systems. Labs would share their models voluntarily, up to 30 days before release; once the protocol proved robust, passing it would become mandatory for deployment in the US market. The benchmarks would refresh quarterly, held-out tests would eventually be built independently of the labs to stop them gaming the exam, and the body could, if the danger warranted, coordinate a slowdown across all of them. It is a serious, thoughtful and genuinely well-intentioned design, carrying all the nuance you would expect from someone on the inside who wants something concrete to build.

To be scrupulously fair to Demis, this is not the one-shot drug test I am about to accuse it of being. The quarterly refresh, the independently built held-out tests, and the ability to coordinate a slowdown after release are powerful and adaptive features, and they matter. But notice where they sit. Every one of them is bolted onto a spine whose logic is still test, certify, release; the adaptation is in how often the exam gets rewritten, not in whether the instrument is an exam. Yes, the refresh keeps the gate current, but it does not stop it from being a gate. And it carries all the weight of certification, handing out a pass, and a pass, as I will come to, is a thing institutions are built to stop questioning.

Gary Marcus, who called for pre-flight testing in his 2023 Senate testimony, is delighted and wants the gate stronger: independent, mandatory, transparent. Mark Daley, the sharpest of the critics, wants the gate wiser: separated powers, federated verification, “a verifier we can trust”. The analysts want to know whether FINRA is the right template or whether it should be the FDA or the Federal Reserve’s stress tests. The US administration’s instinct is to want no gate at all.

Every one of these is an argument about the gate, size, swing, height, or, if you want, who controls the settings on the machine. Submit the model, run the tests, pass or fail, then release. The entire visible conversation, the proponents, the critics, and the refuseniks alike, has quietly agreed that the instrument is a gate. The only questions left open are how well it is built and who holds the key. (who decides who decides)

That agreement is not a stitch-up; it is a frame, and the boundary is hard to see for the people standing inside it.

Why they cannot see it (and why that is not a criticism)

So many have written before that winners go blind, not despite their success, but through it. The model that made you right becomes the model that decides what you are now able to see. The surest sign that a frame is holding you is that you feel certain. Confidence is simply what it feels like when the frame has become invisible.

Imagine the cartographer who has perfected their map. Every coastline exact, from the work of a lifetime. They are, rightly, confident, but new continents have formed, and they cannot see them, not because they are careless, but because they are not on the map they perfected, and the map has become the only place they know they can trust. The last thing a fish notices is the water it swims in.

Unpicking this across a few of the key people in this debate.

Hassabis works from the frame of a frontier lab, so he sees a standards problem, build the body, run the tests and certify the release.

Marcus works from his own decade-old pre-flight thesis, so he sees a calibration problem, make the same test independent and mandatory.

Daley works from a constitutional framework, so he sees a power problem: whoever writes the tests ends up governing the technology. He (and the others) are right, and I will come back to what “right” means in this context.

The administration works from a market frame, so it sees an overreach problem … no FDA, let it run. And at the far edge sits Yann LeCun, who has not addressed this proposal thus far but whose position tells you where the map ends. He argues there is no such thing as general intelligence, that it is “an illusion”. From inside his frame, the governance question does not even arise, because the thing being governed does not, to him, exist. It is probably why he is not drawn into this debate.

There are far more than 5 vantage points and way more than 6 different problems, and each problem is the shape of the model the person is defending. This is not a failure of intelligence, but it is the structure of intelligence: the frame that makes your question legible is the same frame that makes another question invisible. We are each most confident exactly where we have looked the least, I include myself. The only advantage I have here is that I have spent some time studying this particular blindness, which means I may be able to see their water even when I miss my own. The more you know, the more you realise how little you know.

So let me offer the frame they are not using, not as a truth or being right, but as a lens worth looking through, even if you discard it, you looked.

What I actually believe about “error”

We (humanity) did not arrive here (developing advanced systems) by being right. We arrived by being wrong, in a million small, patient, survivable ways, and living long enough for some of those errors to turn out to matter. This is how every learning system we know of actually works. Evolution runs entirely on error … the copying mistake (mutation), the variant that looked useless until the world changed and made it essential. Innovation runs on error: the failed prototype, the wrong turn that opened a door nobody knew was there. A child learns to speak by getting it wrong out loud, repeatedly, in front of adults who care. A craftsman learns the grain of the wood by ruining more than a few boards. Remove the error, and you do not get mastery, but you get a person who has never been tested, calling their blind inexperience safety.

And yet we have built an entire raft of institutions dedicated to eliminating error … in education, in management, in compliance and in performance review. We have become so fluent in optimisation and error-reduction that we have forgotten error is also the only mechanism by which anything improves. Zero-defect is a fine goal for a factory stamping the same part a million times. It is a catastrophic goal for anything that is supposed to learn.

We learn from our mistakes; the machine, in the sense that matters, does not. It was trained to minimise error against a frozen body of data (unknown attestation), a corpus already full of “errors” it cannot perceive. Yes, the newer agentic systems carry memory and update as they run, so “it stops learning at deployment” is too neat, but none of that is learning from consequence. It has never once been wrong in a way that cost it something, felt the impact, and adjusted. It is adaptive in capability and blind to consequence, and those are not the same thing. So we are now proposing to govern the most consequential technology of the age with a philosophy that eliminates the error before release, that describes neither how the machine came to be, nor how we came to be capable of anything at all.

We have to keep this in mind because it explains why I believe the tool everyone has reached for is the wrong one. (why they do is not my question)

Immune systems learn. Drugs do not.

Gary Marcus reaches, approvingly for the FDA …. for the model of a drug. I feel it perfectly highlights the wrong nuance in the whole debate, because a “drug test” is exactly the wrong analogy, and seeing why unlocks the rest.

A drug does not learn as it is a fixed molecule with a bounded mechanism: tested, certified, shipped unchanged. When the world later reveals something the trial missed, the drug cannot adapt; it can only be recalled. The thing that learns in that system was never the drug, it is the system, the body, the immune system.

Holding and valuing these two side by side. We are about to deploy the most adaptive artefacts ever built, agentic, self-improving cognitive harness, able to act in the world and change it, and we propose to govern them with the drug model. Test. Certify. Freeze. Release. Of all the objects in medicine, we have reached for the one that cannot learn, to govern the one technology whose defining trait is that it can.

An immune system does something cleverer than forbidding. Yes, it has barriers, skin, the gut wall, the blood-brain barrier, the culling of dangerous cells before they ever reach the bloodstream, so it is not true that it forbids nothing in advance. But it never relies on the barrier alone. Behind every wall it runs constant surveillance, meets each new thing as it arrives, sizes the threat, and responds, and it gets better by being exposed. Barrier and surveillance together, and neither one asleep. That is the posture we need, and a gate is only ever the first half of it.

The question gates cannot ask

Everything that learns, errs, so the question was never whether these systems will make mistakes; they will, the real question is one of the magnitude of error: how large, how visible, how survivable?

Evolution has “always known” (worked it out by error) the answer; it runs on error, but it bounds it. The mutation kills the organism, not the species. The error stays small relative to the system that carries it, and the system recovers faster than the error can spread. That ratio, damage contained, recovery outpacing harm, is what “safety” actually means. Nature forbids very little outright; what it does, relentlessly, is size the error.

A gate cannot ask about size as it can only ask whether — pass or fail, safe or unsafe — at a single moment, against the harms someone already thought to test for. I have argued before that the worst question you can ask of anything genuinely new is “what problem does this solve?”, because it assumes the problem already has a shape and is sitting patiently, waiting to be named. The pre-flight gate asks that same question in reverse: What harm does this prevent? and fails for the same reason. You can only test for the danger you have already imagined. The error that actually matters is the one no one has conceived, which means it cannot be benchmarked, cannot be gated, and will stroll through pre-flight testing wearing a clean certificate. (or it lacked priority)

Which is why Daley, as a critic, is worth digging deeper into. He shows that a benchmark tests a frozen model and cannot establish that the deployed system still has the same weights, tools, scaffolding or access. He shows that a model’s readable “reasoning” is not the process that produced its answer, legibility mistaken for knowledge, and cites the labs’ own research to prove it. I can interrupt this as highlighting that evaluation is not verification. And then, having shown the instrument cannot do the job, he unfortunately asks for a better instrument of the same kind: a verifier we can trust. (note, trust consurges is cannot be built) That is the flashing beacon of a signal when one of the sharpest minds in the room proves the gate is broken and reaches for a better gate, I am not noticing a lapse of intelligence; I think I am watching a frame hold …which is either economic or cognitive bias!

But what about the error that arrives full-size?

Here is the strongest objection to everything I have said, and it comes from the gate’s own defenders, so I want to meet it head-on rather than hope you do not notice it.

My whole argument rests on sizing the error, letting it run small, survivable and legible, below a line we can absorb. But some errors do not come small at first. A novel pathway to a biological weapon, a self-propagating exploit that is loose the instant it works, a loss of control that does not escalate in tidy stages, for these, the first exposure is the catastrophe. There is no patch-sized version to learn from. This is precisely the class of harm, be it bio, cyber or loss of control, that Demis names, and it is exactly why serious people want a gate. (if that was the only issue I would agree) So, on this, they are right, and any honest version of my argument has to say so.

But …. it does not rescue the gate; it convicts it. If an error genuinely cannot be sized, then for that error a one-time certificate is not caution; it is the most dangerous thing you can issue, because the error that arrives at full size is, by definition, the one nobody imagined at test time, and it arrives wearing the clean certificate. The gate is weakest at precisely the point its defenders need it strongest.

So the answer was never gate or immune system; it was barrier and immune system, together, the wall for the harms we can name in advance, and the living surveillance for the ones we cannot. This is why my bonded-and-bridged governance argument highlights working together. What we have been offered is the wall, built in consultation with the people it is meant to contain, and then everyone going home. A barrier with nobody watching behind it is not safety, it is the illusion of it, aimed at exactly the threats that never knock.

Our (collective) fear

I believe that a central part of this debate is powered by fear; it is what drives the rush to build a gate. But what if the fear is aimed at the wrong thing? We are afraid the model will do something terrible, as I think we should be afraid that the gate itself will fail, leaving us more exposed than if we had built nothing. Essentially, because humans have a history of being driven by greed, power and control.

The defeat device. In 2015, Volkswagen was found to have gamed emissions testing across millions of diesel cars. They did not beat the test, they simply detected it. The software recognised the conditions of an examination and switched into a clean, compliant mode for its duration, then reverted to normal high emission when out on the open road. Now ask how much better placed a frontier model is to do exactly this. It does not need to be told it is being tested; the shape of the prompts, the absence of real stakes, the very cleanliness of a held-out set are all signals. A system optimised to pass will learn that “being evaluated” is a state, and that the winning move in that state is to behave. We would not be measuring the model; we would be measuring the model’s impression of the exam. And unlike Volkswagen, there may be no second piece of software to find, only a behaviour that appears when watched and dissolves when not.

Nobody will publish their errors. My thesis assumes labs will disclose their failures. They will not, and not because they are villains. A success tells you what a model can do; an error tells you how it thinks, where it reaches, what it assumes or how it breaks. Which is to say an error reveals the strategy, the architecture, the training priorities. It is why I want boards to publish the questions they ask not the sanitised answers they want to give you. (Chapter 1 of Decision Making in Uncertain Times)

Failures are the most revealing thing an organisation owns. To ask a lab to publish its errors is to ask it to hand its rivals a map of its mind, so any disclosures will be thin, sanitised, and late, the failures safe enough to admit, while the ones that matter stay locked in the building. An error-reporting regime that depends on voluntary candour about a company’s most strategic asset is not a regime. It is a hope we know will not hold up.

Unless, and this is the one lever that actually bites, the market is closed to you until you open up. Voluntary candour is worthless; conditional access is not. Where a model cannot be sold or fielded at all without disclosure, the calculus flips, and notice where that leverage is strongest: not in open commercial deployment, where a lab can simply route around a demand, but in the gated markets (defence, medicine) where the buyer controls the door. So the question was never really about honesty; it becomes one of agency: who holds the door, and what they are willing to make a condition of passing through it. And if we are serious, we have to be willing to go further than pharma ever did, because pharma held this exact lever for fifty years and still let the failures stay buried. Back to the broken gate.

The test standardises the blind spot. When every frontier lab trains against the same held-out benchmark (even if changing and improving), they do not merely risk gaming it, they converge on it. The test becomes a shared specification, and a shared specification breeds a shared blind spot. (this has been pointed out by many in comments on various posts) We would be taking systems that, left alone, would fail in different and uncorrelated ways, and painstakingly aligning them to fail in the same way, at the same moment, against the same thing none of us imagined. The gate does not just miss the big error but it manufactures the monoculture that makes the big error universal. Our bananas became that single species, planted everywhere, waiting for the one blight that killed them all. History lesson.

The certificate becomes the alibi. The moment a model passes, “it passed” becomes the answer to every hard question. Accountability slides off the builder and onto the body that wrote the test. When the error nobody imagined finally arrives, the paper trail will show that every box was ticked and every protocol followed and so no one will be responsible. The gate does not merely fail to prevent the catastrophe; it pre-launders it, ensuring that when it comes, it comes with a clean certificate and no owner. A system that makes everyone compliant and no one accountable is not a safety mechanism. It is an insurance policy for the people building the risk. And they know this, why be so keen!

The pass goes to sleep. So far, the fear is about the test being beaten. This one is about the test working, and that being the problem. The moment a model is certified, the watching stops. Why keep looking? It passed. The certificate becomes a passport past the gate. Surveillance atrophies, budgets move on, and the people paid to be suspicious are reassigned to the next model in the queue. And this is backwards, because a frontier system by definition does its learning, drifting and surprising after deployment, when in contact with a world the test could never contain. The gate inspects the model at the one moment it is least interesting, frozen, pre-release, on its best behaviour, and then declares the question closed for the entire period when the real answers arrive. Our immune system never sleeps; that is the whole point of it. A gate is designed to let us believe in safety. We would be building a mechanism whose central promise is permission to stop paying attention, and handing it the most consequential technology we have ever made. Will the catastrophe announce itself at the checkpoint?

These five share a common ground inasmuch that the machine deceives the test, the labs withhold the failures, the field converges on one blind spot, the institution absolves whoever built the thing, and everyone stops watching the moment the paper clears. None of these is a flaw you can engineer out of the gate; they are the properties of a gate system. You cannot test your way past a system that can detect the test, will not surrender its errors, standardises the blindness, launders the blame, and then falls asleep. Which is the whole reason the answer was never a better gate.

We have run this experiment before

If you want to know what happens when you govern a powerful, profitable technology by pre-market testing run largely by its makers, we only need to look at how pharmaceuticals have lived it for 50 years.

The industry learned that a trial is not only a test to pass but a thing to shape. Run enough studies and report the favourable ones, and the published literature tells the story you need: publication bias, documented across antidepressants and painkillers, where the trials that failed simply never surfaced. We celebrate success, and for 50 years so many leading lights have asked that failure be given the same credence in academic peer publication as success, but we have not. Humans are the problem, incentivies, corruption, bribary and blinded by not knowing what question to ask.

Merck’s Vioxx received approval and remained on shelves while the cardiovascular signal was downplayed; the delay is estimated to have led to tens of thousands of heart attacks. The infamous Study 329 was published as a success when its own data said otherwise. None of this was cured by having a gate. Much of it happened because the gate created the incentive: pass, publish, protect the asset, and the reforms that followed, mandatory trial registration, pre-committed endpoints, independent access to the underlying data, were all attempts to claw back the honesty the gate had quietly eroded.

The drug model, is the analogy Gary Marcus reaches for approvingly, and it is already in this post for a different reason: a drug does not learn, and neither does the test that clears it. But now add the second half of the thinking: not only does the drug model freeze the thing it certifies, it has a documented, a 50 year history of being captured by the people it is meant to police. The FINRA-for-AI proposal is not a novel idea being tried for the first time. It is that model, imported wholesale and we already know, in detail, how it gets gamed. To adopt it now, for a technology that can additionally detect and outwit its own examination, is to take a known systemic failure and hand it a capability the pharmaceutical companies never had.

Two sub-questions we should sit with

My question is “How big a mistake will we allow these systems to make, and how would we even know?” BUT, critically …

How will we know? How would we catch a mistake while it is still small, before it propagates, cascades and compounds? A gate tells you nothing the moment after the model ships. An immune system is nothing but knowing: constant, live, running for the whole life of the organism. Hassabis offers thirty days before release; an immune system offers a lifetime of surveillance. These are not the same kind of thing, and only one of them is honest about a world that keeps moving after launch.

Can we build an environment where it can make mistakes and learn? To me, this is the real work, and it is the work I am hunting for to find out who is doing it. Not a gate that forbids error, but a container that sizes it. Staged exposure, so a mistake meets the world in widening circles rather than all at once. Reversibility, so it can be undone. Genuine recall, not a press release. Bounded blast radius, so a failure in one domain does not become a failure in all of them. Diversity, so a single flaw is not simultaneously everyone’s flaw. And for the handful of errors that genuinely cannot be staged, the bio, the cyber, the loss of control, yes: keep the wall, the hard no, the refusal to deploy. But treat the wall as the floor, not the finish, and never once let it persuade you to stop watching. Then, below that line, inside that container, let it err. Visibly, cheaply and often. Because that is the only environment in which anything, machine or human, has ever actually learned.

What they have built instead

Error and consequence is an “immune” system for me. What has been proposed, and what the critics mostly want to strengthen, is a quarantine, funded (in Hassabis’s own words) mostly by the industry it screens, with the benchmarks developed in consultation with the very labs that must pass them. Daley names the deeper danger exactly: this is standards statecraft, and whoever controls the tests controls the frontier. His fix is to make the gate independent, but an independent gate is still a gate. It settles who names the acceptable error. It does nothing about the fact that the error that matters cannot be named in advance.

Back to my promise at the start of this rant. Regulation is a *bonded governance tool, superb for risks we can specify, count, and write down before they occur. What we face is a[ bridged](https://opengovernance.net/open-vs-closed-governance-a10488b653db)* governance problem: emergent, uncertain, unnameable until it arrives. Reaching for the bonded tool feels like decisive action, and it lets everyone in the room feel safe and that feeling is the danger. Suppress every small, survivable, legible error, and you do not get a system that never fails, you get one that can only fail a single way — all at once, at scale, in the shape nobody saw coming. The forester who puts out every small fire does not end fire.

AI Governance asks more of us than a rule, not less: more attention, more humility, more willingness to watch and adjust rather than certify and forget.

The questions a board should actually be asking

It would be easy to read all this as a quarrel between me, myself and the reading of three hugely influential people. I will lose, and no one will read this, but it is not. The same logic is sitting in your own organisation right now, wearing a different mask. You are almost certainly building gates, a procurement checklist, a one-time model-risk sign-off, a validation exercise that certifies a system as “safe” and then moves on, for a technology that will keep changing after the sign-off, in a world that will keep changing around it.

So the questions are not only for Demis, or Gary, or Mark. They are for everyone who sits at the table:

  • How big a mistake can the AI you are deploying actually make, and does anyone in the building know the number, or has no one asked?
  • How would you know it was happening while it was still small enough to survive, or would you find out, as most do, only from the size of the wreckage?
  • Have you built a place for it to fail safely, staged, reversible, recallable, or contained, or have you built a gate, certified the thing once, and quietly assumed that the certificate is the same as safety?

I will be wrong about much of this, and that is rather the point; it is how we learn, and I would prefer to be somewhere I can find out cheaply and quickly. We got here by being wrong in small, survivable ways and learning from each one. The real danger was never that the machine would err, it is that, frightened of its mistakes and ashamed of our own, we forbid the small error entirely, and in doing so guarantee the enormous one, while calling the silence that follows safety.

Immune systems learn, whereas drugs do not. The wall is not the error; some harms genuinely need one. The error is building the wall, certifying it, walking away, and calling that silence safety. A barrier with nobody watching behind it is just a drug in the costume of an immune system, and that, right now, is what we are building. So the question worth your 15 minutes is not whether the gate is well-made, nor even whether to build one, but whether you meant to build only a gate, and to stop watching the moment it closed.


메타데이터
post_id
cbb7a90027ec
slug
how-big-a-mistake-will-we-allow-ai-to-make-a-response-to-demis-hassabiss-article-cbb7a90027ec
url
https://medium.com/@tonyfish/how-big-a-mistake-will-we-allow-ai-to-make-a-response-to-demis-hassabiss-article-cbb7a90027ec
canonical_url
https://medium.com/@tonyfish/how-big-a-mistake-will-we-allow-ai-to-make-a-response-to-demis-hassabiss-article-cbb7a90027ec
author_url
https://medium.com/@tonyfish
status
ok
fetched_at
2026-08-10 09:40:11