← Back to list

Forbidden to err

Being wrong is critical to (for) progress. Demand a machine that cannot make a mistake, be wrong or err, and we will not build…

Tony Fish · 2026-07-06 20:35 · 0 claps · 19.5 min read
#leadership #ai #decision-making #mistakes #learning
Open on Medium ↗
Wiki topics: AI · AI · General BIZ · Business Strategy EDU · Education & Learning

Forbidden to err

Being wrong is critical to (for) progress. Demand a machine that cannot make a mistake, be wrong or err, and we will not build “intelligence”; we have created “god”

Straight up, this is a provocation for anyone who governs, funds, or deploys AI and believes that “an error” is a defect to be removed rather than the engine that drives change and growth. It is written to make you think …. not to argue or agree.

42 seconds. We have enough evidence to be certain that we learn by “being wrong” or “through errors and mistakes”; evolution, children, athletes, and scientists all thrive on lived mistakes. Yet in compute and AI, we have made “error” a sin, and we have also penalised errors in our education systems, and now consider it a fundamental wrong that a “machine” should ever make one. But “the machine” never learned our way at all; it is optimised to close the gap to a target we set, which is refinement, not becoming. Eliminate the idea of error, and you have not built intelligence; you have built a god, and a god, being finished as it cannot make an error …, can no longer learn. The question we need to explore is “how big an error or mistake will you allow the machine to make … to learn?”

how big an error or mistake will you allow the machine to make … to learn?

Football — well, it is the World Cup 2026

Watch any world Cup football game, and you already know how it will be lost, not by the better side being outplayed across 90 + 6 minutes for advertising (don’t go there) but by one touch, one misjudged backpass, one keeper who comes for a ball they should have left, a single “error” (as judged by the losing sides commentators) that settles everything and if in England it will be replayed until the player who made it can barely walk down his own street, unlike the Lionesses who won. At all elite levels in any sport, this is probably a law of nature, because two individuals or teams arrive so evenly matched, so drilled and so near to flawless, that the contest is no longer who is better but who “errs” less, and the loser is simply the one who made the most mistakes. (often forced by the other team)

What global sport teaches supporters is that an “error” is how you lose. We learn it on the pitch and in the exam hall long before we can vote, a mistake is a defeat and a shame to be carried, as a neurominority I, like so many others, have the scars.

However, is that not the reverse of the truth? Every player on that pitch reached the final by erring 10,000 times where no one was watching, through the miss-hit passes and the missed touches and the failures on the training ground that were never optional but mandatory, the only way anyone becomes good enough to win or lose a competition. The same word, error, is doing two jobs, separated only by who is watching and when: the one that built the athlete and the one that is blamed for the result.

Question: “Was it possible to learn before we had a word for wrong/ error/ mistake?”

The reality is that learning was there as the fire got out of control, the animal escaped, the slipped hand, the heavy touch, the missed guess, the copy of a gene that came back imperfect. What the words “error/ fault/ wrong/ mistake” added was a “category”, and the category is where the power lies.

Once you can name “error” you can forbid it, mark against it, optimise it away, and build an institution on the supply of forgiveness. Schools do it with red ink, sport does it with the post-match inquest, regulators do it with frameworks, and we are now doing it, at scale and with extraordinary conviction, to the AI machine we have decided is too important to be wrong.

However, learning appears to have been learned by being wrong, making errors, creating mistakes, and the machine we are creating may be a powerful thing we do not want to get wrong, because we fear the consequences.

The mechanism we should not be ashamed of

Evolution does not select for correctness but operates on creating and copying errors. The majority of mutations are mostly useless, sometimes fatal, and just occasionally waiting for the day the world changes and turns yesterday’s defect into today’s survival. Reproduction and innovation are “error machines” that pay off. The COVID immune quirk, the exaptation, the ancient error that becomes the essential strength forty thousand years later, none of it arrives from a process that got things right. Life appears to have arrived from a process permitted to get things wrong, repeatedly, at volume and without judgement or shame.

We accept this in biology, but economics and business ideology means we spend our energy focused on removing errors. We have built an entire business, media and political circus around “error reduction” and called it progress. Six Sigma, zero defects, and marking schemes that reward the reproduction of the agreed truth/fact, whilst penalising the deviation that might have been the invention or discovery. I am not saying these methods are wrong or do damage, but it is a mindset. We have quietly stripped error out of the places it used to live, the workshop, the kiln, the kitchen, the pitch, the studio and the laboratory bench, the messy physical domains where you learned because the clay collapsed, the weld cracked, the sauce split or the experiment failed. Remove the error and you remove the teacher. In this line of thinking, we appear to have mistaken the absence of errors and mistakes for the presence of being better.

I am standing on the shoulders of others, and I want to show some respect before we reach the part that is critical for AI, governance, and decision-making.

Background

Philosophers such as Peirce called it fallibilism, and Karl Popper who built it into an entire theory of knowledge, conjecture and refutation, where a claim counts as knowledge only if it can be proven wrong, got here a long time before me. They realised that to forbid error, and the claim is no longer falsifiable, which in Popper’s own terms means it is no longer knowledge at all. David Deutsch, a British physicist at the University of Oxford (who is often described as the “father of quantum computing”), presents error correction as the single process behind all knowledge and behind open societies too, because the ones that survive are those that make their mistakes cheap to find and to fix. Nassim Taleb (Black Swan) argues that suppressing the small errors does not remove risk but stores it, transferring it into the rare and catastrophic blow-up, so that the system, which never wobbles, is not the strong one but the one quietly accumulating fragility until it shatters.

Those leading in the study of the brain now believe that the brain looks less like a camera (old model) and more like a prediction machine (probability), able to fill in where there is an “error” in the forecast. Predictive processing theory holds that learning is proportional to surprise (errors in the forecast). We update exactly as much as the world violated our expectation, so without error, there is no update at all. Error has value, but we want to eliminate it. Smell the paradox?

But the machine does not learn the way we/you do

It is tempting to say that we built the machine on error and then forbade it to err, which is both tidy AND false. The machine is not trained on error; it is trained on data that contains known and unknown errors. Data is a vast semi-frozen historical corpus riddled with mistakes, biases and falsehoods that a machine cannot perceive as mistakes, and it is then “tuned” to minimise prediction error against that corpus. LLMs are not learning from their mistakes; they are learning to reproduce a target that is itself full of mistakes that nobody has flagged as an error, never corrected by the world it describes, but tuned towards a frozen record of it. Then, at deployment, the (training) weights stop moving. Inside a single conversation, the model can still be corrected, and it will behave differently for the rest of that exchange, a real adaptation, but one that dissolves the moment the session ends and never touches the weights inside the beast, a fast loop with a human inside it, not the slow one that trains tomorrow’s model. When a deployed model hallucinates and no human is there to correct it, it learns nothing from the event, because there is no consequence, no update and no becoming. The biological mutation and the model’s hallucination are not the same kind of object at all, since one is the raw material of becoming and the other is inert.

So the asymmetry is that humans learn from errors and mistakes and the machine we have does not. (new model relearn but old models don’t) We are the system that grows by being wrong, feeling the consequence and updating, whilst the machine is a system that minimises a loss against a flawed snapshot and then stops. We have spent this whole anxious moment getting the two backwards, pouring our fear into making the frozen machine never wrong, whilst quietly removing real error, the actual teacher, from the one learner that genuinely runs on it, which is us.

Here is some thinking exploring Agentic AI + Harness (self-learning cognitive), where this can learn, not the Agentic AI but the Harness.

So, where is the AI provocation? The machine as a single artefact does not learn by inference, which is true, but the system you are governing is not one frozen artefact; it is a loop. Today’s deployment generates tomorrow’s training signal, the failures, the surprises and the places it broke against a world it forecast wrongly. A “zero error” deployment regime does not sit harmlessly downstream of learning; it starves the loop of exactly the surprising failures the next model needs in order to become anything more than this one. You are not preventing mistakes; you are foreclosing the next generation’s only curriculum and calling that prudence. See the irony?

Three errors, one word

“error” is carrying at least 3 different ontologies, and we collapse them into one, which is where both the bad governance and the bad theology come from.

ONE “error” as gap, the engineer’s error. In a control loop, the error is simply the distance between where you are and where you said you wanted to be, the setpoint minus the measured output, written e = r − y, and the entire loop exists to drive that distance to zero. This is “negative feedback”, a highly successful idea in engineering, running the thermostat, the autopilot, the cruise control and the controller in every plant in manufacturing. It is goal-oriented (teleological) to the core, because the target is fixed in advance and the apparatus converges towards it, so the error is not a mistake to be learned from but a measurement that tells the loop which way to push.

What is outside of the loop is who set the setpoint, and when, and what each attempt was allowed to cost, because the order you tune yourself to was never a fact about you, it was a fact about the price of being wrong in the world that trained you, and that price moves. A loop shaped by trials that were rare and expensive learns to be heavily damped, cautious, slow to close the gap, because a large overshoot might be the one you do not recover from. A loop shaped by trials that are cheap and constant learns the opposite, to run underdamped on purpose, to overshoot fast and correct faster, because by the time caution has finished being careful the world has already moved the setpoint again.

The behaviour of that loop depends on its “order”, and the orders are worth describing clearly because the “harnesses” now wrapped around AI are built from them. A first-order design responds in proportion to the present gap and approaches the target smoothly, governed by a single time constant, never overshooting but often never quite arriving, leaving a small steady-state offset it cannot close. A second-order design adds memory and inertia, the accumulated integral of past error, and now it can close that offset completely, but it pays for the privilege with overshoot, going past the target so that the gap reverses sign and the reversed gap pulls it back, ringing around the setpoint before it settles. How much it rings is governed by a single number, the damping ratio ζ, and the size of the first overshoot follows from it through exp(−ζπ/√(1−ζ²)), so that an underdamped loop oscillates whilst a critically damped one arrives as fast as it can without ever crossing the line.

A third-order design and above adds anticipation, the derivative of the error, a model of where the gap is heading rather than only where it sits, which lets the loop pre-empt, and at the same time gives it the capacity to go unstable, to amplify rather than damp, to chase its own tail. In cybernetics this is the climb from a loop that corrects towards a fixed setpoint, to a loop that can move its own setpoint, to a loop that can change how it moves its setpoints, what Chris Argyris called single and double loop and Roger Bateson called “deutero learning”, the learning of how to learn.

Every one of these, 1st, 2nd and 3rd order, are all “error as gap.” Even the overshoot, which looks the most like productive wrongness, is nothing of the kind, because it is a gap with a reversed sign, the loop going too far on its way to a target it already held, producing nothing the setpoint did not already specify. It converges, it refines, and it arrives, but it does not become.

TWO “error” as mutation, the biologist’s error, and here there is no setpoint at all. The copying error in the gene is not measured against a target and corrected, because there is no target, only variation thrown into a world that will decide, retrospectively, whether any of it was worth keeping. This error does not converge but diverges, opening options that were never specified in advance, most of them useless, and the rare one becoming a target only after the fact (has an advantage in this context). This is the error the whole of this article has been about, the generative one, the deviation that creates rather than corrects.

THREE “error” as transgression, the moral error, the thing that is named, forbidden and blamed, the hallucination reclassified as sin and answered from the pulpit. It is neither a measurement nor a mutation but a category laid over the top of both, and whoever holds it decides which deviations are faults to be confessed and which are simply the cost of a system that is still alive.

One word, three jobs. The gap converges towards a point you already labelled, the mutation discovers points no one had thought of, and the transgression judges both against a point someone in authority chose. The damage begins the moment a single word is allowed to carry multiple definitions, because then a generative mutation can be judged as a transgression and corrected away as though it were merely a gap, whilst the machinery doing the correcting gets called learning when it is doing nothing of the sort.

This is exactly what is now happening to the language of self-learning around large language models. The harnesses we are wrapping around these models, the self-critique loops, the agentic correct-and-retry, the verifier and reward loops, the reflect-and-revise scaffolds, are control loops, and the error they run on is “error as gap.” They generate a candidate, measure its distance from a target, whether a reward model, a test that passes or fails, or a critic’s judgement, and refine towards that target, overshooting and correcting much as a 2nd or 3rd order feedback design does. They are very good at it, but they are converging rather than mutating. A model that refines itself towards your reward signal is not becoming something new; it is arriving more precisely at something you already specified, so to call it self-learning is to smuggle the biologist’s error in through a door marked with the engineer’s.

The danger is real as the harness genuinely corrects, and we will mistake its fluency at closing gaps for the generative error this essay is defending. We will point at the self-improving agent and say that it learns from its mistakes, when what it does is drive a measurement to zero against a target it cannot question. The mutation, the deviation that would create a setpoint nobody was holding, stays exactly as absent as it was before, and we carry on forbidding it in the only two places it actually lives, the surprising failure at deployment and the lived mistake in ourselves. So when you are told that a system learns from its errors, the first question worth asking is which error, the gap or the mutation, because almost always it is the gap, and the gap, however elegantly it is closed, was never the thing that made anything new.

The pulpit moves …

This section picks up from this article, “Is Probability a new pulpit?In short, it exploited the fact thatnaming” has always been how power works.

We call the model’s failures hallucinations, and we have made their elimination the central moral project to debate. Notice what that framing does, because it takes the surprising deviation, the thing that, fed back, is the engine, and reclassifies it as sin. Whoever gets to define “sin” holds the pulpit. The body that sets the threshold for acceptable error in an AI system is not making a technical decision; it is naming the category, drawing the jurisdiction and deciding what may and may not arise, which is not safety engineering but theology with a confusion matrix.

Prof Sidney Dekker has spent years dismantling the corporate cult of “zero harm”, arguing that the vision of “zero is not the absence of failure” but the absence of honesty about failure, a theatre of clean dashboards that drives the real errors underground where they cannot be learned from. Amy Edmondson draws attention to the fact that the zero crowd rejects the distinction between preventable failure in known territory, which you should indeed reduce, and intelligent failure at the frontier of knowledge, which is the only way the frontier ever moves. Lump them together under one word, whether error or hallucination or defect, and you will optimise away the second whilst congratulating yourself on the first. How many times have I and others said this ….

In the language I have been using for a while now, you can tend the conditions from which capability consurges, or you can optimise those conditions away in pursuit of a clean number, and the clean number is a desurgence, the state where conditions collapse and nothing new arises any more. Most of the AI governance I read is a beautifully documented blueprint for desurgence, all mountains of assurance that everything is fine, precision about exactly the wrong things, and the quiet removal of the one capability, to be wrong in a generative direction, that made the thing worth having.

The “god” problem

This brings me to the oldest question hiding underneath all of it, and it turns out the theologians have been fighting about for a while … a long while

“God”, in some classical ideals, never learns and never errs, as in Aristotle’s unmoved mover and the scholastic actus purus, pure actuality with no unrealised potential and therefore no change. Is that a comforting attribute? A being with no potential left to realise cannot update, cannot be surprised and cannot become, which means perfection and learning are mutually exclusive at this level. To be incapable of mistake (different from changing your mind) is to be finished, complete and static, a closed book in a universe that keeps writing new pages, and the “god” who never errs avoided having to learn at the price of the end of becoming.

This is precisely the fault line that process theology, through Whitehead and Hartshorne, opened against the classical picture, a god genuinely affected by the world, responsive, in process and able to be moved. The argument that a perfect being must be a static one, and that a living one must therefore be in some sense unfinished, is not my heresy; it is a live dispute, and we, with our new shiny AI tools, are wandering into it carrying a GPU.

Here is a question: When we demand a machine that does not “err,” we are not asking for intelligence; we are asking for the classical god, the frozen, finished, (absolute perfection) thing, and then expressing surprise that it cannot do the one thing that made us, which is to get it wrong, feel it and become something we were not yesterday. If improvement comes from trial, error, correction and optimisation, and it does, since that is the whole of the case above, then the governance question is not how we stop it from making mistakes. It is how big an error we are willing to let it make, because the size of the error you permit is the size of the learning you allow, and an organisation that permits none has chosen, quite deliberately, not to grow, however carefully it dresses the decision up as prudence.

do you realise how IMPORTANT that question is?

The err you learned is not the err you need

There is a second asymmetry, and it has nothing to do with machines. It is the gap between the errors you learned to make and the errors the world now prices, and if we are honest, most of us are still running a loop tuned to a world that no longer exists.

The Wright brothers, and every steam engine builder before them, learned through years of trial and error because each trial cost a life, a limb, a machine, or a season, so the skill that generation actually built was not flight; it was how to survive being wrong when wrong was catastrophic, patient, heavily damped, and suspicious of speed. We look back and call it craft, but it was calibrated caution wearing craft’s clothes.

Raising money and building a business in the 1990s ran the same way. A bad meeting could cost you a year, a round could take eighteen months, so the skill was patience, rationed against how rarely you got to try again. Today, an idea can become a business with £10k a day in revenue inside a week, because the tools and the access have collapsed the cost and the cycle time of a trial while multiplying how many you get to run. That is the paradox worth naming plainly: more tools and more access did not make us err less, it made erring cheaper and more frequent at once, and cheap, frequent error does not feel like error at all, it feels like iteration, which is exactly why this moment reads as more creative and more innovative than any before it whilst the underlying mechanism, being wrong and correcting, has not changed in the slightest.

I used to fix the bugs before going live. Now we go live and fix as we go, shipping behind a feature flag, running the A/B test as an institutionalised error nobody is ashamed of, selling a product that does not exist yet from a landing page to see who pays, building in public and letting a timeline correct us in hours rather than a board correcting us in quarters. And a version of all of it is the one this article already called out: correcting a model mid-conversation and watching it adapt, visibly, for the rest of the session, the very same asymmetry, now happening to the person doing the correcting, all day, at speed.

We built our frameworks, patience, caution, and the long apprenticeship, out of a world where trials were scarce and expensive, and we love those frameworks because we paid for them in years, which is precisely why they are so hard to put down. But the frameworks were never timeless; they were contextual, tuned to a cost of error that has since moved twice in 40 years while we stayed exactly where the framework left us.

So the skills worth building now are not the old list; they are the old list plus the two variables the old apprenticeship never had to separate because it never had enough trials to tell them apart. Errors, refinement, performance, mistakes, thinking, dwelling and time were always on it. Cost belongs there as well, what each attempt is allowed to cost you (risk), because that is the “number” that actually sets how cautious a person needs to be, and repetition belongs there, how many trials you get to run inside a unit of time, which is the variable that moved furthest between 1985 and 2026, further than error itself. Time itself deserves splitting into different meanings rather than being left as one word: feedback latency, how long between acting and finding out you were wrong, which the tools have collapsed towards zero, and dwell time, how long you sit with the result before it actually changes you, which nothing has collapsed at all. Generate data faster than you can turn it into learning, and you get noise dressed as progress, which may be the modern failure mode we see but don’t have a name for — what was broadband before we had the word?

What was to err in 1980 is not what is to err in 2026, and the frameworks we propagate, love and attach ourselves to were built to solve a problem we are no longer solving, however hard-won they were, however much we owe them. I need to take this on

The questions we did not know we had to ask

These are not questions to agree with but questions to remain uncomfortable with, and they are best asked slowly enough that the discomfort of not knowing the answer is held rather than resolved.

The first is, “what actually is your tolerance for error?”, stated as a number, and whether you understand that you have just set your ceiling for learning at exactly the same height, because if the answer is zero then you should say plainly what you have chosen, a system that performs and does not improve, which may be right for the ledger and the cockpit where wrong is simply wrong, but is catastrophic for the strategy.

The second turns away from the machine and back towards your people, “who are the learners that actually run on error?”, and asks where you have already stripped it out of them, in the graduate scheme that rewards reproducing the model answer, the review that punishes the productive failure, and the craft and the workshop and the bench that you made efficient, because whilst you obsess over the machine’s mistakes it is worth knowing what learning has already left the building with theirs.

The third asks, “are you starving the loop?” since your deployed model does not learn from today’s failure, but the next one is trained on it, so a zero-error deployment regime quietly removes tomorrow’s curriculum, and freezing the artefact is not the same as governing the system.

I would argue that the third matters most when the world changes, as it will, and asks which of your current errors is the ancient mutation that saves you, because you are under pressure right now to eliminate the redundant, the inefficient, the deviant and the off strategy, and some of that is precisely the immune quirk you will need when the shock you cannot yet name arrives. Efficiency is the removal of options, so the only thing worth knowing is how many you are removing, and which.

Skill was never a certificate

So should you defend the skill you worked years to earn, or hand it over to whatever the tools now make possible? ouch

That question already contains its own mistake, because it treats the skill as a destination rather than as one already-spent trial in a loop that was always supposed to keep running. ouch

A person who holds a hard-won instinct from 1995 as identity rather than as one input still waiting to be updated is doing exactly what this article accused the machine of doing earlier: training once on a corpus, tuning to a target that has since moved, and reproducing that target forever, mistaking the absence of new error for the presence of continued competence.

The “god or wicked problem” was never only about the machine. It was about anyone, human or otherwise, who stops being wrong on purpose and calls the stillness mastery, because perfection does not require silicon, it only requires deciding, at some point, that you are finished learning and are now simply right, and the frozen skill defended past its context is that decision wearing a CV.

What was to “err” in 1980 is not what is to err in 2026, so the discipline is not to keep the old errors sacred; it is to go looking for the new ones, on purpose, in public, at whatever price the world is currently charging for them. The alternative is not safety. It is a slower, more personal version of the exact freezing this essay has been warning about the whole way through.

I did not arrive here by being right, nor am I ever right. I arrived by being wrong in a million small, patient, generative ways and surviving long enough for some of them to matter, and the machine we have built does not work like that, because it never tasted a mistake in its or someone else’s life. The danger, then, is not that it will err, but that, frightened of its failures and ashamed of our own, we forbid the one thing that still makes us capable of anything, the lived mistake, and call the silence that follows safety.

Let’s debate “How big a mistake will you allow the machine to make ?”


메타데이터
post_id
9007ca6cae6c
slug
forbidden-to-err-9007ca6cae6c
url
https://medium.com/@tonyfish/forbidden-to-err-9007ca6cae6c
canonical_url
https://medium.com/@tonyfish/forbidden-to-err-9007ca6cae6c
author_url
https://medium.com/@tonyfish
status
ok
fetched_at
2026-07-10 04:31:59