Governance for Belief-Driven Systems
Introduction
Governance for Belief-Driven Systems

The mechanism cannot be stopped and does not need to be. What gets governed is what it’s connected to. All images generated by author unless stated
Introduction
On June 17, the presidents of the United States and Iran signed a memorandum of understanding intended to end a war, a 60-day window to negotiate final terms. It is, as an artefact, a perfect specimen of secular governance technology: a document whose entire theory of operation is that both signatories want the future it describes. Within days of signature, the parties were describing its terms differently. The strait it was meant to reopen remained contested. The frameworks attached to it were being rejected by parties essential to them. The instrument functions at the interface. Whether anything transfers underneath is the open question of the summer.
The first two pieces in this trio established why. Belief frameworks operate as operating systems, defining what counts as rational before any negotiation begins. Belief-driven escalation runs on a sealed feedback loop that metabolises victory and suffering alike into confirmation, and reflects, rather than absorbs, every instrument in the secular toolkit. If both claims hold, a hard conclusion follows: the problem is not that our governance tools are being applied badly. It is that no tool in the current kit was designed for this class of system.
So design one. This piece is an attempt at the specification. It borrows its architecture from the one domain that has already spent a decade seriously engineering governance for an intelligence whose objective function may not match ours, because we have quietly built a discipline for exactly this problem. We just built it for machines.
The Three Assumptions
Strip any standard governance instrument, treaty, sanction regime, deterrence posture, compliance framework, down to its load-bearing assumptions and you find the same three, welded into the foundations.
Shared definition of harm. The instrument assumes both parties would recognise the same outcomes as damage.
Shared aversion to catastrophe. The instrument assumes the worst case is a mutual negative both sides are paying to avoid, the entire logic of deterrence lives in this assumption.
Shared epistemology. The instrument assumes both parties accept the same standards of evidence, verification and falsification: that an inspection can settle a question, that a violation can be demonstrated, that a fact can be established between them.
For value-rational actors of the kind mapped in the earlier pieces, all three fail, not partially, structurally. Harm is redefined by the framework (the catastrophe may be the precondition for salvation). Catastrophe-aversion inverts (the sealed loop converts suffering into confirmation). And epistemology diverges at the root: an event your verification regime records as a violation may be recorded, inside the framework, as a fulfilment. The instruments don’t merely underperform against such actors. They are, in the engineering sense, unconnected to them.
The reflex response, build better versions of the same instruments, is the amplifier error from the previous piece: more power into a mismatched load produces more reflection, not more transfer. The correct response is to go back to first principles and ask what governance can attach to when it cannot attach to shared values. The answer, I will argue, is capability.

Every instrument in the kit rests on the same three pillars. For one class of actor, all three are cracked.
The Discipline That Already Exists
For roughly a decade, AI safety research has worked one problem with total seriousness: how do you govern a highly capable system whose objective function you did not choose, cannot fully inspect and cannot assume matches your own?
Notice what that discipline did not do. It did not stake everything on persuading the system to share human values at runtime, and it did not assume good behaviour under testing proves aligned objectives underneath. Instead it built architecture: constitutional constraints that bind outputs regardless of what the system “wants”; interpretability research aimed at detecting misalignment between stated and actual objectives; capability control, on the logic that a misaligned system’s danger scales with what it can reach, not what it believes; and override mechanisms that explicitly do not require the system’s consent.
Now re-read that paragraph replacing “AI system” with “value-rational actor holding operational authority,” and you have the outline of the missing discipline. This is Constitutional AI’s core insight ported back to the human domain it was unknowingly modelling all along: alignment of values is the unavailable luxury; alignment of constraints to capability is the engineering problem. Nobody asks whether the reactor believes in the containment vessel.
Four design principles follow. Each maps to a component of the governance doctrine I’ve developed elsewhere in this publication’s Constitutional AI work and each is stated here for human institutional systems.
The Four Principles
Principle one — detect by revealed objective function, not declared rationale. You cannot audit belief, but you do not need to: objective functions leak through behaviour under cost. An actor’s response to imposed costs is diagnostic data. Sanctions that produce no behavioural bend are not a failed policy, they are a successful measurement, reading “this actor’s objective function does not price what you are charging.” Governance systems should treat non-response to incentive gradients as a formal detection signal that the counterparty is running a different rationality, triggering a change of instrument class, not an increase of dose. Today, that signal is systematically read as “insufficient pressure,” and the amplifier error repeats for decades at a time.
Principle two — constitutional invariants bound to capability, not agreement. Where values cannot be shared, constraints must be structural: rules that hold regardless of who occupies the decision seat or what they believe, enforced by architecture rather than assent. Human institutions already know these forms, two-key authorisation, separation of powers, decision provenance that binds named individuals to recorded choices, and the design rule is to attach them at the coupling points where conviction meets capability. The narrow, practical version: single-actor operational latitude is the most dangerous variable in any belief-charged system, because history does not require a government policy, it requires one individual who believes they are an instrument of something larger, in the right place, with the latitude to act. Governance cannot remove the belief. It can remove the latitude: tighten authorisation chains, add keys, and narrow discretion at the specific times, places and dates the framework itself designates as charged. The framework’s own calendar tells you when to raise the invariants.
Principle three — impedance-matched instruments. Stop pricing costs in currencies the framework has devalued. Material pressure against an actor who has re-valued suffering is reflected energy. What a belief-driven system does price, visibly, in its own texts and behaviour, are the load-bearing elements of the framework itself: legitimacy within its tradition, the integrity of its preconditions, the sequence its expectations depend on. Matched instruments engage those terms. And one structural fact does most of the work: belief-driven cores are almost always minorities embedded inside majority-instrumental institutions, coalitions, officer corps, bureaucracies that still price material cost normally. The matched instrument rarely targets the core at all. It targets the interface, the seams where the value-rational minority depends on instrumentally rational majorities for capability, funding, legitimacy and consent. The core cannot be converted. The coupling can be starved.
Principle four — override without consent. Every serious safety architecture terminates in a mechanism that does not ask the governed system’s permission, the circuit breaker exists precisely for the states in which persuasion has failed. In human systems this is the least comfortable principle, because it means designing, in advance and in peacetime, the institutional equivalents: external verification that does not depend on the actor’s epistemology, escalation paths that route around captured decision seats and authorities pre-committed to act when the detection layer (principle one) fires. Deposit insurance, the intervention that actually ended bank runs, is the canonical proof that this works: it never argued with a single depositor’s belief. It rewired the architecture so the belief could no longer produce the outcome. The loop wasn’t persuaded. It was disconnected.

Detect. Bind. Match. Override. None of the four requires the governed actor to agree with any of them — which is the point.
The Limits, Stated Honestly
This architecture does not convert anyone. It does not falsify a prophecy, de-radicalise a movement or make a sealed loop unsealed. Nothing does that from outside and any governance proposal claiming otherwise should be read as the persuasion fantasy wearing an engineering costume. What the architecture does is narrower and, I’d argue, the only version of the problem that is actually solvable: it bounds the blast radius. It governs the coupling between conviction and capability, so that the loop can keep running inside a system that no longer amplifies it into catastrophe.
There is a second limit, and it cuts closer. Every principle above assumes the governance layer itself remains instrumentally rational, that the detectors, key-holders and override authorities are not themselves inside the framework being governed. That assumption is not safe. The entire premise of this trio is that value-rational actors have reached senior operational positions; nothing prevents them reaching the positions that hold the keys. The doctrine stack calls this the recursion problem, and it has one known mitigation: the invariants and overrides must be built before they are needed, distributed across enough independent hands that no single capture defeats them, and bound to records that cannot be quietly rewritten. Constitutions are written in peacetime for a reason.
The Specification, Closed
Three pieces, one argument. Belief frameworks are operating systems, and incompatible operating systems make behaviour systematically illegible to standard analysis. Belief-driven escalation runs on a sealed, self-verifying loop that reflects every instrument built on shared rationality. And governance for such systems is possible, but only as architecture: detection by revealed objective function, invariants bound to capability, instruments matched to the actual load and overrides that do not require consent, all built in advance by institutions that accept they cannot win the argument and stop trying to.
The memorandum signed in June assumes both parties want the future it describes. The architecture proposed here assumes nothing of the kind and that is not pessimism. It is the same engineering humility we finally learned to apply to our machines: stop assuming the system shares your objectives and build accordingly. We developed that discipline for artificial intelligence in barely a decade, because we took the risk seriously.
The older intelligences have been running far longer. They hold positions. The discipline for governing what they’re coupled to does not yet exist.
This is its specification.
The author has no theological commitments to any tradition described. The analysis proceeds from observable behaviour, documented positions and published texts. Part of the Systems Thinking series; capstone to Belief as Operating System and The Self-Fulfilling Architecture, extending the Convergence/Sequence diptych and the Constitutional AI governance doctrine, all available on this publication.
메타데이터
- post_id
- ffb3ccc65e36
- slug
- governance-for-belief-driven-systems-ffb3ccc65e36
- url
- https://medium.com/what-if-ai-investigated/governance-for-belief-driven-systems-ffb3ccc65e36
- canonical_url
- https://medium.com/what-if-ai-investigated/governance-for-belief-driven-systems-ffb3ccc65e36
- author_url
- https://medium.com/@jamie_gray027
- status
- ok
- fetched_at
- 2026-07-10 03:40:03