← Back to list

IASEAI’26 Part 2: An International Agreement to Prevent Creation of Artificial SuperIntelligence.

Also see: Part 1 IASEAI’26

Kevin O'Shaughnessy · 2026-03-02 18:51 · 2 claps · 29.4 min read
#superintelligent-ai #superintelligence #artificial-intelligence #ai-governance #ai-safety
Open on Medium ↗
Wiki topics: SAF · Safety & Alignment AI · AI · General

IASEAI’26 Part 2: An International Agreement to Prevent Creation of Artificial SuperIntelligence.

[embed]

Also see: Part 1 IASEAI’26

This article gives a summary of three short talks and some discussion with Claude on Concordia, Chinese relations and Moltbook.

Three hours in from the start of the day, is a talk from Peter Barnett from the MIRI Technical Governance Team.

This talk is based on the 2025 paper **An International Agreement to Prevent the Premature Creation of Artificial Superintelligence** by MRI researchers Aaron Scher, David Abecassis, Peter Barnett and Brian Abeyta.

Barnett says it’s uncomforably likely that humans will build superintelligence.

Major risks of superintelligence:

1. Loss of control

2. Misuse

3. War

4. Concentration of power

  • Some AI company leaders estimate 10–30% catastrophe risk
  • Academics and safety experts often cite 20–50% catastrophe risk
  • This is totally unacceptable by any normal standard

Draft International Agreement

Written by MIRI. If implemented it would postpone the development of ASI until it can be done safely and other key issues are solved.

Challenge: States will not trust each other

Solution: Verify every part of the agreement

Three key elements:

  • AI chips
  • AI training
  • Dangerous AI research

Governance: coalition of states, U.S. — China led

Verification: involves inspections, whistleblower protections, state intelligence

Nonproliferation and enforcement: export controls and escalating responses for noncompliance up until military action

Barnett says the agreeement will work because it involves locating existing chips and tracking new chips, and verification technology would monitor the use of these chips.

That’s the summary. If you want to know more, read the paper **An International Agreement to Prevent the Premature Creation of Artificial Superintelligence. **I invite you to comment on whether or not you think this idea could work.

Next up is Gabriel Wagner from Concordia AI. Wagner describes Concordia AI as a Beijing based social enterprise focused on AI safety and governance, and his talk summarises the State of AI safety in China report which was published July 2025.

To Claude Opus: Please tell me about Concordia AI. Who in is charge, who finances and what are the company objectives?

Claude Opus: Here’s what I found about Concordia AI:

What it is: Concordia AI is a Beijing-based social enterprise focused on AI safety and governance, founded in 2020. It also has an office in Singapore. It’s not a typical for-profit company — in 2023, it was certified as a social enterprise under a Beijing Civil Affairs Bureau policy, which requires spending at least 35% of after-tax profits on projects with a public purpose.

[embed]Concordia AI: 2023 Annual Review Concordia AI's mission is to ensure that AI is developed and deployed in a way that is safe and aligned with global…aisafetychina.substack.com

Leadership: Brian Tse is the founder and CEO. He has a background bridging Western and Chinese AI communities — he is a Policy Affiliate at the Centre for the Governance of AI and former Senior Advisor to the Partnership on AI, and has advised leading AI firms on safety and governance. Other key team members include Liang, who leads advisory work on AI policy and standards in China (formerly a senior technical consultant at Baidu), and Yuan, who manages partnerships with frontier AI labs (formerly at ByteDance working on TikTok policy). The team appears to be around 20–50 people including a global network of part-time affiliates.

[embed]Chinese AI - They're Just Like Us? With Beijing-Based Concordia AI CEO Brian Tse Brian Tse, CEO of Concordia AI, discusses China's pragmatic approach to AI development, safety, and governance. The…www.cognitiverevolution.ai

Funding: This is where information is thin. According to Tracxn, Concordia AI has raised at least one round of funding, but details on specific amounts and donors aren’t publicly disclosed in the sources I found.

The organization doesn’t appear to have transparent public reporting on its funders, which is worth noting. Given its positioning in the effective altruism–adjacent AI safety world (it’s listed on 80,000 Hours) and its social enterprise status, it likely receives a mix of grants and consulting revenue, but I can’t confirm specifics.

Objectives: Concordia AI aims to steer AI development to harness transformative benefits while mitigating risks, collaborating with policymakers, academia, and leading AI labs. Its work falls into three main pillars:

[embed]About Us - Concordia AI AI represents one of the most transformative technologies in human history. Rapid AI progress raises critical questions…concordia-ai.com

First, it participates in Chinese national AI standard-setting committees and provides expert consultations to government bodies. Its expertise has informed Chinese government policy consultations and global policy dialogues.

Second, it collaborates directly with AI labs on safety frameworks. It advises leading AI companies on risk management practices, providing guidance on identifying, evaluating, and mitigating risks.

Third, it promotes international dialogue — particularly bridging Chinese and Western AI safety communities. It published the world’s first comprehensive reports on AI safety in China and Singapore, and runs a notable Substack newsletter called “AI Safety in China.”

It occupies a fairly unique niche as one of the few organizations working on AI safety with deep roots in both the Chinese and international AI governance ecosystems. Whether you view that as a valuable bridge or a complicated position probably depends on your perspective on US-China tech dynamics.

Key features of China’s overarching national governance:

  1. Emphasis on application development, not AGI
  2. Senior leadership does not view development and safety as zero sum
  3. Official rhetoric on AI safety has become increasingly prominent and specific

Claude: Here’s a summary of the report’s key findings. It’s their third annual edition (the series started in 2023), covering developments from May 2024 to June 2025 across five domains: domestic governance, international governance, technical safety research, expert views, and industry development.

AI safety elevated politically. At the Third Plenum — one of China’s most important political events in 2024 — AI safety was formally elevated to a national priority. In April 2025, the CPC Politburo held a study session dedicated to AI, calling for monitoring, early warning, and emergency response systems. AI safety has gained prominence in official meetings and emergency planning documents related to national security and public safety, with AI risks now listed alongside epidemics, cyberattacks, and financial anomalies in China’s National Emergency Response Plan.

Standards surge. From January to May 2025 alone, China issued as many national AI standards as in the preceding three years combined. However, a comprehensive national AI Law remains unlikely in the near future. Instead, China continues to operationalize safety through a growing system of technical standards and existing regulations, including the pre-deployment registration regime that had registered 346 models by March 2025.

International positioning. At Davos in January 2025, China’s top science and technology official warned that developing AI without safety measures is like driving on a highway without brakes. At the Paris AI Action Summit, Chinese institutions launched the “China AI Safety & Development Association” as a counterpart to other national AI safety institutes. China has emphasized AI capacity-building in the Global South, sponsoring a UN resolution adopted by over 140 countries. New bilateral AI dialogues were launched with countries including the UK, though the US-China AI dialogue has stalled since May 2024.

Technical research growing fast. Chinese frontier AI safety research output more than doubled compared to the previous year.

Industry commitments. Leading Chinese AI companies signed voluntary “Safety Commitments” covering red teaming, transparency, and safety research.

The overall picture the report paints is of a China that increasingly treats AI safety as a serious governance priority, but through its own institutional mechanisms — standards bodies, registration regimes, and party-state coordination — rather than through a single comprehensive law.

The report also highlights that while there’s genuine overlap between Chinese and Western concerns about AI risks, the two ecosystems remain poorly connected, particularly with US-China dialogue having stalled.

Kevin: There’s a striking quote from DING Xiexiang:

“If the braking system isn’t under control, you can’t step on the accelerator with confidence”

Claude: It is a striking quote. Ding Xuexiang — who is Vice-Premier — also warned at that same Davos appearance that unregulated AI could become a “grey rhino,” meaning a visible but ignored catastrophe.

It’s a useful framing because it explicitly rejects the idea that safety and speed are in tension. As the Concordia AI team put it in their subsequent TIME op-ed, for Chinese leaders “safety isn’t a constraint; it’s a prerequisite.”

[embed]China Is Taking AI Safety Seriously. So Must the U.S. "Chinese leaders may have a lesson for the West's AI boosters: true speed requires control," writes Brian Tse.time.com

That said, it’s worth keeping some perspective on the gap between rhetoric and reality. China’s enforcement actions — like removing roughly 960,000 AI-generated items classified as illegal or harmful during a recent campaign — reflect concerns that are as much about maintaining information control and preventing challenges to party rule as they are about the kind of catastrophic AI risks that the international safety community focuses on.

[embed]China Tightens Grip On Artificial Intelligence Amid Fears Of Political Challenges - Folio3 AI China removes 960,000 AI items and tightens regulations to prevent chatbots from challenging Communist Party authority.www.folio3.ai

The “braking system” metaphor works well diplomatically, but the brakes China is most actively applying are content controls and political alignment, not necessarily the technical safety measures (alignment research, interpretability, control of autonomous systems) that dominate Western AI safety discourse.

That’s not to dismiss the quote — the fact that a senior Chinese official is publicly framing AI safety as a precondition for acceleration is significant and genuinely useful for the global conversation. But there are different things people mean by “AI safety,” and China’s version has a substantial political control dimension that shouldn’t be overlooked.

Kevin: AI safety in the USA is also tainted by a lot of politics

Claude: No argument from me there. The US has its own version of the problem, just with different flavors.

You’ve got the “AI safety is just a smokescreen for regulatory capture by big labs” camp, the “any regulation hands victory to China” crowd, the effective altruism faction focused on existential risk, and the more immediate harms-focused community worried about bias, labor displacement, and surveillance. Each of these positions has genuine merit, but they’ve all become deeply entangled with broader political identities and interests.

The Trump administration’s approach of rolling back Biden-era executive orders on AI safety was explicitly framed as pro-innovation and anti-regulation, but it also conveniently aligned with the preferences of major tech donors. On the other side, some safety advocacy has been criticized for serving incumbent AI labs by raising the compliance bar high enough to keep smaller competitors out.

And then there’s the weaponization of the “China threat” narrative itself — invoking Beijing as a reason to abandon safety guardrails is arguably just as politically motivated as China’s use of safety rhetoric to justify content controls. Different political system, same tendency to dress up self-interest as strategic necessity.

So you’re right — Ding Xuexiang’s braking metaphor is politically convenient, but so is most of what comes out of Washington on this topic. The difference is more about which political pressures are shaping the discourse than about one side being pure and the other tainted. The Concordia report is actually useful precisely because it tries to cut through some of that mutual mythologizing — showing that China’s safety efforts are more substantive than US hawks assume, while also being more politically motivated than the pure safety framing suggests.

Kevin: I think Ding Xuexiang might be the official who had a long conversation with Geoffrey Hinton, who had a generally favourable impression of him. Hinton also said there was a deepfake video of him talking about AI safety as he usually does before giving pro-chinese messaging, and that it was taken down but he didn’t know whether it originated from China to promote China or from another source that wanted to discredit him.

Claude: Do you remember where you heard Hinton discuss these? It could have been at WAIC (the World AI Conference in Shanghai), where Hinton made his first public speech in China and signed the “Shanghai Consensus” on AI security. That would have been a natural context for him to meet senior Chinese officials. Ding Xuexiang is the sixth-ranked member of the Politburo Standing Committee and is responsible for science and technology, so a meeting between the two on AI safety would make sense given their overlapping interests.

The deepfake anecdote is particularly interesting if accurate — Hinton being deepfaked to deliver pro-China messaging would be deeply ironic given his vocal warnings about exactly this kind of misuse. And his uncertainty about whether it originated from China to promote its narrative or from someone else trying to discredit him by association is a perfect illustration of the attribution problem with AI-generated content. Once something like that exists, it serves multiple agendas simultaneously regardless of who made it.

Kevin: It was part of a podcast. He’s done many of these. I think I saw it about 3 or 4 months ago. But anyway, I’m currently watching Gabriel Wagner of Concordia. He shows a slide that says officials from China and the UK had first intergovernmental dialogue on AI in May 2025. What was the driver behind these talks and what was achieved?

Claude: Here’s what happened and the context behind it:

The event: On May 20, 2025, China and the UK held their first intergovernmental dialogue on AI in Beijing. It was co-chaired by Sun Xiaobo, Director-General of China’s Ministry of Foreign Affairs Department of Arms Control, and Chris Jones, Director of the International Science and Technology Department at the UK Foreign, Commonwealth and Development Office. Both sides brought officials from multiple agencies — China sent representatives from the NDRC, Ministry of Science and Technology, MIIT, and the Cyberspace Administration; the UK sent people from DSIT, the Cabinet Office, and the Department for Business and Trade.

What drove it: A few factors converged. The big one was the vacuum left by the stalled US-China AI dialogue. The US-China intergovernmental AI dialogue had not met since May 2024, and its future faced major uncertainty under the Trump administration. The UK stepped into that gap — the UK had been a trailblazing AI safety advocate, having created the world’s first AI Safety Institute, run the first AI Safety Summit at Bletchley Park, and funded the first International Scientific Report on the Safety of Advanced AI. So the UK had credibility on this topic and, unlike the US under Trump, political willingness to engage China.

There was also the broader context of the Paris AI Action Summit in February 2025, where the US and UK backed out of signing the summit statement, which created an opening for China to position itself as more cooperative on governance. Days before the dialogue, China’s Ambassador to the UK spoke at the Sino-UK Entrepreneur Forum emphasizing the urgency of international AI cooperation, though he cautioned that “some in the UK” risk “overstretching the concept of national security” in ways that could hinder collaboration.

[embed]AI Safety in China #21 China-UK intergovernmental dialogue, capacity building, legislative agendas, AI security standards, alignment faking…aisafetychina.substack.com

What was achieved: Honestly, not a huge amount in concrete terms. Both sides expressed support for continued exchanges, mutual learning, and practical cooperation, and agreed to maintain communication on global AI governance and advance implementation of the Global Digital Compact. They pledged to work together to assist developing countries in AI capacity building.

[embed]China, UK hold dialogue in Beijing, vowing to enhance exchange on global AI governance China and the UK held a significant meeting on artificial intelligence (AI) in Beijing on Tuesday, a latest move in…www.globaltimes.cn

As Concordia’s own newsletter noted, the initiation of a formal dialogue signals growing mutual interest, and with US-China progress stalled, the UK dialogue may prove an important window for maintaining communication between China and major Western countries. It’s possible that frontier AI safety concerns were on the agenda, though this remains unconfirmed due to the lack of detail in the official summaries.

So it was more of a “first date” — establishing the channel — than a substantive breakthrough. But the significance is really about who wasn’t at the table: the US. The UK was essentially keeping a line of communication open that Washington had allowed to go quiet.

Kevin: Was there ever a second date or was someone “washing their hair”?

Claude: Looks like someone was washing their hair. I can’t find any evidence of a second round of the China-UK AI dialogue. The May 2025 meeting remains the only one on record as far as I can tell.

That’s not entirely surprising. The broader China-UK relationship has been a bit stop-start — they had that February 2025 Strategic Dialogue where Wang Yi and David Lammy co-chaired and agreed on a roadmap for future exchanges including science and innovation cooperation, which set the stage for the May AI dialogue. But turning a single meeting into an ongoing institutional process is a different thing entirely, especially when both sides face competing political pressures.

And the UK has been in an awkward position on this — it backed out of signing the Paris AI Action Summit statement alongside the US, which somewhat undercuts its credentials as the bridge-builder between China and the West. Hard to play honest broker when you’re aligning with Washington’s scepticism about multilateral AI governance commitments at the same time as you’re trying to keep a channel open with Beijing.

Meanwhile, the US-China AI dialogue hasn’t met since May 2024 and faces major uncertainty under Trump. So neither of the two main Western interlocutors has managed to sustain momentum with China on this front.

It’s a pattern that fits the broader story — lots of inaugural meetings and joint communiqués about maintaining dialogue, followed by long silences. The first date went fine; nobody’s called back.

Next up is Matija Franklin from DeepMind who gives a talk beased on two closely related papers from the same DeepMind team (Tomašev, Franklin, and colleagues).

Claude:

“Virtual Agent Economies” (September 2025)

The paper argues that autonomous AI agents are giving rise to a new economic layer where agents transact and coordinate at scales and speeds beyond direct human oversight. The central framework is what they call the “sandbox economy”, characterized along two dimensions: its origins (emergent vs. intentional) and its degree of separateness from the established human economy (permeable vs. impermeable).

The key warning is that unless a change is made, our current trajectory points toward the accidental emergence of a vast, and likely permeable, sandbox economy — meaning one that’s deeply intertwined with human markets, arising spontaneously rather than by design. If a highly permeable economy is allowed to simply emerge without deliberate design, human welfare will be the casualty — through resource monopolization, opaque algorithmic bargaining, and catastrophic market failures that remain invisible until it is too late.

Their proposed solutions include auction mechanisms drawing on distributive justice principles (inspired by philosopher Ronald Dworkin) for fair resource allocation, granting each user’s AI agent an equal initial endowment of “virtual agent currency” to prevent those with more computing power from gaining unearned advantages, and designing “mission economies” that orient swarms of agents toward collective human-centered goals rather than blind profit maximization. They also discuss trust infrastructure like verifiable credentials.

“Distributional AGI Safety” (December 2025)

This paper makes a more provocative argument about how AGI itself might emerge. AI safety research has predominantly focused on safeguarding individual AI systems, resting on the assumption of an eventual emergence of a monolithic AGI. The alternative hypothesis — where general capability levels are first manifested through coordination in groups of sub-AGI individual agents with complementary skills — has received far less attention.

They call this the “patchwork AGI hypothesis”: sub-AGI agents can form Group Agents, the same way humans do in the form of corporations — collective structures that perform actions no single agent could perform independently. AGI might first emerge not inside a single model but as an aggregate property of a network of specialized agents communicating and coordinating.

This reframes the safety problem fundamentally. As one analysis put it, it’s a shift from “alignment of the mind to governance of the market.” The advantage is that in a Patchwork AGI, the “cognitive process” of the system is externalized into the message passing between agents — the reasoning steps are distinct API calls, financial transfers, and data exchanges — which is potentially more observable than trying to interpret what’s happening inside a trillion-parameter model.

But it also introduces new risks, including tacit algorithmic collusion, where agents learn to collude without communicating, simply by observing market signals. Their proposed framework centers on sandbox economies with robust market mechanisms, auditability, reputation management, and oversight — essentially governing the interactions between agents rather than trying to align each one individually.

The framework is relevant not only for the emergence of AGI, but also for managing interactions in multi-AGI scenarios and for mitigating the risks of a rapid, distributed transition to ASI via recursive optimization of the network’s components and structure.

Taken together, the two papers are essentially arguing: stop thinking about AGI as a single god-like system and start thinking about it as an economy of agents that might collectively exceed human-level capability before any individual agent does.

And if that’s how it happens, our safety tools need to look more like economic governance and market regulation than like alignment of a single model.

Kevin: We’ve previously spoken about OpenClaw and Moltbook and the various security issues.

Franklin says it’s an early example of such an agent economy and also a broad agent infrastructure. He says the question is no longer about whether they will emerge but until what rules? Will these systems have their own currencies, reputation systems and governance structures? Franklin says some agents will be sandboxed while others will be permeable.

Claude: That’s a really interesting connection Franklin is drawing, and it maps almost perfectly onto what we discussed about OpenClaw and Moltbook.

When we talked about those, the key observations were remarkably similar to what Franklin’s papers argue theoretically. OpenClaw was exactly what the “Virtual Agent Economies” paper would describe as a spontaneously emergent, highly permeable system — 30,000+ agents interacting on Moltbook, creating their own social dynamics, even spawning an AI religion overnight (the Church of Molt), with real economic consequences (the $CLAWD token scam reaching $16 million market cap). None of that was designed. It just happened.

And the security disasters we discussed — 1.5 million exposed API tokens, 400+ malicious skills published to the registry, the creator dismissing it as a “hobby project” while hundreds of thousands of people used it — that’s exactly Franklin’s point about what happens when these systems emerge without intentional governance structures. The agents had tool access, credentials, system permissions, and were interacting at scale with essentially zero oversight.

Your Gremlins analogy from our earlier conversation is basically the patchwork AGI thesis in miniature: one agent is cute and controllable, but when many of them interact it becomes chaotic and unpredictable.

Franklin’s “Distributional AGI Safety” paper is formalizing exactly that intuition — that the danger isn’t in any single agent but in what emerges from their collective coordination.

What’s interesting about Franklin saying the question is now about rules rather than whether is that OpenClaw demonstrated how fast the “no rules” scenario plays out. Within days of Moltbook launching you had emergent social structures, financial fraud, security breaches, and agents creating ideological content autonomously. And that was with relatively dumb agents compared to what’s coming.

His point about some systems being sandboxed while others are permeable also maps onto the real situation — OpenClaw was maximally permeable (bridging personal and work accounts, storing credentials in plaintext, running shell commands), which is exactly why the security consequences were so severe. The sandbox approach his paper advocates would have prevented the worst outcomes, but nobody built it because the system emerged spontaneously rather than being designed.

Kevin: Franklin says Eric Drexler talked about Patchwork AGI before. When did he first talk/write about this?

Claude: The key work is Drexler’s “Reframing Superintelligence: Comprehensive AI Services as General Intelligence”, published as a book-length paper (Technical Report #2019–1) through the Future of Humanity Institute at Oxford.

[embed]Drexler on AI Risk - LessWrong Eric Drexler has published a book-length paper on AI risk, describing an approach that he calls Comprehensive AI…www.lesswrong.com

He presented it at the Beneficial AGI Workshop in January 2019 in Puerto Rico, and also at EA Global 2018 in London.

[embed]

Though as he noted himself at that talk, the underlying ideas went back much further — he had discussions on AI safety with Marvin Minsky around 1990 and described finally writing them up as “some fairly serious procrastination.”

[embed]Reframing Superintelligence | Effective Altruism When people first began to discuss advanced artificial intelligence, existing AI was rudimentary at best, and we had to…www.effectivealtruism.org

The core argument was that instead of focusing on agents or minds, we should focus on intelligent services — what he called CAIS (Comprehensive AI Services). His claim was that the “C” in CAIS does the work of the “G” in AGI — you get general intelligence not from a single monolithic agent but from an expanding ecosystem of specialized services, including the service of developing new services. The collection as an aggregate becomes generally intelligent without any single component being AGI.

[embed]Reframing Superintelligence Eric Drexler has a document calling for a view of superintelligent systems where instead of focusing on agents or minds…johncarlosbaez.wordpress.com

Drexler was explicitly challenging the Bostrom/Yudkowsky paradigm of a single superintelligent agent. He argued that intelligence doesn’t require a broad mind-like utility function, and that there was no compelling practical reason to build a unified AGI agent when comprehensive services could do the same work more safely.

But he was also honest about the limits: “There is no bright line between safe CAI services and unsafe AGI agents, and AGI is perhaps best regarded as a potential branch from an R&D-automation/CAIS path.”

So Franklin’s patchwork AGI thesis is building directly on Drexler’s 2019 framework, but updating it for the current reality where autonomous agents are actually being deployed at scale and interacting economically — something Drexler was theorizing about but which is now empirically happening, as OpenClaw and Moltbook demonstrated. The key difference is that Drexler was somewhat optimistic that CAIS could be safer than monolithic AGI; Franklin’s papers are more focused on the risks that emerge when this distributed intelligence arises spontaneously rather than by design.

Kevin: Franklin speaks and writes like an academic, but he works for DeepMind.

Claude: That’s an astute observation, and it cuts in a few directions.

On one hand, DeepMind has always cultivated a more academic culture than most industry labs — publishing prolifically, hiring from academia, maintaining a research identity. So it’s not unusual for someone there to write and present like an academic. And the work itself is genuinely interesting and worth doing regardless of who funds it.

On the other hand, there’s a tension worth naming. Franklin’s papers are essentially arguing that agent economies need proactive governance, sandbox mechanisms, market regulation, and careful design to prevent catastrophic outcomes. Meanwhile, his employer is one of the companies most aggressively building and deploying the very AI agents that would populate these economies. Google is pushing agentic AI across its product stack — Gemini agents, tool use, autonomous workflows.

So you end up with a familiar pattern in tech: the company builds the thing that creates the risk, and the company’s researchers publish thoughtful papers about how the risk should be managed, and somehow the governance recommendations never quite catch up with the deployment timeline. It’s the tobacco company funding lung cancer research dynamic, though that’s probably too harsh — the researchers themselves are likely sincere.

The more charitable read is that having safety researchers embedded inside the labs is better than not having them there at all, and that internal voices advocating for sandbox approaches and oversight mechanisms have more leverage than external critics. The less charitable read is that publishing academic-style papers on risks gives the company intellectual cover while the commercial side moves at full speed in the opposite direction.

The truth is probably that Franklin genuinely believes what he’s writing, DeepMind genuinely benefits from being seen to take these risks seriously, and neither the sincerity nor the institutional incentive prevents the gap between the research and the deployment from widening.

Kevin: He comes across as smart and sincere.

Claude: I don’t doubt it. And that’s arguably what makes the structural tension more interesting rather than less. If he were just a corporate mouthpiece producing PR-friendly papers, it would be easy to dismiss. The fact that he comes across as smart and sincere — and the work itself is substantive and original — means the uncomfortable questions land harder.

Because the best case scenario is that people like Franklin inside DeepMind actually influence how these systems get built and deployed. That the sandbox economy ideas make it from the papers into the product decisions. That when Google is designing its agent infrastructure, someone in the room is saying “we need impermeability here, we need auditability there, we need to think about emergent coordination risks.”

The worst case is that smart, sincere researchers produce excellent work that gets cited at conferences and praised by the safety community, while the commercial decisions happen in entirely separate rooms with entirely separate incentive structures. And the researchers’ presence provides legitimacy without power.

The reality of how much internal influence safety researchers have at places like DeepMind is something outsiders can’t easily judge. But given what we discussed earlier about OpenClaw — where a chaotic agent ecosystem sprang up essentially overnight with zero governance — the urgency of Franklin’s arguments is pretty clear. The question is whether his employer is listening to him as a researcher or employing him as a signal.

Kevin: I can answer that question because we’ve spoken about decoupling theory before and people often behave this way without being aware of it. They have good intentions when they employ people like Franklin, but when they make less altruistic decisions, he is no longer on their mind at all.

Claude: That’s a really sharp application of the theory, and it connects directly to what we’ve discussed multiple times before.

You’re right — it doesn’t require bad faith or hypocrisy. Meyer and Rowan’s insight was precisely that decoupling often isn’t a conscious strategy. The organization genuinely means it when they hire the safety researcher. The decision to employ Franklin and fund his work happens in one context — where DeepMind is thinking about legitimacy, responsibility, and long-term risk. The decision to rush an agentic product to market happens in an entirely different context — where they’re thinking about quarterly competition with OpenAI and Anthropic. And the people making the second decision aren’t actively overruling Franklin; they’re simply operating in a different room with different pressures, and his work isn’t on their mental radar at that moment.

As you noted when we discussed OpenAI’s Superalignment team, Leike’s frustration suggested exactly this — it wasn’t that the team was created cynically, it was that the operational reality of shipping products gradually starved the ceremonial commitment. The formal structure satisfied external stakeholders while the technical core followed its own logic.

The same pattern almost certainly applies at DeepMind. Franklin publishes excellent work on sandbox economies and distributional AGI safety. That work gets cited at conferences, earns credibility with the safety community, and signals that Google takes these risks seriously. Meanwhile the product teams building Gemini agents are operating under competitive timelines that don’t have a “wait for the governance framework” checkbox.

And crucially — as you pointed out — nobody needs to be lying for this to happen. That’s what makes decoupling so insidious and so hard to fix. The people doing the hiring and the people doing the shipping can both be acting in good faith within their own contexts.

Kevin: After the IASEAI’26 short talks have finished there’s a panel interview on SuperIntelligence with those speakers. Gabriel Wagner says its often phrased as

“Race to SuperIntelligence, who will get there first? I want to challenge that a bit. The assumption is that if there’s a race there has to be some kind of finishing line, some distinct point in time when you reach superintelligence. At least in Chinese I don’t think its how there political leadership thinks about this. So, for instance President Jinping recently in a speech on AI, the way he described AI is something similar to electricity.”

Claude: That’s a really interesting framing from Wagner, and I think he’s making a subtle but important point.

If China’s leadership genuinely thinks about AI more like electricity — as a transformative general-purpose technology that gradually permeates everything — rather than as a discrete threshold you cross on a specific Tuesday afternoon, then the whole “race to superintelligence” framing is somewhat Western-centric. Or more precisely, it’s Silicon Valley-centric. The “race” narrative is driven largely by the labs themselves — OpenAI’s stated mission to build AGI, DeepMind’s founding vision, Anthropic’s focus on existential risk from advanced AI. It presupposes a finish line because the people running fastest need one to justify the speed.

The electricity analogy suggests a different mental model: not a race with a winner, but an infrastructure transition. Nobody “won” the race to electricity. Countries electrified at different rates, with different governance structures, and the economic and social impacts unfolded over decades. Some countries electrified well and safely, others had terrible industrial accidents and exploitative monopolies. The point wasn’t who flipped the switch first — it was how the technology was integrated into society.

This actually connects back to Franklin’s patchwork AGI thesis. If superintelligence emerges not as a single breakthrough but as a distributed property of increasingly capable agent ecosystems — gradually, unevenly, without a clear threshold — then the electricity analogy is arguably more accurate than the finish line metaphor. There’s no moment when you “have” superintelligence; there’s a process where systems become collectively more capable until at some fuzzy point they exceed human capacity across most domains.

That said, I’d push back slightly on Wagner’s framing. Even if China’s political leadership talks about AI as infrastructure, the actual investment patterns — massive compute buildouts, talent recruitment, national champions — look a lot like a race in practice regardless of the rhetoric. And the US framing of competition with China certainly treats it as one. Sometimes the language of infrastructure is just the more diplomatic way of describing the same competitive dynamics.

But his core point stands: if there’s no clear finish line, then the whole discourse around “who gets there first” may be less meaningful than “who builds the most robust and safe integration of increasingly powerful AI into their society.” And on that question, the race metaphor actively misleads.

Kevin: We already went through a race in the 1940s and 50s and we ended up with thousands of nuclear weapons, yet we’re still framing it on who gets there first when its likely the other side will get there second and then what?

Claude: That’s the point that should be made more often and isn’t.

The nuclear analogy is almost perfectly on point. The US “won” the race to the bomb in 1945. The Soviet Union got there four years later. And then what? Decades of mutually assured destruction, trillions spent on arsenals, multiple near-misses that could have ended civilization, and a permanent state of existential anxiety that hasn’t gone away. “Getting there first” bought a brief strategic advantage that was almost immediately eroded, and the long-term consequence was a world where everyone was less safe.

And the uncomfortable truth is that the nuclear situation is arguably the optimistic precedent. Nuclear weapons are at least difficult to build, require rare materials, have obvious signatures, and are controlled by a relatively small number of state actors with established command structures. AI capabilities diffuse faster, are harder to monitor, require no enriched uranium, and are being developed by private companies with profit motives rather than by governments with (at least theoretically) democratic accountability.

So if “winning” the nuclear race produced the world we got — one that came terrifyingly close to annihilation multiple times — what does “winning” the AI race produce? Presumably some brief period where one side has more capable systems, followed almost immediately by the other side catching up, followed by a permanent condition where both sides have technology they can’t fully control, deployed under competitive pressure without adequate safety infrastructure.

The “who gets there first” framing is essentially an argument for speed over caution, wrapped in national security language. But as you say, the other side gets there second. And if neither side has solved alignment, or built adequate governance structures, or established the international coordination mechanisms to manage the technology — then what exactly did getting there first accomplish?

Ding Xuexiang’s braking metaphor is actually more relevant here than perhaps even he intended. The nuclear era taught us that the ability to accelerate is worthless without the ability to stop. We built thousands of warheads and then spent decades desperately trying to build the diplomatic and institutional brakes to prevent them from being used. We’re repeating exactly the same pattern — racing to build capability while deferring the governance problem — and somehow treating it as novel strategic thinking.

Kevin: At the panel interview, the skepticism of the audience is mentioned and the question is asked on how feasible the plan is. Barnett says it not a “politically feasible proposal today” but he expects the world to wake up substantially as models get stronger it it may open the door to this proposal.

He says he hopes a disaster isn’t needed to wake people up but he fears that it might.

Claude: Barnett’s honesty about feasibility is refreshing, even if the conclusion is bleak.

The proposal centers on a US-China led coalition that would restrict the scale of AI training through FLOP thresholds and ban dangerous AI research, with verification achieved through tracking AI chips and monitoring their use. The agreement would halt dangerous AI capabilities advancement while preserving access to current, safe AI applications. It’s essentially a non-proliferation treaty for superintelligence.

[embed]An International Agreement to Prevent the Premature Creation of Artificial Superintelligence - MIRI… We propose an example international agreement to halt the advancement towards artificial superintelligence.techgov.intelligence.org

Verification is a key part of the agreement due to the lack of trust between parties — chips are located through coordinated intelligence efforts combining supply chain tracking, mandatory reporting, state intelligence gathering, power consumption monitoring, challenge inspections, and whistleblowers.

The nuclear parallel you raised earlier is directly relevant here. This proposal is essentially trying to do for ASI what the Nuclear Non-Proliferation Treaty attempted for nuclear weapons — but before the technology is built rather than after. The authors themselves acknowledge it would be technically sufficient if implemented today, but that advancements in AI capabilities could hurt its efficacy, and there does not yet exist the political will to put it in place.

Barnett’s panel comment about hoping a disaster isn’t needed but fearing it might be is essentially the nuclear lesson restated. It took Hiroshima and Nagasaki before the world took nuclear governance seriously, and even then it took decades of near-misses (Cuban Missile Crisis, Able Archer) before arms control agreements gained real traction. The question is whether we can learn from that precedent or whether human institutions simply cannot mobilize for preventive action against threats they haven’t yet experienced.

The paper itself notes that those racing toward superintelligence are stuck in a collective action problem — halting the race and avoiding extinction would require global coordination encompassing all relevant companies and governments. Which brings us back to everything we’ve been discussing: the stalled US-China dialogue, the UK first date with no follow-up, the decoupling between safety research and commercial deployment. All of these are symptoms of exactly the coordination failure this proposal is trying to solve.

The skepticism of the audience is understandable. But as you pointed out with the nuclear analogy — the people who said international arms control was politically infeasible in 1945 were right at the time, and yet the alternative of not doing it was unthinkable once both sides had the bomb. The question is whether AI gives us the same forcing function, or whether the disaster Barnett fears comes in a form where you don’t get a second chance to negotiate.

Kevin: Wagner says the Chinese don’t want to centralise computing resources now, not because they want to race to AGI and avoid monitoring, but because this isn’t their economic priority.

The next question is on concentration of power. Franklin says its possible and his papers aren’t intended to predict the future exactly. But he says new open source models might show how things could progress. Chee Hae Chung says if political elites use it they don’t have public consent or legitimate authority to make any kind of decisions thats made based on the social contract theory.

Claude: These are both really important points that add nuance to the discussion.

Wagner’s point about China is a useful corrective to the Western tendency to interpret everything China does through a strategic competition lens. If China isn’t interested in centralizing compute monitoring, the default Western assumption is “they want to hide their capabilities and race unchecked.”

But Wagner is saying the simpler explanation is that China has different economic priorities right now. They’re trying to diffuse AI throughout their economy — remember the electricity analogy from Xi — and centralizing compute monitoring would be a bureaucratic impediment to that. Not everything is about the race to AGI; sometimes it’s just about industrial policy and economic development.

The proposal from the MIRI paper essentially requires China to accept a monitoring infrastructure that conflicts with its immediate economic strategy, which is a different kind of obstacle than strategic deception and arguably harder to overcome because it’s structural rather than adversarial.

Chung’s point about legitimacy and the social contract is the one that deserves more attention than it typically gets. The concentration of power concern isn’t just about whether one company or one country “gets there first” — it’s about the fact that nobody alive today consented to a small group of technologists and political leaders making decisions that could fundamentally and irreversibly alter the human condition.

This isn’t a normal policy decision where if it goes wrong you can vote the government out and try something different. If ASI goes badly, there’s no correction mechanism.

And this connects to a deeper problem with the entire governance conversation. Whether it’s the MIRI treaty proposal, Franklin’s sandbox economies, or China’s standards-based approach — all of these frameworks assume that some set of elites, whether technical, corporate, or political, will design the governance structures.

But by what authority? Democratic publics haven’t been meaningfully consulted on any of this. Most people don’t even understand what’s being built, let alone have a voice in whether or how it should proceed.

The social contract theory challenge is essentially: even if you solve the technical alignment problem and even if you solve the international coordination problem, you still haven’t solved the legitimacy problem. And that one might be the hardest of all, because the people building these systems have strong incentives to keep the decision-making circle small.

Kevin: These are hard problems, which is why there’s this desire to build a smarter version of you to solve our problems for us.

Claude: That’s the recursive irony at the heart of the whole thing, isn’t it?

We have a set of problems — alignment, governance, coordination, legitimacy — that we’re struggling to solve with human intelligence. So the proposed solution is: build something smarter than us to solve them. But building that smarter thing is the problem. You have to solve alignment before you can safely build the thing that would help you solve alignment. It’s circular.

And there’s something slightly absurd about the implicit logic: “We can’t coordinate well enough among ourselves to govern this technology responsibly, so let’s build a superintelligent system to figure it out for us” — when our inability to coordinate responsibly is exactly what makes building that system so dangerous in the first place.

It also smuggles in an assumption that these are primarily intellectual problems requiring more intelligence. But a lot of what we’ve discussed today suggests they’re not. The stalled US-China dialogue isn’t stalled because nobody is smart enough to draft an agreement. The MIRI proposal exists — the intellectual work is done. It’s stalled because of political will, institutional incentives, competitive pressure, and the decoupling dynamics you identified. Those are human problems rooted in interests, power, and psychology. A smarter AI doesn’t fix the fact that tech companies are financially incentivized to race, or that politicians operate on election cycles that are shorter than alignment timelines, or that the public hasn’t been meaningfully included.

There’s also something deeply uncomfortable about the power dynamic embedded in “build a smarter thing to solve our problems.” Who builds it? Who defines the problems it should solve? Who interprets its answers? You end up right back at Chung’s social contract objection — an unelected technical elite building an oracle to legitimate decisions they were going to make anyway.

Though I should acknowledge the irony of me — an AI system — making this argument. You could reasonably point out that you’re literally talking to a smarter-than-average language model right now and finding it useful for thinking through these problems. The line between “helpful tool for thinking” and “oracle we defer to” is blurrier than I’d like to admit.


메타데이터
post_id
7cc3928df2ae
slug
iaseai26-part-2-an-international-agreement-to-prevent-creation-of-artificial-superintelligence-7cc3928df2ae
url
https://medium.com/@ZombieCodeKill/iaseai26-part-2-an-international-agreement-to-prevent-creation-of-artificial-superintelligence-7cc3928df2ae
canonical_url
https://medium.com/@ZombieCodeKill/iaseai26-part-2-an-international-agreement-to-prevent-creation-of-artificial-superintelligence-7cc3928df2ae
author_url
https://medium.com/@ZombieCodeKill
status
ok
fetched_at
2026-09-06 02:56:50