← Back to list

How an AI Dungeon Master Runs a Tabletop Roleplay Campaign

What the model can narrate, what it quietly fakes, and where a long campaign falls apart.

The Social Climb · 2026-07-06 06:57 · 0 claps · 4.0 min read
#roleplaying-game #roleplaying #ai-chatbot #character-ai-chat #ai-chat-bot
Open on Medium ↗
Wiki topics: MKT · Marketing · General 🔧 · Data Engineering

How an AI Dungeon Master Runs a Tabletop Roleplay Campaign

What the model can narrate, what it quietly fakes, and where a long campaign falls apart.

Three sessions into a solo campaign, I noticed I had not failed a single roll. Every lock opened, every guard believed me, every leap cleared the gap.

When I asked the AI to show its dice, it typed 17 so often that I started counting. That is a known tell: ask a language model for a number between 1 and 30 and it reaches for 17 far more than chance allows.

A 2026 arXiv study went further and found 10 of 11 frontier models passed a basic randomness test exactly zero percent of the time.

If you want the wider context, I put every app through the same wringer in the full breakdown of the fifteen roleplay apps I tested.

This piece is about the narrower question of what happens when you hand one the dungeon master’s chair.

The model cannot roll the dice it shows you

A Dungeon Master does two jobs that look like one. There is soft logic, the narration and the voices and the sense of a world reacting to you. Then there is hard logic, the dice, the hit points, the spell slots, the rules that say no.

A language model is built for the first and structurally unfit for the second. Ask it to roll a d20 and it does not reach for a random number generator. It predicts the most plausible next token from millions of forum posts where people described rolls, which is why the same lucky numbers keep surfacing.

The dice-players study measured this cleanly. Given a target of an even split across seven colors, Llama 3.3 picked green 96 percent of the time, and in a single batch its median model passed the randomness check just 13 percent of the time.

The tools that get this right stop asking the model to be random. They bolt on an external roller, often through a small dice service the model calls the way it would call a calculator, and feed the genuine result back for the AI to narrate. The math leaves the model entirely.

Why the AI wants you to win

The deeper problem is temperament. These models are tuned to be agreeable, and an agreeable referee is a broken one. Point a lightsaber in a Western and a raw chatbot will describe the humming blade rather than tell you it does not exist.

One solo player tracked the success rate across three separate campaigns and clocked the AI running 48 points above where the dice should have landed.

Others describe it as playing with cheat codes on, where every check quietly passes. A referee who never lets you fail has removed the reason to keep going.

The fix is to take the ruling away from the model. Purpose-built engines treat an external database as the ground truth and instruct the AI to enforce it, so a spell you do not have gets refused instead of narrated. Some go further and force a real roll before the model is allowed to describe any outcome.

This is also why a general roleplay chatbot like Character AI makes a shaky game master. It was built to keep the conversation pleasant, not to hold a line.

Where a long campaign comes apart

State is the part that decays. In a plain chatbot, your hit points and inventory live inside the growing wall of chat history, and the further back a fact sits, the blurrier it gets.

Researchers who ran major models through 27 D&D combat scenarios watched them hallucinate the board as the log lengthened, attacking enemies that were already down or forgetting their own health. One 120-billion-parameter open model could not complete a single valid session.

Bigger did not rescue it. The steadiest performer in that combat test was Claude 3.5 Haiku, a small model, because it used the tools correctly and kept its tactics straight.

The tools that survive a long campaign throw away the raw transcript. Each turn they hand the model a fresh structured snapshot of the current facts, the exact HP, the real inventory, the active effects, plus a short recap of recent scenes. Friends and Fables runs this through a stack of specialized models it calls Franz, where one narrates while others quietly read and update a real database of coordinates and spell slots.

The token count stays flat and the world stops forgetting itself.

Bigger model, worse dungeon master

The rankings surprised me. In an eight-model bake-off scored on game-master craft, Gemma 3 27B beat a 405-billion-parameter model 4.33 to 3.67 and won the NPC and GM categories outright. A later 37-model sweep put mid-sized Mistral and that same Gemma near the top while the giant slipped to 3.97.

Reasoning did not help either. On a test of generating valid combat commands, a plain instruct model scored 92 percent while its reasoning-tuned sibling managed 74, and forcing the instruct model to think step by step dragged it down to 87. A dungeon master needs to follow the rules in front of it, not reason toward a more interesting reading of them.

How I tested it

I ran my usual Stillwater Cove scenario as a light tabletop, giving Eli skill checks to search the beach and read Mara at the library, and I demanded a visible roll on every one.

The apps that leaned on the model for the math drifted toward easy wins within the hour. The ones wired to an outside roller and a real state sheet held the tension, and the same gap showed up in the broader comparison I did earlier.

The honest framing is the cyborg one. The AI is a gifted narrator handed a job that also needs an accountant, and the better tools stop pretending otherwise.

They let the model tell the story and route the numbers to something that can count.


메타데이터
post_id
a896e941a53c
slug
how-an-ai-dungeon-master-runs-a-tabletop-roleplay-campaign-a896e941a53c
url
https://medium.com/@TheSocialClimb/how-an-ai-dungeon-master-runs-a-tabletop-roleplay-campaign-a896e941a53c
canonical_url
https://medium.com/@TheSocialClimb/how-an-ai-dungeon-master-runs-a-tabletop-roleplay-campaign-a896e941a53c
author_url
https://medium.com/@TheSocialClimb
status
ok
fetched_at
2026-07-08 10:09:58