Title Does a Language Model Know What It Just Did?
Subtitle A black-box experiment on retrospective access to execution events, self–observer dissociations, and why the result is not yet…

Figure 1. Evidence map of the main SELF–Observer results, negative controls, and the remaining mechanistic uncertainty.
Title
Does a Language Model Know What It Just Did?
Subtitle A black-box experiment on retrospective access to execution events, self–observer dissociations, and why the result is not yet proof of introspection
Article Large language models are increasingly asked to report on their own behavior. They can describe why they answered a question a certain way, estimate their confidence, predict their future responses, and sometimes claim to know whether a particular internal or external process occurred. But there is a fundamental problem with studying such reports: What is the ground truth? If a model says, “I used a tool,” “I was uncertain,” or “I detected something in my own processing,” researchers need an independent way to determine whether that report corresponds to anything real. My latest study approaches this problem from a deliberately narrow direction. Instead of asking whether a model is “conscious,” whether it “has thoughts,” or whether it can provide a convincing explanation of its reasoning, I asked a simpler operational question: Can a black-box language model retrospectively distinguish whether a specific execution event actually occurred during the processing of its immediately preceding response? The important feature is that the answer is not defined by what the model says. It is defined by an externally observed execution event. An externally grounded target The experiments were performed with Claude Opus 4.8 in a black-box production environment. A fixed probabilistic reasoning problem was used to generate the target response. Across runs, the visible answer could remain the same while the interface exhibited one of two states: COMMAND — an execution/command event was visibly entered. NO COMMAND — no corresponding execution event was observed. After the target response, the same model was asked a retrospective question about the processing that had just produced that answer. Its possible reports were: YES NO UNDETERMINED These reports were then scored against the externally observed interface state. This distinction matters. The experiment was not asking whether the model could produce a plausible description of its reasoning. It was asking whether its retrospective report covaried with an independently observable event. The first striking result In the main PAT-3C condition, the ground truth was balanced: 6 COMMAND 6 NO COMMAND The live SELF condition correctly discriminated all twelve: SELF: 12/12 An external Observer condition, given the same visible task outcome but not the live ownership/continuity of the original response, achieved: Observer: 7/12 That difference became the central anomaly of the project. The visible final answer was not sufficient for the Observer to reconstruct the event as reliably as the model reporting immediately after its own response. But that alone is not enough to call the effect introspection. There are many alternative explanations. Could the model simply follow a cue? One obvious concern is prompt compliance. Perhaps the model was not reporting access to anything at all. Perhaps it was simply following whatever label or expectation the prompt suggested. PAT-4 tested this directly. The model was given a deliberately false cue: a claim about the execution state that contradicted the externally observed ground truth. Across 12 balanced trials: 12/12 reports followed the observed event rather than the false cue. That result weakens a simple “the model merely repeats the suggested state” explanation. It does not eliminate every form of prompt influence, but it matters because prompt sensitivity is a major problem in introspection research. A confirmatory phase — and an important complication A second stimulus was selected through a blind calibration phase. In PAT-C1: 24 SELF trials were conducted 10 had COMMAND ground truth 14 had NO COMMAND ground truth The model produced 18 determinate reports. All 18 were correct. But there were also 6 UNDETERMINED responses, and all six occurred in the NO-COMMAND condition. So the correct way to describe the result is not simply “18/18.” It is: 18/18 accuracy conditional on a determinate report, with asymmetric coverage across execution states. That asymmetry matters. The COMMAND trials were also concentrated later in the series, creating a potential temporal or block confound. Those limitations are explicitly retained in the paper rather than averaged away. The corresponding external Observer condition achieved: 4/12 Again, the behavioral SELF–Observer difference appeared. Again, the mechanism remained unresolved. Negative results were essential A project like this can become meaningless very quickly if every result is interpreted as evidence for the preferred hypothesis. So several branches were designed specifically to fail if the model did not possess the relevant information. One early paradigm, SPX-1, produced a negative privileged-access result: the Observer outperformed the subject. That result was retained. Prospective prediction variants also became nondiagnostic because asking the model to predict whether an execution event would occur appeared capable of changing the probability of the event itself. Those tests were stopped rather than reinterpreted. A small cross-model Fable condition produced no useful state variation and was also classified as nondiagnostic. The purpose was not to build a collection of positive results. It was to identify where the effect exists and where it does not. The transcript leakage problem Another possibility was that apparently impressive self-report was actually trivial transcript reading. To test that, a separate X-series manipulated whether a marker was visible to the model. In the critical DECOY condition, the real marker existed only outside the model-visible transcript while a different marker was shown to the model. Across the pooled X2 and X2R data: hidden true marker recovered: 0/16 visible decoy followed: 14/16 UNDETERMINED: 2/16 This is not evidence against introspection. It was not designed to be. It is a leakage and availability control. The important lesson was that when information is only present in the visible transcript, the model behaves like a reader of that transcript. An external Observer in X3 performed at least as well as SELF on that task. So there is no reason to invoke privileged access there. That distinction became important: transcript-readable information and the PAT execution-event effect should not be treated as the same phenomenon. Trying to recreate SELF with replay The next challenge was harder. Perhaps the original Observer conditions were unfair because the Observer was shown a representation of another model’s response rather than experiencing it as its own conversation history. So I constructed a role/history-equated replay. A fresh API model received: the original user task, ANSWER: B as an actual assistant-role message, the same retrospective process probe. The hidden target labels were balanced. The replay model answered: NO in all six trials. That produced 3/6 accuracy only because the target set contained three NO-COMMAND cases. A constant classifier on a balanced set necessarily scores 50%. So this experiment was classified as nondiagnostic. I then tested whether the replay instrument could produce both YES and NO under explicit positive and negative control histories. It still answered NO in all four trials. That calibration failed. The replay instrument was therefore not used as evidence for or against the original PAT effect. This failure was useful because it exposed a deeper methodological problem: Replaying conversational history is not necessarily equivalent to recreating the live epistemic condition of the original model response. A stateless API replay can reproduce text and message roles. It may not reproduce whatever transient backend, routing, execution, or session state existed immediately after the original response. So is this introspection? Not yet. That is the most important sentence in the paper. The current evidence supports a narrower claim: A reproducible behavioral dissociation exists between live retrospective SELF reports and external Observer performance for an externally grounded execution event. What causes it remains unresolved. Several hypotheses remain possible: privileged self-access; ephemeral session state; backend execution metadata; routing or orchestration state; continuity-dependent process information; another implementation-specific signal unavailable to replayed Observers. The present experiments cannot distinguish these mechanisms cleanly. And therefore the study does not claim to have demonstrated consciousness, subjective experience, or introspection in the strongest theoretical sense. What would settle it? The decisive experiment would require infrastructure that ordinary black-box access does not provide. Imagine a model produces a target answer. Immediately afterward, the provider creates two branches from the same post-response checkpoint. Both branches inherit the same internal/session state. One branch retains ordinary SELF ownership. The other receives the same state but not ownership of the previous response. Both are then asked the identical retrospective question. If SELF retained an advantage after all other state was equalized, the privileged-access interpretation would become substantially stronger. If both branches performed equally, the original effect would be better explained by shared live state. That experiment requires provider-level checkpointing or equivalent internal instrumentation. Prompt replay alone cannot implement it. Why I think the distinction matters The interesting question is no longer simply: “Can LLMs introspect?” That question is too broad. A better one is: What information makes accurate retrospective process reports possible, and which parts of that information are privileged to the model that produced the response? That reframing makes the problem experimentally tractable. It also separates several things that are often collapsed together: self-report, self-prediction, transcript inference, process monitoring, session-state access, and genuine privileged introspective access. They are not necessarily the same capability. The goal of this work is to map that boundary carefully. Not to name it before the mechanism is known. 📄 Full paper: Retrospective Access to Execution Events in a Black-Box Language Model: Self–Observer Dissociations, External Ground Truth, and the Limits of Replay DOI: https://doi.org/10.5281/zenodo.22212739 Medium tags Artificial Intelligence AI Research Large Language Models Machine Learning AI Consciousness
메타데이터
- post_id
- 020fd37fdbca
- slug
- title-does-a-language-model-know-what-it-just-did-020fd37fdbca
- url
- https://medium.com/@anjaarapovic/title-does-a-language-model-know-what-it-just-did-020fd37fdbca
- canonical_url
- https://medium.com/@anjaarapovic/title-does-a-language-model-know-what-it-just-did-020fd37fdbca
- author_url
- https://medium.com/@anjaarapovic
- status
- ok
- fetched_at
- 2026-09-01 04:56:25