← Back to list

The 4 Tests That Could Show AI Models Have Beliefs

There is a lot of debate today about whether AIs truly have beliefs or not. Some argue that they’re just stochastic “parrots” that have…

Strad Slater · 2026-02-12 05:50 · 0 claps · 4.6 min read
#ai-safety #ai-interpretability #philosophy #ai-belief
Open on Medium ↗
Wiki topics: SAF · Safety & Alignment PHI · Philosophy

The 4 Tests That Could Show AI Models Have Beliefs

There is a lot of debate today about whether AIs truly have beliefs or not. Some argue that they’re just stochastic “parrots” that have gotten really good at outputting legible text without any actual internal understanding of what it’s saying. Others think that models are developing internal beliefs about the world that help aid in its ability to produce accurate and useful outputs.

As of now, the debate is still largely unresolved. One of the most useful techniques for detecting beliefs in AI is interpretability which aims to decode the internal activations in a model and map them to interpretable features and concepts that humans can understand. A rising direction in this field is the search for a truth/belief feature in a model that can be directly tied to some internal activations.

However, the field lacks a unanimously agreed upon standard for what should count as a belief with respect to LLMs. It’s hard to find something inside a model when people aren’t exactly sure what they should be looking for.

Press enter or click to view image in full size

In the paper, “Standards for Belief Representations in LLMs”, a standard for LLM-based beliefs is proposed in order to help better guide interpretability researchers in their search for belief representation in LLMs. The paper divides the standards into 4 criteria; accuracy, coherence, uniformity and use, based on philosophical and practical motivations.

The Motivations for the 4 Criteria

The researchers wanted the standards for LLM-based belief to match closely to the philosophical concept of belief in humans, because if it didn’t then calling these things beliefs might not make sense.

However, because models are structurally different from humans and we rely more on their internal processes rather than observable behavior to assess beliefs (which is the opposite of how we evaluate human beliefs), the criteria used for models may differ slightly from those used for humans.

One motivation is that, the representation of LLM-based belief should be action-guiding. These beliefs should be somewhat utilized in the actions an LLM makes. This motivation heavily contributes to the standard accounts of human belief as well.

The researchers are also motivated by the need for these representations of belief to explain why LLMs are successful. This is useful as it helps settle the debate above about whether model’s are successful due to belief formation or some other means such as stochastic processes.

Finally, the representations must be measurable and interpretable. This comes from the motivation that the representations be practical which in part means useful to humans. It’s possible that LLMs might hold some type of belief representation that is completely uninterpretable to humans. The researchers argue that these cases should not be treated as beliefs in the same way we use them for humans as they hold no practical use to us.

The Four Criteria for Belief Representation in LLMs

Accuracy

A model’s beliefs must be accurate with respect to it’s training data where accuracy is expected. Given that these beliefs should help explain the success of a model, it makes sense that these beliefs be accurate in cases where accurate beliefs would help a model achieve its goals.

Having accurate beliefs is very useful for the discovery of beliefs within LLMs since the detection of beliefs will be easiest on datasets with clear true or false statements that we already know the ground truth on.

However, this does point to a limitation of accuracy as the sole criteria in that it’s hard to use it for statements where a clear ground truth has not been established such as “humans have free will” and “most people won’t benefit from taking multi-vitamins.” This makes it difficult to generalize the use of accuracy as a criteria to beliefs outside a model’s dataset.

Coherence

The beliefs of a model should be logically consistent. For example, a belief in one statement should mean the model does not believe in it’s negation. Or if a model believes in A and B separately, than it should also believe the statement A and B. Beliefs should be consistent with all these different logical variations.

This consistency also applies to semantic variations of the belief. For example, if a model believes “this cup is to the right of this dog,” then it should also believe “this dog is to the left of this cup.” These are semantically different sentences, yet the rational belief in one automatically determines the belief in the other.

This criteria also stems from the motivation of having a model explain it’s successes, since coherence of belief is useful for a model to say accurate things about the world. It’s also a very practical criteria as knowing one coherent belief allows us to infer all the other logical variations of the belief such as its negation.

Uniformity

A model’s representation of belief must be uniform across various domains. It’s not enough for a model’s representation of belief to be accurate and coherent if it only applies to a very narrow set of facts and domains. Oftentimes, when this is the case, the representation might not even be representing belief but rather something more specific to the domain that correlates with truth.

Uniformity is a very practical criteria as it allows the representation of belief to generalize to other domains. This means we could use it to find beliefs about other domains that the model hasn’t even trained on.

Use

A model must use their beliefs as a way to actually form their final response. A belief is useless for us to know if it’s completely useless to the model. We can’t make reliably predictions about what a model will do based on an idle belief that it never uses.

The researchers explain how “use” can and has been determined by simply taking a section of the model one predicts is associated with truth and altering it to see if it reliably changes the output of the model. If so, then it is very likely that that section is used for producing the final output.

Takeaway

While the field of AI interpretability will be heavily reliant on technical work, it’s important to remember that the human-like nature of this technology makes philosophy a very applicable field as well. This paper acts a great example to how theoretical work on philosophical concepts such as the nature of belief can be practically useful for technical work in AI.

Hopefully this standard of belief presented by the researchers in the paper can help those actually trying to detect beliefs do so in a structured way. It would be interesting to see if any standards exists of other types of propositional attitudes in LLMs such as desires and goals. This would be very useful for the development of “Thought Logging” in LLMs, a tool that would log propositional attitudes such as beliefs and desires from LLMs in real time. You can read more about this topic in an article I wrote here.


메타데이터
post_id
1bfb62e22feb
slug
the-4-tests-that-could-show-ai-models-have-beliefs-1bfb62e22feb
url
https://medium.com/@stradslater/the-4-tests-that-could-show-ai-models-have-beliefs-1bfb62e22feb
canonical_url
https://medium.com/@stradslater/the-4-tests-that-could-show-ai-models-have-beliefs-1bfb62e22feb
author_url
https://medium.com/@stradslater
status
ok
fetched_at
2026-07-20 23:38:49