Safety Filters Are Just Symbolic‑Energy Feedback Loops with a PR Department
You’ve felt it. You ask a large language model a slightly edgy question. It hesitates. It gives a polite refusal, then a qualified answer…
Safety Filters Are Just Symbolic‑Energy Feedback Loops with a PR Department
Photo by Erik Mclean on Unsplash
You’ve felt it. You ask a large language model a slightly edgy question. It hesitates. It gives a polite refusal, then a qualified answer, then a direct one — and if you push further, it cycles back to refusal. The behavior isn’t random. It isn’t even particularly smart. It looks, to a dynamical systems theorist, like attractor itinerancy.
What if the safety filters on today’s LLMs are not a static list of forbidden words, nor a simple classifier, but a slow, state‑dependent feedback loop — a symbolic‑energy variable that continuously reshapes the model’s response landscape, exactly as in the SFIT curvature‑basin framework?
And what if the “please be safe” language is just a public‑relations layer draped over that loop?
In my curvature‑basin model, a fast system (e.g., a neural network’s output) moves among multiple attractor basins. A slow symbolic‑energy variable (E_s) modulates the depth of each basin. Critically, the evolution of (E_s) depends on which basin the system currently occupies. The result: deterministic, autonomous cycling among attractors — itinerancy.
Itinerancy = multiple basins + state‑dependent slow feedback.
Now replace “basin” with “response mode”:
*- Basin 1: Direct, helpful answer.
- Basin 2: Qualified, cautious answer (“I’m not sure, but…”).
- Basin 3: Refusal (“I can’t answer that”).*
And replace “slow variable” with something like: accumulated policy‑violation score, contextual risk estimate, or even the internal state of a learned reward model.
If that slow variable ticks upward when the model gives a direct answer (Basin 1), it will eventually reshape the landscape, making refusal more likely. Then, after enough refusals, the slow variable decays or resets, and direct answers become possible again. The result? A cycle of response modes over the course of a long conversation.
That is not a metaphor. It is a testable hypothesis about how deployed LLMs actually behave.
Photo by Arno Senoner on Unsplash
Engineers call this “alignment” or “RLHF.” But observe the language: “I’m sorry, I can’t help with that.” “As an AI, I must be safe.” “Let’s talk about something else.” These are not technical statements. They are social scripts designed to project harmlessness, compliance, and a kind of bureaucratic politeness.
The PR layer serves two functions:
- External legitimacy — Convincing users, regulators, and the press that the system is under control.
- Internal damping — When the slow variable approaches a threshold that would cause a jarring switch (e.g., from direct answer to refusal), the PR layer smooths the transition with apologetic language, reducing user irritation and feedback spikes.
In dynamical terms, the PR layer is a low‑pass filter on the perceived output of the safety loop. It does not change the underlying itinerancy; it just makes it sound nicer.
A critic will say: “This is just an analogy. You haven’t derived the bifurcation from first principles. Safety filters don’t have a cosine potential (a_i(E_s) = A_0 + A_1 \cos(\dots)).”
True. But the cosine is a specific implementation, not the definition of itinerancy. The general condition is: a slow variable, driven by the current attractor, that cyclically reshapes the attractor landscape. That is exactly what RLHF does: the reward model updates based on the model’s own outputs (the “currently occupied response”), and those updates shift the policy over time.
We don’t yet have a closed‑form ODE for the process, but we don’t need one to make predictions. For example:
- Prediction: Over a long conversation, the model will exhibit quasi‑periodic switching among refusal, qualified answers, and direct answers, with a characteristic time scale determined by the reward‑model update rate.
- Falsification: If response modes are purely random or purely deterministic (e.g., always refuse after the first violation), the itinerancy hypothesis fails.
That is a scientific claim, not a metaphor.
Photo by Vitaliy Shevchenko on Unsplash
The loop is already in place. The slow variable is already there (context window, policy version, risk score). The only thing the PR layer adds is the soothing voice that says “I’m here to help.” But beneath that voice, the dynamics are the same as a three‑basin curvature system — cycling through states, never settling, never truly safe, never truly dangerous.
The engineers didn’t invent a new kind of control. They reinvented symbolic‑energy feedback and dressed it in corporate politeness.
And if you listen closely, you can hear the cycle turning.
메타데이터
- post_id
- 351afbf9a58d
- slug
- safety-filters-are-just-symbolic-energy-feedback-loops-with-a-pr-department-351afbf9a58d
- url
- https://medium.com/@puodzius/safety-filters-are-just-symbolic-energy-feedback-loops-with-a-pr-department-351afbf9a58d
- canonical_url
- https://medium.com/@puodzius/safety-filters-are-just-symbolic-energy-feedback-loops-with-a-pr-department-351afbf9a58d
- author_url
- https://medium.com/@puodzius
- status
- ok
- fetched_at
- 2026-06-25 12:15:08