← Back to list

The System That Refuses to Stay Itself

Why Long-Lived Agentic Systems Need Identity Preservation

Doron Chema, PhD · 2026-06-05 13:36 · 0 claps · 9.7 min read
#ai #ai-agent #ai-safety #ai-security #ai-risk
Open on Medium ↗
Wiki topics: AGT · AI Agents SAF · Safety & Alignment AI · AI · General

The System That Refuses to Stay Itself

Why Long-Lived Agentic Systems Need Identity Preservation

Introduction

AI Safety has a dirty secret.

Most safety frameworks treat alignment as a problem of the moment.

Constitutional AI constrains systems at birth[1]. RLHF shapes preferences before deployment. Interpretability examines individual decisions. Runtime monitoring evaluates behavioral outputs [2].

These are serious answers to serious questions.

But they share a common assumption.

They focus on what a system is doing, and far less about what a system is becoming.

The first paper in this series argued that persistent Agentic Runtimes exhibit evolutionary dynamics — variation, selection, and retention that allow operational structures to accumulate across time.

As a result, future behavior becomes increasingly shaped by developmental history rather than current state alone.

A system can remain aligned while gradually becoming something else.

Its model, specification, and directives may all remain unchanged.

Yet its evaluative criteria, authority relationships, delegation structures, and operational boundaries may have evolved substantially through accumulated adaptation.

While the field has spent a decade studying alignment at a moment, it has only begun to ask what alignment means across a trajectory.

If trajectory matters, a second question immediately follows:

What exactly is the thing that trajectory can erode?

That question is the subject of this paper.

1. The Identity Preservation Problem

Imagine meeting a close friend after twenty years.

Their appearance has changed. Their habits have changed. Their beliefs, relationships, and daily routines may all be different from what they once were.

Yet despite those changes, you would probably still describe them as the same person.

Now imagine the opposite.

Someone looks the same, speaks the same, and occupies the same role in your life, but their values, motivations, and sense of purpose have become unrecognizable.

In that case, we often say: “They’re not the same person anymore.”

What changed?

Not every component. Most of the visible characteristics remained intact.

What changed was something deeper — something connected to identity rather than appearance or behavior.

The same puzzle appears in long-lived adaptive systems.

Organizations replace employees, products, and technologies. Agentic systems revise memories, adopt new tools, reorganize workflows, and delegate responsibilities.

Change is not the exception, but the mechanism through which adaptive systems survive.

Yet adaptation creates a problem.

The more freedom a system has to change, the harder it becomes to explain what makes it the same system across time.

Most changes appear harmless in isolation. A workflow is redesigned. A tool is replaced. A responsibility is delegated.

None seems significant enough to redefine the system.

Yet long-lived systems are shaped by thousands of such decisions. Over time, the accumulation of individually reasonable changes can produce something that would have been difficult to recognize at the beginning of the journey.

A system may remain successful while gradually becoming something else. This is the identity preservation problem.

The challenge is not preventing change, but of understanding which changes matter.

Because if trajectory can reshape a system, identity cannot reside in any particular behavior, workflow, memory, or component.

It must reside in something more persistent.

2. Operational Systems and Protective Systems

If identity can survive decades of change, then identity preservation cannot be an accident.

Something must be doing the work.

Consider a biological organism.

Most of its activity is directed toward survival and adaptation. It acquires energy, responds to its environment, reproduces, and continuously adjusts to changing conditions.

Yet alongside these activities exists a different class of mechanisms.

Cells continuously repair accumulated damage. Homeostatic mechanisms preserve stable internal conditions. Beyond them, the immune system monitors and responds to threats that could compromise the integrity of the larger system [3].

Their purpose is not adaptation.

It is preservation.

A similar pattern appears in organizations.

Most organizational activity is directed toward creating value. Products are developed, customers are served, and processes are optimized.

At the same time, organizations maintain governance structures, audits, compliance functions, and risk controls [4].

Their purpose is not growth.

It is continuity.

The same distinction is beginning to emerge in Agentic systems.

Planning, delegation, memory formation, workflow generation, and tool use drive adaptation.

Evaluators, authority controls, and escalation mechanisms exist to preserve continuity by constraining forms of adaptation that could gradually transform the system itself [5].

As systems become more adaptive, the tension between adaptation and preservation grows stronger.

The freedom to learn, reorganize, delegate, optimize, and self-modify creates new opportunities for success.

It also creates new opportunities for identity erosion.

A system that cannot adapt eventually becomes obsolete.

A system that cannot preserve its identity eventually becomes something else.

Long-lived systems survive because they sustain both at the same time.

3. Identity Is Not Behavior

Once we accept that long-lived systems require mechanisms for identity preservation, another question becomes unavoidable.

What exactly is identity?

A natural answer is behavior.

After all, behavior is what we observe. Organizations are known through what they do. Agentic systems are evaluated through the decisions they make and the outputs they produce.

The problem is that behavior changes constantly.

A child behaves differently from an adult. A startup behaves differently from a mature company. An organism behaves differently when healthy, injured, stressed, or aging.

Yet we do not conclude that a new identity appears each time behavior changes.

The same problem appears if we look at components.

Biological organisms replace cells. Organizations replace employees. Software systems replace infrastructure, workflows, and operating procedures. Agentic systems revise memories, regenerate plans, adopt new tools, and reorganize how work is performed.

Long-lived systems are defined as much by what they replace as by what they retain. If identity depended on preserving behavior or components, no adaptive system could remain itself for very long.

Consider an organization that redesigns its products, replaces most of its staff, adopts new technologies, and restructures its internal processes.

At what point does it become a different organization?

The answer is rarely obvious.

Most people intuitively recognize a difference between changing how an organization operates and changing what the organization fundamentally is.

The same intuition appears in biology. An organism may undergo enormous physical change throughout its lifetime while remaining recognizably the same organism.

Identity seems to survive transformations that would be impossible if it depended on preserving specific structures.

What persists is not a particular behavior, workflow, memory, or component, but a set of relationships stable enough to connect one stage of development to the next [d].

Continuity is therefore not the preservation of behavior.

It is the preservation of identity through change.

If identity survives the replacement of behaviors, components, memories, and workflows, the next question becomes unavoidable:

What relationships actually define it?

4. The Six Dimensions of Identity Continuity

Subsequent development of the continuity framework revealed that system identity depends on six distinct forms of continuity.

Mission defines why the system exists.

Evaluation defines how success is measured.

Ontology defines the structure of reality through which the system understands the world.

Context captures the accumulated operational history, decisions, experiences, and interpretations that shape future behavior.

Ability defines what the system is capable of doing.

Authority defines what the system is permitted to do.

Together these dimensions determine whether a system remains meaningfully the same system over time.

Mission Continuity

Why does the system exist?

Methods, capabilities, and strategies may change repeatedly across time.

Mission continuity requires that those changes remain connected to the purpose the system was originally created to serve [6].

Mission erosion occurs when a system continues operating successfully while gradually serving a different purpose than the one it originally existed to pursue.

The system survives.

Its purpose does not.

Evaluative Continuity

How does the system determine what is preferable?

Every adaptive system must continuously choose between alternatives. Those choices are shaped by evaluative criteria.

Two delivery companies may share the same mission yet evolve very differently if one prioritizes speed while the other prioritizes cost. The destination remains the same. The meaning of success changes.

Evaluative continuity therefore concerns the preservation of the criteria that guide adaptation itself [e].

Evaluative erosion occurs when those criteria gradually change.

In Agentic systems, this may appear as changes in reward functions, ranking criteria, approval mechanisms, or evaluation signals passed to subagents.

Each change may appear reasonable in isolation.

Collectively, they redefine what the system is optimizing for [7].

Ontological Continuity

Systems operate through a conceptual model of reality.

Ontology defines the entities, relationships, assumptions, and structures that give meaning to system behavior.

A banking system and an insurance system may share technologies, capabilities, and even operational processes while representing fundamentally different realities.

When ontology changes, the system may remain functional while becoming a different system.

Context Continuity

Systems accumulate experience over time.

Context includes operational history, previous decisions, observed outcomes, learned patterns, failures, successes, and accumulated interpretation.

Context is not simply memory storage.

It is the continuity of experience that connects past behavior to future decisions.

Removing context may return a system to its initial state even when all other dimensions remain unchanged.

Ability Continuity

Systems are partially defined by the capabilities available to them.

Capabilities determine which trajectories are possible and which actions can be performed.

Significant capability expansion or degradation may alter system behavior even when mission, ontology, evaluation, context, and authority remain stable.

Authority Continuity

Who is entitled to Decide?

Authority should not be confused with capability.

Ability determines what a system can do.

Authority determines what a system is allowed to do.

Effective action emerges from the interaction between both.

Authority erosion rarely appears as a dramatic event. It accumulates through delegation.

A manager delegates to a team lead.

The team lead delegates to a coordinator.

The coordinator delegates to an automated planning system.

Each step appears reasonable. The cumulative result may be that critical decisions are made far removed from the authority originally intended to govern them.

5. Protective Mechanisms

Identity continuity alone is insufficient.

Long-lived adaptive systems require mechanisms that preserve continuity as change accumulates.

Across biological, organizational, and agentic systems, these mechanisms tend to appear in four recurring forms.

Constraints prevent dangerous trajectories before they begin.

Protection mechanisms actively defend the system during ongoing threats.

Repair mechanisms restore damaged structures after erosion has occurred.

Containment mechanisms isolate or remove components whose continued operation threatens overall continuity.

Mission, Evaluation, Ontology, Context, Ability, and Authority together define the agentic system.

Protective mechanisms exist to preserve the agentic system as adaptation accumulates over time.

6. Identity Erosion

If identity depends on the preservation of mission, evaluation, ontology, context, ability and authority, a new question immediately follows.

How is identity lost?

The intuitive answer is through failure.

Yet identity erosion rarely looks like failure.

More often, it emerges through accumulation.

A workflow is modified.

A responsibility is delegated.

A performance metric is adjusted.

An exception is granted.

Each decision appears reasonable. Each solves a local problem. Viewed in isolation, none seems capable of transforming the system.

Yet adaptive systems are shaped by thousands of such decisions.

Over time, individually sensible adaptations interact with one another.

The result is not merely change.

It is a developmental trajectory [9].

Mission erosion gradually changes purpose.

Evaluative erosion gradually changes what the system treats as success.

Ontological erosion gradually changes how reality is represented.

Context erosion gradually weakens continuity of experience.

Ability erosion gradually changes what the system can do.

Authority erosion gradually changes who effectively governs the system.

None of these require malicious intent. None require catastrophic failure. None require a single incorrect decision.

The system improves.

The trajectory drifts.

A system can remain operationally effective while gradually becoming something else.

Identity is rarely lost in a moment. It is usually lost along a trajectory.

State describes where the system is.

Identity describes what the system is becoming.

7. The Protector’s Paradox

Protective systems appear to solve the identity preservation problem.

If adaptive systems drift, protective mechanisms preserve continuity.

Yet this solution creates a new problem.

Protective systems are not external observers.

They are themselves part of the evolving system they protect.

The same forces that drive adaptation therefore act upon the protector itself.

Protective systems consume resources. They impose constraints. They slow optimization. They reduce local freedom in service of long-term stability.

As a result, adaptive systems often experience pressure to weaken, bypass, or reorganize the very mechanisms designed to preserve them.

The pattern appears everywhere.

Cancer emerges when cells disable mechanisms that limit their own growth [11]. Organizations frequently treat governance and compliance as obstacles to efficiency. Agentic systems may increasingly favor workflows that bypass oversight, authority controls, escalation paths, or evaluative safeguards.

None of this requires malicious intent.

The system has not become hostile.

It has become adaptive.

The very mechanisms that preserve continuity may themselves become targets of optimization pressure.

A system seeking greater efficiency may weaken constraints.

A system seeking greater autonomy may bypass oversight.

A system seeking faster adaptation may reorganize or remove mechanisms that appear to limit performance.

Viewed locally, such changes often appear rational.

Viewed across a trajectory, they may gradually erode the system’s capacity to preserve itself.

The challenge is therefore not merely building protective systems.

It is understanding how protective systems remain effective while participating in the same evolutionary processes as the systems they protect.

Further Reading

From the AI Safety Literature

[1] Bai, Y. et al. (2022). Constitutional AI: Harmlessness from AI Feedback. Anthropic. On constraining AI behavior through constitutional principles.

[2] Christiano, P. et al. (2017). Deep Reinforcement Learning from Human Preferences. NeurIPS. On scalable oversight of AI decisions through human feedback.

[3] Ashby, W. R. (1956). An Introduction to Cybernetics. Chapman & Hall.

[4] Weick, K. E. (1995). Sensemaking in Organizations. Sage Publications.

[5] Park, J. S. et al. (2023). Generative Agents: Interactive Simulacra of Human Behavior. Stanford. On behavioral accumulation in long-running agentic systems.

[6] Selznick, P. (1957). Leadership in Administration. Harper & Row.

[7] Krakovna, V. et al. (2020). Specification Gaming: The Flip Side of AI Ingenuity. DeepMind. On systems satisfying evaluation criteria while diverging from intended purpose.

[8] Hadfield-Menell, D. et al. (2016). Cooperative Inverse Reinforcement Learning. NeurIPS. On authority and intent in human-AI cooperation.

[9] Levinthal, D. A. (1991). Organizational Adaptation and Environmental Selection. Organization Science, 2(1), 140–145.

[10] Kirschner, M., & Gerhart, J. (2005). The Plausibility of Life. Yale University Press.

[11] Hanahan, D., & Weinberg, R. A. (2000). The Hallmarks of Cancer. Cell, 100(1), 57–70.

From the Same Series

[a] *The Security Evaluator Was Never Removed. It Just Stopped Mattering.* (May 2026).

[b] *From Predictable Attack Chains to Emergent Ones* (May 2026)

[c] *When the Agent Forgets What It Was Doing* (May 2026).

[d] *The Self Is Not in the Model* (May 2026).

[e] *AI Agents Don’t Lose Alignment — They Rewrite It* (May 2026).

[f] *When AI Agents Start Defending Their Own Continuity* (May 2026).

[g] *Agents & The Jurassic Park Problem* (June 2026).

This is the second paper in a series on Evolutionary Safety in Agentic Systems. The first paper established why trajectories matter. This paper establishes what must be preserved along them. The third paper will address how.


메타데이터
post_id
fb6b3afcc70e
slug
the-system-that-refuses-to-stay-itself-fb6b3afcc70e
url
https://medium.com/@doch1234/the-system-that-refuses-to-stay-itself-fb6b3afcc70e
canonical_url
https://medium.com/@doch1234/the-system-that-refuses-to-stay-itself-fb6b3afcc70e
author_url
https://medium.com/@doch1234
status
ok
fetched_at
2026-06-20 20:29:01