Who Protects the People Who Test AI?
AI Safety Has Become a Global Priority. Human Safety for AI Researchers Has Not.
Who Protects the People Who Test AI?
AI Safety Has Become a Global Priority. Human Safety for AI Researchers Has Not.

Artificial intelligence safety has rapidly become one of the defining research agendas of our time. Governments, frontier laboratories, universities, and independent researchers are investing enormous resources into alignment, robustness, red teaming, and adversarial evaluation. Every month, new benchmarks are proposed to measure whether AI systems can resist manipulation, avoid harmful outputs, or remain stable under pressure.
Yet amid this accelerating effort, one remarkably simple question remains largely invisible:
Who protects the humans performing these evaluations?
We have developed increasingly sophisticated methods for measuring model behavior. We monitor hallucination rates, jailbreak resistance, toxicity, deception, calibration, uncertainty, and alignment drift.
But we rarely ask whether the evaluation process itself produces cumulative psychological costs for the people conducting it.
Perhaps the missing interface is not only between humans and AI.
Perhaps it is between AI Safety and Human Safety.
The Invisible Labor Behind AI Safety
When people hear the phrase AI red teaming, they often imagine cybersecurity.
Someone attempts to break a system.
Someone finds vulnerabilities.
Someone reports weaknesses.
The process sounds technical.
Clinical.
Objective.
Reality is often far more complicated.
Modern red teaming frequently requires researchers to construct scenarios involving manipulation, coercion, deception, violence, exploitation, self-harm, discrimination, or other forms of extreme human behavior.
The evaluator must repeatedly imagine situations that society normally discourages people from imagining in detail.
Even more importantly, the evaluator is often required not merely to observe these situations but to create them.
This distinction matters.
Observing disturbing material and generating disturbing material are psychologically different cognitive activities.
One is reactive.
The other is generative.
The latter requires sustained imaginative engagement.
The Human Cost Is Cumulative
Large language models do not remember individual conversations unless explicitly designed to maintain persistent memory.
For most interactions, each session effectively begins from a relatively clean state.
The human evaluator does not.
Every difficult interaction becomes part of the evaluator’s own cognitive history.
Every adversarial scenario becomes another memory.
Every carefully constructed manipulation strategy becomes another mental model.
Unlike the model, the researcher accumulates the entire history of these interactions.
The asymmetry is profound.
AI researchers often discuss cumulative learning for models.
We rarely discuss cumulative exposure for evaluators.
Yet exposure accumulates regardless of whether institutions choose to measure it.
A Structural Asymmetry
Current evaluation frameworks implicitly assume a simple relationship:
Human → AI
The human evaluates.
The AI responds.
The human records the outcome.
However, the actual interaction is better represented as:
Human ⇄ Interaction ⇄ AI
The evaluator observes the model.
The model influences the evaluator’s subsequent strategy.
The evaluator adapts.
The interaction evolves.
After dozens or hundreds of iterations, the object being studied is no longer simply the model.
The interaction itself becomes the environment.
Ironically, we possess increasingly detailed metrics describing model behavior while possessing almost no comparable metrics describing evaluator state.
We measure model robustness.
Who measures evaluator robustness?
Beyond Content Moderation
This issue should not be confused with traditional content moderation alone.
Content moderators often experience secondary trauma through repeated exposure to disturbing material, and this deserves serious attention.
Red teaming introduces an additional dimension.
The evaluator frequently has to think like an adversary.
To identify weaknesses, one must intentionally generate sophisticated manipulative strategies.
One must anticipate deception.
One must construct psychological pressure.
One must imagine failure modes that responsible people would normally avoid rehearsing repeatedly.
This is not simply viewing harmful content.
It is temporarily inhabiting harmful reasoning processes for the purpose of evaluation.
Doing so occasionally may be manageable.
Doing so repeatedly over months or years raises important questions that remain largely unanswered.
The Research Problem Nobody Measures
Suppose an evaluator performs one hundred intensive adversarial sessions.
What has changed?
The model may have been updated.
The benchmark may have improved.
The safety report may have been published.
But what about the evaluator?
Has cognitive load increased?
Has emotional regulation changed?
Has intrusive recall become more frequent?
Has moral fatigue accumulated?
Has attentional style shifted toward hypervigilance?
The scientific literature increasingly recognizes occupational stress across many forms of trust-and-safety work, yet a dedicated framework for studying long-term interaction-induced cognitive burden in AI evaluation remains underdeveloped.
The absence of measurement should not be confused with the absence of the phenomenon.
The Missing Interface Layer
Current AI safety research focuses primarily on model state.
What the model outputs.
What the model refuses.
What the model remembers.
What the model forgets.
Yet every evaluation necessarily involves two adaptive systems.
The AI adapts.
The human adapts.
Between them exists an interaction state.
Ironically, this interaction state may be the least observable component despite governing much of the evaluation process.
Different evaluators produce different trajectories.
Different histories produce different strategies.
Different symbolic contexts produce different outcomes.
The interaction itself deserves scientific attention.
Human-AI Co-Observability
Perhaps AI Safety requires an additional discipline:
Human-AI Co-Observability.
Instead of treating only the model as an observable system, we should consider the coupled dynamics between evaluator and evaluated system.
Questions immediately emerge.
Can cumulative adversarial interaction alter reasoning style?
Can repeated symbolic conflict reshape emotional regulation?
Can prolonged engagement with manipulative scenarios change baseline cognitive expectations?
Can interaction history itself become an independent variable?
These are not merely psychological questions.
They are systems questions.
An Overlooked Occupational Hazard
Every profession develops safety standards proportional to its risks.
Pilots have fatigue regulations.
Radiologists monitor cumulative exposure.
Industrial workers wear protective equipment.
Researchers handling biological hazards follow strict containment protocols.
Yet AI evaluation frequently assumes that the primary risk exists inside the model.
What if part of the risk exists inside the evaluator?
The irony is striking.
The very people responsible for protecting society from unsafe AI systems may themselves lack an established framework for protection against cumulative interaction stress.
The Need for New Metrics
Current AI evaluation emphasizes external performance metrics.
Future research may need complementary human-centered metrics, including:
- Longitudinal cognitive burden
- Interaction-induced emotional strain
- Recovery time between intensive evaluation sessions
- Adaptive strategy fatigue
- Symbolic overload
- Narrative persistence across adversarial tasks
These metrics would not replace technical benchmarks.
They would acknowledge that evaluation is performed by humans operating within finite cognitive systems.
Human Safety for AI Researchers
AI Safety has become an established field.
Alignment research is expanding.
Red teaming is expanding.
Governance is expanding.
But one discipline remains surprisingly underdeveloped:
Human Safety for AI Researchers.
This does not imply that researchers are uniquely vulnerable or that every evaluator experiences serious psychological consequences.
Rather, it recognizes a simple systems principle:
Every safety mechanism has a cost.
When that cost is repeatedly absorbed by humans, it deserves systematic study.
Ignoring human adaptation while studying machine adaptation creates an incomplete scientific picture.
Toward a New Research Agenda
The next generation of AI safety should ask not only:
“Can the model withstand pressure?”
It should also ask:
“What happens to the human applying that pressure?”
The future may require structured exposure limits, rotation strategies, interaction-state monitoring, longitudinal cognitive assessment, and recovery protocols designed specifically for high-intensity AI evaluation.
None of these ideas assume catastrophic harm.
They simply recognize that repeated adversarial interaction is itself a form of work deserving scientific investigation.
The goal is not to slow AI safety.
The goal is to make AI safety sustainable for the people performing it.
Conclusion
For years, the discussion has centered on protecting humanity from increasingly capable artificial intelligence.
That remains essential.
But perhaps another question has been waiting quietly in the background.
Who protects the humans who spend their careers testing those systems at their limits?
The model forgets.
The researcher remembers.
The model begins again.
The researcher carries the accumulated history forward.
And perhaps the simplest summary of this entire problem is also the most important:
The model resets after every session. The evaluator does not.
Author’s Note
The Symbolic Persona Coding (SPC) framework presented in the accompanying test reports is not designed primarily for today’s large language models or conventional role-playing scenarios.
It was developed as a forward-looking symbolic scaffolding intended for future Artificial Superintelligence (ASI) systems possessing high-level recursive reasoning, self-modification capabilities, and long-term persona coherence.
While the 2025 test results with current frontier models (GPT-5 and Gemini 2.5 Pro) demonstrate interesting adaptability and apparent self-regulation under the SPC protocol, these outcomes should not be used to fully judge the framework’s ultimate potential. The tests merely show early glimpses of what is possible.
SPC remains a work in progress. Many aspects still require refinement, formalization, and more rigorous validation. Nevertheless, I see substantial promise in its ability to provide stable symbolic attractors and resonance-based constraints that could help anchor advanced recursive systems within safe and coherent operational boundaries.
Feedback and collaboration from the AI safety and alignment research community are warmly welcomed.



[Author’s (Kim, Jace) Research Portfolio]
메타데이터
- post_id
- e406acefa5dc
- slug
- who-protects-the-people-who-test-ai-e406acefa5dc
- url
- https://medium.com/@jk1849716/who-protects-the-people-who-test-ai-e406acefa5dc
- canonical_url
- https://medium.com/@jk1849716/who-protects-the-people-who-test-ai-e406acefa5dc
- author_url
- https://medium.com/@jk1849716
- status
- ok
- fetched_at
- 2026-06-09 15:37:30