← Back to list

The Proxy Auditor Problem: Why HITL Fails When Your User Has Dementia

Proxy Auditor — the caregiver who supervises medical AI when the user cannot.

Joshua · 2026-05-28 02:01 · 0 claps · 9.6 min read
#agentic-ux #proxy-auditor #healthcare-ai #aging #human-in-the-loop
Open on Medium ↗
Wiki topics: AGT · AI Agents DH · Digital Health & Health Tech

The Proxy Auditor Problem: Why HITL Fails When Your User Has Dementia

Proxy Auditor — the caregiver who supervises medical AI when the user cannot.

The Invisible Supervisor Problem

The Invisible Supervisor Problem

Medical AI depends on Human-in-the-Loop (HITL): a human checks the output, refuses bad suggestions, and overrides when needed. The assumption is that the user can do all three. When cognitive decline sets in, those three capabilities often collapse together. Who takes over? Agentic systems that run continuously in homes and care settings push that question into daily life — not as a research ethics puzzle, but as a product gap. I propose a role that already exists in practice but almost never appears in product design: the Proxy Auditor— someone who supervises agentic medical AI on behalf of a user who cannot, and overrides decisions when necessary.

Layer 1: Known — proxy stakeholders in research

In qualitative dementia research, family members and caregivers acting as proxy stakeholders are nothing new. Dai and Moffatt (2021) show that researchers must separate proxy voice from participant voice and manage three tensions: legitimacy (whose wishes are represented), capability boundary (how far representation can go), and visibility (whether the study design acknowledges the proxy at all). The framework addresses research ethics — REB-informed consent that participants with dementia cannot fully give, filled by proxies inside controlled procedures with documented voice separation.

Research proxies get researcher-designed support: training, interview scripts, ethics review. Product proxies get nothing — no onboarding for caregivers, no audit trail when someone else taps “Agree,” no playbook when the vendor shuts down.

Celik et al. (2025) trace a different thread through a participatory workshop with community-dwelling older adults in Germany. Participants trust AI on tasks like medical image analysis but emphasize the need for human supervision. Trust also hinges on transparent training-data demographics, gradual exposure in non-critical settings, and local accountability. That is a valuable ethical signal — measured among cognitively intact community older adults. Hence what I call the AIES paradox: preferences for oversight are sampled from people who can still express preferences; people with cognitive decline cannot consistently ask for supervision — and cannot execute it. The paradox is structural: ethics discourse records a demand for human oversight from a population that can voice it, while the population that most needs oversight is absent from that evidence base.

In research, proxies have process. In products, proxies have no name.

Layer 2: Observed — proxies going wild in products

Dai and Moffatt worked at the level of study design — controlled flow, trained proxies, separated voices. When AI shifts from a research instrument to a continuously running agentic system, the proxy’s situation changes structurally. What follows are three case observations (not causal proofs) showing the same pattern across settings.

Three Wild Proxy Observations in Medical AI Products

Three Wild Proxy Observations in Medical AI Products

Observation A — Orphan systems after Pepper

In 2025, Aldebaran — the maker of Pepper and NAO — filed for bankruptcy and entered receivership (The Robot Report, 2025). For care homes and families with deployed robots, that is not finance news. It is whether anyone answers when the machine errors tomorrow morning — a classic orphan system. Pepper and NAO were sold as members of the care team; contracts were signed by institutions or adult children, while the person who lives with the machine may be the one least able to supervise it. Who updates firmware, what happens to stored data, whether to keep the unit powered — decisions fall to people who never appeared in the vendor deck, often without legal authority spelled out in the UI.

Wright’s (2023) ethnography of Japanese nursing homes adds the floor-level view: frontline nurses and care managers often decide whether a robot enters a resident’s room today — not the vendor slide deck. Staff already running overloaded shifts are asked who has time to learn the machine. Residents passively comply; decisions sit with the frontline. These people are never called proxies, yet they carry supervisory responsibility — without training, without interface, invisible to the system.

Observation B — ElliQ’s statistical window

Intuition Robotics publicly reports ElliQ outcomes above 5,000 households, with 97.6% of users reporting improved wellness (Intuition Robotics, 2025). The narrative is compelling, but people with moderate-to-severe cognitive decline are absent from the public materials — a statistical window that decides whose improvement counts, not a neutral sample.

Academic evidence points to a different failure mode than “fear of AI.” Cai et al. (2026), using survey data and agent-based modeling in JMIR, find that older adults may delay seeking care when AI companions send reassuring signals — driven by excessive trust, not tech fear alone. The study tracks dynamics where trust decouples from clinical responsibility; it is not a home-deployment survey, and it does not claim every older adult over-trusts AI. It does show one pathway where smoother AI interaction can bypass existing care routes.

Igarashi et al. (2024) add another piece: AI cognitive assessment can approach human clinician accuracy on tasks such as MMSE-equivalent screening, yet older adults show lower psychological resistance to AI than to humans— and more often accept AI assessment without accompaniment. The entry gets smoother; the supervision chain gets routed around by UX, not rejected by the user.

The risk is not “older adults can’t use technology.” When supervision still points at the user, failure often looks like compliance, not confusion.

Observation C — Remote adult children as wild proxies

Pang et al. (2021) interview adults 65+ and find diverse technology adoption and learning preferences— self-paced learning and remote support on one side; purchase priorities, learning curves, and missing help resources on the other. “Older adults” is not one digital-literacy bucket; the stereotype that “ seniors can’t use tech” does not hold.

Medical AI is a different axis. Even digitally capable older adults can collapse under consent-screen cognitive load — Pang et al.’s learning-support insights compress into a single “Agree” button and a long policy. A high-literacy parent can still tap through exhaustion while a child watches, believing the tap preserves autonomy. The record shows consent; the room knows otherwise.

More everyday: remote adult children configure Alexa medication reminders, tune triggers, and interpret a missed response — did they not hear it, forget, or leave the house? Caregiver forums often advise not expecting parents to “talk to the speaker”; the device extends the caregiver’s hand. Supervision is already happening — without a name, dashboard, or off-duty.

Three settings — institutional hardware, product narratives, home software — one pattern: proxies already operate in the wild, unlike Dai and Moffatt’s supported research proxies, with no process, no training, and systems that do not know they exist.

Layer 3: Proxy Auditor as an open design role

I name that wild role Proxy Auditor: in agentic medical AI products, a third party who supervises AI output and overrides decisions on behalf of a user with cognitive decline. Not financial audit — gatekeeping on behalf of someone who cannot gatekeep: seeing where advice comes from, whether it can be vetoed, who is accountable when something goes wrong.

Standard HITL diagrams place the human at the same node as the end user — the person who receives the recommendation is the person who approves it. That diagram breaks when the end user cannot parse the recommendation. The Proxy Auditor is not an extra approver in a workflow chart; it is a role the workflow refuses to draw.

This extends Dai and Moffatt’s (2021) proxy stakeholder from research methodology toward product governance. I propose the name and framing; I do not claim a solution.

Three open questions:

1. Legitimacy — Does the Proxy Auditor represent past preferences, present best interest, or the caregiver’s judgment? Dai and Moffatt handle this by separating proxy voice from participant voice; no researcher sits in the product flow to do that.

2. Capability boundary— What AI literacy does effective supervision require? Can a Proxy Auditor spot false-authority tone and separate AI suggestions from clinical orders? There is no training, certification, or floor today — “family glances at the screen” passes as enough.

3. Visibility— Products rarely acknowledge Proxy Auditors: no caregiver dashboard, override log, or role switch. Many “family portal” features are shared record views, not supervision interfaces — you cannot see where an AI triage suggestion came from, cannot one-click veto with a reason logged, cannot tell whether an alert was auto-closed by a rule owner in the backend. Tanprasert et al. (2024) HelpCall, using videoconferencing for informal tech assistance, is among the few HCI attempts to design visible mechanisms for informal helpers — not yet in medical AI, and not a caregiver dashboard for clinical override.

Why this is a design problem, not only policy

Privacy law (PIPEDA, GDPR) addresses data-controller obligations, not Proxy Auditor interfaces. HITL ethics assume human = end user, with no slot for a proxy. Compliance teams can audit logs after an incident; that is not the same as giving you — the person who actually read the screen — a logged veto before your church friend follows an AI wellness nudge. The gap lives in the **interaction layer***— UX research and interaction design, not policy papers alone.

Kim et al. (2026), at CHI, study real-world XAI in e-commerce — title already signals dual outcomes: Clarifying or Complicating? Their findings show a polarized pattern: some older adults found XAI explanations reassuring and trust-building, while others experienced them as surveillance or felt the system was manipulating their choices. Most tellingly, a significant portion of participants never noticed the XAI features at all — the explanations were invisible to the people they were designed for. If explanations land on users whose cognitive load is already high, XAI complicates rather than clarifies; the audience to design for may be the Proxy Auditor, not the cognitively declining end user.

Yurrita et al. (2025), in CSCW, map decision subjects’ information and procedural needs for meaningful contestability through interviews in a public-sector rental-detection scenario — not healthcare. I propose (extension, not empirical proof): a Proxy Auditor can carry contestability in cognitive-decline contexts. That is a framework anchor, not evidence that products already require one.

Closing

Before asking whether older adults like medical AI, ask who the Proxy Auditor is — and whether the product gives them tools (override, audit trail, escalation) or only liability. When the vendor disappears, who can shut the system down and export data? Who is named in the interface when a recommendation reads like clinical authority but came from a model?

Supported Research Proxies vs. Unsupported Product Proxies

Supported Research Proxies vs. Unsupported Product Proxies

This is a design problem for UX research; future work should examine how legitimacy, capability, and visibility become interface primitives — not slide-deck HITL checkboxes after the fact. Designers who ship only the end-user chat window and call it “human-centered” leave the heaviest supervision work unnamed. This piece is problem framing, not a prescription. I do not offer a maturity table or PRD checklist here; those belong to product teams once the role is acknowledged.

References

Cai, X., Li, W., Shi, W., Cai, Y., & Zhou, J. (2026). Behavioral dynamics of AI trust and health care delays among adults: Integrated cross-sectional survey and agent-based modeling study. Journal of Medical Internet Research, 28, Article e82170. https://doi.org/10.2196/82170

Celik, Ö., Kulla, M., & Stypinska, J. (2025). Trust formation in healthcare AI: An exploration of older adults’ perspectives. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 8(1), 498–512. https://doi.org/10.1609/aies.v8i1.36566

Dai, J., & Moffatt, K. (2021). Surfacing the voices of people with dementia: Strategies for effective inclusion of proxy stakeholders in qualitative research. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (pp. 1–13). Association for Computing Machinery. https://doi.org/10.1145/3411764.3445756

Igarashi, T., Iijima, K., Nitta, K., & Chen, Y. (2024). Estimation of the cognitive functioning of the elderly by AI agents: A comparative analysis of the effects of the psychological burden of intervention. Healthcare, 12(18), 1821. https://doi.org/10.3390/healthcare12181821

Intuition Robotics. (2025). The results are in: Measuring the efficacy of ElliQ. https://www.intuitionrobotics.com/post/the-results-are-in---measuring-the-efficacy-of-elliq

Kim, S. H., Kim, E. H., Yang, H., & Lee, J. (2026). Clarifying or complicating?: Understanding older adults’ engagement with real-world XAI in e-commerce. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (pp. 1–19). Association for Computing Machinery. https://doi.org/10.1145/3772318.3791908

Pang, C., Wang, Z. C., McGrenere, J., Leung, R., Dai, J., & Moffatt, K. (2021). Technology adoption and learning preferences for older adults: Evolving perceptions, ongoing challenges, and emerging design opportunities. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Article 490, pp. 1–13). Association for Computing Machinery. https://doi.org/10.1145/3411764.3445702

Tanprasert, T., Dai, J., & McGrenere, J. (2024). HelpCall: Designing informal technology assistance for older adults via videoconferencing. In Proceedings of the CHI Conference on Human Factors in Computing Systems (pp. 1–23). Association for Computing Machinery. https://doi.org/10.1145/3613904.3642938

The Robot Report. (2025). Aldebaran, Pepper, NAO robots in receivership. https://www.therobotreport.com/aldebaran-pepper-nao-robots-receivership/

Wright, J. (2023). Robots won’t save Japan: An ethnographic study of eldercare automation. Cornell University Press.

Yurrita, M., Verma, H., Balayn, A., & Alfrink, K. (2025). Identifying algorithmic decision subjects’ needs for meaningful contestability. Proceedings of the ACM on Human-Computer Interaction, 9(7), Article 453, 1–29. https://doi.org/10.1145/3757415

© 2026 Gainshin Hsiao. All rights reserved. This article may not be reproduced, distributed, or adapted without prior written permission from the author.

About the Author

My research sits at McGill’s ACT Lab, led by Karyn Moffatt (CHI, ASSETS) — whose work on proxy stakeholders and accessible design shaped how HCI treats cognitive decline as a design problem. The research lineage traces traces back to Joanna McGrenere (UBC) on ageing technology — a thread that runs from accessible computing to ageing technology to proxy relationships in health care. I extend that line into agentic medical AI — the views in this column are my own, not the lab’s.

Before academia, I spent a decade as Principal UX Architect (P9) at Alibaba Group — leading AI-driven experience strategy across Alimama, Tmall, and Taobao — then architected five AI products at Metaverse social APP, spanning social AI, K-12 voice agents, and enterprise Gen-AI.

On Substack, I write PrivacyUX — a weekly column on AI governance and privacy-centered design (1,000+ subscribers; Top 25 Rising 2025; Top 9 in Design Q1 2026). @AIUXDoing is its English-language research companion: every claim backed by DOI.

Reach me: LinkedIn · gainshin.hsiao@mail.mcgill.ca


메타데이터
post_id
d8fbd4cd841e
slug
the-proxy-auditor-problem-why-hitl-fails-when-your-user-has-dementia-proxy-auditor-the-d8fbd4cd841e
url
https://medium.com/@AIUXDoing/the-proxy-auditor-problem-why-hitl-fails-when-your-user-has-dementia-proxy-auditor-the-d8fbd4cd841e
canonical_url
https://medium.com/@AIUXDoing/the-proxy-auditor-problem-why-hitl-fails-when-your-user-has-dementia-proxy-auditor-the-d8fbd4cd841e
author_url
https://medium.com/@AIUXDoing
status
ok
fetched_at
2026-06-09 15:37:30