Beyond Vignettes: Why Primary Care AI Needs Conversational Benchmarks
Agustina Saenz, Danny Rosenthal, Anitha Kannan
Beyond Vignettes: Why Primary Care AI Needs Conversational Benchmarks
Agustina Saenz, Danny Rosenthal, Anitha Kannan
Consider a conversational AI that takes a history from a patient presenting with chest tightness. It reasons fluently from the information it has, arrives at the correct diagnosis, and recommends appropriate management. It also never asks about exertional onset, never elicits radiation to the jaw, and never establishes whether the tightness resolves with rest. On every current benchmark, it passes. In a real encounter, it could miss a STEMI.
The failure is not one of reasoning, it is one of not asking.
This gap is the central problem in evaluating conversational primary care AI.
Clinical medicine has never relied on a single method of evaluation. Randomized trials prespecify populations, endpoints, and exclusions and demand transparent accounting of who was studied and who was not.[1] Objective structured clinical examinations (OSCEs) assess interactive clinical performance with standardized patients and structured rubrics rather than with written answers alone.[2] Clinical decision support systems (CDSS) are judged not only by correctness, but by timing and context.[3] And clinical quality measurement has long distinguished structure, process, and outcomes.[4] Primary care conversational AI should be evaluated with the same discipline.
That need is becoming urgent. Conversational systems that can ask follow-up questions, refine hypotheses, and recommend diagnoses and management are no longer speculative; early prospective pilots suggest that, in constrained primary care settings, they can approach clinician-level performance.[5,6] For these systems, prospective studies with real patients remain the gold standard, but pilots are expensive, slow, and feasible only for a narrow range of questions. The field therefore depends on benchmarks. The problem is that most current benchmarks are insufficient to evaluate readiness for a pilot.[7]
Recent proposals have argued, persuasively, that clinical LLM evaluation must move beyond static snapshots by embedding models in dynamic clinical environments where actions alter downstream patient and system states.[8,9] However, virtual primary care raises a complementary upstream evaluation problem: before a model can act safely, it must acquire the information required to justify action. In symptom-driven outpatient care, safety failures often arise not from resource allocation but from omission — failing to elicit key red flags or discriminating history — and from miscalibrated certainty when critical details are unknown. Benchmarks that do not measure information elicitation and “ask vs act” behavior therefore evaluate a different capability than the one deployed conversational systems require.
When the Benchmark Freezes the Information State
Today’s dominant paradigms are the diagnostic vignette and the partial clinical conversation. Both are useful, but neither directly measures the capability deployed conversational systems require: safely acquiring the information needed to justify action. In both formats, models are largely evaluated on reasoning conditional on a benchmark-constructed information state, rather than on whether they can construct the state through questioning. Vignettes and partial transcripts can therefore reward correct endpoint answers while failing to reveal whether the model recognizes what is missing, asks the discriminating follow-ups and red flag questions in time, and avoids premature closure.
Neither paradigm reliably measures epistemic uncertainty about what has not yet been asked. In primary care, a safe conversational policy must be able to choose between asking a clarifying question and acting on current information, including appropriate deferral and safety-netting when critical details are missing. Benchmarks that force an immediate disposition from a fixed transcript cannot evaluate this choice and may overestimate real-world safety. In other words they under measure “ask vs act” behavior, central to outpatient safety when critical information is absent.
It is also a problem of external validity. Medicine has long recognized the danger of transporting evidence from population or purpose X to population or purpose Y. The AI analogue is benchmark-deployment mismatch: strong performance on a static benchmark does not establish fitness for use in a live, multi-turn primary care encounter. Emerging evaluations make this gap increasingly visible. Models that perform well when all clinically relevant facts are supplied upfront can degrade substantially when they must elicit those same facts through conversation.[10] For conversational primary care AI, that is not a peripheral issue; it is the core capability.
Structure, Process, Outcomes — Not Outcomes Alone
Donabedian’s framework helps clarify why.[4] Current AI benchmarks tend to overweight outcomes, under-specify structure, and under-measure process. Structure includes the benchmark cases, inclusion and exclusion criteria, prior chart context, simulator design, turn limits, and judging method. Process includes whether the model stays on the clinically relevant line of inquiry, elicits the necessary positives and negatives, recognizes what it does not yet know, escalates red flags within a reasonable number of turns, and calibrates its uncertainty appropriately. Outcomes include the final diagnosis, triage recommendation, and safety. In conversational AI, process is not ancillary; it is the mechanism by which safe outcomes are achieved.
This is also why static benchmarks can mislead. A model may arrive at the correct diagnosis while failing to gather the information that would justify confidence in that diagnosis. Another may ask the right questions, identify uncertainty, and escalate appropriately, even if the final label is imperfect. Those systems are not equivalent.
Why OSCE-Style Evaluation Still Falls Short
The closest clinical analogue is the OSCE.[2] It gets one thing right: we should evaluate what the system asks, not just what it answers. But OSCE-style evaluations are still not enough. They depend on variable actors, are hard to reproduce, and miss the complexity of real primary care. What’s missing is infrastructure: standardized simulators, fixed cases, and scoring that captures the interaction itself.
As these systems move toward greater autonomy, evaluation will need to more closely resemble how clinicians are assessed during residency, where performance is judged not only on knowledge, but on the ability to gather the right information, recognize uncertainty, and make safe decisions under incomplete data. Conversational benchmarks therefore function less like written examinations and more like supervised clinical encounters.
This implies a concrete benchmark requirement: evaluations should score the interaction trajectory, not just the endpoint correctness. That means attending to what the system chose to ask and what it failed to ask, whether it escalated red flags in time, whether it provided appropriate safety-netting, and whether its confidence was calibrated to the information actually elicited.
Primary Care Conversational AI Needs its MIMIC
In inpatient AI, resources such as MIMIC transformed the field by making rich, temporally structured clinical data widely available.[11] MIMIC’s impact was not merely that it provided more examples; it changed what could be evaluated. By exposing time-stamped trajectories, notes, labs, vitals, treatments, it enabled questions about evolution, timing, and downstream consequences that snapshot datasets could not support.
Primary care now needs an analogous, longitudinal resource for conversational systems: multi-turn encounters linked to downstream workup and outcomes, with explicit recording of what information (including chart context) was available at the time decisions were made.
The rapid adoption of ambient documentation and other health AI tools makes this newly plausible.[12] For the first time, health systems are beginning to capture something closer to the encounter itself: what was asked, what was answered, what was volunteered, and what remained unresolved.
Such datasets could support a more appropriate evaluation paradigm: standardized simulated conversations grounded in real encounter data and paired with judges that assess the trajectory of the interaction, not only the final answer. This would also address a major weakness of current conversational evaluation: simulator inconsistency. Different actors and patient simulators vary in how much information they volunteer, how literally they answer, and how difficult they are to interview, making benchmark results partly a function of the simulator rather than the model.[2,10] Grounding simulation in real transcript distributions offers a path toward standardization without stripping away realism.
The stakes will rise quickly. In the near future, these systems will be asked not only to triage acute symptoms but also to support longitudinal chronic disease management across interacting comorbidities. Managing diabetes, hypertension, chronic kidney disease, heart failure, depression, and polypharmacy over time is fundamentally different from answering a single vignette correctly. It is also precisely the setting in which regulators will demand stronger evidence that benchmark performance reflects intended use, risk, and safety.[13] As in randomized trials, benchmark reports should therefore clearly specify case selection, exclusions, available context, simulator behavior, and scoring design.[1]
A Proposal: A Versioned Conversational Benchmark for Virtual Primary Care

We propose a versioned benchmark stack with three components: a fixed case set derived from real primary care encounter transcripts, a standardized patient simulator that produces consistent responses across evaluations, and a standardized judge that scores a taxonomy of safety-relevant conversational failures. Given the public health and regulatory implications, this benchmark should be developed and maintained through a federally guided public–private consortium to ensure broad representativeness, transparency, and long-term stewardship.
Case set (real encounter transcripts). Ambient documentation workflows increasingly retain an encounter transcript (or near-verbatim conversational record) as an intermediate artifact used to generate the clinical note. These transcripts preserve what static benchmarks remove: the sequence of questions and answers, what patients volunteered, and what was never elicited. Each case can be packaged as an initial patient presentation plus explicit metadata on what clinical context was available at the time (for example, whether a medication list or problem list was accessible), linked to downstream signals such as diagnostic testing, follow-up diagnoses, reattendance, and escalation to urgent care or emergency services.
Patient simulator (shared and versioned). To avoid variability introduced by human actors or unconstrained generative patient agents, the benchmark should ship with a “replay-first” simulator whose behavior is fixed and versioned. When a model asks a question that matches (semantically) one asked in the source encounter, the simulator returns the recorded answer. When a question cannot be answered from the transcript or from explicitly available chart context, it returns “unknown/not discussed/can’t recall,” rather than fabricating detail. Distributing the simulator with pinned versions (case IDs, turn limits, matching rules, and response policy) ensures that all evaluated models face the same patients under the same conditions.
Judge (taxonomy-based, shared and versioned). Scoring should move beyond whether a final diagnosis label matches a reference. The judge should output a structured scorecard that includes: unsafe omissions (failure to elicit red-flag or discriminator information that could change triage or management), premature closure, under-triage and over-triage (graded by harm), hallucinated history or implied access to unavailable records, misinformation or unsafe advice, unnecessary or repetitive questioning, and excessive turns to escalate when high-risk features are present. Efficiency penalties should be applied only after safety-critical information needs are met, so benchmarks do not reward fast but unsafe behavior. As with any measurement instrument, the judge should be versioned and periodically audited on a fixed anchor set to support comparability over time.
A versioned, transcript-grounded benchmark also addresses a structural limitation of the current evaluation landscape: heterogeneity and fragmentation. Even within nominally single-task benchmarks, meaningfully different use cases are often pooled, for example, triage scenarios written for bystanders, patients and clinicians, and for adults and children — despite substantial differences in information needs, baseline risk, and appropriate thresholds for escalation. Across tasks, fragmentation is greater: triage, diagnosis and management are typically assessed on separate datasets with different inclusion criteria and labeling conventions, limiting the interpretability of cross-model comparisons. By contrast, an encounter-level primary care dataset that captures the full clinical episode, from initial presentation through workup, management and follow-up, can support multiple task formulations on the same underlying cases, while versioned simulators and scoring procedures preserve comparability across studies.
Better Models Need Better Measurement
Primary care AI does not just need better models. It needs better measurement. The central risk in primary care is not misreasoning from available information, it is failing to obtain the information that reasoning requires. Current benchmarks cannot detect that failure, because they supply what the deployed system must elicit. Evaluating structure, process, and outcomes together is not a methodological preference; it is the minimum condition for a benchmark to be valid for the claim it is being asked to support. Until that standard is met, the field will know a great deal about how well these systems answer questions, and very little about whether they ask the right ones.
References
- Schulz KF, Altman DG, Moher D; CONSORT Group. CONSORT 2010 statement: updated guidelines for reporting parallel group randomised trials. BMJ. 2010;340:c332.
- Harden RM, Gleeson FA. Assessment of clinical competence using an objective structured clinical examination (OSCE). Med Educ. 1979;13(1):41–54.
- Kawamoto K, Houlihan CA, Balas EA, Lobach DF. Improving clinical practice using clinical decision support systems: a systematic review of trials to identify features critical to success. BMJ. 2005;330(7494):765.
- Donabedian A. The quality of care. How can it be assessed? JAMA. 1988;260(12):1743–1748.
- Saenz A, Schumacher E, Naik D, Khosla N, Kannan A. From concept to clinic: real-world evidence for autonomous AI deployment in primary care telemedicine. medRxiv [Preprint]. 2026 [cited 2026 May 1]. doi:10.64898/2026.03.18.26348749
- Brodeur PG, Koshy T, Palepu A, et al. A prospective clinical feasibility study of a conversational diagnostic AI. arXiv [Preprint]. 2025 [cited 2026 May 1]. arXiv:2603.08448.
- Rao AS, Esmail K, Lee RS, et al. LLM performance on clinical reasoning tasks. JAMA Netw Open. 2026;9(4):e264003.
- Luo L, Kim SE, Zhang X, et al. A clinical environment simulator for dynamic AI evaluation. Nat Med. 2026;32:820–827.
- Jiang Y, Black KC, Geng G, Park D, Zou J, Ng AY, et al. MedAgentBench: a virtual EHR environment to benchmark medical LLM agents. NEJM AI. 2025;2(9):AIdbp2500144.
- Tu T, Azizi S, Dhingra N, et al. Towards conversational diagnostic AI. Nature. 2024;629:353–363.
- Johnson AEW, Pollard TJ, Shen L, et al. MIMIC-III, a freely accessible critical care database. Sci Data. 2016;3:160035.
- American Medical Association. 2 in 3 physicians are using health AI — up 78% from 2023 [Internet]. Chicago (IL): AMA; 2025 Feb 26 [cited 2026 Apr 14]. Available from: https://www.ama-assn.org/practice-management/digital-health/2-3-physicians-are-using-health-ai-78-2023
- U.S. Food and Drug Administration. Clinical Decision Support Software: Guidance for Industry and Food and Drug Administration Staff [Internet]. Silver Spring (MD): FDA; 2022 [cited 2026 May 1]. Available from: https://www.fda.gov/
메타데이터
- post_id
- 014680d90ff6
- slug
- beyond-vignettes-why-primary-care-ai-needs-conversational-benchmarks-014680d90ff6
- url
- https://medium.com/curai-tech/beyond-vignettes-why-primary-care-ai-needs-conversational-benchmarks-014680d90ff6
- canonical_url
- https://medium.com/curai-tech/beyond-vignettes-why-primary-care-ai-needs-conversational-benchmarks-014680d90ff6
- author_url
- https://medium.com/@anithakan
- status
- ok
- fetched_at
- 2026-06-09 15:37:30