← Back to list

The Average Clinician Doesn’t Exist

What an AI scribe study’s headline number hides

Grace Ann Hansen in Towards AI · 2026-06-12 10:01 · 0 claps · 10.0 min read paywalled
#artificial-intelligence #healthcare #data-science #health-informatics #electronic-health-record
Open on Medium ↗
Wiki topics: RAG · RAG & Retrieval ML · Machine Learning AI · AI · General 🔬 · Science · General

The Average Clinician Doesn’t Exist

What an AI scribe study’s headline number hides

Photo by Vitaly Gariev on Unsplash

Photo by Vitaly Gariev on Unsplash

If you like my article,👏 Clap 50 times (yes, you can!). Medium’s algorithm favors this, increasing visibility to others who then discover the article.

🔔 Follow me on Medium, LinkedIn, and subscribe to get my latest article. You can also subscribe to me on Substack!

Right now, in some hospital board meeting, a single number is on a slide. Sixteen minutes. That is how much documentation time the largest study yet of ambient AI scribes found that adopting clinicians saved per eight-hour day, and it is the wrong number to be staring at.

The study, published in *JAMA on April 1, 2026, by Lisa Rotenstein and colleagues across five academic health systems*, did the careful work. The slide does not. What gets carried out of a paper like this one is a lone figure, stripped of the distribution it came from, and a lone figure is exactly what this finding cannot survive.

The number, and what it actually is

Start with the figure itself, since it is not what most people think it is. Sixteen minutes is not a median. It is not the typical clinician’s day. It is a difference-in-differences estimate: the change in documentation time among the people who turned the scribe on, minus the change among a control group who never did, measured over the same window. That design is the right one. It strips out the trends that would have moved documentation time anyway, the EHR upgrades and staffing shifts, and seasonal patterns, and isolates what the scribe added. The catch is that what it isolates is still an average effect across everyone who adopted, and “everyone who adopted” turns out to be a crowd of people who used the tool in nothing alike ways.

The arithmetic matters here. The study tracked 8,581 ambulatory clinicians from June 2023 through August 2025. Of those, 1,809 adopted an ambient scribe. The remaining 6,772 are the comparison group. So the sixteen-minute figure (95% CI 13.7 to 18.3) is a gap between two averages, drawn from a population where roughly four in five clinicians never adopted the tool at all. The companion figure, 13.4 fewer minutes in the electronic health record overall (95% CI 9.1 to 17.7), works the same way. In the study’s own framing, those amount to a 10% drop in documentation time and a 3% drop in total EHR time. Modest, and reported as modest. Visit volume rose by about half a visit per week.

The tools under study were not toys. They were the three products most health systems are actually buying: Ambiance, Nuance DAX Copilot, and Abridge, all wired into Epic, with EHR time measured through Epic’s own usage logs rather than a survey. This is the first published result from a five-system research consortium built to answer exactly the question a board wants answered. They answered it. The slide just quotes the wrong line.

The average describes almost no one.

Here is what the slide leaves off.

Adoption was opt-in at four of the five systems, and most clinicians declined. That alone should set off an alarm, since opt-in selects for the clinicians most likely to benefit: the early adopters, the documentation-heavy, the ones already hunting for relief. The adopter average sits tilted upward before a single minute gets counted. And even inside that self-selected group, the usage splits hard. Only about a third ran the scribe in at least half their visits, the threshold the study tied to the largest gains. The rest used it now and then, or tried it and drifted off. Multiply the two filters together. One in five clinicians adopted, and a third of those were heavy users, so the people the tool was built for, the ones running it through half their day, are something like seven in every hundred clinicians the study watched.

Now split the adopters by who they were and what they did, and the sixteen minutes come apart in your hands.

The heavy users, that top third, saved 21.3 minutes of total EHR time and 27.3 minutes of documentation time, per Mass General Brigham’s summary of its own study: roughly twice and three times the pooled number. Primary care adopters saved about 25 minutes in the record and 26.9 minutes in documentation. Surgical specialists saved nothing the study could tell apart from zero. Female clinicians saved about 19 minutes of EHR time; male clinicians, about 6 minutes. Medical specialists landed near 10.

Read that spread again. The same tool, in the same systems, in the same study, bought one clinician half an hour of her evening back, and another clinician no measurable change at all. Sixteen minutes is not the middle of that range. It is the place the math settles once you fold in a heavy-using primary care doctor and a surgeon who never opened the app, and then average both against 6,772 people who were not using it.

Why is the mean the wrong summary for this shape

I spent six years building clinical analytics platforms inside a health system, and I can tell you what happens to a figure like sixteen minutes once it leaves the methods section. It becomes a bar on a slide. The bar gets set against a vendor’s promise, the promise gets set against a per-seat license, and a purchasing decision falls out of the bottom. None of those steps carries the shape of the data forward. The bar is all that survives the trip.

The shape is the finding. When uptake is voluntary, and usage is long-tailed, with a small group of heavy users carrying most of the benefit, the mean no longer describes anyone’s actual day. It turns into an artifact of the mixing. Move thirty more surgeons into the sample, and the mean drops. Run a mandatory rollout instead of an opt-in, and every per-clinician minute estimate moves, not since the software changed, but since the denominator did.

Make it concrete with a hundred clinicians and the study’s own subgroup numbers. Hand the scribe to all hundred. About twenty turn it on. Of those twenty, seven run it hard and save something near 25 minutes a day; the other thirteen dip in and out and save maybe a third of that. The eighty who declined save nothing, since they never adopted. Add up the minutes across the people who actually used it, divide, and you land in the neighborhood of the published figure. Now look at what that average just did. It took a clinician who saves 25 minutes, a clinician who saves zero, and a clinician who saves neither. The mean is not a compromise between them. It is a fictional clinician standing in for a room full of real ones who behaved nothing like each other.

This is the place where the tired median-versus-mean argument gets it wrong, so let me be exact. The trouble is not that the study reported a mean when it should have reported a median. The trouble is reporting any single point of central tendency for a group this varied. A median would bury the structure just as deep. What this finding needs is the distribution behind it: the histogram, the quartiles, the share of adopters who saved nothing, the share who saved an hour. The study has that structure in its subgroup tables. The slide files it in the trash.

None of this is a pedantic complaint about statistics. It has a price. A system that buys, on average, deploys, on average: it hands the scribe to a flat slice of every department, watches the pooled minutes come back modest, and concludes the tool underdelivered. What actually underdelivered was the targeting. The seven-in-a-hundred who would have run it hard got the same nudge as the surgeon who was never going to touch it, and the budget got spread thin enough that nobody’s day changed much. Then, when the contract comes up for renewal, the average looks unimpressive, and a tool that gave its heavy users back half an hour every evening gets shelved for the wrong reason. The flattening does more than mislead the slide. It misroutes the money, and then it lets the technology take the blame.

The same tool, read three different ways.

Pull back from one study and the variance gets wider, not narrower. A perspective in *npj Digital Medicine*, led by Joshua Ohde at Mayo Clinic, aligned results across health systems and found the same category of tool posting numbers that do not appear related. Mass General Brigham’s earlier work clocked a median of 5.6 minutes saved per appointment. Kaiser Permanente Medical Group measured something closer to 18 seconds per appointment. Intermountain found no productivity change it could call real.

Eighteen seconds and five and a half minutes are not the same product behaving erratically. They are one product that meets different schedules, specialties, note cultures, EHR builds, and definitions of what counts as time saved. The npj authors catalog the ways a scribe stalls when it scales: speaker attribution that fails in a crowded exam room, notes that bloat instead of tighten, EHRs that will not take the draft cleanly, clinicians who trust the output too readily and stop checking it, two-party-consent states where recording a visit is its own legal question, and regulators who have not settled what these tools legally are. Not one of those problems shows up in a sixteen-minute headline, and every one of them changes the number for the system that hits it.

Illustrative distribution built from the study’s reported subgroup figures. Image by the author.

Illustrative distribution built from the study’s reported subgroup figures. Image by the author.

What stayed consistent was not the minutes.

Here is the turn. Across all of these studies, the variable that moved the least reliably was time. The number that moved most reliably was how clinicians felt about their work.

The new JAMA study found that after-hours documentation, the “pajama time” clinicians log from home once the kids are down, did not budge. Its senior author, Rebecca Mishuris, the chief health information officer at Mass General Brigham, said plainly that the modest time savings were unlikely to fully account for changes in burnout. She is right, and the evidence on burnout backs her up from a different angle. A separate multicenter study in *JAMA Network Open* followed 263 clinicians who used an ambient scribe and observed a reduction in burnout from 51.9% to 38.8% within a month. Kaiser Permanente group, after more than 2.5 million scribe-assisted visits, reported that 84% of clinicians felt the tool improved how they interacted with patients, even where the clock barely moved.

That is the signal hiding behind the noisy minutes. The benefit these tools hand over most dependably is not on the stopwatch. It shows up as lower cognitive load at the end of a visit, fewer notes carried home in the head, and the plain fact of whether a doctor looks at you or at a screen as you talk. Time savings are the metric everyone reaches for, since time is the easy thing to put on a slide. Attention and exhaustion are harder to chart, and they are where the change actually lives.

What a better dashboard shows

So what would honest reporting look like? Segment before you average. Break the savings out by usage intensity, since a tool used in 5% of visits and one used in 80% are not the same deployment wearing one label. Break it out by specialty, since primary care and surgery live on different documentation planets. Report the share of adopters who saved nothing or lost time, not just the winners at the top. And pair every minute figure with an adoption figure, since a 27-minute saving that reaches seven clinicians in a hundred is a different purchase than a 5-minute saving that reaches every one of them.

A dashboard built that way is harder to screenshot into a triumphant board slide. It refuses to resolve into a single number that a vendor can quote back to you. That is not a flaw in the dashboard. That is the dashboard telling the truth about a tool whose value is real, uneven, and concentrated, and refusing to launder the unevenness into a clean average that points the budget at the wrong people.

The money makes the same point.

Follow the dollars and the distribution surfaces again. The JAMA study estimated that adopters generated roughly $167 per clinician per month in additional visit revenue from that small bump in volume, and the authors called the figure a conservative floor. Set it next to the price tag. A companion commentary in *JAMA Network Open* notes that systems pay between $200 and $600 per clinician per month to license these tools.

On average, revenue does not cover the cost of the license. The math only closes if you stop averaging: put the tool in the hands of the heavy users, the primary care doctors, the people whose practice shapes, let it run all day, and the return reads nothing like it does when smeared evenly across a roster where most people will use it twice and quit. A flat per-seat number across an unsorted population is the sixteen-minute bar again, wearing a finance hat.

The same trap catches the marketing from the other side. The vendor pitch for ambient scribes has long promised to hand back an hour or two of a clinician’s day. You can hear the echo of it in a widely shared figure that Kaiser’s Permanente physicians saved 15,791 hours of documentation. That is a true number and an enormous one. Divide it across 7,260 physicians and more than 2.5 million visits, and it comes to something near 18 seconds per appointment. A giant aggregate and a trivial per-visit saving are the same dataset, told to flatter at one scale and to deflate at another. The hour-a-day promise and the sixteen-minute headline are both averages doing a magic trick.

Report the distribution

So here is the ask, for anyone who builds these dashboards, signs these contracts, or writes these headlines. Stop shipping the average. The average clinician saved sixteen minutes. The average clinician does not exist. What exists is a primary care doctor who clawed back half an hour of her evening, a surgeon whose tool did nothing measurable for, and thousands more sorted across every point in between.

Put the histogram on the slide. Put the quartiles next to the mean, and name the share of clinicians for whom the tool did nothing. The distribution is harder to read in fourteen-point font at a board meeting, and it is the only honest thing in the room.

Sixteen minutes is not the finding. The spread is the finding.

Author Note. Grace Ann Hansen is an independent researcher and writer, and an MBA & PhD graduate student in health informatics and artificial intelligence. She is also a published author, a professional musician, a gymnastics coach, and a queer transgender woman living in Sioux Falls, South Dakota. All interpretation, argument, and prose are her own. Correspondence concerning this article should be addressed to Grace Ann Hansen at grace@graceannhansen.com.


메타데이터
post_id
1e64ed05c79f
slug
ai-scribe-average-myth-1e64ed05c79f
url
https://medium.com/@graceannhansen/ai-scribe-average-myth-1e64ed05c79f
canonical_url
https://medium.com/@graceannhansen/ai-scribe-average-myth-1e64ed05c79f
author_url
https://medium.com/@graceannhansen
status
ok
fetched_at
2026-06-12 22:02:08