Thought leadership
Evidence is a lagging indicator: what the largest AI scribe study actually tells us
The biggest real world study of ambient AI scribes found modest time savings and no measurable change downstream. Why we think the evidence is lagging the technology, not contradicting it.
Published

Two years. Five hospitals. 1,800 clinicians using ambient AI scribes, measured against 6,770 who were not.
It is the largest real world look at ambient documentation published so far, co led by Mass General Brigham and UCSF, and the first set of results out of the Ambient Clinical Documentation Collaborative. If you have been waiting for a study big enough to settle the argument about AI scribes, this is the closest thing yet.
The findings are modest. That is worth sitting with rather than spinning.
What the study found
Clinicians using AI scribes got back roughly 16 minutes a day on documentation and 13 minutes a day in the EHR overall. In relative terms that is a 10% reduction in documentation time and a 3% reduction in total EHR time.
The revenue picture is similar in shape. Seeing slightly more patients produced a statistically significant increase of about $167 per clinician per month. Real, measurable, and not the number anyone was modelling in their business case.
Time spent in the EHR outside working hours did not differ meaningfully between the two groups. The pyjama time stayed.
Three details matter more than the headline figures:
- Intensity changed everything. Clinicians who used the tool in more than half of their visits saw roughly twice the reduction in total EHR time and three times the reduction in documentation time. But only about 32% of users reached that threshold. Most of the cohort was using the tool occasionally, and the average reflects that.
- The gains were not evenly distributed. Primary care physicians, advanced practice providers and heavy EHR users saw the clearest benefit. A single system average flattens a wide spread of individual experience.
- Burnout and minutes are not the same measure. Earlier work has linked ambient documentation to meaningful reductions in burnout. The senior author on this study made the point plainly: reductions this modest are unlikely to explain those burnout findings on their own. Something else is happening in how clinicians work with these tools, and we do not yet have a clean way to measure it.

The emergency department version of the same story
Separate analysis from the Brigham emergency department lands in the same place from a different angle. There, an AI scribe saved about 1.6 minutes per note. A human scribe saved roughly double that. Neither changed how many patients a physician could see in a shift, and neither changed what the hospital collected per patient.
The minutes were real. They just did not reach the patient in the waiting room.
That is the honest tension in all of this. Time saved at the keyboard is not the same as time saved in the department. In a system where the constraint is beds, boarding, staffing and test turnaround, handing a clinician back ninety seconds per note does not move the queue. The bottleneck downstream absorbs it.
Why we are still optimistic
Two reasons, and neither is wishful thinking.
- Absence of evidence is not evidence of absence. These are observational studies of early deployments, measuring the outcomes that happen to be easy to instrument. Keystrokes and EHR timestamps are easy. Cognitive load, attention during the consultation, decision quality, the number of times a clinician does not have to reconstruct a story from scratch: much harder. The measurement is lagging the experience, and clinicians keep reporting a benefit the timers are not capturing.
- The category is genuinely early. The models, the workflows, the integration points and the habits around them are all a year or two old. Judging the ceiling of ambient AI from 2025 deployments is like judging laparoscopic surgery from the first year of the learning curve.
So our instinct is this: in clinical AI, evidence is going to be a lagging indicator.
The uncomfortable part
That is a deeply uncomfortable sentence in healthcare, and it should be.
For pharmaceutical agents and physical devices, the cycle runs one way. Decades of testing, trials and evidence generation, and only then widespread adoption. The evidence leads, adoption follows, and the whole regulatory apparatus is built around that order.
Clinical AI has bucked that trend. Rightly or wrongly, adoption arrived first and the evidence is now trying to catch up with tools that are already in daily use across thousands of clinicians. Studies like this one are the catching up.
Which leaves the sector with a real gap. We do not have a shared framework for evaluating what these tools are worth today, and we have an even weaker one for what they might be worth downstream to patients, clinicians and health systems. Documentation minutes became the default metric largely because they were the easiest thing to count, not because they were the right thing to count.
What this means if you are the one buying
None of this argues against buying. It argues for buying with your eyes open.
- Assume the system average will disappoint you. Plan for the distribution instead. The value sits with heavy users in the right specialties, so adoption depth is the variable to manage, not licence count. A tool used in 20% of visits will return roughly nothing.
- Do not build the business case on throughput. Two independent analyses now show no measurable change in patients seen or revenue captured at the department level. If the finance case rests on seeing more patients, it will not survive contact with the data.
- Decide in advance what a win looks like. If the goal is retention, burnout or the quality of the consultation, measure those directly and accept that they are slower and messier to evidence. If the goal is minutes, measure minutes and set realistic expectations.
- Look at where the time actually goes. Documentation is one source of friction. Finding the right protocol, tracking down a current guideline, working out which version of policy applies on this ward tonight: that is another, and it is one that lands squarely in clinical decision making rather than after the fact paperwork. Different tools, different constraint, different measurement.
Where we land
A study this size finding modest results is not a failure of the technology. It is the evidence base doing its job, at the pace evidence moves, on tools that are still early.
The right response is neither to declare ambient AI overhyped nor to wave the findings away. It is to get much better at asking what we are actually trying to improve, and then measuring that thing honestly, including when the answer is uncomfortable.
We would rather be part of an industry that publishes a modest result than one that never checks.
Eolas Medical gives clinicians one tap access to their own institution's protocols and guidelines, at the point of care. Better Knowledge. Safer Care.
Sources: Rotenstein et al., JAMA (2026), first published results of the Ambient Clinical Documentation Collaborative; Mass General Brigham newsroom, April 2026; Annals of Emergency Medicine (2026); STAT, September 2026.
Related reading

Thought leadership
The Golden Triangle of Clinical AI
We've been diving into the three streams of synthesis converging at the moment of clinical decision and have written an article exploring each corner... Conversation Synthesis i.e what is being said in the room, Chart Synthesis i.e what is known about this patient and Knowledge Synthesis i.e both what medicine knows and how this health system practises. Each corner has its own category leaders. Each is at a different stage of maturity. Each has real limits. The integrated future of Clinical AI runs on all three.
Published

Thought leadership
Zero Antimicrobial Prescribing Errors: What a Nature-Published Study Means for the AI and the Future of Antimicrobial Stewardship
Antimicrobial resistance is accelerating. The tools clinicians use to fight it haven’t kept pace — until now
Published

Thought leadership
Delivering Clinical Knowledge at the Point of Care: Engaging HCPs with Timely, High-Quality Content
Timely, embedded content boosts HCP engagement and patient care.
Published