AI Digital Assistants in Healthcare: What the ROI Case Studies Actually Show
Every vendor pitch for an AI digital assistant in healthcare comes with a number attached: hours saved, calls deflected, dollars recovered. Some of those numbers hold up under scrutiny. Others come from a single case study written by the company selling the product. This piece pulls together the ROI case studies that have a named organization, a stated method, and (where possible) independent verification, and separates them clearly from vendor-only claims.
Quick Facts
- A 2025 randomized trial in NEJM AI of 238 UCLA Health physicians found DAX Copilot improved burnout scores but produced no statistically significant reduction in documentation time.
- OSF HealthCare reported more than $2.4 million in ROI in one year from its AI virtual assistant, Clare.
- CommonSpirit Health’s internally built assistant Insightli generated more than $10 million in value as part of over $100 million in system-wide AI and automation savings in fiscal year 2025.
- An MGMA Stat poll found only 19% of medical group practices use any chatbot or virtual assistant for patient communication.
What Counts as an AI Digital Assistant in Healthcare Right Now
“AI digital assistant” gets used loosely, so it helps to split the category into what’s actually being deployed. There are ambient documentation scribes that listen to a visit and draft a note. There are patient-facing chatbots and voice agents that handle scheduling, FAQs, and intake on a hospital website or phone line. There are internal, staff-facing assistants that help employees write, summarize, or search internal knowledge. And there are newer voice AI agents that place outbound calls to patients for chronic care follow-up. Each category has a different evidence base, and conflating them is the fastest way to walk into a purchase you can’t defend later. For the basics on how these tools fit into the broader AI-in-healthcare picture, see our simple guide to what AI in healthcare actually means.
| Assistant Type | Strongest Evidence Found | Financial ROI Evidence |
|---|---|---|
| Ambient documentation scribes (DAX Copilot, Nabla, Suki) | Randomized trial (NEJM AI) on burnout/task load | Mixed; documentation time savings not statistically significant in the RCT |
| Patient-facing chatbots/virtual assistants | Named health system case studies (OSF, Weill Cornell) | Reported by individual organizations; adoption still low overall |
| Staff-facing enterprise assistants | Health system self-reported operational data | CommonSpirit’s Insightli: $10M+ reported value |
| Voice AI care-management agents | Vendor-published case studies only | Vendor-reported multiples, not yet independently replicated |
Clinical Documentation Assistants: The Best-Studied ROI Case Studies So Far
Ambient AI scribes have the most rigorous evidence of any digital assistant category, largely because health systems and academic centers have run controlled studies rather than relying only on vendor testimonials.
The UCLA NEJM AI Trial: Burnout Improved, Documentation Time Didn’t
The strongest study to date is a pragmatic randomized trial published in NEJM AI in late 2025. Researchers at UCLA Health randomly assigned 238 outpatient physicians, representing 14 specialties, to either one of two AI scribe applications, Microsoft Dragon Ambient eXperience (DAX) Copilot or Nabla, or a usual-care control group from November 4, 2024, to January 3, 2025. This trial matters because, unlike most vendor case studies, it was not commissioned by either scribe company. Independent secondary reporting on the trial found the DAX arm showed a +2.8 point improvement on the Mini-Z burnout scale, with physician task load dropping by 39.9 points. Documentation time, however, told a different story: time-in-note fell by only 1.7% compared to control, a result that was not statistically significant (P = 0.66), and DAX was used in only about a third of eligible visits during the trial. The takeaway for buyers: burnout relief and documentation time savings are separate claims, and a vendor’s marketing may lean on the one that tested better.
The Peer-Matched Cohort Study: Engagement Up, After-Hours Work Up Too
A separate, independent cohort study published in the Journal of the American Medical Informatics Association tracked DAX use across an integrated health system from March to September 2022. It enrolled 99 providers representing 12 specialties, with 76 matched control group providers included for analysis, and found positive trends in provider engagement, while non-participants saw worsening engagement and no practical change in productivity, alongside a statistically significant worsening of after-hours EHR use. The study also reported no quantifiable effect on patient safety, which is worth noting whenever a vendor implies a safety benefit without data behind it.
Independent Result at Atrium Health: No Significant Improvement
Not every large-scale rollout has produced a positive result. Coverage in Healthcare IT News on research into Dragon Ambient eXperience’s general availability at Atrium Health (now part of Advocate Health) reported that the general availability of Nuance’s Dragon Ambient Experience copilot in Atrium Health’s electronic health records did not reveal significant improvements in key metrics for the organization. CommonSpirit’s own CIO has said something similar about ambient tools system-wide: one area that hasn’t delivered on the promised financial benefit is ambient listening technology, according to Becker’s Hospital Review. This doesn’t mean ambient scribes don’t work; it means results vary by organization, rollout maturity, and how “success” is measured.
Patient-Facing AI Digital Assistants: Call Deflection and Access ROI
Chatbots and virtual assistants aimed at patients, rather than clinicians, have a smaller but more financially concrete evidence base, largely because the metric (calls handled, appointments booked) is easier to count than a clinical outcome.
OSF HealthCare’s $2.4 Million Year With Its Virtual Assistant Clare
OSF HealthCare partnered with Fabric to deploy an AI virtual care navigation assistant named Clare on its website. According to the published case study, the software functioned as an AI virtual care navigation assistant, guiding patients to the best resources for their inquiry, acting as a single point of contact and available 24 hours a day to help patients during and outside of business hours. By handling symptom checks, scheduling, and resource navigation directly, Fabric diverted calls from the call center, which is where the reported $2.4 million in one-year ROI came from. This is a vendor-published case study naming a real health system, so it’s stronger than an anonymous statistic, but it’s still the vendor’s own account of its own customer, not an independent audit.
MGMA Data: Adoption Is Still Low, But Where It Exists, It Moves Volume
Broader adoption numbers put OSF’s result in context. An April 2025 MGMA Stat poll found only about one in five (19%) medical group practices use some version of chatbot or virtual assistant for patient communication, while 81% do not, based on 375 applicable responses. Where practices have implemented these tools well, the volume shift can be substantial: MGMA reports that at Weill Cornell Medicine, shifting appointment scheduling to a 24/7 chat interface led to a 47% increase in appointments booked digitally via an AI chatbot. That gap between low overall adoption and strong results where it’s used well is a pattern worth watching if your organization is deciding whether to pilot a patient-facing assistant; our guide to AI patient intake and digital check-in covers the operational side of that decision in more depth.
Staff-Facing Enterprise Assistants: CommonSpirit’s $10 Million Internal Tool
A less-discussed category is the internal, staff-facing assistant, an LLM interface employees use for drafting, summarizing, or searching internal content rather than anything patient-facing. CommonSpirit Health built its own version, called Insightli, and has published some of the clearest financial figures in this article. According to Becker’s Hospital Review, CommonSpirit’s generative AI assistant, Insightli, has generated more than $10 million in total value through productivity gains and subscription savings, and it’s part of a system where CommonSpirit generated more than $100 million in value through AI and robotic process automation in fiscal year 2025, with 242 applications now live across its hospitals. What makes this case study more credible than most is that CommonSpirit’s own CIO has publicly pushed back on hand-wavy AI ROI claims, saying leaders should be able to show what have your clinicians, your finance people and your operators actually said you were going to achieve, and you actually achieved afterwards. That discipline is rare enough to be worth flagging. For a broader look at how operational AI compares to clinical AI investment, see our piece on operational efficiency versus clinical AI.
Voice AI Care-Management Assistants: Promising, But Vendor-Reported Only
The newest category is outbound voice AI agents that call patients directly for chronic care check-ins, pre-procedure prep, or care gap closure. Hippocratic AI is the most visible vendor here, and its own published customer results include figures like a 360% increase in team capacity for chronic care management, a 30% reduction in readmission rates, and a 12x average ROI across different use cases, according to the company’s own site. These are notable numbers, and they’re attached to named use cases rather than vague marketing copy, which is more than many competitors offer. But they come entirely from the vendor’s own reporting, not an independent, peer-reviewed evaluation of the kind that exists for ambient scribes. Treat these as an early signal worth tracking, not settled evidence, until an outside researcher publishes a comparable study. Our related piece on AI patient engagement tools looks at how this category fits alongside more established patient outreach methods.
Where AI Digital Assistant ROI Claims Fall Apart
A few patterns show up across every category above. First, a financial ROI claim and a clinician or patient experience claim are genuinely different pieces of evidence, and a vendor citing only one should prompt a question about the other, since the NEJM AI trial showed strong burnout improvement alongside a null result on time savings. Second, sample size and independence matter enormously: a 238-physician randomized trial with no vendor funding carries more weight than a single-customer case study written by the company selling the tool. Third, some widely repeated statistics in this space, like specific call-deflection percentages in the 65% to 85% range that circulate in vendor marketing blogs, don’t trace to any independently published methodology or sample size, so they should be treated as unverified marketing claims rather than settled facts until a named, checkable source appears. Our companion page on independent AI-in-healthcare ROI case studies goes deeper on how to separate vendor-reported numbers from independently verified ones across the wider AI landscape, not just digital assistants specifically.
How to Vet an AI Digital Assistant ROI Case Study Before You Buy
Ask any vendor presenting an ROI case study three things: who funded the study, how large and how independent was the sample, and what happened to the secondary metrics they didn’t lead with. If a scribe vendor only shows a documentation-time chart, ask about burnout survey results and after-hours EHR time, since the evidence above shows those can move in opposite directions. If a chatbot vendor only shows a dollar figure, ask for the underlying call volume data. And if the case study is a single customer with no stated sample size or timeframe, treat it as a data point, not a benchmark. Pairing this with a broader look at AI triage and scheduling systems for hospitals can help ground a digital assistant purchase in the operational metrics your organization already tracks.
FAQ
Do AI digital assistants in healthcare actually save money?
Some do, with real numbers attached. CommonSpirit Health reported its internally built assistant Insightli generated more than $10 million in value, part of over $100 million in AI and automation savings system-wide in fiscal year 2025. OSF HealthCare reported over $2.4 million in ROI in one year from its virtual assistant Clare, though these remain individual, named results rather than guaranteed outcomes.
Do ambient AI scribes reduce physician documentation time?
The evidence is mixed. A 2025 randomized trial in NEJM AI involving 238 UCLA Health physicians found no statistically significant reduction in documentation time for DAX Copilot, but did find meaningful improvements in burnout and task load scores. A separate peer-reviewed cohort study found DAX use came with a statistically significant increase in after-hours EHR time.
How many medical practices actually use a patient-facing chatbot?
An MGMA Stat poll from April 2025 found only 19% of medical group practices used any version of a chatbot or virtual assistant for patient communication, with 81% not using one at all.
Are vendor-reported ROI numbers for AI digital assistants reliable?
They should be read as a claim about that vendor’s own customers, not an industry average. Figures like Hippocratic AI’s self-reported 12x average ROI or 30% reduction in readmission rates come from the company’s own published case studies, not an independent, peer-reviewed evaluation, so they’re informative but shouldn’t be weighed the same as a randomized trial.




0 Comments