Generative AI in Healthcare: Real Examples and Limits

generative AI in healthcare

Written by ai-healthcare

July 26, 2026

Generative AI in Healthcare Means Something Narrower Than the Hype Suggests

When people say “generative AI in healthcare” they usually mean one specific thing: a large language model that writes something (a note, a message, a summary) rather than a model that classifies an X-ray or predicts a readmission. That distinction matters because most of the older, well-validated “AI in healthcare” examples, like mammography screening algorithms, are predictive models, not generative ones. Generative tools are newer, less proven at the bedside, and adopted for a narrower set of jobs: documentation, patient communication, and chart summarization.

This piece sticks to real, named deployments in US and Canadian health systems, the adoption numbers behind them, and the specific failure modes that peer-reviewed research has already documented. If you want the fuller catalog of AI use cases beyond generative tools, our 12 applications of AI in healthcare piece covers imaging, triage, and operations separately.

Tool Primary Job Where It’s Deployed Notable Adoption Fact
Nuance DAX Copilot (Microsoft) Ambient clinical documentation from patient-clinician conversations Epic and Oracle Health integrations across large US systems DAX users cut documentation time by 50% per Nuance/Microsoft
Abridge Real-time visit transcription into SOAP notes plus patient-facing summaries UPMC, Emory Healthcare, and Yale New Haven Health Supports 55+ medical specialties and 28 languages as of mid-2025
Mayo Clinic “Inpatient Insights” EHR-embedded AI summarization of inpatient charts Three academic tertiary hospital sites on Epic 512 pilot clinicians (physicians, APPs, therapists, pharmacists) in an Aug-Oct 2025 evaluation
Mass General Brigham patient-message drafting LLM-drafted replies to patient portal messages Pilot across ambulatory practices system-wide Published in The Lancet Digital Health with documented safety limitations
AIwithCare / RECTIFIER (Mass General Brigham spinout) Retrieval-augmented generation for clinical trial eligibility screening Spun out to license to other health systems Built to outperform manual trial-eligibility screening

Generative AI Adoption Is Real But Shallower Than the Headlines Imply

The topline adoption numbers look impressive until you separate “used somewhere” from “embedded in core clinical work.” A national survey of 2,174 nonfederal acute care hospitals, published in JAMA Network Open, found 31.5% of hospitals were already using generative AI integrated with their EHR by 2024, with 24.7% planning to adopt it within a year and 43.7% delayed or uncertain. That same study found major teaching hospitals (53.9%) and system-affiliated hospitals (38.5%) were far more likely to be early adopters than independent hospitals (16.3%), a gap that tracks resources more than clinical need.

McKinsey’s survey of US healthcare leaders tells a similar story at the organization level: generative AI adoption rose from 25% of organizations in late 2023 to 47% in 2024 and reached 50% by the end of 2025, with more than 80% of surveyed leaders saying they’d deployed at least one generative AI use case to end users.

A separate executive survey of 120 US health systems found the pattern holds at the platform level: clinical note-taking sits at 68% adoption with 62% year-over-year growth, while AI-based clinical documentation improvement sits at 43% adoption with 59% growth. Documentation, not diagnosis, is where the money and attention are going. That’s consistent with what we’ve covered in our deep dive on AI medical scribes, which remains the single most mature generative AI use case in US clinics today.

Real Generative AI Examples Already Running Inside US Health Systems

Ambient scribing at scale: Nuance DAX and Abridge

Nuance, acquired by Microsoft in 2022, built DAX Copilot on its decades-old Dragon dictation technology and layered generative AI on top to turn full visit audio into a structured note. Nuance reports that DAX users cut documentation time by 50%, though the same coverage notes that accuracy and reliability topped the list of concerns in a KLAS Research survey of health system executives. Abridge, its closest large-scale competitor, is deployed at systems including UPMC, Emory Healthcare, and Yale New Haven Health, and as of mid-2025 supported over 55 medical specialties and 28 languages.

Chart summarization: Mayo Clinic’s inpatient pilot

Mayo Clinic tested an EHR-embedded generative AI summarization tool called Inpatient Insights across three academic tertiary hospital sites running a single integrated Epic instance, giving 512 pilot users, including physicians, advanced practice providers, respiratory therapists, and pharmacists, access between August and October 2025. Mayo has also partnered with Google Cloud on a broader generative AI application aimed at clinical workflows and research, part of a wider pattern of health systems pairing hyperscaler cloud partnerships with in-house pilots rather than buying one packaged product.

Patient messaging: Mass General Brigham’s cautionary pilot

Mass General Brigham ran one of the more closely watched pilots: using an LLM to draft replies to patient portal messages inside the EHR, across a set of ambulatory practices. The results, published in The Lancet Digital Health, found the approach may help reduce physician workload and improve patient education, but also found limitations that could affect patient safety. One of the study’s authors put it directly: “keeping a human in the loop is an essential safety step when it comes to using AI in medicine, but it isn’t a single solution”. That single line captures the whole adoption story better than most vendor marketing does.

Generative AI Adoption in Canadian Health Systems Looks Different

Canada’s adoption story is shaped less by competitive vendor pressure and more by capacity strain. A 2026 Canadian Standards Association policy report notes that more than six million Canadians lack a primary care physician, with wait times rising and demographic change intensifying pressure for decades, and argues the balance of risk now favors supervised generative AI adoption over caution.

On the ground, that’s showing up as clinician-driven pilots rather than top-down mandates. At North York General Hospital, AI-assisted mammography screening improved detection through a machine-generated second reading, alongside a separate rollout of an AI clinical scribe and an internal policy chatbot that gives staff instant access to frequently updated guidelines. The same reporting found that across successful Canadian pilots, the common thread was a clinical champion who led the work from the start.

Canada also has a data infrastructure advantage most US systems don’t: the Vital network, built on its predecessor GEMINI, is the largest multi-institutional hospital data-sharing network in Canada, holding data on more than 3 million hospitalizations and covering roughly 60% of Ontario’s hospital patients, giving federal AI strategy planners a shared foundation that most fragmented US EHR environments lack. If you’re tracking events where these Canadian rollouts get discussed in detail, our AI in Healthcare Conferences Canada guide covers the Toronto event focused specifically on moving hospital AI from pilots to real implementation.

Where Generative AI in Healthcare Still Falls Short

Hallucinations remain unsolved, not just rare

A cross-industry 2025 survey found 44% of organizations reported experiencing negative consequences from generative AI use, with average financial losses of $4.4 million per incident. In healthcare specifically, a review of ChatGPT-focused clinical studies documented hallucinations or fabricated information in three of nine studies examined, which is a meaningfully high rate for a technology already being piloted inside EHRs.

Diagnostic reasoning is not where generative AI is strong

General-purpose generative AI is fundamentally different from the narrow, image-trained models used in radiology. Across a broad set of studies, general-purpose generative AI averaged 52.1% accuracy across 83 studies on open-ended diagnosis, close to a non-expert clinician’s performance, compared to narrow diagnostic models that reach about 96% for diabetic retinopathy detection. That gap is exactly why nearly every large-scale deployment covered above is documentation or communication, not diagnosis. For a fuller breakdown of failure categories, our 7 real risks of AI in healthcare piece walks through bias and error patterns specifically.

Regulation is still catching up to the technology

The FDA issued comprehensive draft guidance in January 2025 covering AI-enabled devices across their total product lifecycle, and held a second Digital Health Advisory Committee meeting in November 2025 specifically on generative AI-enabled digital mental health devices. But post-market oversight is thin: one analysis found only about 5% of AI-enabled devices had reported adverse-event data by mid-2025, including device malfunctions and one death. Most documentation and messaging tools, meanwhile, don’t go through device clearance at all, since they’re marketed as workflow software rather than diagnostic devices, which leaves a regulatory gray zone that hospital compliance teams are still working through.

What This Means for Health IT Teams Evaluating Generative AI Now

Three practical takeaways stand out from the evidence. First, adoption correlates with resources, not necessarily clinical need. Academic and system-affiliated hospitals lead, independent hospitals lag, which means a purchasing decision should weigh your IT capacity as heavily as vendor claims. Second, the safest, best-validated use cases right now are narrow: ambient documentation, chart summarization, and drafting (not sending) patient messages, all with a clinician reviewing before anything reaches a chart or patient. Third, don’t treat “generative AI” and “AI in healthcare” as interchangeable when evaluating a tool. If a vendor’s pitch depends on diagnostic accuracy claims for an LLM-based product, that’s the exact area where the published evidence is weakest. Our broader pros and cons guide is a useful next stop if you’re building a board-level case either way.

FAQ

What is generative AI in healthcare, exactly?

Generative AI refers to large language models that produce new text, summaries, or conversation rather than just flagging a pattern in an image or dataset. In healthcare, that mostly means tools that draft clinical notes, patient messages, and chart summaries rather than tools that detect tumors or predict readmission risk, which are usually built on older predictive AI methods.

How many US hospitals are actually using generative AI?

A JAMA Network Open study of 2,174 acute care hospitals found 31.5% were already using generative AI integrated with their EHR in 2024, with another 24.7% planning to adopt within a year. Separately, McKinsey found half of US healthcare organizations had implemented generative AI by the end of 2025.

Is generative AI reliable enough for diagnosis?

Not on its own for open-ended diagnostic reasoning. General-purpose generative AI has averaged around 52.1% accuracy across 83 studies on open-ended diagnosis, close to a non-expert clinician, which is why nearly all deployed use cases are documentation and communication rather than diagnosis.

Is generative AI in healthcare regulated by the FDA?

Some tools that qualify as medical devices fall under FDA oversight, and the agency issued draft guidance in January 2025 covering the total product lifecycle of AI-enabled devices. Most documentation and messaging tools, however, operate outside formal device clearance for now.

You May Also Like…

0 Comments

Submit a Comment

Your email address will not be published. Required fields are marked *