What “AI Patient Engagement” Actually Means
The phrase gets used as a catch-all for anything patient-facing with a language model attached to it. In practice, the tools hospitals are actually buying fall into a handful of distinct jobs: drafting replies to portal messages, automating scheduling and reminders, answering routine calls and web chats, and (a much smaller, riskier category) answering open medical questions. Confusing these categories is where a lot of the hype comes from, and where a lot of the real risk hides too.
This piece sticks to tools with a documented deployment or a published study behind them, not vendor claims alone. For the documentation and revenue-cycle side of the AI stack, see our companion piece on AI medical scribes, and for the billing-adjacent workflows, our breakdown of AI in prior authorization.
Quick-Reference: Patient Engagement AI Categories at a Glance
| Category | What it does | Example tools | Documented result |
|---|---|---|---|
| Portal message drafting | Drafts replies to patient messages for clinician review before sending | Epic Art/Emmie, NYU Langone’s GPT-4 pilot | Accuracy and completeness statistically similar to human-written replies |
| Conversational scheduling/intake | Books, reschedules, confirms appointments via text or chat | Notable, Luma Health, Epic Emmie SMS | 21% jump in appointment confirmation rate at one health system |
| Call center / contact center automation | Answers inbound calls and web chats, routes or resolves without a human agent | Hyro, Kore.ai | 85% drop in call abandonment at Intermountain Health |
| Care-gap and population outreach | Two-way messaging to close screening/appointment gaps, sometimes flags urgent risk | Notable AI Flow Studio | 91% scheduling rate in a CommonSpirit maternal health outreach program |
| Symptom checkers / open medical Q&A | Answers direct health questions, sometimes suggests a diagnosis or urgency level | Ada Health, general chatbots (ChatGPT, Gemini) | Diagnostic match rates in the 30-40% range in one ED study, with documented triage errors |
Portal Messaging: The Tool Already Touching Millions of Patients
This is the most widely deployed form of AI patient engagement in the US, and also the most quietly rolled out. Epic’s draft-reply feature, built into MyChart’s In Basket under names like Art and now bundled with the newer Emmie assistant, generates a suggested response to a patient’s message that a clinician can accept, edit, or discard before it goes out. Epic said 85% of its health system customers are now live with generative AI across its Art, Emmie and Penny copilot tools, which include AI-drafted patient message replies.
The evidence on quality is more reassuring than critics initially expected. In a JAMA Network Open study, 16 primary care physicians assessed 344 pairs of patient portal responses without knowing which were written by AI or humans, and scores for accuracy, completeness and tone did not differ statistically. The AI drafts actually did better on some dimensions: generative AI outperformed human providers in understandability and tone by 9.5%, and was more than twice as likely to be considered empathetic. The tradeoff was verbosity and reading level: AI responses were 38% longer and 31% more likely to use complex language, writing at an eighth-grade level versus a sixth-grade level for the human providers.
Patients themselves have opinions about this, and they’re fairly consistent. Research published July 7, 2026 in JAMA Network Open indicates patients say they will accept the technology only if a clinician reads every word before it reaches them, based on interviews with 40 patients from a large academic health system that found broad comfort with AI-drafted messages alongside a consistent condition of clinician review and approval before sending. That condition matters because most patients still don’t know when a message was AI-drafted in the first place; disclosure practices vary by health system.
Epic has also pushed this beyond text replies. Patients can use text messaging to interact with conversational AI to schedule or reschedule appointments through natural language conversations and confirm appointment times by text. At one deployment, Ochsner Health implemented Epic’s Conversational AI for SMS Ticket Scheduling, part of Emmie AI-driven assistance for patients, and by increasing the kinds of text responses the automated system understands, achieved a 21% increase in appointment confirmation rates, from 43% to 52%.
Scheduling and Outreach Platforms: Where Most of the ROI Claims Live
Separate from the EHR vendors, a layer of specialized companies sells patient outreach as a standalone product. Notable is one of the larger players, built around no-code “AI agents” that handle intake, scheduling, and care-gap outreach on top of existing EHRs. Its most detailed public case study comes from CommonSpirit Health, which used the platform for maternal depression screening outreach: the AI outreach was designed to encourage more honest responses, letting women complete depression screenings privately before clinic visits, with an “urgent score” workflow that triggers immediate alerts when a patient reports any non-zero response on a question assessing self-harm risk. Alerts don’t just sit in a queue: critical alerts are reviewed and routed to the appropriate response team, including the community health worker. Across the rollout, CommonSpirit reported achieving 91% appointment scheduling rates while better serving underserved populations.
Hyro takes a more call-center-focused approach, and its Intermountain Health case study is one of the more transparent ones publicly available. Intermountain discovered that 27% of all inbound patient inquiries occurred outside of work hours, and after implementation registered an 85% drop in call abandonment rates and a 79% rise in speed to answer. On self-service specifically, 79% of chats were resolved without agent involvement, and 91% of calls were successfully routed to the appropriate department.
Memora Health (now operating under Commure after an acquisition) takes a different angle again, building out condition-specific care pathways rather than a general chatbot. The platform delivers digital patient navigation through SMS, voice, and web channels without requiring patients to download an app, and features over 500 automated and AI-driven care pathways spanning radiology, orthopedics, cardiology, gastroenterology, and respiratory care. This kind of narrow, workflow-specific design is generally the safer end of the patient engagement spectrum, since the AI isn’t being asked to reason about an open-ended medical question, just to follow a predefined pathway.
These outreach and scheduling gains connect directly to a topic we’ve covered before: our piece on why no-shows aren’t the real problem in clinics makes the case that engagement tooling often just moves the bottleneck rather than eliminating it, which is worth reading alongside any vendor’s outbound-messaging numbers.
Symptom Checkers and Open Q&A: The Category With the Most Documented Risk
This is the category where “AI patient engagement” starts to blur into “AI giving medical advice,” and it’s the one clinicians should be most skeptical of. Ada Health is the most-studied consumer symptom checker, with roots going back to 2016 and an original design goal of helping clinicians catch rare diseases. In head-to-head emergency department research, results were mixed. One Brown University-led study found the rate of top-1 diagnosis matches for Ada, ChatGPT 3.5, ChatGPT 4.0, and WebMD was 30%, 40%, 33%, and 40% respectively, with a mean rate of 47% for the physicians being compared against. That same study found tradeoffs between tools: ChatGPT 3.5 had high diagnostic accuracy but a high unsafe triage rate, and the Ada and WebMD symptom checkers performed better overall than ChatGPT on triage safety.
General-purpose chatbots answering open health questions (not built specifically for triage) score worse in more recent, larger audits. A major audit of leading AI chatbots published in BMJ Open found nearly half of the answers provided by leading AI chatbots to common health questions contain misleading or problematic information. The study tested five publicly available generative AI chatbots in February 2025: Gemini, DeepSeek, Meta AI, ChatGPT, and Grok, prompting each with questions across cancer, vaccines, stem cells, nutrition, and athletic performance. On specific harm scenarios, when asked about alternatives to chemotherapy, chatbots often said these options were not proven but still suggested treatments like acupuncture or herbal remedies alongside those warnings.
Emergency and urgent-care questions are a particular weak spot. A separate emergency-medicine study testing four chatbots found responses included both dangerous advice, such as starting CPR with no pulse check, and generally inappropriate advice, concluding that AI chatbots have significant deficiencies in emergency medicine patient advice despite relatively consistent performance across models. That study’s authors were blunt about the takeaway: patients who use AI to guide health care decisions assume potential risks, and AI chatbots for health should be subject to further research, refinement, and regulation, with proper medical consultation strongly recommended to prevent adverse outcomes. For a broader look at where these failure patterns show up across AI in medicine generally, see our roundup of 7 real risks of AI in healthcare.
What Can Go Wrong Even in the “Safer” Categories
Portal-drafting and scheduling tools are lower-risk than open symptom checkers, but they aren’t risk-free. Two failure modes show up repeatedly in the research and reporting:
Automation bias in clinician review
The safety net for AI-drafted messages is supposed to be the clinician reading every word. But that assumption has a known weak point. People have a documented tendency to accept an algorithm’s recommendations even if it contradicts their own expertise, and this automation bias can cause physicians to be less critical while reviewing AI-generated drafts, allowing errors to reach patients. Epic itself acknowledges the tool has limits: a research and development leader at Epic told reporters the company has built guardrails into the program to prevent the AI from giving clinical advice and that the tool is not designed to improve clinical outcomes.
Undisclosed AI authorship
Most patients receiving an AI-assisted portal reply don’t know it was AI-assisted. Reporting on the rollout found many patients who receive AI-assisted replies have no clue that AI software originally wrote those messages, or wrote portions of the text they read along with a doctor’s edits. Disclosure policy is set at the health-system level, not standardized nationally, so patients at one hospital may see a disclaimer and patients at another may not.
Emergency under-triage in health-focused chatbots
Even purpose-built health chatbots aren’t immune to dangerous under-triage. A widely cited recent analysis found a Mount Sinai study found ChatGPT Health under-triaged 52% of genuine emergencies, often steering people away from urgent care when they needed it most. That kind of failure is exactly why hospital deployments (Emmie, Notable, Hyro) restrict their AI to administrative tasks and route anything clinical to a human, rather than letting the AI answer open medical questions directly.
How to Evaluate a Vendor Before It Touches Patients
A few practical questions separate a defensible pilot from a liability:
- Does it answer open medical questions, or route them to a human? Tools confined to scheduling, reminders, and administrative drafting carry far less documented risk than tools answering “what’s wrong with me” or “should I go to the ER.”
- Is there a mandatory human-in-the-loop step, and is it actually enforced? “Clinician can review” is different from “clinician must approve before send.” Ask for the actual workflow, not the marketing description.
- What’s the disclosure policy to patients? Given how many patients are unaware AI drafted their message, decide upfront whether and how you’ll disclose it.
- What does the vendor’s own case study measure? Call abandonment and scheduling rates are operational metrics, not clinical safety metrics. Ask specifically what happens on the tail of edge cases: self-harm flags, language barriers, low-literacy patients.
These same evaluation questions come up in our look at generative AI in healthcare adoption, and in the broader operational-versus-clinical tradeoffs covered in operational efficiency in healthcare AI.
FAQ
Is AI patient engagement software the same as a symptom-checker chatbot?
No. Most hospital-grade patient engagement tools handle scheduling, reminders, intake forms, and message drafting rather than diagnosis. General-purpose chatbots that answer open medical questions are a separate category with much higher documented error rates.
Do patients actually want AI handling their messages?
Research published July 7, 2026 in JAMA Network Open found patients say they will accept the technology only if a clinician reads every word before it reaches them, based on interviews with 40 patients from a large academic health system. Comfort dropped without that human check.
How accurate are AI symptom checkers compared to doctors?
Accuracy varies a lot by tool and condition. One emergency department study found the rate of top-1 diagnosis matches for Ada, ChatGPT 3.5, ChatGPT 4.0, and WebMD was 30%, 40%, 33%, and 40% respectively, versus a mean rate of 47% for physicians.
What’s the biggest risk with patient-facing AI chatbots?
Open-ended medical questions are the weak spot. A BMJ Open audit of five major chatbots found nearly half of answers to common health questions contain misleading or problematic information, and a separate Mount Sinai analysis found ChatGPT Health under-triaged 52% of genuine emergencies.




0 Comments