ChatGPT Prompts for Doctors: What Actually Works (and What’s Risky)
Physicians are no longer debating whether to use ChatGPT. They’re debating how.A 2026 survey by the American Medical Association found that 72% of physicians reported using AI in clinical practice, up from 62% the year before. Other international physician surveys tell a similar story: generative AI adoption in clinical practice has been climbing quickly across specialties in recent polling, though the exact figures and methodologies vary from survey to survey.
That adoption curve has outpaced most hospitals’ policies. Some prompts genuinely save time and produce useful drafts. Others quietly create a HIPAA problem or a false sense of confidence in wrong information. This piece separates the two, using the actual studies and vendor documentation behind the claims, not general reassurance.
Quick Facts
- 72% of US physicians reported using AI in clinical practice in 2026, up from 62% the year before, per the AMA
- Hallucination rates ranged from 50% to 82% across six different AI chatbots when a single fabricated medical detail was slipped into a clinical prompt, per a Mount Sinai study
- Standard ChatGPT is not HIPAA compliant: once patient information is entered, it sits on OpenAI’s servers, which is technically a data breach
- OpenAI’s physician advisors tested 6,924 conversations in their daily work across clinical care, documentation, and research before launch, and overall physicians rated 99.6% of responses as safe and accurate
What ChatGPT Prompts for Doctors Work Well in Practice
The clearest wins are administrative, not diagnostic. Physicians use ChatGPT and similar tools to draft notes, write prior auth letters, create patient education materials, and cut admin time, with use cases like prior auth letters drafted at midnight, patient-friendly discharge summaries, and differential lists on tough cases. None of that requires the model to be right about a diagnosis; it requires the model to be a fast, competent writer that a clinician then edits.
Drafting Prior Authorization and Referral Letters
This is the single most commonly cited practical use. Pre-built skills and clinician-specific starter prompts support common clinical tasks, including drafting notes, referrals, prior authorization letters, and patient instructions for clinician review. The template stays the same; you supply the clinical reasoning and the specifics, and the model handles formatting and phrasing. If your practice is also trying to cut administrative load more broadly, our piece on AI data entry automation in clinics covers where that time actually goes.
Plain-Language Patient Education Handouts
Common use cases include reviewing care pathways, synthesizing evidence with citations, reasoning through differentials, drafting prior authorization letters, summarizing patient charts, and creating plain-language explanations for patients. A prompt like “rewrite this discharge instruction at a 6th-grade reading level, keep all medication names and dosages exact” tends to work well because it’s a translation task, not a clinical judgment task.
Literature Orientation, Not Literature Conclusions
ChatGPT can be a reasonable starting point for scanning a topic, but it is a poor final source for citations. One JMIR study found hallucination rates stood at 39.6% for GPT-3.5, 28.6% for GPT-4, and 91.4% for Bard when generating references for systematic reviews. Use it to get oriented on a topic, then verify every citation independently, ideally through a tool built specifically for sourced clinical evidence.
Where ChatGPT Prompts for Doctors Get Genuinely Risky
The risky prompts aren’t obviously risky. They’re the ones where a clinician trusts an authoritative-sounding answer without checking it, or pastes in more patient detail than they realize.
The Misinformation Vulnerability Mount Sinai Documented
Researchers created fictional patient scenarios, each containing one fabricated medical term such as a made-up disease, symptom, or test, and in the first round, without extra guidance, the chatbots routinely elaborated on the fake medical detail, confidently generating explanations about conditions or treatments that do not exist. Hallucination rates ranged from 50% to 82% across six different AI chatbots, according to the Mount Sinai research team. As one of the senior researchers put it “Even a single made-up term could trigger a detailed, decisive response based entirely on fiction.” This matters directly for prompting: if a resident types a slightly wrong lab name or a mis-transcribed term from a chart, the model is more likely to invent a plausible-sounding answer than to flag the error. For more on this class of failure across AI tools generally, see our overview of real risks of AI in healthcare.
The HIPAA Problem With Standard ChatGPT
This is the risk most likely to create real legal exposure, not just an embarrassing note. Once something is entered into ChatGPT, it is on OpenAI’s servers, and they are not HIPAA compliant; the protected health information is no longer internal to the health system, and that is, technically, a data breach. Opting out of model training doesn’t fix this: physicians can opt out of having OpenAI use the information to train ChatGPT, but regardless of whether you’ve opted out, you’ve just violated HIPAA because the data has left the health system. Data entered into public AI tools may be stored, processed, or used in ways that do not meet HIPAA requirements, and even partially anonymized data still carries re-identification risk. If you’re evaluating AI tools for your EHR workflow specifically, our guide to AI EHR integration costs and compliance walks through the BAA question in more depth.
Where Diagnostic Accuracy Actually Sits
ChatGPT is not uniformly bad at clinical reasoning, but it is uneven. One study using published clinical vignettes from the Merck Manual found ChatGPT achieved an overall accuracy of 71.7% across all 36 clinical vignettes, with the highest performance in making a final diagnosis at an accuracy of 76.9% and the lowest performance in generating an initial differential diagnosis at an accuracy of 60.3%. That gap matters: the task doctors most want help with, a broad differential on an ambiguous presentation, is exactly where the model is weakest. Separately, a systematic review and meta-analysis found ChatGPT-4 achieved a pooled accuracy of 81.8% on medical licensing exams compared to ChatGPT-3.5’s 60.8%, and an accuracy rate of 72.2% on in-training residency exams compared to 57.7% for ChatGPT-3.5, a reminder that model version matters enormously and older screenshots of “ChatGPT can’t do X” may already be outdated.
A Safer Prompting Framework for Clinical Use
The good news from the Mount Sinai research is that prompting technique itself measurably changes the error rate. In a second round, researchers added a one-line caution to the prompt, reminding the AI that the information provided might be inaccurate, and with that added prompt, errors were reduced significantly. The instruction told the model to use only clinically validated information and acknowledge uncertainty instead of speculating further, aiming to encourage it to flag dubious elements rather than generate unsupported content. Practically, that means building a habit of appending something like “if any detail here seems inconsistent or unverifiable, say so rather than explaining it” to clinical prompts.
Two other habits follow from the evidence above. First, never enter identifiers, patient names, dates of service, or anything from the 18 HIPAA identifiers as listed by the Department of Health and Human Services into standard ChatGPT; if those are included, it would be a HIPAA violation, but if you don’t have any of those, you’re fine. Second, treat every diagnostic or reference-heavy output as a draft that requires independent verification, not a finished answer, particularly given hallucinated reference rates as high as 28.6% to 91.4% in one systematic-review analysis. Readers interested in how clinical decision support tools are supposed to handle this differently may want our piece on AI clinical decision support in US hospitals.
ChatGPT for Clinicians: What Changed and What Didn’t
OpenAI’s answer to the adoption numbers above was a dedicated product. OpenAI made ChatGPT for Clinicians free for verified U.S. physicians, nurse practitioners, and pharmacists, supporting clinical care, documentation, and research. It includes trusted clinical search with citations based on medical sources, and documentation support for drafting notes, referrals, prior authorization letters, and patient instructions for clinician review. Before release, physician advisors tested 6,924 conversations in their daily work across clinical care, documentation, and research, and overall physicians rated 99.6% of responses as safe and accurate.
That number deserves a caveat that a clinician-focused blog put well: OpenAI’s physician advisors tested 6,924 conversations before launch and rated 99.6% of responses safe and accurate, but those numbers are worth remembering for what they are, the company’s own internal testing, not independent verification. The HIPAA picture also hasn’t fully changed by default. Entering any protected health information into the free tiers is a HIPAA violation, full stop, and turning off chat history or using temporary chat doesn’t fix this; treat this tool exactly like standard ChatGPT unless your organization has a specific BAA authorization in place. HIPAA support is available through a business associate agreement for eligible accounts, and conversations will not be used to train models, OpenAI said. The upgrade is real, but it doesn’t erase the need for the same discipline described above.
Quick Reference: What Actually Works vs. What’s Risky
| Prompt Use Case | Risk Level | Why |
|---|---|---|
| Drafting prior auth / referral letter templates (no patient identifiers) | Low | Formatting task, clinician reviews before sending |
| Rewriting discharge instructions in plain language | Low | Translation task, not a diagnostic judgment |
| General literature orientation on a topic | Moderate | Reference hallucination rates of 28.6% to 39.6% for GPT models mean every citation needs independent verification |
| Pasting real patient notes, names, or chart excerpts into standard ChatGPT | High | Data leaves the health system and is technically a HIPAA breach |
| Asking for a differential diagnosis as a final answer | High | Weakest documented performance area at 60.3% accuracy |
| Trusting an answer built on a possibly-wrong detail you typed | High | Hallucination rates of 50% to 82% when fabricated details are present |
FAQ
Is ChatGPT HIPAA compliant for use with patient information?
Standard, public ChatGPT is not HIPAA compliant because OpenAI does not sign a Business Associate Agreement for that version, and once patient information is typed in, it has left the health system’s control, which is technically a data breach. OpenAI’s newer ChatGPT for Clinicians and ChatGPT for Healthcare products offer a BAA pathway for eligible accounts, but that coverage doesn’t start automatically; you need specific authorization to sign a BAA before coverage begins, and until then you should treat it exactly like standard ChatGPT.
How often does ChatGPT give wrong medical information?
It depends heavily on the task and framing. A Mount Sinai study found hallucination rates ranged from 50% to 82% across six different AI chatbots when a fabricated medical detail was slipped into a clinical vignette, while a separate JMIR study found ChatGPT invented 28.6% to 39.6% of academic references depending on the model version.
What’s the safest way to prompt ChatGPT for clinical use?
Mount Sinai researchers found that adding a one-line caution telling the model the input might contain errors, and instructing it to flag uncertainty rather than speculate, significantly cut hallucination rates. Combine that with never entering identifiable patient details and always verifying clinical claims against a primary source before acting on them.
Should doctors use ChatGPT for differential diagnosis?
Use it as a brainstorming aid, not a decision-maker. A study using Merck Manual clinical vignettes found ChatGPT’s overall accuracy across diagnosis and management tasks was 71.7%, with its weakest performance, 60.3%, on generating an initial differential diagnosis.




0 Comments