Home  |   Subscribe  |   Resources  |   Reprints  |   Writers' Guidelines

E-News Exclusive

CDI’s Battle Against AI Hallucinations, Context Errors

By Elizabeth S. Goar

Following a checkup with his primary care physician, who used ambient listening tools to capture encounter notes, Medicomp System CMO Jay Anders, MD, was surprised to see diagnoses of atherosclerotic heart disease and diabetic complication listed in the clinical documentation. While he does have high blood pressure, he has never been diagnosed with heart disease, nor is there anything in his history that would account for the AI scribe inserting anything related to diabetes.

It’s a real-world example of how AI-generated documentation can be polished, structured, and compliant on its face, while silently misrepresenting what happened during the encounter. “That gap,” Anders says, “is where hallucinations live.”

Making the Distinction
Hallucinations occur when an AI system generates information that was never present in the source material. Clinical context errors can be more subtle. The information may exist somewhere in the record, but the AI interprets or applies it incorrectly.

“In my experience, clinical context errors can be particularly difficult to detect during human review because the individual facts may appear accurate on the surface. The problem is not necessarily that the AI invented information; it’s that it misunderstood the clinical context surrounding it,” says David Cohen, Greenway Health’s chief product and strategy officer.

An underlying issue is AI’s inability to account for all atypical presentations, adds Core-CDI CEO Glenn Krauss, RHIA, CCS, CCDS, C-CDI. It fills gaps using the flotsam and jetsam that may be floating around the patient’s medical history or by creating a diagnosis from whole cloth. This puts clinical documentation integrity (CDI) specialists in the uncomfortable position of being asked to infer clinical intent.

“Coder and CDI rules say we’re not allowed to infer a diagnosis, and that we can’t code from an inferred diagnosis. We must contact the physician. That's our code of ethics. So why are we letting software do it?” he asks.

The Data Disconnect
Part of the problem is the trust clinicians place in ambient AI. A study on AI documentation accuracy by SmallPDF found that 74% of health care professionals act on AI summaries without checking the source, and 91% consider AI document summaries accurate.

That trust may be misplaced, says Melinda McGuire, MS, MAS, RHIA, founder of Highland eHealth and program director of HIT/HIM with Baker College. “When vendors tout ‘98% accuracy,’ health system executives often assume that means 98% of finished notes are error-free. Peer-reviewed research shows a clear split: Sentence-level error rates look low, but document-level error occurrence is high.”

One 2025 study evaluating large language model (LLM)-generated clinical note summarization found a sentence-level hallucination rate of 1.47% and classified 44% of detected hallucinations as “major” errors. Another test of two commercial ambient scribe products found that 70% of draft notes contained at least one error and averaged 2.9 errors per note.

“Consider what a 1.47% sentence-level hallucination rate looks like at scale: A health system generating 10,000 AI-assisted notes a week would see a meaningful, recurring stream of erroneous clinical data points entering the EHR, well before accounting for clinician catch rates at sign-off,” McGuire says.

Noting that it’s difficult to establish a single hallucination rate, Stephanie Smith, MD, vice president of clinical intelligence at Accuity, cites a 2026 study that found omissions in 18% of reviewed notes, hallucinations in 11.5%, erroneous inclusions in 9.3%, and major errors in 5.3%. Meanwhile, 94.7% of the evaluated notes were considered free of significant errors.

“That combination is important,” she says. “AI-generated documentation can perform very well, overall, while still producing a relatively uncommon error with potentially significant consequences. In medicine, the denominator matters, but so does the severity of the miss.”

Catching Errors
AI-assisted documentation can create coding and compliance risks, including misinterpretation of historical conditions, unsupported high-severity diagnoses, missing clinical context, and inappropriate query prompts. To counter these risks, CDI specialists must become compliance gatekeepers, verifying clinical support, confirming diagnoses were actively managed during the encounter, comparing notes against objective data, and ensuring AI-generated recommendations remain compliant and audit defensible.

“Catching these failure modes requires moving beyond informal spot checking,” McGuire says. “Health systems must equip their CDI teams with updated audit tools, specific AI error taxonomies, and clear governance protocols.”

Krauss adds that AI-generated documentation must be viewed with a critical eye so the right questions can be asked. “If you … just rely on the output of a suggested query, that’s a problem. That’s going to cause denials,” he says.

Smith suggests CDI professionals “focus less on whether the sentence sounds right and more on whether the record actually supports the statement in context.”

She adds, “As AI-generated documentation becomes more fluent and more accurate overall, the human reviewer may be asked to read an increasing volume of plausible, mostly correct text to find an increasingly uncommon error. The error does not necessarily announce itself. It may be one incorrect sentence surrounded by 10 accurate ones.”

Governance Is Key
CDI professionals are being asked to govern content they weren't trained to evaluate and with little support, according to Anders, who says mandatory provider review policies are still catching up, and auditing frameworks for AI-authored notes “barely exist.”

“The LLMs themselves are not self-correcting. Training on user input can entrench errors as readily as it resolves them,” he adds.

Krauss notes that LLMs have improved significantly, but still fall prey to garbage-in, garbage-out. “It’s only as good as the questions [clinicians] ask … If you don’t ask the right questions, you’re not going to get the right answers,” he says.

That’s why Yaroslav Dokuchaev, CEO and founder of DICO by Expert Radiotech, considers AI hallucinations primarily a governance problem. He says trust comes not from convincing AI-generated text, but “from knowing exactly which parts of the record were generated by AI, which were verified by a clinician, and who is ultimately accountable for the final document.”

An effective governance framework begins with a simple principle: AI may assist with documentation, but accountability remains human, Cohen says. Effective AI documentation governance should also include defined use cases, mandatory human oversight, routine auditing for errors and hallucinations, and clear responsibility among clinical, compliance, and technology stakeholders. Organizations should also assess vendor validation practices, performance measures, system limitations, and processes for identifying, reporting, and resolving documentation issues.

“The goal is to create a system of controls that makes errors detectable and manageable, while ensuring that documentation remains accurate, clinically trustworthy, and compliant,” Cohen says. “That balance is what allows health care organizations to realize the productivity benefits of AI without sacrificing the integrity of the medical record.”

— Elizabeth S. Goar is a freelance health care writer based in Wisconsin.