Quick answer: the sharpest evidence on AI clinical documentation errors and mental health risk is not about the software making things up. Across 20,302 primary care notes in JAMA Psychiatry (January 21, 2026), visits documented by an ambient AI scribe recorded more neuropsychiatric symptoms than unscribed visits, and were associated with lower odds of any psychiatric intervention (adjusted odds ratio 0.83; 95% CI, 0.72 to 0.95). Depression diagnosis codes fell to 9% of AI-scribed visits versus 12% of unscribed ones. Visits documented by a human scribe showed no such drop. The record got fuller while the care got quieter.

By Matthew Sexton, LCSW, NATC — a Licensed Clinical Social Worker and Certified Narcissistic Abuse Treatment Clinician who runs a fee-for-service practice and designed the documentation side of VibeCheck.luxury.

An ambient AI scribe listens to a visit and drafts the note. That pitch is obvious and mostly honest. A clinician stops typing and starts looking at the person across from them. I have argued before that the EHR turned clinicians into unpaid transcriptionists, and any tool that hands that hour back is doing something real. I am not here to smash it.

I am here about a finding almost nobody is discussing, because it does not fit the story we have all agreed to tell about AI. That story is about a robot that makes things up. What the researchers found is stranger, and worse.

Fuller notes came with less clinical action

Researchers compared four matched groups of roughly 5,076 primary care notes each: visits documented by an ambient AI scribe, visits documented by a human virtual scribe, contemporaneous visits with no scribe, and prior-year visits with no scribe. Then they looked at what the clinician actually did afterward.

AI-scribed notes captured more symptom language across all six of the research domains the team measured. On the negative valence domain, which covers the fear, loss, and distress vocabulary that carries most depressive presentation, AI-scribed notes scored 2.05 against 1.57 for unscribed visits. They also ran close to twice as long, 13,629 characters versus 7,932 for contemporaneous unscribed notes.

By any documentation metric a health system tracks, those are better notes. They are more complete and more defensible in an audit.

And the depression diagnosis codes came in at 447 of the AI-scribed visits, 9%, against 604 of the contemporaneous unscribed visits, 12%. On the composite measure of any psychiatric intervention, AI-scribed visits landed at 14% against 17% for both the human-scribed and unscribed groups. Human scribes, doing the same job with a person instead of a model, showed no significant difference from unscribed visits at all. Whatever this is, it is specific to the AI.

Documentation methodNotes (matched)Depression diagnosis codeAny psychiatric intervention
Ambient AI scribe~5,076447 (9%)14%
Human virtual scribe~5,076No significant difference from unscribed17%
Contemporaneous, no scribe~5,076604 (12%)17%
Prior-year, no scribe~5,076

Adjusted odds of any psychiatric intervention for AI-scribed visits versus unscribed visits: 0.83 (95% CI, 0.72 to 0.95). An em dash marks a figure this article does not carry. Source: JAMA Psychiatry, January 21, 2026.

Two bar charts of matched primary care visits. Depression diagnosis code: AI-scribed visits 9 percent, unscribed visits 12 percent. Any psychiatric intervention: AI-scribed visits 14 percent, human-scribed visits 17 percent, unscribed visits 17 percent. The AI-scribed arm is the only one that falls below the others on both measures, while human-scribed visits showed no significant difference from unscribed visits. The chart reports documentation and treatment rates only and makes no claim about patient outcomes.
Figure 1. Across both measures, the AI-scribed arm is the only one that drops — and the human-scribed arm, doing the same job with a person, does not move. These are documentation and treatment rates, not patient outcomes. Source: JAMA Psychiatry, “Psychiatric Documentation and Management in Primary Care With Artificial Intelligence Scribe Use,” published online January 21, 2026.

Two honest caveats. This is a correlational cohort study rather than a randomized trial, so the right verb is associated with and not caused, and nobody has established the mechanism. Researchers did not find a broken note. They found a documentation system that produced a more impressive record alongside less recorded action.

I have a hypothesis, and I will label it as exactly that. A note that already contains the symptoms can feel finished. When the low mood, the sleep disruption, and the appetite change are all sitting there in tidy paragraphs, the sensation of having addressed something may arrive without the addressing. That is a guess about attention, not a finding. What is not a guess is the gap between what got documented and what got done.

Every accuracy study was run somewhere else

Three rigorous evaluations of ambient AI scribe accuracy have been published in the last two years, and they are genuinely good work. A Frontiers in Artificial Intelligence evaluation (October 22, 2025) put 97 encounters through 388 paired blinded reviews and found hallucinations in 31% of AI notes against 20% of physician-authored notes (p = 0.01). A UC Davis pragmatic pilot in JMIR Medical Informatics (April 17, 2026) ran 7,545 AI-generated notes across 31 physicians and scored 356 in detail. A UC Berkeley and UCSF team analyzing clinician safety feedback, published at Machine Learning for Health (December 1, 2025), worked from 50,123 encounters across 470 physicians.

Now look at the specialties. Family practice. Internal medicine. Cardiology. OB/GYN. Orthopedic surgery. Otolaryngology. Dermatology. Pediatrics. Endocrinology. Sleep medicine.

Not one of those studies included psychiatry, psychology, or psychotherapy.

Meanwhile the American Psychological Association’s 2025 Practitioner Pulse Survey of 1,742 practitioners, reported in the APA Monitor (March 1, 2026), found 22% of psychologists already using AI for note-taking and dictation, 29% using AI at least monthly, and 56% having used it at least once. Two thirds, 67%, remain concerned about data breaches, and at least 60% are worried about inaccurate outputs.

So behavioral health is documenting with tools whose error profile has never been tested on a behavioral health encounter. Clinicians using them are not being reckless. They are doing what the evidence appeared to license, and the evidence was collected in a cardiology clinic.

Vendors sell these tools horizontally, as if a knee exam and a trauma intake were the same recording with different words in it. A knee exam has a right answer sitting in the room. A therapy hour has silences that mean something, a person circling a topic three times before landing on it, and one sentence about not wanting to be here that arrives sideways and gets walked back within ten seconds.

The AI clinical documentation error you cannot catch is the one that is not there

Public conversation about AI documentation fixates on fabrication. In the data, fabrication is not the main event.

In the UC Davis pilot, accidental omissions were the most frequent error at 18% of scored notes, 64 of 356. Hallucinations came second at 11.5%, 41 of 356. Accidental inclusions ran 9.3%, and bias was rare at 1.1%. Omission outran fabrication by roughly 1.6 to 1.

Hold both halves of that study, because the calibration matters. That same pilot found 94.7% of notes, 337 of 356, free of significant errors. These tools are not broadly broken. It also found a tail. In 5.3% of notes, 19 of 356, reviewers rated an error as posing a risk of serious harm if left uncorrected. The authors’ conclusion was blunt, that careful clinician review of notes remains imperative.

That conclusion is the whole safety model, and it is worth saying out loud what it asks of a person. A 5.3% serious-harm-if-uncorrected rate is tolerable only if the correction reliably happens. And the hardest error to correct is the missing one. Catching a hallucination means noticing something wrong on the page, which is a reading task. Catching an omission means recalling what was said in a fifty-minute conversation and registering that it is absent, which is a memory task performed at the end of a full day by someone who has already had six of those conversations.

In general medicine, an omission is often a missing lab value that resurfaces on the next panel. In mental health, the omission is the risk statement. It is the passive ideation mentioned once and never returned to, the collateral detail from a family member, the sentence that would have changed the disposition. Nothing in the record shows you it is gone.

I could not find a single study measuring whether the catch actually happens. Every safety argument for these tools in behavioral health rests on a step nobody has tested.

— Matthew Sexton, LCSW, NATC

Who said that, and how were they saying it

Two smaller findings deserve more attention than they got.

The UC Berkeley and UCSF team categorized the safety concerns clinicians voluntarily submitted while using an ambient scribe. Of the concerns clinicians did report, 18.5% involved medication names, dosages, or instructions recorded incorrectly, and 9.5% involved speaker misattribution, meaning confusing who said what or attributing a caregiver’s account to the patient. Physicians separately described name confusion, misgendering, and speaker-recognition failures when more than one person attended the visit.

An important guardrail on that paper. Only 0.93% of encounters carried any free-text feedback at all, and 415 of the 470 physicians submitted none. That is a reporting rate, not an error rate, and the authors say plainly that further study is needed to contextualize the absolute degree of risk. Nobody should read it as “18.5% of notes had medication errors.” It describes the texture of what worried clinicians, not how often they were right.

Even so, look at where misattribution lands. In a knee exam it is a formatting nuisance. In a family session, a couples session, a collateral call, or an intake where a parent fills in the history, who said a thing is the clinical content.

A second finding is about how the underlying transcription behaves. A peer-reviewed study at the ACM Conference on Fairness, Accountability, and Transparency, “Careless Whisper” (June 2024), found that 1.4% of audio segments produced an entirely fabricated hallucination sequence, and that 38% of those hallucinations contained explicit harms, including 19% that perpetuated violence and 8% that invented false authority. One disparity stayed with me. Hallucination rates ran 1.7% for speakers with aphasia against 1.2% for controls. Models fabricate more when speech is disfluent, halting, or full of pauses.

Be precise about what that study is. It analyzed recorded speech from a research corpus, not clinical encounters and not therapy sessions. Its relevance here is an inference, and I want it labeled as one. Transcription models of this class sit underneath commercial medical documentation, and their worst failure mode appears on exactly the speech patterns that fill a psychiatric hour. Long pauses. Halting delivery. Flat or pressured speech. Thought that arrives out of order. Nobody has measured what that does to a therapy transcript, and that is the point.

Adoption finished before governance started

By 2025, nearly two-thirds (62%) of the 2,784 US hospitals running Epic had adopted an ambient AI documentation tool, according to research from Emory’s Rollins School of Public Health published in The American Journal of Managed Care and summarized in the institutional release (February 4, 2026). Adoption was substantially higher at nonprofit hospitals than at for-profit ones.

The first national governance framework for clinical AI arrived on September 17, 2025. The Joint Commission and the Coalition for Health AI released initial guidance calling for policies, local validation, and monitoring “to be flexibly interpreted and integrated into existing or new processes as deemed appropriate for the context of any organization.” A voluntary AI certification for the Joint Commission’s more than 22,000 accredited and certified organizations was promised in future playbooks.

Read those two dates together. The guidance is non-binding, it defers validation to each organization’s discretion, its operational playbooks were still forthcoming when it was announced, and it landed in the same year the adoption survey was measuring. Governance did not fail to keep up. It was never in the race.

State law is moving faster than professional standards here, and it lands on the license rather than on the vendor, which I have written about at length in the state AI patchwork post. The professional body closest to therapists has issued ethical guidance, not enforceable standards. So the practical answer to “who is responsible when the note is wrong” is the clinician, every time, using a tool whose failure modes were characterized in a different specialty.

That is the system critique, and it needs no villain. It needs a procurement process that bought thoroughness by word count, a regulatory floor that arrived after deployment, and a research literature that studied the specialties with the loudest budgets first.

I have written separately about who controls the scribe and who keeps the recording; that post is about control of the tool, and this one is about whether the tool was ever tested on a mental health encounter in the first place.

What a real safety net for mental health documentation would look like

Not “review the note.” Everyone already says that, and it is doing no work. None of what follows should land on the individual clinician, and every bit of it currently does.

  • Test the tool on the encounter type you actually run. If a vendor cannot show you accuracy data from a behavioral health session, they have not tested it there. Ask, and write down the answer.
  • Audit for omission specifically. Pull ten sessions where you remember a risk statement, and check whether it made it into the note. Fabrication is visible. Absence is not, and only a deliberate check finds it.
  • Treat multi-speaker sessions as a separate risk category. A tool validated on one voice in a room has not been validated on three.
  • Watch the gap between documentation and action. If your notes are getting longer while your diagnoses and interventions are not tracking with them, that is the JAMA Psychiatry pattern showing up in your own practice. It is measurable without a research team.
  • Keep the clinician as the author, not the approver. There is a real difference between drafting a note and signing off on one, and the entire error-catching argument depends on which of those is happening at six o’clock. That is the whole of what clinician-in-the-loop is supposed to mean.

None of this is a reason to throw the tools out. I would rather a clinician look at a person than at a keyboard, and the burden these tools relieve is not imaginary. It is a reason to stop pretending the evidence base covers ground it has never touched. When we designed the documentation side of VibeCheck.luxury, the rule was that the clinician owns and reviews every note, because the safety model has no other floor.

Covered was never the same as cared for. A complete note was never the same as a treated patient. The 20,302-note study is the first real measurement of the distance between those two things in mental health, and the distance was wider than anyone expected. Somebody should measure it in a therapy room next. Until then, the clinician reading a draft at the end of a long day is the only safety system these tools have. If you want to talk through what that looks like in a specific practice, book a call.

FAQ

Do AI scribes make more errors than clinicians writing their own notes? On hallucination, yes, in the one head-to-head study available. A Frontiers in Artificial Intelligence evaluation (October 22, 2025) ran 388 paired blinded reviews across 97 clinical encounters and found hallucinations in 31% of AI ambient-scribe notes versus 20% of physician-authored notes (p = 0.01). AI notes also scored higher on thoroughness and lower on succinctness. The specialties studied were general medicine, pediatrics, OB/GYN, orthopedic surgery, and adult cardiology. No mental health or behavioral health specialty was included, so this figure does not describe therapy or psychiatric documentation.

What is the most common type of AI documentation error? Omission, not fabrication. In a UC Davis pragmatic pilot published in JMIR Medical Informatics (April 17, 2026), physicians scored 356 AI-drafted notes and found accidental omissions in 18% (64 of 356), hallucinations in 11.5% (41 of 356), and accidental inclusions in 9.3% (33 of 356). Bias was rare at 1.1%. That matters because omission is the error a reviewer is least equipped to catch: spotting it requires remembering what was said and noticing it is absent, rather than seeing something visibly wrong on the page.

How risky are the errors that do get through? The same UC Davis pilot found that 5.3% of AI-drafted notes (19 of 356) contained an error rated as posing a risk of serious harm if left uncorrected. The authors concluded that careful clinician review of notes remains imperative. Read the other side of that figure too: 94.7% of notes were free of significant errors. The honest reading is that these tools are mostly accurate with a tail that matters, and that the entire safety model depends on the review actually happening.

Has anyone tested AI scribes on therapy or psychiatric sessions? Not in the peer-reviewed accuracy literature. The rigorous studies of ambient AI scribe error rates were conducted in family practice, internal medicine, cardiology, OB/GYN, orthopedics, otolaryngology, dermatology, pediatrics, endocrinology, and sleep medicine. None included psychiatry, psychology, or psychotherapy. Meanwhile the American Psychological Association’s 2025 Practitioner Pulse Survey of 1,742 practitioners found 22% already using AI for note-taking and dictation, and 29% using AI at least monthly in practice.

Sources

  1. Psychiatric Documentation and Management in Primary Care With Artificial Intelligence Scribe Use, JAMA Psychiatry, published online January 21, 2026. The 20,302-note cohort, the adjusted odds ratio of 0.83 for any psychiatric intervention, depression codes at 9% versus 12%, the composite intervention outcome at 14% versus 17%, note length of 13,629 versus 7,932 characters, and the absence of any effect for human scribes.
  2. Quality of Clinical Notes Created by Ambient Listening Generative AI: Pragmatic Prospective Pilot Study, JMIR Medical Informatics, UC Davis, April 17, 2026. The 5.3% serious-harm-if-uncorrected rate, the 94.7% error-free rate, and the error-type breakdown of 18% omissions, 11.5% hallucinations, 9.3% accidental inclusions, and 1.1% bias.
  3. Assessing the quality of AI-generated clinical notes: validated evaluation of a large language model ambient scribe, Frontiers in Artificial Intelligence, October 22, 2025. The 31% versus 20% hallucination comparison across 388 paired blinded reviews, the thoroughness and succinctness scores, and the specialty list.
  4. Patient Safety Risks from AI Scribes: Signals from End-User Feedback, Machine Learning for Health 2025 Findings Track, UC Berkeley and UCSF, December 1, 2025. The 50,123-encounter corpus, the 0.93% feedback-rate caveat, and the cluster breakdown including 18.5% medication errors and 9.5% misattribution.
  5. Careless Whisper: Speech-to-Text Hallucination Harms, ACM Conference on Fairness, Accountability, and Transparency, June 2024. The 1.4% segment hallucination rate, the 38% harm rate within hallucinations, and the 1.7% versus 1.2% disparity for disfluent speech.
  6. Ambient AI Tool Adoption in US Hospitals and Associated Factors, The American Journal of Managed Care, reported via the Emory institutional release, February 4, 2026. Nearly two-thirds (62%) of the 2,784 US hospitals using Epic had adopted an ambient AI documentation tool by 2025.
  7. AI in the therapist’s office: Uptake increases, caution persists, APA Monitor on Psychology reporting the 2025 Practitioner Pulse Survey (n = 1,742), March 1, 2026. The 22% note-taking figure, 29% monthly use, 56% ever-use, and 67% data-breach concern.
  8. Joint Commission and Coalition for Health AI (CHAI) Release Initial Guidance to Support Responsible AI Adoption Across U.S. Health Systems, Coalition for Health AI, September 17, 2025. The voluntary, locally interpreted framework and the forthcoming certification for more than 22,000 accredited organizations.

Disclaimer

This article is for educational and informational purposes only. It does not constitute medical, clinical, legal, or therapeutic advice, and reading it does not create a therapist-client relationship with Matthew Sexton, LCSW or Mental Wealth Solutions, Inc. Although the author is a licensed clinical social worker, the content in this article is not clinical assessment, diagnosis, or treatment.

The studies described here measured groups of clinicians and notes, mostly outside behavioral health, and their findings describe populations rather than any individual practice, product, or reader. AI documentation tools differ substantially from one another, change between software versions, and are governed by rules that vary by state and by professional board and may change after this article is published. Nothing here evaluates a specific product or tells you whether to use one. Before adopting or continuing to use any AI documentation tool in clinical care, confirm the requirements with your licensing board, your malpractice carrier, your compliance team, and qualified counsel, and review the vendor’s own accuracy and data terms directly.

If you are in immediate emotional crisis, you can reach the 988 Suicide & Crisis Lifeline by calling or texting 988 (US). If you are experiencing domestic violence or are in physical danger, contact the National Domestic Violence Hotline at 1-800-799-7233 or visit thehotline.org. In a life-threatening emergency, call 911.

Frequently asked questions.

Do AI scribes make more errors than clinicians writing their own notes?
On hallucination, yes, in the one head-to-head study available. A Frontiers in Artificial Intelligence evaluation (October 22, 2025) ran 388 paired blinded reviews across 97 clinical encounters and found hallucinations in 31% of AI ambient-scribe notes versus 20% of physician-authored notes (p = 0.01). AI notes also scored higher on thoroughness and lower on succinctness. The specialties studied were general medicine, pediatrics, OB/GYN, orthopedic surgery, and adult cardiology. No mental health or behavioral health specialty was included, so this figure does not describe therapy or psychiatric documentation.
What is the most common type of AI documentation error?
Omission, not fabrication. In a UC Davis pragmatic pilot published in JMIR Medical Informatics (April 17, 2026), physicians scored 356 AI-drafted notes and found accidental omissions in 18% (64 of 356), hallucinations in 11.5% (41 of 356), and accidental inclusions in 9.3% (33 of 356). Bias was rare at 1.1%. That matters because omission is the error a reviewer is least equipped to catch: spotting it requires remembering what was said and noticing it is absent, rather than seeing something visibly wrong on the page.
How risky are the errors that do get through?
The same UC Davis pilot found that 5.3% of AI-drafted notes (19 of 356) contained an error rated as posing a risk of serious harm if left uncorrected. The authors concluded that careful clinician review of notes remains imperative. Read the other side of that figure too: 94.7% of notes were free of significant errors. The honest reading is that these tools are mostly accurate with a tail that matters, and that the entire safety model depends on the review actually happening.
Has anyone tested AI scribes on therapy or psychiatric sessions?
Not in the peer-reviewed accuracy literature. The rigorous studies of ambient AI scribe error rates were conducted in family practice, internal medicine, cardiology, OB/GYN, orthopedics, otolaryngology, dermatology, pediatrics, endocrinology, and sleep medicine. None included psychiatry, psychology, or psychotherapy. Meanwhile the American Psychological Association's 2025 Practitioner Pulse Survey of 1,742 practitioners found 22% already using AI for note-taking and dictation, and 29% using AI at least monthly in practice.

If you're the therapist here.

Your clients get 4 sessions a month. The other 26 days they're on their own. VibeCheck is the between-session companion that carries those days back to you — clients check in daily, and you walk in already knowing what kind of week it was. Built by Matthew Sexton, LCSW, NATC.