how automated healthcare fails, how you'd know, and what to do at each tier — every claim sourced, reviewed continuously
Researchers and an AP investigation found OpenAI's Whisper speech-to-text model inserting fabricated sentences, and a Whisper-based clinical scribe used by over 30,000 clinicians erased the source audio, removing the way to check.
A peer-reviewed FAccT 2024 study found about 1% of Whisper transcriptions contained whole phrases or sentences absent from the audio, 38% of which included explicit harms such as violent content, false associations or implied false authority; rates were higher for speakers with aphasia and long pauses. In October 2024 the Associated Press reported that over 30,000 clinicians and 40 health systems had adopted a Whisper-based tool from Nabla, used for about 7 million visits, despite OpenAI's warning against use in high-risk domains. Nabla's tool erased the original audio for data-safety reasons, so transcripts could not be compared with the recording; the company said providers must edit and approve notes.[1,2]
No specific patient harm documented in the cited sources; risk is erroneous content in medical records.
All incidents · Models & agents · Human handoff
Information only, not advice. FailSystems is an aggregation and synthesis of published sources. It is not consulting, engineering, legal, regulatory or medical advice, and using it creates no professional relationship. Health systems are complex and no approach fits every organisation: anything you adopt is your own decision, at your own risk, and should be checked against the current official sources and by qualified people who know your setting. Full disclaimer.