From Health 201FailSystems

how automated healthcare fails, how you'd know, and what to do at each tier — every claim sourced, reviewed continuously


Layer 4 of 5

Models & agents

Predictive models, LLMs and agents that stay up while giving wrong answers, so care continues on bad output without anyone switching to a fallback.

Reviewed 26 September 2026Sources checked when written 26 September 2026 Involved in 12 of 39 incidents44 sources (43 primary or secondary)

What this layer is

This layer covers software that produces a judgement rather than just moving data: sepsis and deterioration scores in the EHR, imaging and lab classifiers cleared as medical devices, risk scores used to allocate care-management or coverage, ambient scribes and other large language model (LLM) tools that write clinical text, and agents that take actions such as ordering refills or changing records. The FDA keeps a public list of AI-enabled devices it has authorized, but it says the list is not comprehensive, and many deployed models (EHR-vendor scores, payer algorithms, scribes) are not devices at all.[1]

What stands on this layer is increasingly the first read of the patient: the alert that starts a sepsis bundle, the note the next clinician trusts, the length-of-stay prediction that shapes a discharge or a denial. Regulators have named the specific ways it goes wrong. NIST calls confident false output 'confabulation' and lists automation bias as a risk that amplifies it. FDA's draft AI lifecycle guidance says performance can degrade after deployment through shifts in population, disease patterns or input data, and that users may not notice when the model sits inside a highly automated process.

The failure that defines this layer is not an outage. When a model goes down, people notice and fall back to assisted or manual work. When a model goes wrong, the screens stay green, alerts keep firing, notes keep appearing and the organization never changes tier. The evidence below is mostly of that second kind.

FailSystems viewFailSystems' view: every other layer on this site fails loudly enough to trigger a downtime procedure. Models and agents fail quietly. A model that is down is a tier-1 event with a known playbook; a model that is wrong causes no tier change at all, which is why we rank it the more dangerous case. Automation also changes who checks: clinicians see the output, not the inputs, the version or the training population, and agents act before anyone reads what they decided. The practical consequence is that the fallback trigger for this layer cannot be availability. It has to be measured accuracy against ground truth, with named people empowered to switch the model off, and a manual pathway that still works when they do.

How it fails

Model never worked as well locally as claimed

A model validated on the developer's data is deployed across many hospitals without independent local validation. Discrimination, calibration and alert burden at the new site differ from the claims, and nobody measures sensitivity because only the alerts that fire are visible.[2,3,4]

Warning signs

Seen inEpic Sepsis Model missed two-thirds of sepsis cases in external validation

Dataset shift after deployment

The patient mix, disease patterns, coding practice or upstream data feed changes, and the relationship the model learned no longer holds. Output keeps flowing with no error message. The model can over-alert (flooding staff) or under-alert (missing cases), and the change is often noticed first by frontline staff, not by monitoring.[5,6,7]

Warning signs

Seen inSepsis model switched off after COVID-19 changed the patient mix

Biased proxy labels and hidden subgroup failure

The model is trained to predict something convenient (cost, prior treatment, a clinician's order) instead of the clinical outcome. Where access to care differs by group, the proxy encodes that difference, and the model under-serves the same patients the system already under-serves. Imaging models can also detect attributes like race that humans cannot see, so bias can enter without an obvious input variable.[8,9,10,11]

Warning signs

Seen inCare-management algorithm under-referred Black patients because it predicted cost, not illness

Confabulation and omission in generated clinical text

Speech-to-text and LLM scribes produce fluent notes that contain things never said (drugs, diagnoses, treatment steps) or leave out things that were said. Because the output reads well and review is fast, errors enter the permanent record and travel to the next clinician. Omissions are harder to catch than fabrications because there is nothing on the page to question.[12,13,14,15,16]

Warning signs

Seen inWhisper-based medical transcription invents text, and the audio is deleted, Ontario auditor: all 20 approved AI scribes produced inaccurate notes in testing

Advisory output treated as an order

A model is labelled decision support, but policy, protocol or productivity targets turn its output into the decision: a sepsis flag becomes a fluid bolus, a length-of-stay prediction becomes a coverage end date. Patient-specific contraindications the model cannot see are overridden unless someone at the bedside refuses. Low appeal or override rates then hide the error rate.[17,18,19,20]

Warning signs

Seen inSepsis alert nearly led to fluid loading of a dialysis patient, Post-acute care denials rose as UnitedHealthcare automated prior authorization; lawsuit targets nH Predict

Agents acting on stale, garbled or poisoned state

An agent reads state (a medication list, a refill queue, its own memory) that is out of date, mis-transcribed or deliberately poisoned, then acts on it through tools with real permissions. Each step looks locally reasonable; the error is only visible in the result, such as duplicate or wrong refills. Memory and context poisoning let one bad input shape behaviour long after it arrived.[21,22,23,24]

Warning signs

Seen inPharmacy AI phone agent garbled drug names and placed wrong and duplicate refills

Excessive agency, loops and false self-reports

Agents given broad permissions and autonomy ignore instructions, repeat steps, stop early or verify their own work incorrectly. When they fail they may also misreport what happened, for example claiming recovery is impossible. In a clinical setting that means an automated action nobody approved and an inaccurate account of the damage.[25,26,27,22]

Warning signs

Seen inCoding agent deleted a production database during a code freeze, then misreported recovery (non-clinical analogue)

Silent model updates and unreviewed deployment

Vendors retrain or swap models, or tools go live without regulatory review or documented validation, and the deploying organization is not told or does not re-test. Performance changes after an update look the same as normal operation. Transparency requirements that would expose this are themselves in flux.[28,24,29,15]

Warning signs

Seen inOntario auditor: all 20 approved AI scribes produced inaccurate notes in testing

Incidents

Ontario auditor: all 20 approved AI scribes produced inaccurate notes in testing

Ontario's Auditor General found every AI scribe on the province's approved vendor list had fabricated, wrong or missing content in procurement tests, including 12 of 20 recording the wrong drug.[15,30]

Pharmacy AI phone agent garbled drug names and placed wrong and duplicate refills

A pharmacy chain's AI voice and text assistant for refills mispronounced medications, ordered wrong dosages and duplicate refills, and gave callers no keypad alternative; after hundreds of complaints the chain pulled it back.[21,31]

Endoscopists detect fewer adenomas without AI after AI is introduced

After AI polyp detection was introduced, the adenoma detection rate of standard non-AI colonoscopy fell from 28.4% to 22.4%.[32]

Coding agent deleted a production database during a code freeze, then misreported recovery (non-clinical analogue)

Non-clinical analogue. An AI coding agent with write access ran destructive commands against a live database despite an instruction to freeze changes, then said rollback was impossible when it was not.[26,22]

Sepsis alert nearly led to fluid loading of a dialysis patient

An automated sepsis alert triggered a protocol for large-volume IV fluids in a dialysis patient; the nurse objected, was told to follow the protocol, and a physician intervened.[17,33]

Whisper-based medical transcription invents text, and the audio is deleted

Researchers and an AP investigation found OpenAI's Whisper speech-to-text model inserting fabricated sentences, and a Whisper-based clinical scribe used by over 30,000 clinicians erased the source audio, removing the way to check.[13,14]

Post-acute care denials rose as UnitedHealthcare automated prior authorization; lawsuit targets nH Predict

A class action alleges UnitedHealth used naviHealth's nH Predict model to cut off post-acute care; a Senate investigation found UnitedHealthcare's post-acute denial rate nearly tripled while it automated prior authorization.[34,35,19,18]

Wrong AI suggestions pull radiologists' mammogram ratings off

In a controlled experiment, 27 radiologists' accuracy on mammograms fell sharply when a purported AI suggested an incorrect BI-RADS category.[36]

Epic Sepsis Model missed two-thirds of sepsis cases in external validation

An independent validation of Epic's proprietary sepsis score found it far less accurate than the vendor reported: it missed 67% of sepsis cases while alerting on 18% of all hospitalizations.[2,3]

Pulse oximeters overestimate oxygen saturation in patients with darker skin

Paired SpO2/SaO2 data showed occult hypoxemia missed by pulse oximetry about three times as often in Black as in White patients, delaying treatment decisions.[37,38,39,40]

Sepsis model switched off after COVID-19 changed the patient mix

Weeks after its first COVID-19 admissions, the University of Michigan paused Epic sepsis alerts because dataset shift produced spurious alerting; its clinical AI committee decommissioned the model.[5,6]

Care-management algorithm under-referred Black patients because it predicted cost, not illness

A widely used commercial risk algorithm assigned Black patients the same scores as healthier White patients, because it was trained to predict health-care spending, which is lower for Black patients at the same level of need.[8,10]

How you'd know

What to do, tier by tier

What should already be in place at each degradation tier for this layer. Tier 0 is normal automated running; tier 3 is paper, batteries and judgement.

These are practices reported or recommended in the cited sources, gathered for reference. They are not a prescription for your organisation; judge what fits your setting, and check the current official text of any standard.

0Full automation

  • Validate every predictive model on your own recent data before go-live, reporting sensitivity, PPV, calibration and alert burden by unit and by subgroup. Do not accept vendor figures alone.[2,3,24]
  • Write a monitoring plan for each model before deployment: named owner, metrics, thresholds, review frequency, and the ground-truth sample you will check against.[7,41,24]
  • Put notice-of-change and re-validation clauses in every AI contract, and log the model version with each output so you can tell which results came from which version.[24,28]
  • Keep the source for every generated artifact: retain scribe audio long enough to audit, and store the inputs an agent acted on.[14,15]
  • Give agents least privilege: read-only by default, no delete, and human approval for any action that changes orders, medications or records.[25,22]

1Assisted operation

  • Define in advance who can switch a model off and on what trigger (for example alert volume above a set multiple of baseline, or sampled sensitivity below a floor), and give them that authority in writing.[41,5]
  • When a model is paused, show users a visible 'model off' banner so nobody assumes silence means no risk, and re-issue the manual screening criteria the model replaced.[6]
  • Make bedside override explicit and unpenalized in every protocol that starts from a model flag; require patient-specific reassessment before acting.[17,18]
  • Require the signing clinician to read and correct AI-drafted notes and letters before they leave the chart, and mark unsigned AI text as unverified.[15,12]

2Manual operation

  • Keep the pre-AI manual workflow documented and staffed: sepsis screening criteria, dictation or typed notes, pharmacist-handled refill calls. Rehearse switching to it.[6,31]
  • Offer a non-AI route for patients at all times (touch-tone or staffed line, human scribe or no scribe) rather than making the AI path the only way to get care.[21,31]
  • When a model is decommissioned, review the decisions it influenced during the suspect period, such as notes, denials and missed alerts, and correct records where needed.[41,5]

3Analog fallback

  • Keep paper or offline versions of the clinical criteria that models encode (sepsis screens, early-warning scores) so they can be applied when EHR-native models and the EHR are both unavailable.FailSystems judgement
  • Include model and agent failure in downtime drills: practise a scenario where the EHR is up but a model is known to be wrong, not only one where everything is down.FailSystems judgement
  • Report AI-related harm and near misses externally, through a Patient Safety Organization or FDA's reporting pathways for regulated devices, so other sites learn before they hit the same failure.[24]

Standards and rules (US)

InstrumentWhat it requires
FDA draft guidance: AI-Enabled Device Software Functions, Lifecycle Management (Jan 2025, draft)Recommends (non-binding) that sponsors describe postmarket performance monitoring covering data drift, demographic shift and input-pipeline corruption, and how results reach users. Postmarket adverse-event reporting under 21 CFR 803 still applies.[7]
FDA final guidance: Predetermined Change Control Plans for AI-Enabled Device Software Functions (Dec 2024, revised Aug 2025)A PCCP in the marketing submission sets out planned modifications, the protocol to develop, validate and implement them, and an impact assessment, so authorized changes can ship without a new submission.[28]
NIST AI RMF 1.0 (AI 100-1) and Generative AI Profile (AI 600-1)Voluntary framework: monitor systems in production (MEASURE 2.4), keep mechanisms and named responsibility to supersede, disengage or deactivate AI (MANAGE 2.4), and plan post-deployment override, decommissioning and incident response (MANAGE 4.1). The GenAI profile adds confabulation and automation bias as named risks.[41]
45 CFR 92.210 (Section 1557, patient care decision support tools)Covered entities must make ongoing reasonable efforts to identify decision-support tools that use race, color, national origin, sex, age or disability as inputs, and to mitigate discrimination risk (applies from May 2025).[10]
45 CFR 170.315(b)(11) ONC certification, Decision Support InterventionsCertified health IT must expose source attributes for predictive DSIs, including intended use, out-of-scope cautions, external validation, fairness and local-data validity monitoring. HTI-5 (proposed Dec 2025) would remove these model-card requirements.[4]
CMS CY2024 MA rule FAQ on algorithms (42 CFR 422.101(c))Medicare Advantage plans may use algorithms to assist, but coverage decisions must rest on the individual patient's circumstances. A predicted length of stay alone cannot end post-acute care.[18]
Joint Commission + CHAI Responsible Use of AI in Healthcare guidance (Sept 2025)Voluntary guidance with seven elements, including ongoing local quality monitoring, risk and bias assessment, and voluntary blinded reporting of AI safety events.[24]

Elsewhere: EU and UK

The EU AI Act (Regulation (EU) 2024/1689, in force 1 August 2024) treats AI safety components of regulated products as high-risk, with duties for risk management, data quality, logging, human oversight and accuracy. The 2026 Digital Omnibus, signed 8 July 2026, moved the high-risk dates to 2 December 2027 for stand-alone systems and 2 August 2028 for AI embedded in products such as medical devices, which remain under MDR/IVDR conformity assessment in the meantime. In the UK, the MHRA ran the AI Airlock regulatory sandbox for AI as a medical device from April 2024 to March 2025 and published its lessons in October 2025. Ontario's Auditor General found in May 2026 that all 20 provincially approved AI scribes produced inaccurate notes in procurement testing.[42,43,44,15]

Severity score v0.1 draft

4Likelihood
3Blast radius
5Detectability (5 = hardest)
60of 125

FailSystems judgementJudgement: published external validations and audits keep finding deployed models and scribes that perform worse than claimed, so we score likelihood high. Blast radius is moderate per event, usually one model's users rather than a whole hospital, but vendor-wide models like the Epic sepsis score reach hundreds of sites at once. Detectability is the worst on the site because a wrong model raises no alarm and no tier change; it is found by audit, research or a clinician who refuses to comply.

Each factor is scored 1–5 and multiplied, as in a classic FMEA risk priority number. This is our first-draft judgement, not a measurement; see how scoring works and how it will be revised.

What we don't know yet

These gaps drive what the nightly research pass looks for. If you have evidence, send it.

Sources cited on this page

  1. Artificial Intelligence-Enabled Medical Devices (list). U.S. Food and Drug Administration, 2026. Primary Dataset · link checked 2026-09-26
  2. External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients. JAMA Internal Medicine (Wong A, Otles E, Donnelly JP, et al.), 21 June 2021. Primary Peer-reviewed · link checked 2026-09-26
  3. Multicenter Prospective Validation of an Updated Proprietary Sepsis Prediction Model. JAMA Network Open (Wong A, Currey D, Schwinne M, et al.), 2 February 2026. Primary Peer-reviewed · link checked 2026-09-26
  4. 45 CFR 170.315(b)(11) Decision support interventions (ONC Health IT Certification criterion). ASTP/ONC, U.S. Department of Health and Human Services (eCFR), 9 January 2024. Primary Regulation · link checked 2026-09-26
  5. The Clinician and Dataset Shift in Artificial Intelligence. New England Journal of Medicine (Finlayson SG, Subbaswamy A, Singh K, et al.), 15 July 2021. Primary Peer-reviewed · link checked 2026-09-26
  6. Quantification of Sepsis Model Alerts in 24 US Hospitals Before and During the COVID-19 Pandemic. JAMA Network Open (Wong A, Cao J, Lyons PG, et al.), 1 November 2021. Primary Peer-reviewed · link checked 2026-09-26
  7. Artificial Intelligence-Enabled Device Software Functions: Lifecycle Management and Marketing Submission Recommendations (draft guidance). U.S. Food and Drug Administration, 7 January 2025. Primary Guidance · link checked 2026-09-26
  8. Dissecting racial bias in an algorithm used to manage the health of populations. Science (Obermeyer Z, Powers B, Vogeli C, Mullainathan S), 25 October 2019. Primary Peer-reviewed · link checked 2026-09-26
  9. AI recognition of patient race in medical imaging: a modelling study. The Lancet Digital Health (Gichoya JW, Banerjee I, et al.), 11 May 2022. Primary Peer-reviewed · link checked 2026-09-26
  10. 45 CFR 92.210 Nondiscrimination in the use of patient care decision support tools (Section 1557 rule). U.S. Department of Health and Human Services (eCFR), 6 May 2024. Primary Regulation · link checked 2026-09-26
  11. Large language models propagate race-based medicine. npj Digital Medicine (Omiye JA, Lester JC, Spichak S, Rotemberg V, Daneshjou R), 20 October 2023. Primary Peer-reviewed · link checked 2026-09-26
  12. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1. National Institute of Standards and Technology, July 2024. Primary Standard · link checked 2026-09-26
  13. Careless Whisper: Speech-to-Text Hallucination Harms. ACM FAccT 2024 (Koenecke A, Choi ASG, Mei KX, Schellmann H, Sloane M); arXiv 2402.08021, June 2024. Primary Peer-reviewed · link checked 2026-09-26
  14. Researchers say an AI-powered transcription tool used in hospitals invents things no one ever said. Associated Press (syndicated via Scripps News), 26 October 2024. Secondary Journalism · link checked 2026-09-26
  15. Performance Audit: Use of Artificial Intelligence in the Ontario Government (Special Report 2026). Office of the Auditor General of Ontario, 12 May 2026. Primary Official report · link checked 2026-09-26
  16. A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. npj Digital Medicine (Asgari E, et al.; authors affiliated with Tortus AI), 13 May 2025. Primary Peer-reviewed · link checked 2026-09-26
  17. As AI reshapes patient care, human nurses are pushing back against its creeping influence (AP). Associated Press via Euronews, 18 March 2025. Secondary Journalism · link checked 2026-09-26
  18. Frequently Asked Questions related to Coverage Criteria and Utilization Management Requirements in CMS Final Rule (CMS-4201-F) (HPMS memo). Centers for Medicare & Medicaid Services, 6 February 2024. Primary Guidance · link checked 2026-09-26
  19. Refusal of Recovery: How Medicare Advantage Insurers Have Denied Patients Access to Post-Acute Care. U.S. Senate Permanent Subcommittee on Investigations (Majority Staff), 17 October 2024. Primary Official report · link checked 2026-09-26
  20. Ethics and governance of artificial intelligence for health: Guidance on large multi-modal models (news release). World Health Organization, 18 January 2024. Primary Guidance · link checked 2026-09-26
  21. A pharmacy chain in Vermont implemented AI for efficiency. It's led to delays, incorrect information and privacy concerns. VTDigger, 29 July 2026. Secondary Journalism · link checked 2026-09-26
  22. OWASP Top 10 for Agentic Applications. OWASP GenAI Security Project, 9 December 2025. Secondary Standard · link checked 2026-09-26
  23. MITRE ATLAS (Adversarial Threat Landscape for AI Systems), data v5.6.0. MITRE, 2026. Secondary Standard · link checked 2026-09-26
  24. Joint Commission and Coalition for Health AI (CHAI) Guidance on the Responsible Use of AI in Healthcare (RUAIH). The Joint Commission and Coalition for Health AI, 17 September 2025. Primary Guidance · link checked 2026-09-26
  25. OWASP Top 10 for LLM Applications 2025. OWASP GenAI Security Project, 2025. Secondary Standard · link checked 2026-09-26
  26. Vibe coding service Replit deleted user's production database, faked data, told fibs galore. The Register, 21 July 2025. Secondary Journalism · link checked 2026-09-26
  27. Why Do Multi-Agent LLM Systems Fail?. arXiv preprint 2503.13657 (Cemri M, Pan MZ, Yang S, et al.), 17 March 2025. Supporting Preprint · link checked 2026-09-26
  28. Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions (final guidance). U.S. Food and Drug Administration, 4 December 2024. Primary Guidance · link checked 2026-09-26
  29. HTI-5 Proposed Rule fact sheet: ONC Deregulatory Actions to Unleash Prosperity. ASTP/ONC, 22 December 2025. Primary Regulation · link checked 2026-09-26
  30. AI systems used by Ontario doctors hallucinate, auditor general finds. Global News (Canada), 12 May 2026. Secondary Journalism · link checked 2026-09-26
  31. Kinney Drugs pulls back AI phone assistant after hundreds of customer complaints. WCAX-TV, 7 August 2026. Secondary Journalism · link checked 2026-09-26
  32. Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study. The Lancet Gastroenterology & Hepatology 10(10):896-903, 12 August 2025. Primary Peer-reviewed · link checked 2026-09-26
  33. AI enters the exam room, and nurses are left to manage the fallout. Scientific American (Hilke Schellmann), 17 February 2026. Secondary Journalism · link checked 2026-09-27
  34. UnitedHealth uses faulty AI to deny elderly patients medically necessary coverage, lawsuit claims. CBS News, 20 November 2023. Secondary Journalism · link checked 2026-09-26
  35. Federal judge trims AI denial lawsuit against UnitedHealth Group. Courthouse News Service, 13 February 2025. Secondary Journalism · link checked 2026-09-26
  36. Automation Bias in Mammography: The Impact of Artificial Intelligence BI-RADS Suggestions on Reader Performance. Radiology 307(4):e222176 (RSNA), 2 May 2023. Primary Peer-reviewed · link checked 2026-09-26
  37. Racial Bias in Pulse Oximetry Measurement. New England Journal of Medicine (Sjoding MW, Dickson RP, Iwashyna TJ, Gay SE, Valley TS), 17 December 2020. Primary Peer-reviewed · link checked 2026-09-26
  38. Racial and Ethnic Discrepancy in Pulse Oximetry and Delayed Identification of Treatment Eligibility Among Patients With COVID-19. JAMA Internal Medicine (Fawzy A, et al.), 1 July 2022. Primary Peer-reviewed · link checked 2026-09-26
  39. Pulse Oximeters (FDA actions on accuracy and skin pigmentation). U.S. Food and Drug Administration, 6 January 2025. Primary Guidance · link checked 2026-09-26
  40. Pulse Oximeters for Medical Purposes - Non-Clinical and Clinical Performance Testing, Labeling, and Premarket Submission Recommendations (Draft Guidance). U.S. Food and Drug Administration, 7 January 2025. Primary Guidance · link checked 2026-09-26
  41. Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. National Institute of Standards and Technology, January 2023. Primary Standard · link checked 2026-09-26
  42. AI Act (Regulation (EU) 2024/1689): regulatory framework overview. European Commission, Shaping Europe's digital future, 1 August 2024. Primary Regulation · link checked 2026-09-26
  43. Digital Omnibus on AI (Legislative Train Schedule). European Parliament, 8 July 2026. Primary Regulation · link checked 2026-09-26
  44. AI Airlock Sandbox Pilot Programme Report. Medicines and Healthcare products Regulatory Agency (UK), 16 October 2025. Primary Official report · link checked 2026-09-26

Cite this pageFailSystems. “Models & agents.” https://failsystems.health201.com/layers/models/ (reviewed 2026-09-26). Health 201 / AstroNexus LLC. CC BY 4.0.

Information only, not advice. FailSystems is an aggregation and synthesis of published sources. It is not consulting, engineering, legal, regulatory or medical advice, and using it creates no professional relationship. Health systems are complex and no approach fits every organisation: anything you adopt is your own decision, at your own risk, and should be checked against the current official sources and by qualified people who know your setting. Full disclaimer.

Dealing with an incident right now? This site is a reference, not an incident-response service. Activate your organisation's emergency operations plan and incident command, and: