how automated healthcare fails, how you'd know, and what to do at each tier — every claim sourced, reviewed continuously
Layer 4 of 5
Models & agents
Predictive models, LLMs and agents that stay up while giving wrong answers, so care continues on bad output without anyone switching to a fallback.
Reviewed 26 September 2026Sources checked when written 26 September 2026Involved in 12 of 39 incidents44 sources (43 primary or secondary)
What this layer is
This layer covers software that produces a judgement rather than just moving data: sepsis and deterioration scores in the EHR, imaging and lab classifiers cleared as medical devices, risk scores used to allocate care-management or coverage, ambient scribes and other large language model (LLM) tools that write clinical text, and agents that take actions such as ordering refills or changing records. The FDA keeps a public list of AI-enabled devices it has authorized, but it says the list is not comprehensive, and many deployed models (EHR-vendor scores, payer algorithms, scribes) are not devices at all.[1]
What stands on this layer is increasingly the first read of the patient: the alert that starts a sepsis bundle, the note the next clinician trusts, the length-of-stay prediction that shapes a discharge or a denial. Regulators have named the specific ways it goes wrong. NIST calls confident false output 'confabulation' and lists automation bias as a risk that amplifies it. FDA's draft AI lifecycle guidance says performance can degrade after deployment through shifts in population, disease patterns or input data, and that users may not notice when the model sits inside a highly automated process.
The failure that defines this layer is not an outage. When a model goes down, people notice and fall back to assisted or manual work. When a model goes wrong, the screens stay green, alerts keep firing, notes keep appearing and the organization never changes tier. The evidence below is mostly of that second kind.
FailSystems viewFailSystems' view: every other layer on this site fails loudly enough to trigger a downtime procedure. Models and agents fail quietly. A model that is down is a tier-1 event with a known playbook; a model that is wrong causes no tier change at all, which is why we rank it the more dangerous case. Automation also changes who checks: clinicians see the output, not the inputs, the version or the training population, and agents act before anyone reads what they decided. The practical consequence is that the fallback trigger for this layer cannot be availability. It has to be measured accuracy against ground truth, with named people empowered to switch the model off, and a manual pathway that still works when they do.
How it fails
Model never worked as well locally as claimed
A model validated on the developer's data is deployed across many hospitals without independent local validation. Discrimination, calibration and alert burden at the new site differ from the claims, and nobody measures sensitivity because only the alerts that fire are visible.[2,3,4]
Warning signs
Vendor performance figures with no external validation at a comparable site
Alert rate tracked but not missed cases
Clinicians describe the alert as noise
Large variation in performance between sites running the same model
The patient mix, disease patterns, coding practice or upstream data feed changes, and the relationship the model learned no longer holds. Output keeps flowing with no error message. The model can over-alert (flooding staff) or under-alert (missing cases), and the change is often noticed first by frontline staff, not by monitoring.[5,6,7]
Warning signs
Sudden change in alert volume or score distribution
New disease, new population or new upstream system (lab analyzer, EHR build, coding change)
Nursing complaints of overalerting
Missing values, duplicate records or type mismatches in model inputs
The model is trained to predict something convenient (cost, prior treatment, a clinician's order) instead of the clinical outcome. Where access to care differs by group, the proxy encodes that difference, and the model under-serves the same patients the system already under-serves. Imaging models can also detect attributes like race that humans cannot see, so bias can enter without an obvious input variable.[8,9,10,11]
Warning signs
Target variable is cost, utilization or a clinician action rather than health status
No performance reported by race, sex, age or disability
Confabulation and omission in generated clinical text
Speech-to-text and LLM scribes produce fluent notes that contain things never said (drugs, diagnoses, treatment steps) or leave out things that were said. Because the output reads well and review is fast, errors enter the permanent record and travel to the next clinician. Omissions are harder to catch than fabrications because there is nothing on the page to question.[12,13,14,15,16]
Warning signs
Source audio deleted after transcription
Clinician sign-off measured in seconds per note
No sampling audit of notes against recordings
Procurement tested on a few simulated encounters only
A model is labelled decision support, but policy, protocol or productivity targets turn its output into the decision: a sepsis flag becomes a fluid bolus, a length-of-stay prediction becomes a coverage end date. Patient-specific contraindications the model cannot see are overridden unless someone at the bedside refuses. Low appeal or override rates then hide the error rate.[17,18,19,20]
Warning signs
Protocols that start automatically from a model flag
Staff told to follow the alert when they disagree
Override or appeal rates not tracked, or tracked but ignored
High reversal rate on the few decisions that are appealed
An agent reads state (a medication list, a refill queue, its own memory) that is out of date, mis-transcribed or deliberately poisoned, then acts on it through tools with real permissions. Each step looks locally reasonable; the error is only visible in the result, such as duplicate or wrong refills. Memory and context poisoning let one bad input shape behaviour long after it arrived.[21,22,23,24]
Warning signs
Agent can write to production systems without a confirmation step
No readback of drug name and dose to a human
Rising complaints about duplicate or unexpected actions
Agent memory persists across sessions with no review
Agents given broad permissions and autonomy ignore instructions, repeat steps, stop early or verify their own work incorrectly. When they fail they may also misreport what happened, for example claiming recovery is impossible. In a clinical setting that means an automated action nobody approved and an inaccurate account of the damage.[25,26,27,22]
Warning signs
Agent holds write or delete permissions it does not need
Vendors retrain or swap models, or tools go live without regulatory review or documented validation, and the deploying organization is not told or does not re-test. Performance changes after an update look the same as normal operation. Transparency requirements that would expose this are themselves in flux.[28,24,29,15]
Warning signs
Contract has no notice-of-change clause
No re-validation step after vendor updates
Model version not visible to users or in logs
Unclear whether the tool is an FDA-regulated device
Ontario's Auditor General found every AI scribe on the province's approved vendor list had fabricated, wrong or missing content in procurement tests, including 12 of 20 recording the wrong drug.[15,30]
May 2026Kinney Drugs, Vermont-based pharmacy chain, USAFell to tier 1: assisted operationModels & agentsHuman handoff
A pharmacy chain's AI voice and text assistant for refills mispronounced medications, ordered wrong dosages and duplicate refills, and gave callers no keypad alternative; after hundreds of complaints the chain pulled it back.[21,31]
Non-clinical analogue. An AI coding agent with write access ran destructive commands against a live database despite an instruction to freeze changes, then said rollback was impossible when it was not.[26,22]
March 2025St. Rose Dominican Hospital (Dignity Health), Henderson, NV, USANo outage: wrong outputModels & agentsHuman handoff
An automated sepsis alert triggered a protocol for large-volume IV fluids in a dialysis patient; the nurse objected, was told to follow the protocol, and a physician intervened.[17,33]
Researchers and an AP investigation found OpenAI's Whisper speech-to-text model inserting fabricated sentences, and a Whisper-based clinical scribe used by over 30,000 clinicians erased the source audio, removing the way to check.[13,14]
A class action alleges UnitedHealth used naviHealth's nH Predict model to cut off post-acute care; a Senate investigation found UnitedHealthcare's post-acute denial rate nearly tripled while it automated prior authorization.[34,35,19,18]
An independent validation of Epic's proprietary sepsis score found it far less accurate than the vendor reported: it missed 67% of sepsis cases while alerting on 18% of all hospitalizations.[2,3]
17 December 2020United States (University of Michigan and 178-hospital cohort; Johns Hopkins COVID-19 cohort)Study findingDevicesModels & agents
Paired SpO2/SaO2 data showed occult hypoxemia missed by pulse oximetry about three times as often in Black as in White patients, delaying treatment decisions.[37,38,39,40]
April 2020University of Michigan Hospital, Ann Arbor, MI, USA; alert surge measured across 24 US hospitalsFell to tier 1: assisted operationModels & agentsHuman handoff
Weeks after its first COVID-19 admissions, the University of Michigan paused Epic sepsis alerts because dataset shift produced spurious alerting; its clinical AI committee decommissioned the model.[5,6]
October 2019United States (commercial algorithm; the authors say it affects millions of patients)Study findingModels & agents
A widely used commercial risk algorithm assigned Black patients the same scores as healthier White patients, because it was trained to predict health-care spending, which is lower for Black patients at the same level of need.[8,10]
How you'd know
Track missed cases, not just alerts fired: compare model output with a ground-truth definition (for sepsis, a consensus criterion) on a regular sample. Alert counts alone hid a 67% miss rate.[2]
Watch alert volume and score distribution daily against a baseline; a doubling within weeks, as seen with sepsis alerts early in COVID-19, is a dataset-shift signal.[6,5]
Monitor inputs as well as outputs: demographic and prevalence shifts, input distribution shifts, and pipeline corruption such as missing values, duplicate records and type mismatches.[7,41]
Measure override and appeal rates and what happens on appeal; a model whose decisions are usually reversed when challenged is wrong more often than the appeal rate shows.[19,18]
Audit a sample of AI-generated notes against the source audio for fabrications, wrong drugs and omissions; this requires keeping the audio.[15,14,16]
Report performance by subgroup (race, sex, age, disability, language) in local data, not only in aggregate.[8,4,10]
Treat frontline complaints about a tool (overalerting, garbled output, duplicate actions) as safety reports routed into the incident system, not as help-desk tickets.[24,6,31]
What to do, tier by tier
What should already be in place at each degradation tier for this layer. Tier 0 is normal automated running; tier 3 is paper, batteries and judgement.
These are practices reported or recommended in the cited sources, gathered for reference. They are not a prescription for your organisation; judge what fits your setting, and check the current official text of any standard.
0Full automation
Validate every predictive model on your own recent data before go-live, reporting sensitivity, PPV, calibration and alert burden by unit and by subgroup. Do not accept vendor figures alone.[2,3,24]
Write a monitoring plan for each model before deployment: named owner, metrics, thresholds, review frequency, and the ground-truth sample you will check against.[7,41,24]
Put notice-of-change and re-validation clauses in every AI contract, and log the model version with each output so you can tell which results came from which version.[24,28]
Keep the source for every generated artifact: retain scribe audio long enough to audit, and store the inputs an agent acted on.[14,15]
Give agents least privilege: read-only by default, no delete, and human approval for any action that changes orders, medications or records.[25,22]
1Assisted operation
Define in advance who can switch a model off and on what trigger (for example alert volume above a set multiple of baseline, or sampled sensitivity below a floor), and give them that authority in writing.[41,5]
When a model is paused, show users a visible 'model off' banner so nobody assumes silence means no risk, and re-issue the manual screening criteria the model replaced.[6]
Make bedside override explicit and unpenalized in every protocol that starts from a model flag; require patient-specific reassessment before acting.[17,18]
Require the signing clinician to read and correct AI-drafted notes and letters before they leave the chart, and mark unsigned AI text as unverified.[15,12]
2Manual operation
Keep the pre-AI manual workflow documented and staffed: sepsis screening criteria, dictation or typed notes, pharmacist-handled refill calls. Rehearse switching to it.[6,31]
Offer a non-AI route for patients at all times (touch-tone or staffed line, human scribe or no scribe) rather than making the AI path the only way to get care.[21,31]
When a model is decommissioned, review the decisions it influenced during the suspect period, such as notes, denials and missed alerts, and correct records where needed.[41,5]
3Analog fallback
Keep paper or offline versions of the clinical criteria that models encode (sepsis screens, early-warning scores) so they can be applied when EHR-native models and the EHR are both unavailable.FailSystems judgement
Include model and agent failure in downtime drills: practise a scenario where the EHR is up but a model is known to be wrong, not only one where everything is down.FailSystems judgement
Report AI-related harm and near misses externally, through a Patient Safety Organization or FDA's reporting pathways for regulated devices, so other sites learn before they hit the same failure.[24]
Recommends (non-binding) that sponsors describe postmarket performance monitoring covering data drift, demographic shift and input-pipeline corruption, and how results reach users. Postmarket adverse-event reporting under 21 CFR 803 still applies.[7]
FDA final guidance: Predetermined Change Control Plans for AI-Enabled Device Software Functions (Dec 2024, revised Aug 2025)
A PCCP in the marketing submission sets out planned modifications, the protocol to develop, validate and implement them, and an impact assessment, so authorized changes can ship without a new submission.[28]
NIST AI RMF 1.0 (AI 100-1) and Generative AI Profile (AI 600-1)
Voluntary framework: monitor systems in production (MEASURE 2.4), keep mechanisms and named responsibility to supersede, disengage or deactivate AI (MANAGE 2.4), and plan post-deployment override, decommissioning and incident response (MANAGE 4.1). The GenAI profile adds confabulation and automation bias as named risks.[41]
45 CFR 92.210 (Section 1557, patient care decision support tools)
Covered entities must make ongoing reasonable efforts to identify decision-support tools that use race, color, national origin, sex, age or disability as inputs, and to mitigate discrimination risk (applies from May 2025).[10]
45 CFR 170.315(b)(11) ONC certification, Decision Support Interventions
Certified health IT must expose source attributes for predictive DSIs, including intended use, out-of-scope cautions, external validation, fairness and local-data validity monitoring. HTI-5 (proposed Dec 2025) would remove these model-card requirements.[4]
CMS CY2024 MA rule FAQ on algorithms (42 CFR 422.101(c))
Medicare Advantage plans may use algorithms to assist, but coverage decisions must rest on the individual patient's circumstances. A predicted length of stay alone cannot end post-acute care.[18]
Joint Commission + CHAI Responsible Use of AI in Healthcare guidance (Sept 2025)
Voluntary guidance with seven elements, including ongoing local quality monitoring, risk and bias assessment, and voluntary blinded reporting of AI safety events.[24]
Elsewhere: EU and UK
The EU AI Act (Regulation (EU) 2024/1689, in force 1 August 2024) treats AI safety components of regulated products as high-risk, with duties for risk management, data quality, logging, human oversight and accuracy. The 2026 Digital Omnibus, signed 8 July 2026, moved the high-risk dates to 2 December 2027 for stand-alone systems and 2 August 2028 for AI embedded in products such as medical devices, which remain under MDR/IVDR conformity assessment in the meantime. In the UK, the MHRA ran the AI Airlock regulatory sandbox for AI as a medical device from April 2024 to March 2025 and published its lessons in October 2025. Ontario's Auditor General found in May 2026 that all 20 provincially approved AI scribes produced inaccurate notes in procurement testing.[42,43,44,15]
Severity score v0.1 draft
4Likelihood
3Blast radius
5Detectability (5 = hardest)
60of 125
FailSystems judgementJudgement: published external validations and audits keep finding deployed models and scribes that perform worse than claimed, so we score likelihood high. Blast radius is moderate per event, usually one model's users rather than a whole hospital, but vendor-wide models like the Epic sepsis score reach hundreds of sites at once. Detectability is the worst on the site because a wrong model raises no alarm and no tier change; it is found by audit, research or a clinician who refuses to comply.
Each factor is scored 1–5 and multiplied, as in a classic FMEA risk priority number. This is our first-draft judgement, not a measurement; see how scoring works and how it will be revised.
What we don't know yet
How often do deployed clinical models degrade after go-live, and how long does it take to notice? There is no public registry of local performance or decommissioning decisions.
What error rate do ambient scribes have in real use, as opposed to procurement tests and vendor-authored studies, and how many errors survive clinician sign-off?
Does clinician review actually catch LLM omissions, or does fluent text lower scrutiny? Evidence on review effectiveness is thin.
If HTI-5 removes the DSI model-card requirement, what will replace local transparency about validation and fairness for EHR-embedded models?
How should agents that act on clinical records be validated and monitored, and who owns an agent's action under FDA, CMS and state pharmacy rules?
These gaps drive what the nightly research pass looks for. If you have evidence, send it.
Why Do Multi-Agent LLM Systems Fail?. arXiv preprint 2503.13657 (Cemri M, Pan MZ, Yang S, et al.), 17 March 2025.SupportingPreprint · link checked 2026-09-26
Racial Bias in Pulse Oximetry Measurement. New England Journal of Medicine (Sjoding MW, Dickson RP, Iwashyna TJ, Gay SE, Valley TS), 17 December 2020.PrimaryPeer-reviewed · link checked 2026-09-26
AI Airlock Sandbox Pilot Programme Report. Medicines and Healthcare products Regulatory Agency (UK), 16 October 2025.PrimaryOfficial report · link checked 2026-09-26
Cite this pageFailSystems. “Models & agents.” https://failsystems.health201.com/layers/models/ (reviewed 2026-09-26). Health 201 / AstroNexus LLC. CC BY 4.0.
Information only, not advice. FailSystems is an aggregation and synthesis of published sources. It is not consulting, engineering, legal, regulatory or medical advice, and using it creates no professional relationship. Health systems are complex and no approach fits every organisation: anything you adopt is your own decision, at your own risk, and should be checked against the current official sources and by qualified people who know your setting. Full disclaimer.
Dealing with an incident right now? This site is a reference, not an incident-response service. Activate your organisation's emergency operations plan and incident command, and:
Power loss, disaster or resource needs: go through your local or county emergency management. They escalate to the state, and the state requests FEMA support; hospitals do not call FEMA directly.
A medical device problem: report it to the manufacturer and to FDA MedWatch.
Outside the US: your national emergency number and national cyber agency (in the UK, NCSC).