how automated healthcare fails, how you'd know, and what to do at each tier — every claim sourced, reviewed continuously
A safe automated system knows how to become a less automated one. Every clinical workflow needs a defined answer at each tier, decided before the failure.
The tiers describe how much of the automation is still working and trusted, not how serious the incident is. A hospital can be at tier 1 for one layer (a sepsis model switched off) while every other layer runs at tier 0. The layer × tier matrix below is the design question FailSystems asks of every clinical workflow: for each layer, what must already be true if you drop to this tier today?FailSystems judgement
US rules already require practice. The CMS emergency preparedness rule and the Joint Commission's emergency management standards require two exercises a year with an after-action review, but neither requires the scenario to be an EHR downtime or an automation failure. ONC's SAFER guide goes further and recommends unannounced downtime drills at least once a year.[4,3]
The normal state of a digital hospital. Machines acquire and route data (monitors, interfaces, lab analysers with autoverification), analyse it (early-warning scores, results flags), recommend (CPOE decision support, sepsis or deterioration models) and in places act (barcode-gated administration, smart-pump limits, automated dispensing). Clinicians supervise and decide, but much of the checking has moved into software they no longer see. In levels-of-automation terms, most functions sit in the middle of the scale: the computer narrows or suggests and the human approves (Parasuraman, Sheridan and Wickens 2000). Tier 0 is only safe if you can tell when you have left it: the literature shows automation can fail silently while appearing to work (Wright 2016).[1,2,3,4,5,6,7]
What must be true
How you know you're here
How you get back up
The systems are still running, but you can no longer trust them to do their part unsupervised. The EHR is slow enough to cause errors, one interface has stopped, a model or alert is misfiring, or an upgrade changed behaviour. Automation is demoted to an advisory role: humans make each decision and add the checks the software used to do. This is the most dangerous tier to miss, because screens still look normal. Automation bias and complacency push people to keep trusting output that has become wrong (Goddard 2012; Parasuraman, Sheridan and Wickens 2000), and CDS failures often go undetected for long periods (Wright 2016).[8,1,2,3,9,10]
What must be true
How you know you're here
How you get back up
Electronics are up but the cognitive layer is down. Power, monitors, pumps, phones and devices work in standalone mode, but the EHR, CPOE, decision support, barcode verification or lab/pharmacy interfaces are unavailable. Care runs on paper forms, printouts from a read-only viewer, fax, runners and people's memory. This is the tier the downtime literature describes. It is slower and error-prone: lab result reporting was 62% slower on average during downtime in one study (Larsen 2019), a 17-minute results-system outage multiplied clinician read times several-fold (Wang 2016), and in 46% of downtime-related safety reports procedures were not followed or not in place (Larsen 2018). Ransomware can hold a hospital here for weeks. The Joint Commission tells hospitals to prepare for four weeks or longer (SEA 67).[11,9,12,6,3,13,7,14,15,16,17,18]
What must be true
How you know you're here
How you get back up
Power, the network or both are gone, or cannot be trusted. Generators have failed or do not cover the load, UPS batteries are running down, phones and paging may be out, and HVAC and lighting may be lost. What remains is paper, battery-powered devices while their charge lasts, manual techniques (bag-valve ventilation, gravity infusions, cylinder oxygen, direct observation instead of telemetry) and clinical judgement. At this tier the question is often whether to shelter in place or evacuate. NYU Langone evacuated 21 NICU patients in 4.5 hours after Hurricane Sandy's surge cut power in 2012 (Espiritu 2014).[19,13,3,5,4,6,7,20]
What must be true
How you know you're here
How you get back up
What must hold for each layer at each tier. Read a row to see how one layer degrades; read a column to see what a whole hospital needs at that tier. Each row links to the layer's full, sourced defenses; the cells are FailSystems' summary of them.
These are practices reported or recommended in the cited sources, gathered for reference. They are not a prescription for your organisation; judge what fits your setting, and check the current official text of any standard.
| Layer | 0 · Full automation | 1 · Assisted operation | 2 · Manual operation | 3 · Analog fallback |
|---|---|---|---|---|
| Power | The whole power chain (generators, fuel, transfer switches, cooling) is sited above flood level, and production and recovery IT do not share a cloud region or a grid. | Partial outages have clinical triggers: slow labs or order entry open downtime command even while systems are technically up. | Generators are load-tested, the EHR and downtime workstations ride on UPS, there is fuel for two days, and everyone knows which devices are on emergency outlets. | Paper operation for weeks and evacuation without lifts have been rehearsed, and every evacuated patient leaves with a paper summary. |
| Connectivity & data | Phishing-resistant MFA on all remote and vendor access, air-gapped backups, redundant network paths, and contracts that commit third parties to notification and recovery times. | A segmented network, a warm site that can take the whole EHR within hours, and a contracted alternate clearinghouse and reference lab. | A read-only EHR refreshed hourly and printable on backed-up power, a communication channel off the EHR network, and a decision to call downtime within 2 hours. | Current paper forms on every unit, runners to move orders and results, regional mutual aid for cyber incidents, and a planned recovery phase for back-entry. |
| Devices | An inventory with firmware and support dates, SBOMs and patch timelines in contract, staged update rings, and accuracy data by skin tone for oximeters. | Clinicians verify pumps against the current order and question readings that don't fit; clinical devices sit on their own network and can still monitor locally. | Standalone monitors and pumps stocked on critical units, offline recovery kits for endpoints, and manual infusion programming with a double-check. | Paper flowsheets, manual BP cuffs and gravity sets on hand, the skills to use them, and a decided list of what elective work stops. |
| Models & agents | Every model is validated locally before go-live, has a named owner and monitoring plan, logs its version with each output, and agents act with least privilege. | Someone is authorised to switch a model off on a defined trigger, a visible 'model off' banner shows it, and override is explicit and unpenalised. | The pre-AI workflow is documented, staffed and rehearsed, and patients always have a non-AI route to care. | Paper versions of the criteria models encode (sepsis screens, early-warning scores), and drills where the EHR is up but a model is known to be wrong. |
| Human handoff | Clinicians keep unaided performance measured, every alarm and alert has a named responder, and each AI tool's level of automation is written down. | Every mode change is shown on screen and says what the clinician now owns, with takeover procedures for each known failure signature. | Alarm limits are set per patient, and who may silence or widen them is written down and audited; critical alarms are audible wherever the responder is. | Unannounced downtime drills on every unit each year, with every clinician trained on paper ordering and on finding the read-only EHR. |
| Cascades | A dependency map from each clinical service to its suppliers, networks, devices and utilities, with primary and backup never sharing a single point of failure. | Alternates are enrolled for every concentrated supplier, and disconnection criteria are agreed with partners before an incident. | Neighbours and EMS hear early, services are restored in order of clinical criticality, and all staff (not only IT) can run an extended cyber downtime. | Signed transfer agreements, plans for region-wide events where every neighbour is down, and space, power and oxygen for the community's electricity-dependent patients. |
US hospitals are required to exercise. Under the CMS emergency preparedness rule, a hospital must run two exercises a year: an annual full-scale community exercise or facility functional exercise, plus a second one that may be a facilitated tabletop. It must analyse every drill, tabletop and real event, and revise its plan (42 CFR 482.15(d)(2)). The Joint Commission's 2022 EM chapter mirrors this: two exercises a year, with after-action reports reviewed by a committee and sent to senior leaders (EM.16.01.01, EM.17.01.01). CMS guidance says not to test the same scenario every year (SOM Appendix Z). None of these rules names EHR downtime or automation failure as a required scenario. ONC's SAFER guide goes further: it recommends unannounced EHR downtime drills at least once a year, and the Joint Commission suggests drilling annually or quarterly depending on staff turnover (SAFER 2.1; SEA 67).
The evidence that drills work is thin and mostly descriptive. A systematic review of hospital mass-casualty training found drills helped staff learn procedures and exposed weak points in command, communications and patient flow, but study quality was poor and no included study evaluated tabletop exercises (Hsu 2004). A 2025 scoping review found tabletops were the most common hospital training method but could not identify a best one (Malek 2025). Studies of downtime drills are single-site reports: unit drills audited against a checklist (Kashiwagi 2016), a live PACS-offline drill (Dhamija 2022), a quarterly exercise tied to EHR upgrades (Bulson 2024), six-monthly random-unit drills (Lyon 2023), and a nurse escape room with self-reported gains (Rossley 2022). None measures patient outcomes.
Real downtimes show what drills should target. In Larsen 2018, 46% of downtime-related safety reports described procedures that were not followed or not in place. Clinicians kept ordering tests at normal rates during downtime (Larsen 2019). In a simulation, clinicians did not recognise that a patient's deterioration came from a compromised device (Dameff 2018). Staff surveyed after a well-planned downtime still did not know where to find resources (Lyon 2023). FailSystems' reading: exercises should test detection and the declaration decision, the recovery and back-entry phase, and tier 1, not only the switch to paper.
Coming back up is its own hazardous phase, and guidance treats it as part of downtime rather than its end. Restore systems in stages by clinical priority, keep downtime procedures running on each unit until its systems are verified, and for cyberattacks, confirm the attacker is gone before restoring, since restoring without eradication can leave the network compromised (ASPR TRACIE 2022). Restart interfaces in order with empty buffers, because in-transit data can be lost without warning (SAFER 2.5). Before you withdraw manual checks, check that automation behaves as expected. Wright 2016 found that CDS rules can silently stop firing, or fire spuriously after an upgrade. That is why FailSystems suggests running test patients through key rules before announcing uptime.
Back-loading paper is heavy, slow work and needs its own staff. SAFER asks for a process to enter orders as coded data and scan other paper after reactivation, and to merge temporary patient IDs (SAFER 1.3, 1.5). The Joint Commission says to authenticate every transcribed order and to assign staff for data entry (SEA 67). One district needed 2-3 hours of doctor-pharmacist medication reconciliation in high-use areas after an 8-hour planned downtime (Lyon 2023). Paper records made during real downtimes are often incomplete (Larsen 2019), and some manually recorded data may be deliberately left out of the EHR. ASPR advises documenting what that data is and where it is kept (ASPR TRACIE 2022). As early as 2000, LDS Hospital found that recovery from planned downtimes was not smooth, and it wrote down exactly which data had to be re-entered (Nelson 2007).[7,3,2,6,15,11,21,14,32]
The four tiers are FailSystems' own operational framework, not an established standard. They borrow from the levels-of-automation literature. Sheridan and Verplank's 1978 ten-level scale runs from the human doing everything to the computer acting alone. Parasuraman, Sheridan and Wickens (2000) applied it separately to four functions: information acquisition, analysis, decision selection and action. A hospital at tier 0 runs different functions at different levels. Tiers 1-3 describe what happens when those levels are forced down in a failure, rather than chosen at design time. Tier 1 lowers decision and action automation while keeping information automation. Tier 2 removes analysis and decision automation but keeps device-level acquisition and action. Tier 3 also loses most acquisition and action. The same literature warns that high automation brings reduced situation awareness, complacency and skill loss, which matter most when automation fails (Parasuraman et al. 2000). Bainbridge's 'ironies of automation' makes the same point: automating a task can make the human's remaining job harder. Automation bias in clinical decision support is well documented (Goddard 2012). SAE J3016's driving levels 0-5 are a useful vocabulary analogue only. Its 'fallback' and 'minimal risk condition' concepts roughly match a planned drop to a lower tier, but J3016 has no status in healthcare.[33,1,34,8,35]
Cite this pageFailSystems. “Degradation tiers.” https://failsystems.health201.com/tiers/ (reviewed 2026-09-26). Health 201 / AstroNexus LLC. CC BY 4.0.
Information only, not advice. FailSystems is an aggregation and synthesis of published sources. It is not consulting, engineering, legal, regulatory or medical advice, and using it creates no professional relationship. Health systems are complex and no approach fits every organisation: anything you adopt is your own decision, at your own risk, and should be checked against the current official sources and by qualified people who know your setting. Full disclaimer.