how automated healthcare fails, how you'd know, and what to do at each tier — every claim sourced, reviewed continuously
Propagation pattern
Cascades
Not a layer: the way one failure crosses power, connectivity, devices, models and handoff, and spreads from one organisation to many.
Reviewed 26 September 2026Sources checked when written 26 September 2026Involved in 16 of 39 incidents71 sources (68 primary or secondary)
What this layer is
Cascades is not a sixth layer. It is a propagation pattern that runs across the other five: power, connectivity, devices, models and handoff. A cascade starts as a failure in one layer, at one organisation, and turns into a failure in other layers or other organisations. Examples: a remote-access portal is breached and national claims processing stops. A security-software update crashes Windows machines and hospital services go offline. A pathology supplier is hit and three hospital trusts have to call for O-type blood donors.
Safety science has argued for decades that serious accidents in complex systems do not come from one broken part. Perrow called accidents in systems that are both interactively complex and tightly coupled 'normal accidents': they are to be expected, not freak events. Reason's Swiss-cheese model describes harm reaching a patient only when latent conditions and active failures in several defensive layers line up. Leveson's STAMP and CAST treat accidents as a loss of control over the system, not a chain of failed parts. Cook and Rasmussen described how hospitals chasing efficiency 'go solid': the buffers that used to soak up a problem disappear, so that an event in one distant part of the hospital suddenly matters everywhere else. None of these models is settled: professionals disagree on what the parts of the Swiss-cheese model mean, and its critics call it too linear. We use them as lenses, not as laws.
On this site a cascade is described by its path (the order in which it crossed layers, such as devices → connectivity → handoff) and its reach (one department, one hospital, a region, a country). Tracing the path matters because the controls that stop a cascade usually sit at the boundaries between layers and between organisations, and those boundaries are the parts nobody owns.
FailSystems viewFailSystems' view: automation does not add many new ways for a single part to fail. What it adds is coupling. When a hospital runs on a shared EHR, a shared clearinghouse, a shared endpoint agent and a shared pathology network, the same fault reaches every place at once, and the paper workaround has usually withered from lack of use. In our judgement the incidents that do the most harm to patients in automated care will be cascades. Most of them will start outside the hospital, in a supplier the hospital does not control and may not know it depends on. Plan for common-mode failure, not for one component failing on its own.
How it fails
Single shared supplier (concentration risk)
Many hospitals, pharmacies or practices depend on one supplier for one function: claims clearing, pathology, identity, endpoint security. When that supplier fails, every customer loses the function at the same moment, and there is often no quick way to switch because contracts, interfaces and enrolments were built for one path. The customers also cannot see into the supplier, so they cannot judge when service will come back.[1,2,3,4]
Warning signs
One supplier handles a function for most of the organisations in your region or specialty
No alternate supplier is enrolled or tested, or switching would take weeks of EDI or interface work
Contract has no restoration-time commitment or incident-notification clause
You cannot list which clinical workflows stop if that supplier is down for 7 days
One software or content update is pushed to every installation at once and carries a latent defect that testing missed. Because every machine runs the same code, redundancy inside the hospital does not help: the primary and the backup crash together. Updates designed to ship fast, like security content, are the most exposed.[5,6,7]
Warning signs
Vendors can push updates to production endpoints with no staging ring under your control
Primary and backup systems run the same agent, OS build or content version
No inventory of which clinical workstations and servers run each kernel-level agent
Efficiency work strips out slack: spare beds, spare stock, spare staff, manual steps. Activities then depend directly on events elsewhere in the system, so a delay in one place becomes a stoppage in another within hours. In a tightly coupled process there is no time to improvise before the next step needs the output of the failed one.[8,9,10]
Warning signs
Just-in-time supply with no local stock for time-critical items (blood, reagents, drugs)
Bed occupancy routinely near 100%
Paper downtime forms missing, out of date or never drilled
Workflows where the next step cannot start without an electronic result
Hospitals depend on utilities that depend on each other. Loss of electricity can stop water treatment and gas supply, which in turn stops hospital heating, sterilisation and toilets. Patients at home on powered equipment lose it at the same time and arrive at the emergency department. The hospital's generator covers its own electricity but not the water pressure or the patients' home equipment.[11,12,13,14]
Warning signs
Emergency plan assumes municipal water and gas stay up when the grid is down
Boilers or chillers depend on mains water pressure
No register of local patients on home oxygen, dialysis or other electricity-dependent equipment
Regional plan assumes neighbours can take transfers during a region-wide event
A hospital that goes to downtime diverts ambulances and patients to its neighbours. The neighbours have their own systems intact but not the extra capacity, so waits, walk-outs and delays in time-critical care rise there too. One organisation's cyber incident becomes a capacity incident for the region.[15,16,17]
Warning signs
One health system holds a large share of regional inpatient capacity
No regional agreement on diversion, transfers or shared downtime capacity
EMS diversion hours climbing without a known cause
Neighbouring hospitals are told about an outage by the news, not by the affected organisation
Cutting network links to contain an attack is often the right call, but it causes a cascade of its own. Partners that relied on the link lose the service, and organisations that were never infected shut systems down as a precaution because they lack clear central advice. The outage from containment can be larger than the outage from the attack.[1,16]
Warning signs
No pre-agreed criteria for when to disconnect from a partner or supplier
No plan for running the services that ride on that connection while it is cut
Central incident guidance takes hours to reach local sites
Most cascades need several existing weaknesses at once: an unpatched system, a portal without multi-factor authentication, a missing bounds check, an untested backup. Each one is tolerable alone and may sit unnoticed for months. The cascade happens when a trigger finds a path through all of them. Use the Swiss-cheese picture with care. A survey of quality and safety professionals found they read its parts (holes, slices, arrow) in very different ways, and critics argue it is too static and linear. Its defenders still consider it useful because it is systemic.[18,16,1,5,19,20,21]
Warning signs
Known findings (unpatched hosts, missing MFA, failed audits) carried forward year after year
Assessments that were done but gave no one the power to require fixes
Incident reviews that stop at the first human or technical 'root cause'
A race condition in DynamoDB's DNS automation broke a core AWS region for about 15 hours. Some cloud-hosted EHR users slowed or went to paper; others saw nothing.[22,23,24]
A grid collapse cut power to continental Spain and Portugal for about ten hours. Hospitals largely held on generators; care outside them did not.[25,26,27,28,29]
A faulty Rapid Response Content update to CrowdStrike's Falcon sensor crashed about 8.5 million Windows devices worldwide. Outside-in measurement found disrupted services at 759 of 2,232 US hospitals studied.[5,30,31,6,7,32,33]
Ransomware hit Synnovis, the pathology provider for several south-east London NHS trusts and GP practices. Blood testing and matching collapsed, more than 11,000 appointments and procedures were postponed, O-type blood ran short nationally, and one death was later partly attributed to a delayed result.[34,35,36,37,10,38,39,3]
A ransomware attack took Ascension's electronic records offline for about five weeks. Clinicians told KFF Health News of medication errors and delayed lab results, and one said he had no training for the attack; Ascension said its care teams were trained for such disruptions.[40,41]
Attackers used stolen credentials on a Change Healthcare Citrix remote-access portal that had no multi-factor authentication, then deployed ransomware nine days later. Disconnecting the clearinghouse stalled pharmacy claims, medical claims and payments across the US.[1,2,42,43,44]
A month-long ransomware attack on a health system with about 25% of regional inpatient discharges drove patients and ambulances to two unaffected academic EDs, raising their census, waits and stroke activations.[15,47]
Freezing weather knocked out generation and forced the largest controlled load shed in US history. Power loss spread to water systems and hospitals, and to patients at home on powered medical equipment.[48,49,12,11,14,13]
PathPower → Devices → Human handoff
27 September 2020United States (UHS acute and behavioral hospitals)Fell to tier 2: manual operationConnectivity & dataCascades
A security incident led UHS to suspend user access to IT applications across its US operations; facilities ran on offline documentation for up to several weeks.[50,51]
10 September 2017Hollywood, Florida, USAFell to tier 3: analog fallbackPowerCascades
A self-spreading ransomware worm infected 34 English trusts and 603 primary-care and other NHS organisations, and at least 46 more trusts were disrupted. Thousands of appointments were cancelled and five hospitals diverted ambulances.[16,54,55]
Storm surge flooded basements holding fuel tanks and pumps at two Manhattan hospitals whose generators sat on upper floors. Both hospitals evacuated.[56,57,58]
A network loop took down clinical applications at an academic medical centre for about four days, forcing a return to paper it had abandoned years earlier.[64,65,66,67]
How you'd know
Map your dependencies before the event. List every external service a clinical workflow needs (clearinghouse, lab, e-prescribing, identity, endpoint agents, cloud EHR) and mark any service shared with most of your region. If you cannot produce this map, you cannot see a cascade coming.[68,69]
Watch external availability, not only internal alarms. During the CrowdStrike outage, researchers detected disrupted hospital services from outside by scanning network ports and FHIR endpoints every few hours. Hospitals and regional coalitions can use the same kind of outside-in monitoring to spot outages at peers and suppliers.[6]
Track regional load signals: EMS diversion hours, ambulance arrivals, left-without-being-seen rates and stroke-code volume at your own ED. A sudden rise with no local cause may be a neighbour's outage reaching you.[15]
Require suppliers to tell you when they activate their contingency plan. HHS has proposed requiring business associates to report contingency-plan activation within 24 hours. Until that is final, put it in the contract.[44]
Treat a rise in workaround use as a signal. Manual claims, phoned results, O-negative use above baseline and paper orders all show that a coupled system has degraded, often before anyone declares an incident.[10,2]
What to do, tier by tier
What should already be in place at each degradation tier for this layer. Tier 0 is normal automated running; tier 3 is paper, batteries and judgement.
These are practices reported or recommended in the cited sources, gathered for reference. They are not a prescription for your organisation; judge what fits your setting, and check the current official text of any standard.
0Full automation
Build and keep a dependency map that links each clinical service to the suppliers, networks, devices and utilities it relies on. Review it every year and after any major change.[68,69]
Write cybersecurity and restoration requirements into supplier contracts: multi-factor authentication on remote access, a restoration-time commitment, notice within 24 hours of contingency activation, and yearly written evidence of safeguards.[44,1,4]
Require staged rollout for any vendor update that runs with kernel or administrator rights on clinical endpoints. Get control over the rollout ring in writing, and keep a sample of clinical workstations on a delayed ring.[5]
Do not let your primary and backup depend on the same thing. Where a function is life-critical, make sure the backup uses a different supplier, network path or software stack.[9,7]
Close known latent conditions on a deadline: unpatched internet-facing systems, remote-access portals without MFA, unsupported operating systems. Give someone the authority to enforce the deadline.[16,1]
1Assisted operation
Enrol an alternate for each concentrated supplier before you need it (for example, a second clearinghouse EDI enrolment, or a reference-lab agreement), and test the switch once a year.[2,17]
Set pre-agreed disconnection criteria with key partners: who can cut the link, what runs while it is cut, and what evidence brings it back.[1,16]
Keep a read-only copy of the EHR that can print, and test it regularly, so clinicians keep access to records while the primary system or its network is down.[67]
When your blood-matching or lab capacity drops, tell your blood supplier the same day, so that O-type stock can be managed nationally rather than drained locally.[10]
2Manual operation
Keep enough paper downtime forms in every care area for at least 8 hours of ordering, medication administration, lab and radiology, and run unannounced downtime drills at least once a year.[67]
Tell neighbouring hospitals and EMS early when you go on diversion or downtime, and agree in advance how to share load. Regional capacity is a shared resource in a cascade.[15,17]
Rank services by clinical criticality and restore in that order. HHS has proposed a 72-hour restoration target for critical systems. Test whether you could meet it.[44]
Prepare all staff, not only IT, to deliver care during an extended cyber downtime, and prioritise the services that must stay safe.[47]
3Analog fallback
Have signed arrangements with other hospitals to take your patients when your operations are limited or stopped, as the CMS emergency preparedness rule requires. Check that those hospitals do not depend on the same supplier, grid segment or water system as you.[17]
Plan for region-wide events where every neighbour is degraded at once. During the Texas freeze, one Austin hospital found no other hospital could take a large number of transfers.[14,11]
Plan for your community's electricity-dependent patients coming to the ED for oxygen and power when the grid fails. Keep space, outlets and oxygen supply for them.[13,12]
After recovery, review the incident as a control problem (CAST or similar), not a hunt for a root cause. Ask which constraints, feedback loops and decision-makers failed to stop the spread, including at suppliers.[19,70]
Standards and rules (US)
Instrument
What it requires
CMS Conditions of Participation, Emergency Preparedness, 42 CFR 482.15
Hospitals must base their emergency plan on a facility-based and community-based all-hazards risk assessment, have arrangements with other hospitals to receive patients if operations are limited or stop, keep a communication plan with primary and alternate means, and run exercises at least twice a year.[17]
HIPAA Security Rule NPRM, 90 FR 898 (Jan 6, 2025), RIN 0945-AA22 (proposed, not final)
Proposes written procedures to restore critical systems and data within 72 hours, yearly written verification of business associates' technical safeguards, and business-associate notice within 24 hours of activating a contingency plan. It cites the Change Healthcare attack.[44]
NIST SP 800-161 Rev. 1 (May 2022, updated Nov 2024)
Guidance for building cybersecurity supply chain risk management into strategy, policy and risk assessment for the products and services an organisation buys, across organisation, mission and system levels.[69]
ASTP/ONC SAFER Guide: Contingency Planning (2025)
Self-assessment practices for EHR downtime, including paper forms for at least 8 hours, a tested read-only backup EHR, and unannounced downtime drills at least once a year. CMS requires hospitals to attest to the SAFER Guides annually.[67]
Guidance, not a standard. It calls on organisations to prepare all staff to keep care safe through an extended cyberattack downtime.[47]
Elsewhere: EU and UK
EU: the NIS2 Directive (EU) 2022/2555 requires essential and important entities to manage supply-chain security, including the quality and resilience of suppliers' products and services and cybersecurity terms in contracts with direct suppliers (Art. 21). It also provides for coordinated EU risk assessments of critical supply chains (Art. 22). UK: after Synnovis, the Cyber Security and Resilience (Network and Information Systems) Bill would let regulators designate 'critical suppliers' to essential services, and the government's factsheet uses Synnovis as its case study. The bill was still before Parliament in mid-2026. Netherlands: the Dutch Safety Board found in 2020 that hospitals' awareness of IT-failure risk had not kept pace with their dependence on IT. It recommended that hospitals map IT-to-care dependencies, test and drill regularly, and analyse serious outages in depth.[71,3,68]
Severity score v0.1 draft
4Likelihood
5Blast radius
4Detectability (5 = hardest)
80of 125
FailSystems judgementJudgement, first draft. Likelihood 4: five large healthcare cascades in 2017–2024, three of them in 2024 alone, suggest the pattern recurs every year or two somewhere in the US or UK. Blast radius 5: cascades are by definition failures that spread beyond a single organisation; Change Healthcare and CrowdStrike reached national scale. Detectability 4: the triggering weakness is usually latent and often sits inside a supplier the hospital cannot see into, although once a cascade is running it is obvious.
Each factor is scored 1–5 and multiplied, as in a classic FMEA risk priority number. This is our first-draft judgement, not a measurement; see how scoring works and how it will be revised.
What we don't know yet
How much patient harm do cascades cause beyond the first organisation? Only Dameff 2023 measured spillover at neighbours, and it covered one region and one event.
Which healthcare functions are most concentrated in single suppliers nationally (clearinghouses, pathology, e-prescribing, identity, EHR hosting), and what share of hospitals share each one?
Do contract clauses (restoration times, contingency notice, staged rollout) actually shorten outages, or only move liability?
How long does downtime proficiency last after a drill, and how often must paper workflows be practised for a hospital to reach tier 3 safely?
Is disconnecting early to contain an attack net-beneficial for patients once the downstream outage it causes is counted?
These gaps drive what the nightly research pass looks for. If you have evidence, send it.
Cite this pageFailSystems. “Cascades.” https://failsystems.health201.com/layers/cascades/ (reviewed 2026-09-26). Health 201 / AstroNexus LLC. CC BY 4.0.
Information only, not advice. FailSystems is an aggregation and synthesis of published sources. It is not consulting, engineering, legal, regulatory or medical advice, and using it creates no professional relationship. Health systems are complex and no approach fits every organisation: anything you adopt is your own decision, at your own risk, and should be checked against the current official sources and by qualified people who know your setting. Full disclaimer.
Dealing with an incident right now? This site is a reference, not an incident-response service. Activate your organisation's emergency operations plan and incident command, and:
Power loss, disaster or resource needs: go through your local or county emergency management. They escalate to the state, and the state requests FEMA support; hospitals do not call FEMA directly.
A medical device problem: report it to the manufacturer and to FDA MedWatch.
Outside the US: your national emergency number and national cyber agency (in the UK, NCSC).