how automated healthcare fails, how you'd know, and what to do at each tier — every claim sourced, reviewed continuously
Layer 2 of 5
Connectivity & data
The networks, interfaces, vendors and records that carry clinical data, and what happens when data is missing, late or wrong.
Reviewed 26 September 2026Sources checked when written 26 September 2026Involved in 15 of 39 incidents70 sources (65 primary or secondary)
What this layer is
This layer is everything between a clinician's question and the data that answers it: local networks and internet circuits, the EHR and its interfaces to lab, pharmacy and imaging, and the outside services a hospital depends on, such as claims clearinghouses and outsourced pathology. It also covers the data itself: whether it is present, current and correct.
It fails in two different ways. Data can be unavailable: a ransomware attack, a network loop or a vendor outage takes systems down, and staff know they are blind. Or data can be wrong: an order silently goes to a queue no one reads, a downtime copy is hours old, or results entered on paper never make it back. Unavailable data is loud and prompts a switch to backup processes; wrong data is quiet and does not.
Unplanned downtime is common. In one survey, 96% of large US health systems had at least one in three years, and 70% had one longer than 8 hours. Ransomware has made multi-week outages routine: about 44% of ransomware attacks on US care delivery organizations from 2016 to 2021 disrupted care, and in-hospital mortality rises among patients already admitted when an attack begins.
FailSystems viewFailSystems' view: in a paper hospital, losing one department's records was a local problem. In an automated hospital, one identity system, one network core or one shared vendor carries every department's data, so failure is correlated and the backup is a mode of work nobody practises. We think the key distinction is 'unavailable' versus 'wrong'. Most planning targets the first: backups, warm sites, paper forms. The second defeats those plans because nothing tells anyone to use them. Defences against wrong data are reconciliation and monitoring (queues with owners, counts that must match, synthetic transactions), not redundancy.
How it fails
Enterprise ransomware and precautionary shutdown
Attackers encrypt servers and endpoints, often after days of undetected access and data theft. The organization then disconnects everything it cannot yet trust, so the EHR, lab, imaging, pharmacy and communications go dark together. Recovery is a rebuild, not a restart, and takes weeks.[1,2,3,4]
Warning signs
Remote-access portals or VPN accounts without MFA
Unexplained privileged-account activity or large outbound transfers
Backup jobs failing or being deleted
Alerts from CISA/HHS about a group active in the sector
A vendor that many organizations share (claims clearinghouse, pharmacy switch, outsourced pathology) is attacked or fails. Hospitals whose own systems are intact lose a function they cannot perform themselves, and every customer fails at once.[5,6,7,8]
Warning signs
One vendor handles a function with no tested alternative
Contract lacks incident-notification and recovery-time terms
Vendor's remote access into your network is not inventoried
Downtime procedures that decay over days and weeks
The Joint Commission advises hospitals to be prepared to run with life- and safety-critical technology offline for four weeks or longer. Over days, order routing between departments, patient identification and result communication break down; lab turnaround slows and medication checks lapse. Back-entry after recovery creates a second risk period.[9,10,11,12,13]
Warning signs
Downtime drills shorter than a shift or never unannounced
A loop, misconfiguration, carrier cut or failed core switch makes applications unreachable although servers and data are intact. Intermittent 'flapping' is worse than a clean outage because staff cannot tell whether to switch to paper.[14,15]
Warning signs
Single internet path or single carrier
Flat Layer-2 networks spanning buildings
Rising response times and intermittent timeouts
No one owns network lifecycle as a clinical system
Silent data loss or misrouting (data wrong, not absent)
Orders, results or messages are accepted by one system and never reach the next, or land in a queue no one watches. The sender sees success, so no one switches to a backup process. Harm emerges as missed follow-up weeks later.[16,15,17]
Warning signs
Interface error or dead-letter queues without a named owner
Order counts sent vs received that do not reconcile
Clinicians reporting 'I ordered it but nothing happened'
Stale or incomplete record during and after downtime
Read-only downtime copies are snapshots and age from the moment the outage begins. After restoration, data captured on paper is back-entered late or not at all, and results produced during the outage may be absent from the electronic record. Clinicians decide on data that looks current but is not.[15,10,18]
Warning signs
Read-only backup refreshed less than hourly or not tested
No reconciliation owner for paper records after downtime
Attackers target backup systems before encrypting, or backups turn out never to have been restored end to end. The organization then has no clean copy to restore from and must rebuild, or pay.[2,15,3]
When a system diverts ambulances and time-critical patients, nearby EDs absorb the load without extra staff. Waits, walk-outs and time-critical cases rise at hospitals that were never attacked; rural patients face much longer travel.[19,20,21]
Warning signs
A neighbouring system announces diversion or a cyber incident
Sudden EMS arrival increase without a mass-casualty event
A fault inside the provider (DNS automation, internal network congestion) disables core services across a region while the hospital's own building is fine. Impact depends on how each customer and each supplier built on the region: in October 2025 one Epic-on-AWS system slowed and another saw nothing, while NHS trusts using Oracle services went to paper.[22,23,24,25]
Warning signs
No map of which clinical and supplier services run in which region
A race condition in DynamoDB's DNS automation broke a core AWS region for about 15 hours. Some cloud-hosted EHR users slowed or went to paper; others saw nothing.[22,24,25]
A grid collapse cut power to continental Spain and Portugal for about ten hours. Hospitals largely held on generators; care outside them did not.[26,27,28,29,30]
A Class I software correction found that backlogged EHR-to-pump automated programming requests could load stale rate, dose or volume parameters.[31,32]
CISA and FDA reported that a low-cost patient monitor's firmware contained hidden functionality that could allow remote access and sent patient data to an external address; independent researchers later judged it an insecure design rather than an intentional backdoor.[33,34,35,36]
A faulty Rapid Response Content update to CrowdStrike's Falcon sensor crashed about 8.5 million Windows devices worldwide. Outside-in measurement found disrupted services at 759 of 2,232 US hospitals studied.[37,38,39,40,41,42,43]
Ransomware hit Synnovis, the pathology provider for several south-east London NHS trusts and GP practices. Blood testing and matching collapsed, more than 11,000 appointments and procedures were postponed, O-type blood ran short nationally, and one death was later partly attributed to a delayed result.[8,44,45,46,47,48,49,50]
A ransomware attack took Ascension's electronic records offline for about five weeks. Clinicians told KFF Health News of medication errors and delayed lab results, and one said he had no training for the attack; Ascension said its care teams were trained for such disruptions.[12,18]
Attackers used stolen credentials on a Change Healthcare Citrix remote-access portal that had no multi-factor authentication, then deployed ransomware nine days later. Disconnecting the clearinghouse stalled pharmacy claims, medical claims and payments across the US.[5,51,6,52,53]
A month-long ransomware attack on a health system with about 25% of regional inpatient discharges drove patients and ambulances to two unaffected academic EDs, raising their census, waits and stroke activations.[19,13]
PathConnectivity & data → Human handoff
October 2020Mann-Grandstaff VA Medical Center, Spokane, Washington, United StatesNo outage: wrong outputConnectivity & dataHuman handoff
After go-live, the new EHR routed more than 11,000 clinical orders to a hidden queue instead of the intended service, without telling the ordering clinician; VHA identified 149 adverse events.[16]
27 September 2020United States (UHS acute and behavioral hospitals)Fell to tier 2: manual operationConnectivity & dataCascades
A security incident led UHS to suspend user access to IT applications across its US operations; facilities ran on offline documentation for up to several weeks.[56,57]
A self-spreading ransomware worm infected 34 English trusts and 603 primary-care and other NHS organisations, and at least 46 more trusts were disrupted. Thousands of appointments were cancelled and five hospitals diverted ambulances.[58,21,59]
A network loop took down clinical applications at an academic medical centre for about four days, forcing a return to paper it had abandoned years earlier.[14,62,63,15]
How you'd know
Measure EHR response time for key clinical tasks (results review, order entry, patient lookup) continuously; ONC's SAFER guide sets the target at optimally under 2 seconds and defines a functional downtime as any hourly mean response time over 5 seconds, or 3 standard deviations above the mean. Use a synthetic 'test patient' order placed on a schedule to detect silent failure.[15]
Monitor every interface queue, error queue and 'unknown' or dead-letter queue daily, with a named owner and reconciliation of orders sent against orders received.[16,15]
Alert on backup job failures, deletions of backup sets, and failed restore tests; test full restores rather than job completion.[15,2]
Centralize logs and watch for known ransomware tactics (credential abuse on remote access, lateral movement, large outbound transfers) during the days between intrusion and encryption.[7,3,5]
Require third parties to report incidents to you promptly, and subscribe to HHS/CISA advisories and the HHS OCR breach portal so a vendor or neighbour's outage reaches you before patients do.[7,64]
Track EMS arrival and ED census against baseline; a sudden rise without a local cause may mean a neighbouring system has gone down.[19]
What to do, tier by tier
What should already be in place at each degradation tier for this layer. Tier 0 is normal automated running; tier 3 is paper, batteries and judgement.
These are practices reported or recommended in the cited sources, gathered for reference. They are not a prescription for your organisation; judge what fits your setting, and check the current official text of any standard.
0Full automation
Put phishing-resistant MFA on every remote-access portal, VPN and privileged account, starting with vendor access.[3,5,7]
Patch known exploited vulnerabilities promptly and retire unsupported operating systems, including those embedded in diagnostic devices.[3,58,21]
Keep a daily, encrypted, off-site backup separated from normal storage (air gap); keep several generations; back up system configuration monthly and before every upgrade.[15,2]
Build redundant network paths: two internet circuits in different trenches or from different providers, and a routed (not flat Layer-2) core.[15,14]
Inventory every third party that performs a clinical or revenue function you cannot do yourself; write incident-notification and recovery terms into the contract.[7,6]
Complete all nine SAFER Guides every year as a working review, not a yes/no box; CMS accepts 'no' as an answer, so the attestation alone proves nothing.[65,15]
1Assisted operation
Maintain a warm site that can run the whole EHR within 8 hours, more than 50 miles away, and fail over to it at least quarterly.[15]
Segment the network so a compromised zone can be isolated without disconnecting everything; plan in advance which segments stay up.[7,2,21]
Contract and test an alternate clearinghouse or claims submission route, and an alternate reference lab, before an outage.[6,5]
Write and test restoration procedures that bring critical systems and data back within 72 hours, ranked by clinical criticality. The proposed HIPAA Security Rule would require this; do not wait for the final rule.[53,66]
Size interface buffers so data queued during an outage is not lost, and alert users in the EHR when a clinical interface is down.[15]
2Manual operation
Run a read-only backup EHR refreshed at least hourly, tested weekly, printable, and on UPS or generator power at unit level; make sure staff can log in to it.[15]
Keep downtime communication independent of the EHR network (not email, websites or VoIP on the same infrastructure).[15,13]
Call downtime early: activate the warm site or downtime procedures before 2 hours of unplanned outage, not after.[15]
Double-check high-risk medications manually when barcode scanning is unavailable, and use positive patient identification procedures designed for downtime.[15,9,12]
3Analog fallback
Stock current paper forms for orders, medication administration, lab requisitions and results on every unit; keep a paper copy of the downtime policy on units and off-site.[15,67]
Run unannounced downtime drills at least yearly, and at least one exercise that assumes weeks, not hours, without the EHR, lab interface or clearinghouse.[15,13,12]
Assign a runner or courier system for orders and results between departments; paper without a routing method stalls.[12,10]
Agree regional diversion and mutual-aid plans with neighbouring hospitals and EMS for cyber incidents, not only physical disasters.[19,21]
Plan recovery as its own phase: assign owners to back-enter and reconcile paper data, restart interfaces in order, and review harm from delays.[15,13,8]
Standards and rules (US)
Instrument
What it requires
HIPAA Security Rule, 45 CFR 164.308(a)(7) Contingency plan
Covered entities must have a data backup plan, a disaster recovery plan and an emergency-mode operation plan; testing/revision and an applications-and-data criticality analysis are 'addressable'. A January 2025 NPRM would add written procedures to restore critical systems and data within 72 hours, but it was not final as of September 2026.[66]
Hospitals must maintain a system of medical documentation that preserves patient information and keeps records available in an emergency, with the emergency plan and training/testing program reviewed at least every 2 years.[67]
Self-assessment of 13 practices covering disaster recovery, generators, paper forms, tested backups, downtime training, independent communication, interface restart and downtime monitoring. CMS requires hospitals in the Medicare Promoting Interoperability Program to attest annually (yes or no) to completing all nine SAFER Guides.[15]
HHS 405(d) Health Industry Cybersecurity Practices (HICP), 2023 edition
Voluntary, sector-specific: ten practices against five threats including ransomware, scaled for small and large organizations in two technical volumes.[2]
The Joint Commission Sentinel Event Alert 67 (2023)
Not a standard itself; recommends downtime planning committees, response teams, staff training and communication for extended cyber downtime, and points to TJC continuity-of-operations and disaster-recovery requirements.[13]
Elsewhere: EU and UK
In the EU the NIS2 Directive (2022/2555) keeps healthcare within its scope and imposes cybersecurity risk-management and incident-notification duties; ENISA's 2023 health threat landscape found ransomware in 54% of 215 reported health-sector incidents and a dedicated ransomware programme in only 27% of surveyed organisations. In England, DHSC's 2023-2030 cyber strategy aims for all health and social care organisations, including critical suppliers, to be cyber resilient by 2030. WannaCry (2017) and Synnovis (2024) are the reference cases: the first showed how unpatched systems and precautionary disconnection spread disruption, the second how a single pathology supplier can halt a region's diagnostics for months.[68,69,70,58,8]
Severity score v0.1 draft
5Likelihood
5Blast radius
3Detectability (5 = hardest)
75of 125
FailSystems judgementJudgement: likelihood is 5 because unplanned EHR downtime is near-universal and ransomware attacks on care delivery roughly doubled between 2016 and 2021. Blast radius is 5 because shared vendors (Change Healthcare, Synnovis) and precautionary shutdowns take down whole regions or national functions at once. Detectability averages two extremes: outright outages are obvious (about 1), but silent misrouting and stale data can go unnoticed for months (about 5).
Each factor is scored 1–5 and multiplied, as in a classic FMEA risk priority number. This is our first-draft judgement, not a measurement; see how scoring works and how it will be revised.
What we don't know yet
How much patient harm comes from 'wrong data' failures (misrouting, stale copies, lost back-entry) compared with outright outages? No study measures both.
What downtime length should hospitals plan and drill for? Most plans cover 1-3 days, while major ransomware outages last 3-6 weeks.
Do tested read-only backups, warm sites or alternate clearinghouses measurably reduce harm or recovery time? Evidence is mostly self-assessment and expert opinion.
Will HHS finalize the 72-hour restoration requirement, and would recovery-time mandates change outcomes or only paperwork?
How concentrated are US clinical dependencies (clearinghouses, reference labs, hosted EHRs), and which vendors are single points of failure for a region?
These gaps drive what the nightly research pass looks for. If you have evidence, send it.
Cite this pageFailSystems. “Connectivity & data.” https://failsystems.health201.com/layers/connectivity/ (reviewed 2026-09-26). Health 201 / AstroNexus LLC. CC BY 4.0.
Information only, not advice. FailSystems is an aggregation and synthesis of published sources. It is not consulting, engineering, legal, regulatory or medical advice, and using it creates no professional relationship. Health systems are complex and no approach fits every organisation: anything you adopt is your own decision, at your own risk, and should be checked against the current official sources and by qualified people who know your setting. Full disclaimer.
Dealing with an incident right now? This site is a reference, not an incident-response service. Activate your organisation's emergency operations plan and incident command, and:
Power loss, disaster or resource needs: go through your local or county emergency management. They escalate to the state, and the state requests FEMA support; hospitals do not call FEMA directly.
A medical device problem: report it to the manufacturer and to FDA MedWatch.
Outside the US: your national emergency number and national cyber agency (in the UK, NCSC).