how automated healthcare fails, how you'd know, and what to do at each tier — every claim sourced, reviewed continuously
Layer 1 of 5
Power & infrastructure
Electricity, fuel, water, heat and cooling, including the physical data centres everything clinical now runs on; when they fail, the decision support goes too.
Reviewed 26 September 2026Sources checked when written 26 September 2026Involved in 7 of 39 incidents36 sources (36 primary or secondary)
What this layer is
This layer is the physical base of the hospital: utility feeds, generators and automatic transfer switches, uninterruptible power supplies (UPS), fuel, water and heat, HVAC and data-centre cooling, and the physical plant of the off-site data centres that host the EHR and its connected services. Software and control-plane failures inside a cloud provider belong to Connectivity & data. In US hospitals it is regulated mainly through CMS's emergency preparedness rule (42 CFR 482.15) and the NFPA 99 and NFPA 110 codes that CMS enforces.
Everything else on this site stands on it. Monitors, infusion pumps and ventilators need electricity; the EHR, order entry, lab analyzers, PACS, paging and any AI model need electricity, cooling and a working data centre. The documented failures are rarely 'the generator did not exist'. They are a fuel pump in a flooded basement, a cooling unit that trips in record heat, two 'redundant' sites that share the same weather, or a DNS fault in a cloud region hundreds of miles away.
A paper-era hospital that lost power lost light, lifts and life-support devices. A 2026 hospital also loses its records, its order sets, its alerts and its models, IT loads frequently sit on UPS and generator branches that were sized and tested for life-safety loads, not for whole data centres, and a cloud region can fail while the building's lights stay on.
FailSystems viewFailSystems' view: automation moves cognitive work onto the power layer. When the lights go out in 2026 you lose the decision support, the medication checks and the patient's history, not just the monitors, and you lose them at the moment clinicians are also running an evacuation or a surge. Power failures are also where 'redundancy' is most often an illusion: backup systems share a basement, a heatwave, a fuel supplier or a cloud region with the thing they back up. We judge this layer's defining risk to be correlated failure, not single-component failure.
How it fails
Flood-exposed power chain
Generators are raised, but fuel tanks, fuel pumps, transfer switches or switchgear stay at or below grade. Water reaches the lowest component and the whole chain stops. CMS requires flood-free generator placement only for new construction, renovation or new generators, so older sites can remain exposed.[1,2,3,4]
Warning signs
Fuel pumps, day tanks or switchgear in basements or below the design flood elevation
Flood barriers that have never been tested against surge pressure
Staff have warned that 'a little water' would disable the electrical system
On-site fuel covers the design duration, but regional events close roads, knock out fuel pumps at stations and create competition for deliveries. Generators then run in 'constant fear' of stopping. The same fuel shortage keeps staff from getting to work.[1,5,6,7]
Warning signs
Less than two days of on-site fuel
No written priority-delivery agreement, or one supplier shared by every hospital in the region
Emergency power fails on real demand (generator, transfer switch or UPS)
Generators are tested monthly, but a real outage asks for hours or days at full building load. In 2003 multiple New York City hospital generators failed during the blackout, and in 2012 OIG found backup generators unreliable at 28 of the 69 Sandy-area hospitals that lost utility power. NFPA 110 and The Joint Commission set monthly and 36-month load tests to catch this. IT has a further gap: servers and network gear drop in the seconds before generators pick up unless a UPS carries them, and ONC's SAFER guide asks for at least 10 minutes of UPS for the EHR, tested monthly.[8,1,9,10,2,5]
Warning signs
Monthly tests below 30% of nameplate kW or below manufacturer exhaust temperature
No 4-hour test in the last 36 months
Transfer switches never exercised under real building load
Prior surveyor deficiency citations on emergency power
UPS batteries past rated life or no record of monthly UPS tests
Chillers, condensers and air handlers fail in extreme heat or lose their own supply, while the rest of the building still has power. Data-centre equipment overheats and fails within hours; frail patients overheat over days. Both are often seen as facilities problems rather than clinical ones until harm occurs.[11,12,13,14,6]
Warning signs
Condensers sited with poor airflow
End-of-life cooling plant with unfunded replacement
The backup sits in the same flood zone, weather system, grid or cloud region as the primary, so one cause takes out both. Guy's and St Thomas' two data centres backed each other up and failed on the same afternoon.[11,5]
Warning signs
Primary and backup data centres within the same metro area
All production and disaster-recovery workloads in one cloud region
A wide-area grid failure takes out water pressure, heating, fuel supply, telecoms and EMS at once. Hospitals on generators can still lose heat (boilers fed by city water), labs, imaging and records, and receive patients whose home medical devices have stopped.[7,15,16,17,18,19]
Warning signs
Boilers, sterilization or dialysis dependent on municipal water pressure
Fuel-supply sites not on the utility's critical-load list
No plan for electricity-dependent patients in the community
Single telecom carrier for clinical phones and paging
Environmental sensors, alerting and status dashboards run on the same storage, network or region as the systems they watch. When those fail, alerts stop at the worst moment. It happened inside a hospital data centre in 2022 and inside AWS in 2021.[11,20]
Warning signs
Temperature/humidity monitoring hosted on the production SAN or network
Alerts only by email through on-premises servers
Reliance on the provider's status page as your only signal
A grid collapse cut power to continental Spain and Portugal for about ten hours. Hospitals largely held on generators; care outside them did not.[21,22,17,18,23]
Freezing weather knocked out generation and forced the largest controlled load shed in US history. Power loss spread to water systems and hospitals, and to patients at home on powered medical equipment.[7,15,16,24,25,26]
PathPower → Devices → Human handoff
10 September 2017Hollywood, Florida, USAFell to tier 3: analog fallbackPowerCascades
Storm surge flooded basements holding fuel tanks and pumps at two Manhattan hospitals whose generators sat on upper floors. Both hospitals evacuated.[2,27,1]
During the 2003 blackout multiple NYC hospital emergency generators failed. The outage was associated with about 90 excess deaths citywide.[8,30]
How you'd know
Trend every monthly generator test: load as % of nameplate, exhaust temperature, time to transfer. Treat any test below 30% load or below manufacturer exhaust temperature as a failed test.[9,10]
Alarm on data-centre temperature and humidity early (at Guy's the first high-temperature alert, at 26°C, came at 11:29, more than an hour before the main cooling trips at 12:50) and route alerts through a path that does not depend on the data centre.[11]
Compare heat, flood and freeze forecasts against the design limits of cooling plant, flood defences and fuel supply; a first-ever red heat warning was issued four days before the Guy's and St Thomas' failure.[11]
Run your own synthetic checks against cloud-hosted clinical services and DNS; do not rely on the provider's status page, which can itself be impaired.[20,31]
Watch grid-operator emergency notices and water-utility pressure alerts; in Texas, hospitals lost heat when city water pressure dropped.[7,15]
Treat surveyor deficiency citations on emergency power as leading indicators; most Sandy-area hospitals had emergency-related citations before the storm.[1]
Track on-site fuel in hours at current load, not gallons, and confirm delivery contracts before forecast events.[1,5]
What to do, tier by tier
What should already be in place at each degradation tier for this layer. Tier 0 is normal automated running; tier 3 is paper, batteries and judgement.
These are practices reported or recommended in the cited sources, gathered for reference. They are not a prescription for your organisation; judge what fits your setting, and check the current official text of any standard.
0Full automation
Survey the whole power chain against the design flood level: generators, fuel tanks, fuel pumps, transfer switches, switchgear. CMS requires flood-free siting only for new work, so audit existing installations yourself.[3,1,2]
Put a disaster-recovery site outside your weather and grid: SAFER suggests a warm site more than 50 miles away and more than 20 miles from the coast, able to run the whole EHR within 8 hours, tested at least quarterly.[5,11]
Map which cloud regions host your EHR, your suppliers' services and your identity systems. Do not let production and recovery share one region.[31,32,33]
Design data-centre and patient-area cooling for record heat with margin, and fund end-of-life replacements before they become incidents.[11,13]
Host environmental monitoring and alerting on infrastructure independent of what it monitors, with an out-of-band alert path (SMS, pager).[11,20]
1Assisted operation
Set clinical triggers for partial outages: if lab results or order entry slow beyond a set threshold, open downtime command even though systems are technically up.[33,32]
Keep read-only downtime EHR workstations with printers on UPS or generator-backed outlets, and test them on a schedule.[5]
When cooling is failing, start a controlled shutdown of non-critical IT early to protect the clinical core, rather than waiting for hardware to fail.[11]
Give cloud and EHR suppliers a named contact and escalation path in your downtime plan; supplier incidents reached NHS trusts through Oracle and System C.[32]
2Manual operation
Test each generator monthly under load for at least 30 minutes at 30% of nameplate or manufacturer exhaust temperature, and for 4 hours every 36 months (NFPA 110, Joint Commission EC.02.05.07).[9,10,6]
Give the EHR at least 10 minutes of UPS and test the UPS monthly.[5]
Hold at least two days of fuel on site and sign priority delivery agreements that do not depend on the same supplier as every neighbour.[5,1,6]
Know which devices sit on emergency outlets and decide in advance who gets the limited outlets if you lose branches.[1]
Identify every system that needs city water (boilers, sterilizers, dialysis) and plan for loss of pressure.[15]
3Analog fallback
Rehearse paper operation for weeks, not hours; Guy's and St Thomas' ran a 'Paper Hospital' for several weeks.[11]
Rehearse evacuating ventilated, ICU and neonatal patients without elevators or power; NYU Langone moved 21 neonates in 4.5 hours.[27,1]
Send a paper summary with every evacuated patient; receiving hospitals after Sandy got patients with no records.[1,3]
Plan staff transport and fuel for staff vehicles; fuel shortages kept Sandy-area staff at home.[1]
Coordinate with EMS and public health on patients who use home ventilators, oxygen and dialysis; they arrive when the grid fails.[18,16,19]
Hospitals must provide alternate energy for safe temperatures, emergency lighting, fire alarm and sewage; site generators per NFPA 99/101; follow NFPA 99/110/101 emergency power testing and maintenance; have a fuel plan; exercise twice a year and review the plan at least every two years.[6]
NFPA 110, Standard for Emergency and Standby Power Systems (CMS enforces 2010 ed.; 2025 is current)
Sets performance, installation, maintenance and testing of emergency power systems, including monthly load exercise at 30% of nameplate or minimum exhaust temperature and a 4-hour test every 36 months.[34]
NFPA 99, Health Care Facilities Code (CMS enforces 2012 ed.)
Applies electrical and other building-system requirements by risk category; Category 1 covers systems whose failure is likely to cause major injury or death.[10]
The Joint Commission EC.02.05.07 (emergency power testing)
Accreditation standard requiring monthly generator load tests and the 36-month 4-hour test, aligned with NFPA 110.[9]
ONC/ASTP SAFER Guide: Contingency Planning (2025)
Recommended practices: EHR on UPS for at least 10 minutes, generator support for critical EHR functions, 2 days of fuel, flood-safe siting, a remote warm site, and a tested read-only backup EHR.[5]
Elsewhere: EU and UK
In England, Health Technical Memorandum 06-01 (NHS England; last updated April 2017) sets the legal, design, operation and maintenance expectations for hospital electrical infrastructure, including existing sites. The Guy's and St Thomas' review shows those rules did not reach data-centre cooling in practice. In the EU, the Critical Entities Resilience Directive (2022/2557) brings both health and energy into scope and requires designated critical entities to assess all relevant risks at least every four years and keep a resilience plan. The April 2025 Iberian blackout, analysed by the ENTSO-E expert panel, is the reference event for grid-wide failure in Europe.[35,36,11,21]
Severity score v0.1 draft
3Likelihood
5Blast radius
3Detectability (5 = hardest)
45of 125
FailSystems judgementOur judgement: whole-facility power loss is uncommon for any one hospital, but weather and heat events recur often enough across the sector to score 3. When it happens it removes every other layer at once, so blast radius is 5. A blackout itself is obvious, but the causes (a flooded fuel pump, an ageing condenser, a shared failure domain) stay hidden until the event, so detectability scores 3.
Each factor is scored 1–5 and multiplied, as in a classic FMEA risk priority number. This is our first-draft judgement, not a measurement; see how scoring works and how it will be revised.
What we don't know yet
How often do US hospital generators and transfer switches fail on real demand rather than in tests? There is no public, ongoing dataset; the most recent figure found (AP, 2012) is a one-off.
Are cloud-hosted EHRs more or less available than on-premises EHRs during regional events? The October 2025 reports (Tufts vs Baptist) are anecdotes, not a comparison.
How much patient harm do IT-only power and cooling failures cause? Guy's and St Thomas' is a rare published harm review, and it was still open.
How much of the estimated excess mortality after the Iberian blackout came from disrupted hospital and EMS care versus home conditions?
Do NFPA 110 test regimes (30% load monthly, 4 hours every 36 months) predict survival of multi-day outages, and how do CMS-waived microgrid alternatives perform in real events?
These gaps drive what the nightly research pass looks for. If you have evidence, send it.
Cite this pageFailSystems. “Power & infrastructure.” https://failsystems.health201.com/layers/power/ (reviewed 2026-09-26). Health 201 / AstroNexus LLC. CC BY 4.0.
Information only, not advice. FailSystems is an aggregation and synthesis of published sources. It is not consulting, engineering, legal, regulatory or medical advice, and using it creates no professional relationship. Health systems are complex and no approach fits every organisation: anything you adopt is your own decision, at your own risk, and should be checked against the current official sources and by qualified people who know your setting. Full disclaimer.
Dealing with an incident right now? This site is a reference, not an incident-response service. Activate your organisation's emergency operations plan and incident command, and:
Power loss, disaster or resource needs: go through your local or county emergency management. They escalate to the state, and the state requests FEMA support; hospitals do not call FEMA directly.
A medical device problem: report it to the manufacturer and to FDA MedWatch.
Outside the US: your national emergency number and national cyber agency (in the UK, NCSC).