A faulty Rapid Response Content update to CrowdStrike's Falcon sensor crashed about 8.5 million Windows devices worldwide. Outside-in measurement found disrupted services at 759 of 2,232 US hospitals studied.
PathDevices → Connectivity & data → Human handoff
Sources checked when written 26 September 2026
TierAt Mass General Brigham, on-site staff went to downtime procedures and handwritten notes, according to Healthcare Brew; a scan of US hospitals found patient-facing services offline at many sites. How tiers are assigned.
What happened
CrowdStrike's root cause analysis says a new IPC Template Type defined 21 input fields, but the sensor code that called it supplied only 20. Testing and early deployments used wildcard matching for the 21st field, so the mismatch stayed latent. On 19 July 2024 a new Channel File 291 used a non-wildcard criterion for the 21st field. That triggered an out-of-bounds memory read, and Windows hosts that received it crashed. CrowdStrike's remediations include bounds checks, more testing, staged deployment through rings, and customer control over content rollout. Microsoft estimated 8.5 million Windows devices were affected, under 1% of the total, and noted that the damage was outsized because enterprises running critical services use CrowdStrike. CISA confirmed it was not a cyberattack but warned of phishing that took advantage of the outage.
Tully et al. (JAMA Network Open, 2025) scanned hospital IP space and Epic FHIR endpoints from outside. Of 2,232 US hospitals, 759 (34.0%) had detectable disruptions. Of 1,098 disrupted services, 239 (21.8%) were patient-facing. Most services came back within 6 hours, but 43 were down for more than 48 hours. The authors note that network measurement is only a surrogate for clinical impact.[1,2,3,4,5,6,7]
Documented harm
None documented at the patient level. Tully et al. could not confirm patient outcomes. 239 patient-facing services were disrupted at US hospitals.
What it teaches
A security agent with kernel access and vendor-controlled auto-update is a single point of failure across every workstation and device that runs it.
Recovery time is set by hands-on per-machine remediation (disk encryption keys, physical access), not by how fast the vendor rolls back.
Demand staged, customer-controlled update rings for any agent on clinical endpoints.
Redundancy inside a hospital does not help when primary and backup run the same agent and receive the same update.
Security tooling with kernel-level access is itself a common-mode dependency.
Customer control over staged rollout is a contract term worth asking for.
Outside-in monitoring can measure a cascade across hundreds of hospitals within hours.
A single vendor update can push many hospitals down a tier at the same moment, so a peer hospital cannot be assumed to be at tier 0.
External measurement can detect the drop faster than internal reports.
Information only, not advice. FailSystems is an aggregation and synthesis of published sources. It is not consulting, engineering, legal, regulatory or medical advice, and using it creates no professional relationship. Health systems are complex and no approach fits every organisation: anything you adopt is your own decision, at your own risk, and should be checked against the current official sources and by qualified people who know your setting. Full disclaimer.
Dealing with an incident right now? This site is a reference, not an incident-response service. Activate your organisation's emergency operations plan and incident command, and:
Power loss, disaster or resource needs: go through your local or county emergency management. They escalate to the state, and the state requests FEMA support; hospitals do not call FEMA directly.
A medical device problem: report it to the manufacturer and to FDA MedWatch.
Outside the US: your national emergency number and national cyber agency (in the UK, NCSC).