how automated healthcare fails, how you'd know, and what to do at each tier — every claim sourced, reviewed continuously
Layer 3 of 5
Devices & electronics
The monitors, pumps, ventilators, sensors and analyzers that measure and act on patients, and the software and updates that run them.
Reviewed 26 September 2026Sources checked when written 26 September 2026Involved in 13 of 39 incidents58 sources (54 primary or secondary)
What this layer is
This layer is the equipment at the bedside, in the lab and in the patient's home: physiologic monitors, pulse oximeters, infusion pumps, ventilators, continuous glucose monitors, point-of-care and lab analyzers, and the endpoint software (operating systems, security agents, interface adapters) that now runs on or beside them. Nearly all of it is software-driven, networked, and updated by a vendor after installation.
Devices fail in two ways. Loud failures stop the device, raise an alarm or crash the workstation; staff notice and fall back to another device or to manual care. Quiet failures keep producing numbers that look normal and are wrong: a pulse oximeter that reads high on darker skin, a glucose sensor that reads low, a lead analyzer that under-reports, a pump that loads a stale order. The quiet kind never triggers a downtime procedure.
In a paper-era hospital a device error reached one patient through one clinician who could see the device. In an automated hospital device outputs feed EHR flowsheets, early-warning scores, auto-programmed infusions and remote monitoring, and a single vendor update can reach every unit at once. The same connectivity that lets a pump receive an order from the EHR lets a faulty update or a compromised firmware image reach thousands of endpoints in minutes.
FailSystems viewFailSystems' view: a device that stops is a tier-2 problem you can plan for; a device that keeps reporting wrong numbers is the one that hurts people, because nothing in the system tells anyone to change tier. Automation makes this worse in two directions. It amplifies quiet error, because downstream scores, alerts and auto-programming consume the bad value without a human looking at the device. And it synchronizes loud failure, because fleet-wide updates (security agents, firmware, interface software) turn one vendor mistake into simultaneous failure across a hospital or a country. Defense in this layer is less about redundancy of boxes and more about independent cross-checks of values and control over when changes land.
How it fails
Systematic sensor bias in a subgroup
A sensor is accurate on the population it was validated on and biased on others. Pulse oximeters overestimate saturation in patients with darker skin, so hypoxemia is missed and treatment thresholds are crossed later. The device reports normally and nothing alarms.[1,2,3]
Warning signs
SpO2-SaO2 gaps that differ by patient group when paired values are audited
Validation data from the manufacturer that does not report performance by skin tone
Therapy eligibility or escalation rates that differ by group at the same recorded SpO2
Manufacturing or design defect producing plausible wrong values
A batch or design flaw makes a device report values in the normal range that are wrong: glucose sensors reading low, blood lead analyzers reading low. Users act on the number. Detection depends on someone comparing against an independent method, and on the manufacturer reporting promptly, which can fail.[4,5]
Warning signs
Clinical picture that does not match the device value
Discrepancies between point-of-care and reference lab results
Clusters of complaints about one lot or serial range
Changes to instructions for use without a clear safety notice
A vendor pushes a software, firmware or content update to every installed endpoint at once. If the update is faulty, every device or workstation that takes it fails together, and recovery is limited by hands-on remediation per machine. Security agents with kernel access are the extreme case because they update often and without customer staging.[6,7,8]
Warning signs
Endpoint agents or device firmware set to auto-update with no ring or delay
No inventory of which clinical devices run which agents
Recovery runbook that assumes remote management works
Stale or queued commands across a device integration
When EHR-to-device integrations (infusion auto-programming, order interfaces) back up, a queued command can arrive late and be applied to the device as if it were current. The value looks legitimate on the pump screen.[9]
Warning signs
Interface engine queue depth or latency rising
Pump parameters that differ from the current order
Alarms fail to sound (a low-battery alarm that does not fire, wrong priority), sound falsely (spurious power-loss alarms that stop therapy), or sound so often that staff tune them out. The Joint Commission counted 98 alarm-related sentinel events, 80 of them deaths, from 2009 to mid-2012.[10,11,12,13]
Warning signs
High non-actionable alarm rates per bed per day
Alarm limits left at defaults
Vendor corrections that mention alarm behavior
Near misses where an alarm was heard but not acted on
Devices that can connect to a network but no longer receive security updates, or that ship with hidden functions, provide a path to alter device behavior or reach the wider network. The only mitigation may be to disconnect, which removes remote monitoring.[14,15,16,17,18,19]
Warning signs
Devices on end-of-support operating systems
No SBOM or vulnerability disclosure contact from the vendor
Unexpected outbound traffic from device VLANs
Devices on flat networks with clinical workstations
Latent hardware hazard with slow recall remediation
A material or component degrades inside devices already in use, with no alarm. Once found, remediation depends on replacement supply, locating every unit (often in patients' homes) and clear communication, and can take years.[20,21,22]
Warning signs
Recall notices without a tracked list of your affected serial numbers
Recalls are posted, but the notice does not reach the clinician, biomed team or home patient using the device, or arrives without clear action. FDA posting dates reflect classification, which can lag the firm's action.[16,23,20]
Warning signs
No single owner for recall intake and closure
Recalls closed without serial-number reconciliation
Home devices supplied by third parties outside the hospital's view
Some Libre 3 and 3 Plus continuous glucose sensors read lower than actual glucose; FDA classified the correction Class I after reports of 860 serious injuries and 7 deaths.[4,16]
A grid collapse cut power to continental Spain and Portugal for about ten hours. Hospitals largely held on generators; care outside them did not.[24,25,26,27,28]
CISA and FDA reported that a low-cost patient monitor's firmware contained hidden functionality that could allow remote access and sent patient data to an external address; independent researchers later judged it an insecure design rather than an intentional backdoor.[14,15,29,30]
A faulty Rapid Response Content update to CrowdStrike's Falcon sensor crashed about 8.5 million Windows devices worldwide. Outside-in measurement found disrupted services at 759 of 2,232 US hospitals studied.[6,7,31,8,32,33,34]
PathDevices → Connectivity & data → Human handoff
30 June 2021Worldwide (about 15 million devices); USTier 0: automation stayed upDevices
Philips recalled about 15 million breathing devices because PE-PUR foam could degrade into particles and chemicals the patient could inhale; remediation ran years and ended in a consent decree.[20,21,22,12]
Freezing weather knocked out generation and forced the largest controlled load shed in US history. Power loss spread to water systems and hospitals, and to patients at home on powered medical equipment.[35,36,37,38,39,40]
PathPower → Devices → Human handoff
17 December 2020United States (University of Michigan and 178-hospital cohort; Johns Hopkins COVID-19 cohort)Study findingDevicesModels & agents
Paired SpO2/SaO2 data showed occult hypoxemia missed by pulse oximetry about three times as often in Black as in White patients, delaying treatment decisions.[1,2,41,3]
A self-spreading ransomware worm infected 34 English trusts and 603 primary-care and other NHS organisations, and at least 46 more trusts were disrupted. Thousands of appointments were cancelled and five hospitals diverted ambulances.[42,43,44]
PathConnectivity & data → Devices → Human handoff
June 2013United StatesNo outage: wrong outputDevices
Magellan's LeadCare devices, used for more than half of US blood lead tests 2013-2017, gave falsely low results on venous samples; the company delayed telling FDA for 21 months.[5]
Storm surge flooded basements holding fuel tanks and pumps at two Manhattan hospitals whose generators sat on upper floors. Both hospitals evacuated.[45,46,47]
A patient on a cardiac monitor died after the monitor's crisis alarm had been left off; lower-level alarms sounded at the nurses' station but went unheeded.[48,49]
2010Massachusetts, USA (hospital named in the Boston Globe report cited by the Joint Commission)Tier 0: automation stayed upSingle sourceHuman handoffDevices
A 60-year-old ICU patient's monitor alarmed for rising heart rate and falling oxygen saturation; staff responded only after about an hour, when he had stopped breathing.[10]
How you'd know
Audit paired device vs reference values (SpO2 vs SaO2, point-of-care vs lab glucose or lead) at least quarterly and break the gap down by patient group.[1,2]
Subscribe to the FDA Medical Device Recalls database and MAUDE for every device model in your inventory, remembering MDR counts do not establish cause or rate.[23,50]
Subscribe to CISA ICS medical advisories and match them against your device inventory by model and firmware version.[15]
Monitor device integration queues (EHR-to-pump auto-programming, interface engines) for latency and backlog, and alert on growth.[9]
Track alarm load and non-actionable alarm rates per unit; rising rates predict missed actionable alarms.[10,51]
Measure external reachability of your own clinical services; the CrowdStrike study showed internet scanning detected outages at 34% of hospitals.[8]
What to do, tier by tier
What should already be in place at each degradation tier for this layer. Tier 0 is normal automated running; tier 3 is paper, batteries and judgement.
These are practices reported or recommended in the cited sources, gathered for reference. They are not a prescription for your organisation; judge what fits your setting, and check the current official text of any standard.
0Full automation
Require manufacturers to supply accuracy data by skin tone for any oximeter you buy, and prefer devices tested under FDA's 2025 draft protocol.[3,41]
Require an SBOM, a coordinated vulnerability disclosure process and a stated patch timeline in every networked-device contract, mirroring FD&C Act 524B.[17,52,18]
Treat any device with USB, serial, Bluetooth or ethernet ports as internet-capable when you assess it; FDA does.[18]
Put every endpoint agent and device firmware update into staged rings (test group, one unit, then fleet) with a hold period; refuse vendors that cannot support customer-controlled staging.[6,7]
Run IEC 80001-1 risk management before connecting any device to the network, with clinical engineering, IT and the vendor named as owners.[53]
Keep a device inventory with model, serial, firmware version, network location and support end date, and reconcile every recall against it within 10 working days.[54,16]
1Assisted operation
When an oximetry reading does not fit the clinical picture, draw an arterial blood gas before withholding or delaying oxygen-threshold therapy.[1,2]
Verify rate, dose and volume on the pump against the current order before starting any auto-programmed infusion.[9]
Segment clinical devices onto their own networks and block outbound internet by default, so a compromised monitor can still monitor locally.[14,15]
Set alarm limits per patient population and document which alarm signals matter most on each unit, as NPG.01.05.01 requires.[51,10]
2Manual operation
Keep standalone (non-networked) monitors and pumps stocked on each critical unit for use when networked devices or their central stations are down.[16]
Pre-stage offline recovery kits (local admin credentials, disk-encryption recovery keys, bootable media) so endpoints can be restored by hand at scale.[6,31]
Give home patients on recalled sensors a verified fallback (fingerstick meter, strips) and tell them in writing which readings to trust.[4]
Switch to manual infusion programming with independent double-check when the interoperability layer is suspect.[9]
3Analog fallback
Keep paper vital-sign and infusion flowsheets on every unit and drill their use; ECRI ranks digital-darkness unpreparedness the second hazard of 2026.[16]
Maintain manual measurement skills and equipment (manual BP cuffs, gravity infusion sets with drip-rate charts) for when devices cannot be trusted.[16]
Pre-decide which elective procedures cancel when device fleets fail, so the call takes minutes, as it did at Mass General Brigham on 19 July 2024.[31]
Standards and rules (US)
Instrument
What it requires
ISO 14971:2019 (FDA recognition 5-125)
Manufacturers identify hazards, estimate and control risks, and monitor effectiveness of controls across the device life cycle, including post-production information.[55]
IEC 62304:2006+AMD1:2015
Life cycle processes for development and maintenance of medical device software, including software safety classification, change control and problem resolution.[56]
IEC 60601-1-8:2006+AMD1:2012+AMD2:2020
Requirements and tests for medical alarm systems: alarm priority categories, alarm signal characteristics and control states such as pausing and silencing.[13]
IEC 80001-1:2021
The healthcare delivery organization applies risk management for safety, effectiveness and security before, during and after connecting devices or health software to its IT infrastructure.[53]
FD&C Act section 524B (21 U.S.C. 360n-2)
Cyber device sponsors must submit a postmarket vulnerability plan, maintain processes to assure cybersecurity, ship patches on a justified regular cycle and critical fixes out of cycle, and provide an SBOM.[17]
Joint Commission NPG.01.05.01 (2026)
Hospitals identify the most important alarm signals, set policies for managing them, and educate staff; replaces NPSG.06.01.01 from January 2026.[51]
Elsewhere: EU and UK
In the EU, the Medical Device Regulation (EU) 2017/745 makes information security part of the essential requirements: Annex I 17.2 requires software to be built under state-of-the-art life cycle and risk management including information security, and 17.4 requires manufacturers to state minimum hardware, network and IT security requirements, including protection against unauthorised access. In Great Britain, amended post-market surveillance rules in force from 16 June 2025 cut the serious-incident reporting deadline from 30 to 15 days and require manufacturers to submit Field Safety Notices to the MHRA before they go to users, which targets the recall-communication failure mode directly.[57,58]
Severity score v0.1 draft
4Likelihood
4Blast radius
5Detectability (5 = hardest)
80of 125
FailSystems judgementJudgement: device faults reported to FDA are routine (Class I recalls on pumps, ventilators and sensors recur yearly), so likelihood is high. Blast radius is high because fleet updates and population-wide sensors (oximetry, CGMs, a dominant lead analyzer) spread one error across many patients. Detectability is scored hardest (5) because the most harmful mode is a plausible wrong value that triggers no alarm and no downtime procedure.
Each factor is scored 1–5 and multiplied, as in a classic FMEA risk priority number. This is our first-draft judgement, not a measurement; see how scoring works and how it will be revised.
What we don't know yet
How much patient harm does oximetry bias cause at the outcome level (mortality, ICU admission), beyond delayed treatment eligibility?
What share of hospital device fleets run end-of-support software or unmanaged agents, and how does that track with outage and incident exposure?
Do staged, customer-controlled update rings actually reduce fleet-failure blast radius in hospitals, and at what patch-delay cost for security?
How often do EHR-to-device integrations deliver stale or mismatched commands in routine operation, below the threshold of a recall?
Will FDA finalize the 2025 pulse oximeter guidance, and how fast will legacy oximeters in use be replaced?
These gaps drive what the nightly research pass looks for. If you have evidence, send it.
Sources cited on this page
Racial Bias in Pulse Oximetry Measurement. New England Journal of Medicine (Sjoding MW, Dickson RP, Iwashyna TJ, Gay SE, Valley TS), 17 December 2020.PrimaryPeer-reviewed · link checked 2026-09-26
Cite this pageFailSystems. “Devices & electronics.” https://failsystems.health201.com/layers/devices/ (reviewed 2026-09-26). Health 201 / AstroNexus LLC. CC BY 4.0.
Information only, not advice. FailSystems is an aggregation and synthesis of published sources. It is not consulting, engineering, legal, regulatory or medical advice, and using it creates no professional relationship. Health systems are complex and no approach fits every organisation: anything you adopt is your own decision, at your own risk, and should be checked against the current official sources and by qualified people who know your setting. Full disclaimer.
Dealing with an incident right now? This site is a reference, not an incident-response service. Activate your organisation's emergency operations plan and incident command, and:
Power loss, disaster or resource needs: go through your local or county emergency management. They escalate to the state, and the state requests FEMA support; hospitals do not call FEMA directly.
A medical device problem: report it to the manufacturer and to FDA MedWatch.
Outside the US: your national emergency number and national cyber agency (in the UK, NCSC).