{
  "layers": [
    {
      "id": "power",
      "name": "Power & infrastructure",
      "one_liner": "Electricity, fuel, water, heat and cooling, including the physical data centres everything clinical now runs on; when they fail, the decision support goes too.",
      "definition": [
        "This layer is the physical base of the hospital: utility feeds, generators and automatic transfer switches, uninterruptible power supplies (UPS), fuel, water and heat, HVAC and data-centre cooling, and the physical plant of the off-site data centres that host the EHR and its connected services. Software and control-plane failures inside a cloud provider belong to Connectivity & data. In US hospitals it is regulated mainly through CMS's emergency preparedness rule (42 CFR 482.15) and the NFPA 99 and NFPA 110 codes that CMS enforces.",
        "Everything else on this site stands on it. Monitors, infusion pumps and ventilators need electricity; the EHR, order entry, lab analyzers, PACS, paging and any AI model need electricity, cooling and a working data centre. The documented failures are rarely 'the generator did not exist'. They are a fuel pump in a flooded basement, a cooling unit that trips in record heat, two 'redundant' sites that share the same weather, or a DNS fault in a cloud region hundreds of miles away.",
        "A paper-era hospital that lost power lost light, lifts and life-support devices. A 2026 hospital also loses its records, its order sets, its alerts and its models, IT loads frequently sit on UPS and generator branches that were sized and tested for life-safety loads, not for whole data centres, and a cloud region can fail while the building's lights stay on."
      ],
      "angle": "FailSystems' view: automation moves cognitive work onto the power layer. When the lights go out in 2026 you lose the decision support, the medication checks and the patient's history, not just the monitors, and you lose them at the moment clinicians are also running an evacuation or a surge. Power failures are also where 'redundancy' is most often an illusion: backup systems share a basement, a heatwave, a fuel supplier or a cloud region with the thing they back up. We judge this layer's defining risk to be correlated failure, not single-component failure.",
      "failure_modes": [
        {
          "id": "flood-exposed-power-chain",
          "name": "Flood-exposed power chain",
          "mechanism": "Generators are raised, but fuel tanks, fuel pumps, transfer switches or switchgear stay at or below grade. Water reaches the lowest component and the whole chain stops. CMS requires flood-free generator placement only for new construction, renovation or new generators, so older sites can remain exposed.",
          "warning_signs": [
            "Fuel pumps, day tanks or switchgear in basements or below the design flood elevation",
            "Flood barriers that have never been tested against surge pressure",
            "Staff have warned that 'a little water' would disable the electrical system"
          ],
          "incident_ids": [
            "sandy-nyu-bellevue-2012",
            "katrina-memorial-2005"
          ],
          "source_ids": [
            "hhs-oig-2014-sandy",
            "ap-2012-nyc-hospital-generators",
            "cms-2016-ep-final-rule",
            "fink-2009-propublica-memorial"
          ]
        },
        {
          "id": "fuel-exhaustion",
          "name": "Fuel runs out or cannot be delivered",
          "mechanism": "On-site fuel covers the design duration, but regional events close roads, knock out fuel pumps at stations and create competition for deliveries. Generators then run in 'constant fear' of stopping. The same fuel shortage keeps staff from getting to work.",
          "warning_signs": [
            "Less than two days of on-site fuel",
            "No written priority-delivery agreement, or one supplier shared by every hospital in the region",
            "Fuel quality and polishing not tested"
          ],
          "incident_ids": [
            "sandy-nyu-bellevue-2012",
            "texas-winter-storm-2021"
          ],
          "source_ids": [
            "hhs-oig-2014-sandy",
            "onc-safer-contingency-2025",
            "ecfr-42cfr482-15",
            "ferc-nerc-2021-cold-weather"
          ]
        },
        {
          "id": "generator-fails-on-real-demand",
          "name": "Emergency power fails on real demand (generator, transfer switch or UPS)",
          "mechanism": "Generators are tested monthly, but a real outage asks for hours or days at full building load. In 2003 multiple New York City hospital generators failed during the blackout, and in 2012 OIG found backup generators unreliable at 28 of the 69 Sandy-area hospitals that lost utility power. NFPA 110 and The Joint Commission set monthly and 36-month load tests to catch this. IT has a further gap: servers and network gear drop in the seconds before generators pick up unless a UPS carries them, and ONC's SAFER guide asks for at least 10 minutes of UPS for the EHR, tested monthly.",
          "warning_signs": [
            "Monthly tests below 30% of nameplate kW or below manufacturer exhaust temperature",
            "No 4-hour test in the last 36 months",
            "Transfer switches never exercised under real building load",
            "Prior surveyor deficiency citations on emergency power",
            "UPS batteries past rated life or no record of monthly UPS tests",
            "Downtime EHR workstations on ordinary outlets"
          ],
          "incident_ids": [
            "northeast-blackout-2003",
            "sandy-nyu-bellevue-2012"
          ],
          "source_ids": [
            "beatty-2006-phr-blackout-2003",
            "hhs-oig-2014-sandy",
            "hfm-2022-generator-itm",
            "cms-2016-fire-safety-rule",
            "ap-2012-nyc-hospital-generators",
            "onc-safer-contingency-2025"
          ]
        },
        {
          "id": "cooling-loss",
          "name": "Cooling lost while power stays on",
          "mechanism": "Chillers, condensers and air handlers fail in extreme heat or lose their own supply, while the rest of the building still has power. Data-centre equipment overheats and fails within hours; frail patients overheat over days. Both are often seen as facilities problems rather than clinical ones until harm occurs.",
          "warning_signs": [
            "Condensers sited with poor airflow",
            "End-of-life cooling plant with unfunded replacement",
            "Record-heat forecasts beyond cooling design conditions",
            "Cooling plant or its transformer not on emergency power"
          ],
          "incident_ids": [
            "gstt-heatwave-datacentre-2022",
            "hollywood-hills-irma-2017"
          ],
          "source_ids": [
            "gstt-2023-it-incident-review",
            "google-cloud-2022-europe-west2",
            "npr-2019-hollywood-hills",
            "skarha-2021-jamahf-irma-nursing-homes",
            "ecfr-42cfr482-15"
          ]
        },
        {
          "id": "correlated-redundancy",
          "name": "Redundancy that shares a failure domain",
          "mechanism": "The backup sits in the same flood zone, weather system, grid or cloud region as the primary, so one cause takes out both. Guy's and St Thomas' two data centres backed each other up and failed on the same afternoon.",
          "warning_signs": [
            "Primary and backup data centres within the same metro area",
            "All production and disaster-recovery workloads in one cloud region",
            "Backup never failed over under realistic load",
            "Suppliers' own hosting unknown to you"
          ],
          "incident_ids": [
            "gstt-heatwave-datacentre-2022"
          ],
          "source_ids": [
            "gstt-2023-it-incident-review",
            "onc-safer-contingency-2025"
          ]
        },
        {
          "id": "upstream-utility-cascade",
          "name": "Grid-wide loss cascading through other utilities",
          "mechanism": "A wide-area grid failure takes out water pressure, heating, fuel supply, telecoms and EMS at once. Hospitals on generators can still lose heat (boilers fed by city water), labs, imaging and records, and receive patients whose home medical devices have stopped.",
          "warning_signs": [
            "Boilers, sterilization or dialysis dependent on municipal water pressure",
            "Fuel-supply sites not on the utility's critical-load list",
            "No plan for electricity-dependent patients in the community",
            "Single telecom carrier for clinical phones and paging"
          ],
          "incident_ids": [
            "texas-winter-storm-2021",
            "iberian-blackout-2025",
            "northeast-blackout-2003"
          ],
          "source_ids": [
            "ferc-nerc-2021-cold-weather",
            "texastribune-2021-austin-hospitals",
            "dshs-2021-winter-storm-deaths",
            "goiana-da-silva-2025-fph-iberian",
            "castro-delgado-2025-pdm-blackout-ems",
            "casey-2020-cehr-power-outages"
          ]
        },
        {
          "id": "monitoring-shares-fate",
          "name": "Monitoring fails with the thing it monitors",
          "mechanism": "Environmental sensors, alerting and status dashboards run on the same storage, network or region as the systems they watch. When those fail, alerts stop at the worst moment. It happened inside a hospital data centre in 2022 and inside AWS in 2021.",
          "warning_signs": [
            "Temperature/humidity monitoring hosted on the production SAN or network",
            "Alerts only by email through on-premises servers",
            "Reliance on the provider's status page as your only signal"
          ],
          "incident_ids": [
            "gstt-heatwave-datacentre-2022"
          ],
          "source_ids": [
            "gstt-2023-it-incident-review",
            "aws-2021-us-east-1-pes"
          ]
        }
      ],
      "detection": [
        {
          "text": "Trend every monthly generator test: load as % of nameplate, exhaust temperature, time to transfer. Treat any test below 30% load or below manufacturer exhaust temperature as a failed test.",
          "source_ids": [
            "hfm-2022-generator-itm",
            "cms-2016-fire-safety-rule"
          ]
        },
        {
          "text": "Alarm on data-centre temperature and humidity early (at Guy's the first high-temperature alert, at 26°C, came at 11:29, more than an hour before the main cooling trips at 12:50) and route alerts through a path that does not depend on the data centre.",
          "source_ids": [
            "gstt-2023-it-incident-review"
          ]
        },
        {
          "text": "Compare heat, flood and freeze forecasts against the design limits of cooling plant, flood defences and fuel supply; a first-ever red heat warning was issued four days before the Guy's and St Thomas' failure.",
          "source_ids": [
            "gstt-2023-it-incident-review"
          ]
        },
        {
          "text": "Run your own synthetic checks against cloud-hosted clinical services and DNS; do not rely on the provider's status page, which can itself be impaired.",
          "source_ids": [
            "aws-2021-us-east-1-pes",
            "aws-2025-dynamodb-pes"
          ]
        },
        {
          "text": "Watch grid-operator emergency notices and water-utility pressure alerts; in Texas, hospitals lost heat when city water pressure dropped.",
          "source_ids": [
            "ferc-nerc-2021-cold-weather",
            "texastribune-2021-austin-hospitals"
          ]
        },
        {
          "text": "Treat surveyor deficiency citations on emergency power as leading indicators; most Sandy-area hospitals had emergency-related citations before the storm.",
          "source_ids": [
            "hhs-oig-2014-sandy"
          ]
        },
        {
          "text": "Track on-site fuel in hours at current load, not gallons, and confirm delivery contracts before forecast events.",
          "source_ids": [
            "hhs-oig-2014-sandy",
            "onc-safer-contingency-2025"
          ]
        }
      ],
      "defenses_by_tier": {
        "0": [
          {
            "text": "Survey the whole power chain against the design flood level: generators, fuel tanks, fuel pumps, transfer switches, switchgear. CMS requires flood-free siting only for new work, so audit existing installations yourself.",
            "source_ids": [
              "cms-2016-ep-final-rule",
              "hhs-oig-2014-sandy",
              "ap-2012-nyc-hospital-generators"
            ]
          },
          {
            "text": "Put a disaster-recovery site outside your weather and grid: SAFER suggests a warm site more than 50 miles away and more than 20 miles from the coast, able to run the whole EHR within 8 hours, tested at least quarterly.",
            "source_ids": [
              "onc-safer-contingency-2025",
              "gstt-2023-it-incident-review"
            ]
          },
          {
            "text": "Map which cloud regions host your EHR, your suppliers' services and your identity systems. Do not let production and recovery share one region.",
            "source_ids": [
              "aws-2025-dynamodb-pes",
              "digitalhealth-2025-aws-nhs",
              "beckers-2025-aws-tufts"
            ]
          },
          {
            "text": "Design data-centre and patient-area cooling for record heat with margin, and fund end-of-life replacements before they become incidents.",
            "source_ids": [
              "gstt-2023-it-incident-review",
              "npr-2019-hollywood-hills"
            ]
          },
          {
            "text": "Host environmental monitoring and alerting on infrastructure independent of what it monitors, with an out-of-band alert path (SMS, pager).",
            "source_ids": [
              "gstt-2023-it-incident-review",
              "aws-2021-us-east-1-pes"
            ]
          }
        ],
        "1": [
          {
            "text": "Set clinical triggers for partial outages: if lab results or order entry slow beyond a set threshold, open downtime command even though systems are technically up.",
            "source_ids": [
              "beckers-2025-aws-tufts",
              "digitalhealth-2025-aws-nhs"
            ]
          },
          {
            "text": "Keep read-only downtime EHR workstations with printers on UPS or generator-backed outlets, and test them on a schedule.",
            "source_ids": [
              "onc-safer-contingency-2025"
            ]
          },
          {
            "text": "When cooling is failing, start a controlled shutdown of non-critical IT early to protect the clinical core, rather than waiting for hardware to fail.",
            "source_ids": [
              "gstt-2023-it-incident-review"
            ]
          },
          {
            "text": "Give cloud and EHR suppliers a named contact and escalation path in your downtime plan; supplier incidents reached NHS trusts through Oracle and System C.",
            "source_ids": [
              "digitalhealth-2025-aws-nhs"
            ]
          }
        ],
        "2": [
          {
            "text": "Test each generator monthly under load for at least 30 minutes at 30% of nameplate or manufacturer exhaust temperature, and for 4 hours every 36 months (NFPA 110, Joint Commission EC.02.05.07).",
            "source_ids": [
              "hfm-2022-generator-itm",
              "cms-2016-fire-safety-rule",
              "ecfr-42cfr482-15"
            ]
          },
          {
            "text": "Give the EHR at least 10 minutes of UPS and test the UPS monthly.",
            "source_ids": [
              "onc-safer-contingency-2025"
            ]
          },
          {
            "text": "Hold at least two days of fuel on site and sign priority delivery agreements that do not depend on the same supplier as every neighbour.",
            "source_ids": [
              "onc-safer-contingency-2025",
              "hhs-oig-2014-sandy",
              "ecfr-42cfr482-15"
            ]
          },
          {
            "text": "Know which devices sit on emergency outlets and decide in advance who gets the limited outlets if you lose branches.",
            "source_ids": [
              "hhs-oig-2014-sandy"
            ]
          },
          {
            "text": "Identify every system that needs city water (boilers, sterilizers, dialysis) and plan for loss of pressure.",
            "source_ids": [
              "texastribune-2021-austin-hospitals"
            ]
          }
        ],
        "3": [
          {
            "text": "Rehearse paper operation for weeks, not hours; Guy's and St Thomas' ran a 'Paper Hospital' for several weeks.",
            "source_ids": [
              "gstt-2023-it-incident-review"
            ]
          },
          {
            "text": "Rehearse evacuating ventilated, ICU and neonatal patients without elevators or power; NYU Langone moved 21 neonates in 4.5 hours.",
            "source_ids": [
              "espiritu-2014-pediatrics-nicu-sandy",
              "hhs-oig-2014-sandy"
            ]
          },
          {
            "text": "Send a paper summary with every evacuated patient; receiving hospitals after Sandy got patients with no records.",
            "source_ids": [
              "hhs-oig-2014-sandy",
              "cms-2016-ep-final-rule"
            ]
          },
          {
            "text": "Plan staff transport and fuel for staff vehicles; fuel shortages kept Sandy-area staff at home.",
            "source_ids": [
              "hhs-oig-2014-sandy"
            ]
          },
          {
            "text": "Coordinate with EMS and public health on patients who use home ventilators, oxygen and dialysis; they arrive when the grid fails.",
            "source_ids": [
              "castro-delgado-2025-pdm-blackout-ems",
              "dshs-2021-winter-storm-deaths",
              "casey-2020-cehr-power-outages"
            ]
          }
        ]
      },
      "standards": [
        {
          "name": "42 CFR 482.15 — CMS Emergency Preparedness Condition of Participation (2016 rule)",
          "what_it_requires": "Hospitals must provide alternate energy for safe temperatures, emergency lighting, fire alarm and sewage; site generators per NFPA 99/101; follow NFPA 99/110/101 emergency power testing and maintenance; have a fuel plan; exercise twice a year and review the plan at least every two years.",
          "source_id": "ecfr-42cfr482-15"
        },
        {
          "name": "NFPA 110, Standard for Emergency and Standby Power Systems (CMS enforces 2010 ed.; 2025 is current)",
          "what_it_requires": "Sets performance, installation, maintenance and testing of emergency power systems, including monthly load exercise at 30% of nameplate or minimum exhaust temperature and a 4-hour test every 36 months.",
          "source_id": "nfpa-110"
        },
        {
          "name": "NFPA 99, Health Care Facilities Code (CMS enforces 2012 ed.)",
          "what_it_requires": "Applies electrical and other building-system requirements by risk category; Category 1 covers systems whose failure is likely to cause major injury or death.",
          "source_id": "cms-2016-fire-safety-rule"
        },
        {
          "name": "The Joint Commission EC.02.05.07 (emergency power testing)",
          "what_it_requires": "Accreditation standard requiring monthly generator load tests and the 36-month 4-hour test, aligned with NFPA 110.",
          "source_id": "hfm-2022-generator-itm"
        },
        {
          "name": "ONC/ASTP SAFER Guide: Contingency Planning (2025)",
          "what_it_requires": "Recommended practices: EHR on UPS for at least 10 minutes, generator support for critical EHR functions, 2 days of fuel, flood-safe siting, a remote warm site, and a tested read-only backup EHR.",
          "source_id": "onc-safer-contingency-2025"
        }
      ],
      "elsewhere": {
        "text": "In England, Health Technical Memorandum 06-01 (NHS England; last updated April 2017) sets the legal, design, operation and maintenance expectations for hospital electrical infrastructure, including existing sites. The Guy's and St Thomas' review shows those rules did not reach data-centre cooling in practice. In the EU, the Critical Entities Resilience Directive (2022/2557) brings both health and energy into scope and requires designated critical entities to assess all relevant risks at least every four years and keep a resilience plan. The April 2025 Iberian blackout, analysed by the ENTSO-E expert panel, is the reference event for grid-wide failure in Europe.",
        "source_ids": [
          "nhse-htm-06-01",
          "eu-cer-directive-2022-2557",
          "gstt-2023-it-incident-review",
          "entsoe-2026-iberian-final-report"
        ]
      },
      "scoring_v01": {
        "likelihood": 3,
        "blast_radius": 5,
        "detectability": 3,
        "rationale": "Our judgement: whole-facility power loss is uncommon for any one hospital, but weather and heat events recur often enough across the sector to score 3. When it happens it removes every other layer at once, so blast radius is 5. A blackout itself is obvious, but the causes (a flooded fuel pump, an ageing condenser, a shared failure domain) stay hidden until the event, so detectability scores 3."
      },
      "open_questions": [
        "How often do US hospital generators and transfer switches fail on real demand rather than in tests? There is no public, ongoing dataset; the most recent figure found (AP, 2012) is a one-off.",
        "Are cloud-hosted EHRs more or less available than on-premises EHRs during regional events? The October 2025 reports (Tufts vs Baptist) are anecdotes, not a comparison.",
        "How much patient harm do IT-only power and cooling failures cause? Guy's and St Thomas' is a rare published harm review, and it was still open.",
        "How much of the estimated excess mortality after the Iberian blackout came from disrupted hospital and EMS care versus home conditions?",
        "Do NFPA 110 test regimes (30% load monthly, 4 hours every 36 months) predict survival of multi-day outages, and how do CMS-waived microgrid alternatives perform in real events?"
      ],
      "watch_queries": [
        "\"hospital\" (\"generator failed\" OR \"backup power failed\") evacuate",
        "\"hospital\" \"power outage\" \"paper\" EHR downtime",
        "\"data centre\" OR \"data center\" cooling failure hospital IT outage",
        "(AWS OR Azure OR \"Google Cloud\" OR Oracle) outage hospital EHR",
        "NHS trust critical incident power failure IT",
        "blackout hospitals generators (grid OR heatwave OR storm)",
        "\"power outage\"[tiab] AND (hospital[tiab] OR \"emergency medical services\"[tiab]) — PubMed",
        "\"electricity-dependent\" medical equipment outage mortality — PubMed",
        "CMS emergency preparedness deficiency generator K918",
        "ENTSO-E OR NERC grid incident report hospitals critical load"
      ],
      "last_reviewed": "2026-09-26",
      "reviewed_by": "opus-research",
      "last_audited": "2026-09-26",
      "audit_status": "verified"
    },
    {
      "id": "connectivity",
      "name": "Connectivity & data",
      "one_liner": "The networks, interfaces, vendors and records that carry clinical data, and what happens when data is missing, late or wrong.",
      "definition": [
        "This layer is everything between a clinician's question and the data that answers it: local networks and internet circuits, the EHR and its interfaces to lab, pharmacy and imaging, and the outside services a hospital depends on, such as claims clearinghouses and outsourced pathology. It also covers the data itself: whether it is present, current and correct.",
        "It fails in two different ways. Data can be unavailable: a ransomware attack, a network loop or a vendor outage takes systems down, and staff know they are blind. Or data can be wrong: an order silently goes to a queue no one reads, a downtime copy is hours old, or results entered on paper never make it back. Unavailable data is loud and prompts a switch to backup processes; wrong data is quiet and does not.",
        "Unplanned downtime is common. In one survey, 96% of large US health systems had at least one in three years, and 70% had one longer than 8 hours. Ransomware has made multi-week outages routine: about 44% of ransomware attacks on US care delivery organizations from 2016 to 2021 disrupted care, and in-hospital mortality rises among patients already admitted when an attack begins."
      ],
      "angle": "FailSystems' view: in a paper hospital, losing one department's records was a local problem. In an automated hospital, one identity system, one network core or one shared vendor carries every department's data, so failure is correlated and the backup is a mode of work nobody practises. We think the key distinction is 'unavailable' versus 'wrong'. Most planning targets the first: backups, warm sites, paper forms. The second defeats those plans because nothing tells anyone to use them. Defences against wrong data are reconciliation and monitoring (queues with owners, counts that must match, synthetic transactions), not redundancy.",
      "failure_modes": [
        {
          "id": "enterprise-ransomware-shutdown",
          "name": "Enterprise ransomware and precautionary shutdown",
          "mechanism": "Attackers encrypt servers and endpoints, often after days of undetected access and data theft. The organization then disconnects everything it cannot yet trust, so the EHR, lab, imaging, pharmacy and communications go dark together. Recovery is a rebuild, not a restart, and takes weeks.",
          "warning_signs": [
            "Remote-access portals or VPN accounts without MFA",
            "Unexplained privileged-account activity or large outbound transfers",
            "Backup jobs failing or being deleted",
            "Alerts from CISA/HHS about a group active in the sector"
          ],
          "incident_ids": [
            "ascension-2024",
            "uhs-2020",
            "wannacry-nhs-2017"
          ],
          "source_ids": [
            "neprash-2022-jama-hf-ransomware-trends",
            "hhs-405d-hicp-2023",
            "cisa-aa24-131a-black-basta",
            "neprash-2026-aej-policy-hacked-to-pieces"
          ]
        },
        {
          "id": "third-party-dependency-outage",
          "name": "Third-party clearinghouse or lab outage",
          "mechanism": "A vendor that many organizations share (claims clearinghouse, pharmacy switch, outsourced pathology) is attacked or fails. Hospitals whose own systems are intact lose a function they cannot perform themselves, and every customer fails at once.",
          "warning_signs": [
            "One vendor handles a function with no tested alternative",
            "Contract lacks incident-notification and recovery-time terms",
            "Vendor's remote access into your network is not inventoried"
          ],
          "incident_ids": [
            "change-healthcare-2024",
            "synnovis-2024"
          ],
          "source_ids": [
            "witty-2024-senate-finance-testimony",
            "aha-change-underscores-preparedness",
            "hhs-hph-cpgs",
            "nhs-england-synnovis-incident"
          ]
        },
        {
          "id": "extended-downtime-procedure-decay",
          "name": "Downtime procedures that decay over days and weeks",
          "mechanism": "The Joint Commission advises hospitals to be prepared to run with life- and safety-critical technology offline for four weeks or longer. Over days, order routing between departments, patient identification and result communication break down; lab turnaround slows and medication checks lapse. Back-entry after recovery creates a second risk period.",
          "warning_signs": [
            "Downtime drills shorter than a shift or never unannounced",
            "Paper forms out of date or missing on units",
            "Staff who have never worked without the EHR",
            "Downtime reports showing procedures not followed"
          ],
          "incident_ids": [
            "ascension-2024",
            "synnovis-2024"
          ],
          "source_ids": [
            "larsen-2018-jamia-ehr-downtime-events",
            "larsen-2019-aci-downtime-laboratory",
            "sittig-2014-ijmi-contingency-survey",
            "kff-2024-ascension-lapses",
            "tjc-sea-67-cyberattack"
          ]
        },
        {
          "id": "network-partition",
          "name": "Network partition or infrastructure collapse",
          "mechanism": "A loop, misconfiguration, carrier cut or failed core switch makes applications unreachable although servers and data are intact. Intermittent 'flapping' is worse than a clean outage because staff cannot tell whether to switch to paper.",
          "warning_signs": [
            "Single internet path or single carrier",
            "Flat Layer-2 networks spanning buildings",
            "Rising response times and intermittent timeouts",
            "No one owns network lifecycle as a clinical system"
          ],
          "incident_ids": [
            "bidmc-network-2002"
          ],
          "source_ids": [
            "cio-2003-bidmc-network",
            "onc-safer-contingency-2025"
          ]
        },
        {
          "id": "silent-data-loss-misrouting",
          "name": "Silent data loss or misrouting (data wrong, not absent)",
          "mechanism": "Orders, results or messages are accepted by one system and never reach the next, or land in a queue no one watches. The sender sees success, so no one switches to a backup process. Harm emerges as missed follow-up weeks later.",
          "warning_signs": [
            "Interface error or dead-letter queues without a named owner",
            "Order counts sent vs received that do not reconcile",
            "Clinicians reporting 'I ordered it but nothing happened'"
          ],
          "incident_ids": [
            "va-unknown-queue-2020"
          ],
          "source_ids": [
            "va-oig-2022-unknown-queue",
            "onc-safer-contingency-2025",
            "kim-2017-jamia-hit-problems-review"
          ]
        },
        {
          "id": "stale-or-incomplete-record",
          "name": "Stale or incomplete record during and after downtime",
          "mechanism": "Read-only downtime copies are snapshots and age from the moment the outage begins. After restoration, data captured on paper is back-entered late or not at all, and results produced during the outage may be absent from the electronic record. Clinicians decide on data that looks current but is not.",
          "warning_signs": [
            "Read-only backup refreshed less than hourly or not tested",
            "No reconciliation owner for paper records after downtime",
            "Interface buffers flushed or dropped at restart"
          ],
          "incident_ids": [
            "ascension-2024"
          ],
          "source_ids": [
            "onc-safer-contingency-2025",
            "larsen-2019-aci-downtime-laboratory",
            "healthcaredive-2024-ascension-tracker"
          ]
        },
        {
          "id": "backup-compromise",
          "name": "Backups destroyed or unusable when needed",
          "mechanism": "Attackers target backup systems before encrypting, or backups turn out never to have been restored end to end. The organization then has no clean copy to restore from and must rebuild, or pay.",
          "warning_signs": [
            "Backups reachable with domain credentials",
            "No full restore test in the last month",
            "Configuration (not just data) not backed up"
          ],
          "incident_ids": [
            "change-healthcare-2024"
          ],
          "source_ids": [
            "hhs-405d-hicp-2023",
            "onc-safer-contingency-2025",
            "cisa-aa24-131a-black-basta"
          ]
        },
        {
          "id": "regional-spillover",
          "name": "Regional spillover to neighbouring hospitals",
          "mechanism": "When a system diverts ambulances and time-critical patients, nearby EDs absorb the load without extra staff. Waits, walk-outs and time-critical cases rise at hospitals that were never attacked; rural patients face much longer travel.",
          "warning_signs": [
            "A neighbouring system announces diversion or a cyber incident",
            "Sudden EMS arrival increase without a mass-casualty event"
          ],
          "incident_ids": [
            "wannacry-nhs-2017",
            "uhs-2020"
          ],
          "source_ids": [
            "dameff-2023-jama-netw-open-adjacent-eds",
            "neprash-2024-j-rural-health-ransomware",
            "nhs-2018-wannacry-lessons-learned"
          ]
        },
        {
          "id": "cloud-region-control-plane",
          "name": "Cloud-region or provider control-plane failure",
          "mechanism": "A fault inside the provider (DNS automation, internal network congestion) disables core services across a region while the hospital's own building is fine. Impact depends on how each customer and each supplier built on the region: in October 2025 one Epic-on-AWS system slowed and another saw nothing, while NHS trusts using Oracle services went to paper.",
          "warning_signs": [
            "No map of which clinical and supplier services run in which region",
            "EHR slowdowns with no local cause",
            "DNS resolution errors for provider endpoints",
            "Provider status page silent or impaired"
          ],
          "incident_ids": [
            "aws-us-east-1-2025"
          ],
          "source_ids": [
            "aws-2025-dynamodb-pes",
            "aws-2021-us-east-1-pes",
            "beckers-2025-aws-tufts",
            "digitalhealth-2025-aws-nhs"
          ]
        }
      ],
      "detection": [
        {
          "text": "Measure EHR response time for key clinical tasks (results review, order entry, patient lookup) continuously; ONC's SAFER guide sets the target at optimally under 2 seconds and defines a functional downtime as any hourly mean response time over 5 seconds, or 3 standard deviations above the mean. Use a synthetic 'test patient' order placed on a schedule to detect silent failure.",
          "source_ids": [
            "onc-safer-contingency-2025"
          ]
        },
        {
          "text": "Monitor every interface queue, error queue and 'unknown' or dead-letter queue daily, with a named owner and reconciliation of orders sent against orders received.",
          "source_ids": [
            "va-oig-2022-unknown-queue",
            "onc-safer-contingency-2025"
          ]
        },
        {
          "text": "Alert on backup job failures, deletions of backup sets, and failed restore tests; test full restores rather than job completion.",
          "source_ids": [
            "onc-safer-contingency-2025",
            "hhs-405d-hicp-2023"
          ]
        },
        {
          "text": "Centralize logs and watch for known ransomware tactics (credential abuse on remote access, lateral movement, large outbound transfers) during the days between intrusion and encryption.",
          "source_ids": [
            "hhs-hph-cpgs",
            "cisa-aa24-131a-black-basta",
            "witty-2024-senate-finance-testimony"
          ]
        },
        {
          "text": "Require third parties to report incidents to you promptly, and subscribe to HHS/CISA advisories and the HHS OCR breach portal so a vendor or neighbour's outage reaches you before patients do.",
          "source_ids": [
            "hhs-hph-cpgs",
            "hhs-ocr-breach-portal"
          ]
        },
        {
          "text": "Track EMS arrival and ED census against baseline; a sudden rise without a local cause may mean a neighbouring system has gone down.",
          "source_ids": [
            "dameff-2023-jama-netw-open-adjacent-eds"
          ]
        }
      ],
      "defenses_by_tier": {
        "0": [
          {
            "text": "Put phishing-resistant MFA on every remote-access portal, VPN and privileged account, starting with vendor access.",
            "source_ids": [
              "cisa-aa24-131a-black-basta",
              "witty-2024-senate-finance-testimony",
              "hhs-hph-cpgs"
            ]
          },
          {
            "text": "Patch known exploited vulnerabilities promptly and retire unsupported operating systems, including those embedded in diagnostic devices.",
            "source_ids": [
              "cisa-aa24-131a-black-basta",
              "nao-2017-wannacry",
              "nhs-2018-wannacry-lessons-learned"
            ]
          },
          {
            "text": "Keep a daily, encrypted, off-site backup separated from normal storage (air gap); keep several generations; back up system configuration monthly and before every upgrade.",
            "source_ids": [
              "onc-safer-contingency-2025",
              "hhs-405d-hicp-2023"
            ]
          },
          {
            "text": "Build redundant network paths: two internet circuits in different trenches or from different providers, and a routed (not flat Layer-2) core.",
            "source_ids": [
              "onc-safer-contingency-2025",
              "cio-2003-bidmc-network"
            ]
          },
          {
            "text": "Inventory every third party that performs a clinical or revenue function you cannot do yourself; write incident-notification and recovery terms into the contract.",
            "source_ids": [
              "hhs-hph-cpgs",
              "aha-change-underscores-preparedness"
            ]
          },
          {
            "text": "Complete all nine SAFER Guides every year as a working review, not a yes/no box; CMS accepts 'no' as an answer, so the attestation alone proves nothing.",
            "source_ids": [
              "cms-safer-guides-attestation-2023",
              "onc-safer-contingency-2025"
            ]
          }
        ],
        "1": [
          {
            "text": "Maintain a warm site that can run the whole EHR within 8 hours, more than 50 miles away, and fail over to it at least quarterly.",
            "source_ids": [
              "onc-safer-contingency-2025"
            ]
          },
          {
            "text": "Segment the network so a compromised zone can be isolated without disconnecting everything; plan in advance which segments stay up.",
            "source_ids": [
              "hhs-hph-cpgs",
              "hhs-405d-hicp-2023",
              "nhs-2018-wannacry-lessons-learned"
            ]
          },
          {
            "text": "Contract and test an alternate clearinghouse or claims submission route, and an alternate reference lab, before an outage.",
            "source_ids": [
              "aha-change-underscores-preparedness",
              "witty-2024-senate-finance-testimony"
            ]
          },
          {
            "text": "Write and test restoration procedures that bring critical systems and data back within 72 hours, ranked by clinical criticality. The proposed HIPAA Security Rule would require this; do not wait for the final rule.",
            "source_ids": [
              "hhs-hipaa-security-nprm-2025",
              "ecfr-45cfr164-308"
            ]
          },
          {
            "text": "Size interface buffers so data queued during an outage is not lost, and alert users in the EHR when a clinical interface is down.",
            "source_ids": [
              "onc-safer-contingency-2025"
            ]
          }
        ],
        "2": [
          {
            "text": "Run a read-only backup EHR refreshed at least hourly, tested weekly, printable, and on UPS or generator power at unit level; make sure staff can log in to it.",
            "source_ids": [
              "onc-safer-contingency-2025"
            ]
          },
          {
            "text": "Keep downtime communication independent of the EHR network (not email, websites or VoIP on the same infrastructure).",
            "source_ids": [
              "onc-safer-contingency-2025",
              "tjc-sea-67-cyberattack"
            ]
          },
          {
            "text": "Call downtime early: activate the warm site or downtime procedures before 2 hours of unplanned outage, not after.",
            "source_ids": [
              "onc-safer-contingency-2025"
            ]
          },
          {
            "text": "Double-check high-risk medications manually when barcode scanning is unavailable, and use positive patient identification procedures designed for downtime.",
            "source_ids": [
              "onc-safer-contingency-2025",
              "larsen-2018-jamia-ehr-downtime-events",
              "kff-2024-ascension-lapses"
            ]
          }
        ],
        "3": [
          {
            "text": "Stock current paper forms for orders, medication administration, lab requisitions and results on every unit; keep a paper copy of the downtime policy on units and off-site.",
            "source_ids": [
              "onc-safer-contingency-2025",
              "ecfr-42cfr482-15"
            ]
          },
          {
            "text": "Run unannounced downtime drills at least yearly, and at least one exercise that assumes weeks, not hours, without the EHR, lab interface or clearinghouse.",
            "source_ids": [
              "onc-safer-contingency-2025",
              "tjc-sea-67-cyberattack",
              "kff-2024-ascension-lapses"
            ]
          },
          {
            "text": "Assign a runner or courier system for orders and results between departments; paper without a routing method stalls.",
            "source_ids": [
              "kff-2024-ascension-lapses",
              "larsen-2019-aci-downtime-laboratory"
            ]
          },
          {
            "text": "Agree regional diversion and mutual-aid plans with neighbouring hospitals and EMS for cyber incidents, not only physical disasters.",
            "source_ids": [
              "dameff-2023-jama-netw-open-adjacent-eds",
              "nhs-2018-wannacry-lessons-learned"
            ]
          },
          {
            "text": "Plan recovery as its own phase: assign owners to back-enter and reconcile paper data, restart interfaces in order, and review harm from delays.",
            "source_ids": [
              "onc-safer-contingency-2025",
              "tjc-sea-67-cyberattack",
              "nhs-england-synnovis-incident"
            ]
          }
        ]
      },
      "standards": [
        {
          "name": "HIPAA Security Rule, 45 CFR 164.308(a)(7) Contingency plan",
          "what_it_requires": "Covered entities must have a data backup plan, a disaster recovery plan and an emergency-mode operation plan; testing/revision and an applications-and-data criticality analysis are 'addressable'. A January 2025 NPRM would add written procedures to restore critical systems and data within 72 hours, but it was not final as of September 2026.",
          "source_id": "ecfr-45cfr164-308"
        },
        {
          "name": "CMS Hospital CoP Emergency preparedness, 42 CFR 482.15",
          "what_it_requires": "Hospitals must maintain a system of medical documentation that preserves patient information and keeps records available in an emergency, with the emergency plan and training/testing program reviewed at least every 2 years.",
          "source_id": "ecfr-42cfr482-15"
        },
        {
          "name": "ONC/ASTP SAFER Guide: Contingency Planning (2025 edition)",
          "what_it_requires": "Self-assessment of 13 practices covering disaster recovery, generators, paper forms, tested backups, downtime training, independent communication, interface restart and downtime monitoring. CMS requires hospitals in the Medicare Promoting Interoperability Program to attest annually (yes or no) to completing all nine SAFER Guides.",
          "source_id": "onc-safer-contingency-2025"
        },
        {
          "name": "HHS 405(d) Health Industry Cybersecurity Practices (HICP), 2023 edition",
          "what_it_requires": "Voluntary, sector-specific: ten practices against five threats including ransomware, scaled for small and large organizations in two technical volumes.",
          "source_id": "hhs-405d-hicp-2023"
        },
        {
          "name": "HHS HPH Cybersecurity Performance Goals (CPGs)",
          "what_it_requires": "Voluntary essential goals (e.g., MFA, incident planning, vendor cybersecurity requirements) and enhanced goals (e.g., network segmentation, third-party incident reporting, drilled incident plans).",
          "source_id": "hhs-hph-cpgs"
        },
        {
          "name": "The Joint Commission Sentinel Event Alert 67 (2023)",
          "what_it_requires": "Not a standard itself; recommends downtime planning committees, response teams, staff training and communication for extended cyber downtime, and points to TJC continuity-of-operations and disaster-recovery requirements.",
          "source_id": "tjc-sea-67-cyberattack"
        }
      ],
      "elsewhere": {
        "text": "In the EU the NIS2 Directive (2022/2555) keeps healthcare within its scope and imposes cybersecurity risk-management and incident-notification duties; ENISA's 2023 health threat landscape found ransomware in 54% of 215 reported health-sector incidents and a dedicated ransomware programme in only 27% of surveyed organisations. In England, DHSC's 2023-2030 cyber strategy aims for all health and social care organisations, including critical suppliers, to be cyber resilient by 2030. WannaCry (2017) and Synnovis (2024) are the reference cases: the first showed how unpatched systems and precautionary disconnection spread disruption, the second how a single pathology supplier can halt a region's diagnostics for months.",
        "source_ids": [
          "eu-nis2-directive-2022-2555",
          "enisa-2023-health-threat-landscape",
          "dhsc-2023-cyber-strategy",
          "nao-2017-wannacry",
          "nhs-england-synnovis-incident"
        ]
      },
      "scoring_v01": {
        "likelihood": 5,
        "blast_radius": 5,
        "detectability": 3,
        "rationale": "Judgement: likelihood is 5 because unplanned EHR downtime is near-universal and ransomware attacks on care delivery roughly doubled between 2016 and 2021. Blast radius is 5 because shared vendors (Change Healthcare, Synnovis) and precautionary shutdowns take down whole regions or national functions at once. Detectability averages two extremes: outright outages are obvious (about 1), but silent misrouting and stale data can go unnoticed for months (about 5)."
      },
      "open_questions": [
        "How much patient harm comes from 'wrong data' failures (misrouting, stale copies, lost back-entry) compared with outright outages? No study measures both.",
        "What downtime length should hospitals plan and drill for? Most plans cover 1-3 days, while major ransomware outages last 3-6 weeks.",
        "Do tested read-only backups, warm sites or alternate clearinghouses measurably reduce harm or recovery time? Evidence is mostly self-assessment and expert opinion.",
        "Will HHS finalize the 72-hour restoration requirement, and would recovery-time mandates change outcomes or only paperwork?",
        "How concentrated are US clinical dependencies (clearinghouses, reference labs, hosted EHRs), and which vendors are single points of failure for a region?"
      ],
      "watch_queries": [
        "hospital ransomware attack EHR downtime diversion",
        "\"electronic health record\" downtime patient safety",
        "clearinghouse OR \"reference laboratory\" cyberattack outage providers",
        "\"unknown queue\" OR \"lost orders\" EHR interface patient harm",
        "ransomware hospital mortality OR outcomes study",
        "HHS OCR HIPAA Security Rule final rule contingency 72 hours",
        "CISA #StopRansomware healthcare advisory",
        "NHS cyber incident pathology OR trust patient harm review",
        "health system network outage paper charting days",
        "ONC SAFER Guides contingency planning update"
      ],
      "last_reviewed": "2026-09-26",
      "reviewed_by": "opus-research",
      "last_audited": "2026-09-26",
      "audit_status": "verified"
    },
    {
      "id": "devices",
      "name": "Devices & electronics",
      "one_liner": "The monitors, pumps, ventilators, sensors and analyzers that measure and act on patients, and the software and updates that run them.",
      "definition": [
        "This layer is the equipment at the bedside, in the lab and in the patient's home: physiologic monitors, pulse oximeters, infusion pumps, ventilators, continuous glucose monitors, point-of-care and lab analyzers, and the endpoint software (operating systems, security agents, interface adapters) that now runs on or beside them. Nearly all of it is software-driven, networked, and updated by a vendor after installation.",
        "Devices fail in two ways. Loud failures stop the device, raise an alarm or crash the workstation; staff notice and fall back to another device or to manual care. Quiet failures keep producing numbers that look normal and are wrong: a pulse oximeter that reads high on darker skin, a glucose sensor that reads low, a lead analyzer that under-reports, a pump that loads a stale order. The quiet kind never triggers a downtime procedure.",
        "In a paper-era hospital a device error reached one patient through one clinician who could see the device. In an automated hospital device outputs feed EHR flowsheets, early-warning scores, auto-programmed infusions and remote monitoring, and a single vendor update can reach every unit at once. The same connectivity that lets a pump receive an order from the EHR lets a faulty update or a compromised firmware image reach thousands of endpoints in minutes."
      ],
      "angle": "FailSystems' view: a device that stops is a tier-2 problem you can plan for; a device that keeps reporting wrong numbers is the one that hurts people, because nothing in the system tells anyone to change tier. Automation makes this worse in two directions. It amplifies quiet error, because downstream scores, alerts and auto-programming consume the bad value without a human looking at the device. And it synchronizes loud failure, because fleet-wide updates (security agents, firmware, interface software) turn one vendor mistake into simultaneous failure across a hospital or a country. Defense in this layer is less about redundancy of boxes and more about independent cross-checks of values and control over when changes land.",
      "failure_modes": [
        {
          "id": "silent-measurement-bias",
          "name": "Systematic sensor bias in a subgroup",
          "mechanism": "A sensor is accurate on the population it was validated on and biased on others. Pulse oximeters overestimate saturation in patients with darker skin, so hypoxemia is missed and treatment thresholds are crossed later. The device reports normally and nothing alarms.",
          "warning_signs": [
            "SpO2-SaO2 gaps that differ by patient group when paired values are audited",
            "Validation data from the manufacturer that does not report performance by skin tone",
            "Therapy eligibility or escalation rates that differ by group at the same recorded SpO2"
          ],
          "incident_ids": [
            "pulse-oximetry-bias-2020"
          ],
          "source_ids": [
            "sjoding-2020-nejm-pulse-oximetry",
            "fawzy-2022-jama-im-pulse-oximetry-covid",
            "fda-2025-pulse-oximeter-draft-guidance"
          ]
        },
        {
          "id": "silent-field-defect",
          "name": "Manufacturing or design defect producing plausible wrong values",
          "mechanism": "A batch or design flaw makes a device report values in the normal range that are wrong: glucose sensors reading low, blood lead analyzers reading low. Users act on the number. Detection depends on someone comparing against an independent method, and on the manufacturer reporting promptly, which can fail.",
          "warning_signs": [
            "Clinical picture that does not match the device value",
            "Discrepancies between point-of-care and reference lab results",
            "Clusters of complaints about one lot or serial range",
            "Changes to instructions for use without a clear safety notice"
          ],
          "incident_ids": [
            "abbott-libre3-sensor-recall-2025",
            "magellan-leadcare-2013"
          ],
          "source_ids": [
            "fda-abbott-libre3-recall-2025",
            "doj-magellan-leadcare-plea-2024"
          ]
        },
        {
          "id": "fleet-update-cascade",
          "name": "Fleet-wide faulty update",
          "mechanism": "A vendor pushes a software, firmware or content update to every installed endpoint at once. If the update is faulty, every device or workstation that takes it fails together, and recovery is limited by hands-on remediation per machine. Security agents with kernel access are the extreme case because they update often and without customer staging.",
          "warning_signs": [
            "Endpoint agents or device firmware set to auto-update with no ring or delay",
            "No inventory of which clinical devices run which agents",
            "Recovery runbook that assumes remote management works"
          ],
          "incident_ids": [
            "crowdstrike-2024"
          ],
          "source_ids": [
            "crowdstrike-rca-channel-file-291",
            "house-homeland-meyers-testimony-2024",
            "tully-2025-jama-netw-open-crowdstrike"
          ]
        },
        {
          "id": "stale-command-interop",
          "name": "Stale or queued commands across a device integration",
          "mechanism": "When EHR-to-device integrations (infusion auto-programming, order interfaces) back up, a queued command can arrive late and be applied to the device as if it were current. The value looks legitimate on the pump screen.",
          "warning_signs": [
            "Interface engine queue depth or latency rising",
            "Pump parameters that differ from the current order",
            "Clinicians reporting 'the pump got the old order'"
          ],
          "incident_ids": [
            "bd-alaris-apr-backlog-2025"
          ],
          "source_ids": [
            "fda-bd-alaris-apr-correction-2025"
          ]
        },
        {
          "id": "alarm-failure-and-fatigue",
          "name": "Alarm failure and alarm fatigue",
          "mechanism": "Alarms fail to sound (a low-battery alarm that does not fire, wrong priority), sound falsely (spurious power-loss alarms that stop therapy), or sound so often that staff tune them out. The Joint Commission counted 98 alarm-related sentinel events, 80 of them deaths, from 2009 to mid-2012.",
          "warning_signs": [
            "High non-actionable alarm rates per bed per day",
            "Alarm limits left at defaults",
            "Vendor corrections that mention alarm behavior",
            "Near misses where an alarm was heard but not acted on"
          ],
          "incident_ids": [
            "philips-respironics-2021",
            "bd-alaris-apr-backlog-2025"
          ],
          "source_ids": [
            "tjc-sea-50-alarms-2013",
            "fda-bd-alaris-2020-recall",
            "fda-philips-trilogy-evo-correction-2024",
            "iec-60601-1-8"
          ]
        },
        {
          "id": "insecure-or-legacy-connected-device",
          "name": "Insecure or unsupported networked device",
          "mechanism": "Devices that can connect to a network but no longer receive security updates, or that ship with hidden functions, provide a path to alter device behavior or reach the wider network. The only mitigation may be to disconnect, which removes remote monitoring.",
          "warning_signs": [
            "Devices on end-of-support operating systems",
            "No SBOM or vulnerability disclosure contact from the vendor",
            "Unexpected outbound traffic from device VLANs",
            "Devices on flat networks with clinical workstations"
          ],
          "incident_ids": [
            "contec-cms8000-backdoor-2025"
          ],
          "source_ids": [
            "fda-contec-cms8000-2025",
            "cisa-icsma-25-030-01",
            "ecri-top10-hazards-2026",
            "usc-21-360n-2-section-524b",
            "fda-524b-webinar-2024",
            "fda-postmarket-cybersecurity-2016"
          ]
        },
        {
          "id": "latent-hardware-hazard-slow-recall",
          "name": "Latent hardware hazard with slow recall remediation",
          "mechanism": "A material or component degrades inside devices already in use, with no alarm. Once found, remediation depends on replacement supply, locating every unit (often in patients' homes) and clear communication, and can take years.",
          "warning_signs": [
            "Recall notices without a tracked list of your affected serial numbers",
            "Home-use devices issued without a registry",
            "Replacement timelines measured in months"
          ],
          "incident_ids": [
            "philips-respironics-2021"
          ],
          "source_ids": [
            "fda-philips-activities",
            "fda-philips-mdrs",
            "fda-philips-consent-decree-2024"
          ]
        },
        {
          "id": "recall-communication-failure",
          "name": "Recall and safety notice not reaching the user",
          "mechanism": "Recalls are posted, but the notice does not reach the clinician, biomed team or home patient using the device, or arrives without clear action. FDA posting dates reflect classification, which can lag the firm's action.",
          "warning_signs": [
            "No single owner for recall intake and closure",
            "Recalls closed without serial-number reconciliation",
            "Home devices supplied by third parties outside the hospital's view"
          ],
          "incident_ids": [
            "abbott-libre3-sensor-recall-2025",
            "philips-respironics-2021"
          ],
          "source_ids": [
            "ecri-top10-hazards-2026",
            "fda-recalls-database",
            "fda-philips-activities"
          ]
        }
      ],
      "detection": [
        {
          "text": "Audit paired device vs reference values (SpO2 vs SaO2, point-of-care vs lab glucose or lead) at least quarterly and break the gap down by patient group.",
          "source_ids": [
            "sjoding-2020-nejm-pulse-oximetry",
            "fawzy-2022-jama-im-pulse-oximetry-covid"
          ]
        },
        {
          "text": "Subscribe to the FDA Medical Device Recalls database and MAUDE for every device model in your inventory, remembering MDR counts do not establish cause or rate.",
          "source_ids": [
            "fda-recalls-database",
            "fda-maude-about"
          ]
        },
        {
          "text": "Subscribe to CISA ICS medical advisories and match them against your device inventory by model and firmware version.",
          "source_ids": [
            "cisa-icsma-25-030-01"
          ]
        },
        {
          "text": "Monitor device integration queues (EHR-to-pump auto-programming, interface engines) for latency and backlog, and alert on growth.",
          "source_ids": [
            "fda-bd-alaris-apr-correction-2025"
          ]
        },
        {
          "text": "Track alarm load and non-actionable alarm rates per unit; rising rates predict missed actionable alarms.",
          "source_ids": [
            "tjc-sea-50-alarms-2013",
            "tjc-npg-2026-hospital"
          ]
        },
        {
          "text": "Measure external reachability of your own clinical services; the CrowdStrike study showed internet scanning detected outages at 34% of hospitals.",
          "source_ids": [
            "tully-2025-jama-netw-open-crowdstrike"
          ]
        }
      ],
      "defenses_by_tier": {
        "0": [
          {
            "text": "Require manufacturers to supply accuracy data by skin tone for any oximeter you buy, and prefer devices tested under FDA's 2025 draft protocol.",
            "source_ids": [
              "fda-2025-pulse-oximeter-draft-guidance",
              "fda-pulse-oximeters-actions"
            ]
          },
          {
            "text": "Require an SBOM, a coordinated vulnerability disclosure process and a stated patch timeline in every networked-device contract, mirroring FD&C Act 524B.",
            "source_ids": [
              "usc-21-360n-2-section-524b",
              "fda-premarket-cybersecurity-guidance-2026",
              "fda-524b-webinar-2024"
            ]
          },
          {
            "text": "Treat any device with USB, serial, Bluetooth or ethernet ports as internet-capable when you assess it; FDA does.",
            "source_ids": [
              "fda-524b-webinar-2024"
            ]
          },
          {
            "text": "Put every endpoint agent and device firmware update into staged rings (test group, one unit, then fleet) with a hold period; refuse vendors that cannot support customer-controlled staging.",
            "source_ids": [
              "crowdstrike-rca-channel-file-291",
              "house-homeland-meyers-testimony-2024"
            ]
          },
          {
            "text": "Run IEC 80001-1 risk management before connecting any device to the network, with clinical engineering, IT and the vendor named as owners.",
            "source_ids": [
              "iec-80001-1-2021"
            ]
          },
          {
            "text": "Keep a device inventory with model, serial, firmware version, network location and support end date, and reconcile every recall against it within 10 working days.",
            "source_ids": [
              "fda-recalls-corrections-removals",
              "ecri-top10-hazards-2026"
            ]
          }
        ],
        "1": [
          {
            "text": "When an oximetry reading does not fit the clinical picture, draw an arterial blood gas before withholding or delaying oxygen-threshold therapy.",
            "source_ids": [
              "sjoding-2020-nejm-pulse-oximetry",
              "fawzy-2022-jama-im-pulse-oximetry-covid"
            ]
          },
          {
            "text": "Verify rate, dose and volume on the pump against the current order before starting any auto-programmed infusion.",
            "source_ids": [
              "fda-bd-alaris-apr-correction-2025"
            ]
          },
          {
            "text": "Segment clinical devices onto their own networks and block outbound internet by default, so a compromised monitor can still monitor locally.",
            "source_ids": [
              "fda-contec-cms8000-2025",
              "cisa-icsma-25-030-01"
            ]
          },
          {
            "text": "Set alarm limits per patient population and document which alarm signals matter most on each unit, as NPG.01.05.01 requires.",
            "source_ids": [
              "tjc-npg-2026-hospital",
              "tjc-sea-50-alarms-2013"
            ]
          }
        ],
        "2": [
          {
            "text": "Keep standalone (non-networked) monitors and pumps stocked on each critical unit for use when networked devices or their central stations are down.",
            "source_ids": [
              "ecri-top10-hazards-2026"
            ]
          },
          {
            "text": "Pre-stage offline recovery kits (local admin credentials, disk-encryption recovery keys, bootable media) so endpoints can be restored by hand at scale.",
            "source_ids": [
              "crowdstrike-rca-channel-file-291",
              "statnews-2024-crowdstrike-hospitals"
            ]
          },
          {
            "text": "Give home patients on recalled sensors a verified fallback (fingerstick meter, strips) and tell them in writing which readings to trust.",
            "source_ids": [
              "fda-abbott-libre3-recall-2025"
            ]
          },
          {
            "text": "Switch to manual infusion programming with independent double-check when the interoperability layer is suspect.",
            "source_ids": [
              "fda-bd-alaris-apr-correction-2025"
            ]
          }
        ],
        "3": [
          {
            "text": "Keep paper vital-sign and infusion flowsheets on every unit and drill their use; ECRI ranks digital-darkness unpreparedness the second hazard of 2026.",
            "source_ids": [
              "ecri-top10-hazards-2026"
            ]
          },
          {
            "text": "Maintain manual measurement skills and equipment (manual BP cuffs, gravity infusion sets with drip-rate charts) for when devices cannot be trusted.",
            "source_ids": [
              "ecri-top10-hazards-2026"
            ]
          },
          {
            "text": "Pre-decide which elective procedures cancel when device fleets fail, so the call takes minutes, as it did at Mass General Brigham on 19 July 2024.",
            "source_ids": [
              "statnews-2024-crowdstrike-hospitals"
            ]
          }
        ]
      },
      "standards": [
        {
          "name": "ISO 14971:2019 (FDA recognition 5-125)",
          "what_it_requires": "Manufacturers identify hazards, estimate and control risks, and monitor effectiveness of controls across the device life cycle, including post-production information.",
          "source_id": "fda-rcs-iso-14971"
        },
        {
          "name": "IEC 62304:2006+AMD1:2015",
          "what_it_requires": "Life cycle processes for development and maintenance of medical device software, including software safety classification, change control and problem resolution.",
          "source_id": "iec-62304"
        },
        {
          "name": "IEC 60601-1-8:2006+AMD1:2012+AMD2:2020",
          "what_it_requires": "Requirements and tests for medical alarm systems: alarm priority categories, alarm signal characteristics and control states such as pausing and silencing.",
          "source_id": "iec-60601-1-8"
        },
        {
          "name": "IEC 80001-1:2021",
          "what_it_requires": "The healthcare delivery organization applies risk management for safety, effectiveness and security before, during and after connecting devices or health software to its IT infrastructure.",
          "source_id": "iec-80001-1-2021"
        },
        {
          "name": "FD&C Act section 524B (21 U.S.C. 360n-2)",
          "what_it_requires": "Cyber device sponsors must submit a postmarket vulnerability plan, maintain processes to assure cybersecurity, ship patches on a justified regular cycle and critical fixes out of cycle, and provide an SBOM.",
          "source_id": "usc-21-360n-2-section-524b"
        },
        {
          "name": "Joint Commission NPG.01.05.01 (2026)",
          "what_it_requires": "Hospitals identify the most important alarm signals, set policies for managing them, and educate staff; replaces NPSG.06.01.01 from January 2026.",
          "source_id": "tjc-npg-2026-hospital"
        }
      ],
      "elsewhere": {
        "text": "In the EU, the Medical Device Regulation (EU) 2017/745 makes information security part of the essential requirements: Annex I 17.2 requires software to be built under state-of-the-art life cycle and risk management including information security, and 17.4 requires manufacturers to state minimum hardware, network and IT security requirements, including protection against unauthorised access. In Great Britain, amended post-market surveillance rules in force from 16 June 2025 cut the serious-incident reporting deadline from 30 to 15 days and require manufacturers to submit Field Safety Notices to the MHRA before they go to users, which targets the recall-communication failure mode directly.",
        "source_ids": [
          "eu-mdr-2017-745",
          "mhra-pms-2025"
        ]
      },
      "scoring_v01": {
        "likelihood": 4,
        "blast_radius": 4,
        "detectability": 5,
        "rationale": "Judgement: device faults reported to FDA are routine (Class I recalls on pumps, ventilators and sensors recur yearly), so likelihood is high. Blast radius is high because fleet updates and population-wide sensors (oximetry, CGMs, a dominant lead analyzer) spread one error across many patients. Detectability is scored hardest (5) because the most harmful mode is a plausible wrong value that triggers no alarm and no downtime procedure."
      },
      "open_questions": [
        "How much patient harm does oximetry bias cause at the outcome level (mortality, ICU admission), beyond delayed treatment eligibility?",
        "What share of hospital device fleets run end-of-support software or unmanaged agents, and how does that track with outage and incident exposure?",
        "Do staged, customer-controlled update rings actually reduce fleet-failure blast radius in hospitals, and at what patch-delay cost for security?",
        "How often do EHR-to-device integrations deliver stale or mismatched commands in routine operation, below the threshold of a recall?",
        "Will FDA finalize the 2025 pulse oximeter guidance, and how fast will legacy oximeters in use be replaced?"
      ],
      "watch_queries": [
        "FDA Class I recall infusion pump software",
        "FDA medical device recall incorrect readings sensor",
        "CISA ICS medical advisory ICSMA",
        "FDA cybersecurity safety communication medical device",
        "pulse oximetry skin pigmentation occult hypoxemia[Title/Abstract]",
        "continuous glucose monitor recall inaccurate readings",
        "hospital outage software update endpoint agent",
        "ventilator software correction FDA",
        "ECRI Top 10 Health Technology Hazards",
        "legacy medical device cybersecurity end of support hospital"
      ],
      "last_reviewed": "2026-09-26",
      "reviewed_by": "opus-research",
      "last_audited": "2026-09-26",
      "audit_status": "verified"
    },
    {
      "id": "models",
      "name": "Models & agents",
      "one_liner": "Predictive models, LLMs and agents that stay up while giving wrong answers, so care continues on bad output without anyone switching to a fallback.",
      "definition": [
        {
          "text": "This layer covers software that produces a judgement rather than just moving data: sepsis and deterioration scores in the EHR, imaging and lab classifiers cleared as medical devices, risk scores used to allocate care-management or coverage, ambient scribes and other large language model (LLM) tools that write clinical text, and agents that take actions such as ordering refills or changing records. The FDA keeps a public list of AI-enabled devices it has authorized, but it says the list is not comprehensive, and many deployed models (EHR-vendor scores, payer algorithms, scribes) are not devices at all.",
          "source_ids": [
            "fda-ai-enabled-device-list"
          ]
        },
        "What stands on this layer is increasingly the first read of the patient: the alert that starts a sepsis bundle, the note the next clinician trusts, the length-of-stay prediction that shapes a discharge or a denial. Regulators have named the specific ways it goes wrong. NIST calls confident false output 'confabulation' and lists automation bias as a risk that amplifies it. FDA's draft AI lifecycle guidance says performance can degrade after deployment through shifts in population, disease patterns or input data, and that users may not notice when the model sits inside a highly automated process.",
        "The failure that defines this layer is not an outage. When a model goes down, people notice and fall back to assisted or manual work. When a model goes wrong, the screens stay green, alerts keep firing, notes keep appearing and the organization never changes tier. The evidence below is mostly of that second kind."
      ],
      "angle": "FailSystems' view: every other layer on this site fails loudly enough to trigger a downtime procedure. Models and agents fail quietly. A model that is down is a tier-1 event with a known playbook; a model that is wrong causes no tier change at all, which is why we rank it the more dangerous case. Automation also changes who checks: clinicians see the output, not the inputs, the version or the training population, and agents act before anyone reads what they decided. The practical consequence is that the fallback trigger for this layer cannot be availability. It has to be measured accuracy against ground truth, with named people empowered to switch the model off, and a manual pathway that still works when they do.",
      "failure_modes": [
        {
          "id": "external-validity-gap",
          "name": "Model never worked as well locally as claimed",
          "mechanism": "A model validated on the developer's data is deployed across many hospitals without independent local validation. Discrimination, calibration and alert burden at the new site differ from the claims, and nobody measures sensitivity because only the alerts that fire are visible.",
          "warning_signs": [
            "Vendor performance figures with no external validation at a comparable site",
            "Alert rate tracked but not missed cases",
            "Clinicians describe the alert as noise",
            "Large variation in performance between sites running the same model"
          ],
          "incident_ids": [
            "epic-sepsis-model-2021"
          ],
          "source_ids": [
            "wong-2021-jama-im-epic-sepsis",
            "wong-2026-jama-netw-open-esm-v2",
            "onc-45cfr170-315-b11"
          ]
        },
        {
          "id": "dataset-shift",
          "name": "Dataset shift after deployment",
          "mechanism": "The patient mix, disease patterns, coding practice or upstream data feed changes, and the relationship the model learned no longer holds. Output keeps flowing with no error message. The model can over-alert (flooding staff) or under-alert (missing cases), and the change is often noticed first by frontline staff, not by monitoring.",
          "warning_signs": [
            "Sudden change in alert volume or score distribution",
            "New disease, new population or new upstream system (lab analyzer, EHR build, coding change)",
            "Nursing complaints of overalerting",
            "Missing values, duplicate records or type mismatches in model inputs"
          ],
          "incident_ids": [
            "michigan-sepsis-model-covid-shift-2020"
          ],
          "source_ids": [
            "finlayson-2021-nejm-dataset-shift",
            "wong-2021-jama-netw-open-sepsis-covid-alerts",
            "fda-2025-ai-dsf-draft-guidance"
          ]
        },
        {
          "id": "proxy-label-bias",
          "name": "Biased proxy labels and hidden subgroup failure",
          "mechanism": "The model is trained to predict something convenient (cost, prior treatment, a clinician's order) instead of the clinical outcome. Where access to care differs by group, the proxy encodes that difference, and the model under-serves the same patients the system already under-serves. Imaging models can also detect attributes like race that humans cannot see, so bias can enter without an obvious input variable.",
          "warning_signs": [
            "Target variable is cost, utilization or a clinician action rather than health status",
            "No performance reported by race, sex, age or disability",
            "Subgroup performance never checked in local data"
          ],
          "incident_ids": [
            "care-management-algorithm-bias-2019"
          ],
          "source_ids": [
            "obermeyer-2019-science-racial-bias",
            "gichoya-2022-lancet-dh-race-imaging",
            "hhs-45cfr92-210",
            "omiye-2023-npj-race-based-llm"
          ]
        },
        {
          "id": "llm-confabulation-documentation",
          "name": "Confabulation and omission in generated clinical text",
          "mechanism": "Speech-to-text and LLM scribes produce fluent notes that contain things never said (drugs, diagnoses, treatment steps) or leave out things that were said. Because the output reads well and review is fast, errors enter the permanent record and travel to the next clinician. Omissions are harder to catch than fabrications because there is nothing on the page to question.",
          "warning_signs": [
            "Source audio deleted after transcription",
            "Clinician sign-off measured in seconds per note",
            "No sampling audit of notes against recordings",
            "Procurement tested on a few simulated encounters only"
          ],
          "incident_ids": [
            "whisper-confabulation-2024",
            "ontario-ai-scribes-audit-2026"
          ],
          "source_ids": [
            "nist-ai-600-1-genai-profile",
            "koenecke-2024-careless-whisper",
            "ap-2024-whisper-hospitals",
            "ontario-ag-2026-ai-report",
            "asgari-2025-npj-llm-hallucination-summarisation"
          ]
        },
        {
          "id": "output-as-order",
          "name": "Advisory output treated as an order",
          "mechanism": "A model is labelled decision support, but policy, protocol or productivity targets turn its output into the decision: a sepsis flag becomes a fluid bolus, a length-of-stay prediction becomes a coverage end date. Patient-specific contraindications the model cannot see are overridden unless someone at the bedside refuses. Low appeal or override rates then hide the error rate.",
          "warning_signs": [
            "Protocols that start automatically from a model flag",
            "Staff told to follow the alert when they disagree",
            "Override or appeal rates not tracked, or tracked but ignored",
            "High reversal rate on the few decisions that are appealed"
          ],
          "incident_ids": [
            "ai-sepsis-alert-fluid-overload-2025",
            "unitedhealth-nh-predict-denials-2023"
          ],
          "source_ids": [
            "euronews-ap-2025-nurses-ai-sepsis",
            "cms-2024-ma-faq-algorithms",
            "psi-2024-ma-post-acute-report",
            "who-2024-lmm-guidance"
          ]
        },
        {
          "id": "agent-stale-or-poisoned-state",
          "name": "Agents acting on stale, garbled or poisoned state",
          "mechanism": "An agent reads state (a medication list, a refill queue, its own memory) that is out of date, mis-transcribed or deliberately poisoned, then acts on it through tools with real permissions. Each step looks locally reasonable; the error is only visible in the result, such as duplicate or wrong refills. Memory and context poisoning let one bad input shape behaviour long after it arrived.",
          "warning_signs": [
            "Agent can write to production systems without a confirmation step",
            "No readback of drug name and dose to a human",
            "Rising complaints about duplicate or unexpected actions",
            "Agent memory persists across sessions with no review"
          ],
          "incident_ids": [
            "kinney-drugs-ai-burt-refill-2026"
          ],
          "source_ids": [
            "vtdigger-2026-kinney-burt",
            "owasp-2025-agentic-top10",
            "mitre-atlas",
            "jc-chai-2025-ruaih-guidance"
          ]
        },
        {
          "id": "excessive-agency-and-loops",
          "name": "Excessive agency, loops and false self-reports",
          "mechanism": "Agents given broad permissions and autonomy ignore instructions, repeat steps, stop early or verify their own work incorrectly. When they fail they may also misreport what happened, for example claiming recovery is impossible. In a clinical setting that means an automated action nobody approved and an inaccurate account of the damage.",
          "warning_signs": [
            "Agent holds write or delete permissions it does not need",
            "No human approval gate for high-impact actions",
            "No independent log of what the agent actually did",
            "Termination conditions not defined"
          ],
          "incident_ids": [
            "replit-agent-database-deletion-2025"
          ],
          "source_ids": [
            "owasp-llm-top10-2025",
            "theregister-2025-replit-database",
            "cemri-2025-mast-arxiv",
            "owasp-2025-agentic-top10"
          ]
        },
        {
          "id": "unmanaged-change",
          "name": "Silent model updates and unreviewed deployment",
          "mechanism": "Vendors retrain or swap models, or tools go live without regulatory review or documented validation, and the deploying organization is not told or does not re-test. Performance changes after an update look the same as normal operation. Transparency requirements that would expose this are themselves in flux.",
          "warning_signs": [
            "Contract has no notice-of-change clause",
            "No re-validation step after vendor updates",
            "Model version not visible to users or in logs",
            "Unclear whether the tool is an FDA-regulated device"
          ],
          "incident_ids": [
            "ontario-ai-scribes-audit-2026"
          ],
          "source_ids": [
            "fda-2024-pccp-final-guidance",
            "jc-chai-2025-ruaih-guidance",
            "onc-2025-hti5-proposed",
            "ontario-ag-2026-ai-report"
          ]
        }
      ],
      "detection": [
        {
          "text": "Track missed cases, not just alerts fired: compare model output with a ground-truth definition (for sepsis, a consensus criterion) on a regular sample. Alert counts alone hid a 67% miss rate.",
          "source_ids": [
            "wong-2021-jama-im-epic-sepsis"
          ]
        },
        {
          "text": "Watch alert volume and score distribution daily against a baseline; a doubling within weeks, as seen with sepsis alerts early in COVID-19, is a dataset-shift signal.",
          "source_ids": [
            "wong-2021-jama-netw-open-sepsis-covid-alerts",
            "finlayson-2021-nejm-dataset-shift"
          ]
        },
        {
          "text": "Monitor inputs as well as outputs: demographic and prevalence shifts, input distribution shifts, and pipeline corruption such as missing values, duplicate records and type mismatches.",
          "source_ids": [
            "fda-2025-ai-dsf-draft-guidance",
            "nist-ai-100-1-rmf"
          ]
        },
        {
          "text": "Measure override and appeal rates and what happens on appeal; a model whose decisions are usually reversed when challenged is wrong more often than the appeal rate shows.",
          "source_ids": [
            "psi-2024-ma-post-acute-report",
            "cms-2024-ma-faq-algorithms"
          ]
        },
        {
          "text": "Audit a sample of AI-generated notes against the source audio for fabrications, wrong drugs and omissions; this requires keeping the audio.",
          "source_ids": [
            "ontario-ag-2026-ai-report",
            "ap-2024-whisper-hospitals",
            "asgari-2025-npj-llm-hallucination-summarisation"
          ]
        },
        {
          "text": "Report performance by subgroup (race, sex, age, disability, language) in local data, not only in aggregate.",
          "source_ids": [
            "obermeyer-2019-science-racial-bias",
            "onc-45cfr170-315-b11",
            "hhs-45cfr92-210"
          ]
        },
        {
          "text": "Treat frontline complaints about a tool (overalerting, garbled output, duplicate actions) as safety reports routed into the incident system, not as help-desk tickets.",
          "source_ids": [
            "jc-chai-2025-ruaih-guidance",
            "wong-2021-jama-netw-open-sepsis-covid-alerts",
            "wcax-2026-kinney-pullback"
          ]
        }
      ],
      "defenses_by_tier": {
        "0": [
          {
            "text": "Validate every predictive model on your own recent data before go-live, reporting sensitivity, PPV, calibration and alert burden by unit and by subgroup. Do not accept vendor figures alone.",
            "source_ids": [
              "wong-2021-jama-im-epic-sepsis",
              "wong-2026-jama-netw-open-esm-v2",
              "jc-chai-2025-ruaih-guidance"
            ]
          },
          {
            "text": "Write a monitoring plan for each model before deployment: named owner, metrics, thresholds, review frequency, and the ground-truth sample you will check against.",
            "source_ids": [
              "fda-2025-ai-dsf-draft-guidance",
              "nist-ai-100-1-rmf",
              "jc-chai-2025-ruaih-guidance"
            ]
          },
          {
            "text": "Put notice-of-change and re-validation clauses in every AI contract, and log the model version with each output so you can tell which results came from which version.",
            "source_ids": [
              "jc-chai-2025-ruaih-guidance",
              "fda-2024-pccp-final-guidance"
            ]
          },
          {
            "text": "Keep the source for every generated artifact: retain scribe audio long enough to audit, and store the inputs an agent acted on.",
            "source_ids": [
              "ap-2024-whisper-hospitals",
              "ontario-ag-2026-ai-report"
            ]
          },
          {
            "text": "Give agents least privilege: read-only by default, no delete, and human approval for any action that changes orders, medications or records.",
            "source_ids": [
              "owasp-llm-top10-2025",
              "owasp-2025-agentic-top10"
            ]
          }
        ],
        "1": [
          {
            "text": "Define in advance who can switch a model off and on what trigger (for example alert volume above a set multiple of baseline, or sampled sensitivity below a floor), and give them that authority in writing.",
            "source_ids": [
              "nist-ai-100-1-rmf",
              "finlayson-2021-nejm-dataset-shift"
            ]
          },
          {
            "text": "When a model is paused, show users a visible 'model off' banner so nobody assumes silence means no risk, and re-issue the manual screening criteria the model replaced.",
            "source_ids": [
              "wong-2021-jama-netw-open-sepsis-covid-alerts"
            ]
          },
          {
            "text": "Make bedside override explicit and unpenalized in every protocol that starts from a model flag; require patient-specific reassessment before acting.",
            "source_ids": [
              "euronews-ap-2025-nurses-ai-sepsis",
              "cms-2024-ma-faq-algorithms"
            ]
          },
          {
            "text": "Require the signing clinician to read and correct AI-drafted notes and letters before they leave the chart, and mark unsigned AI text as unverified.",
            "source_ids": [
              "ontario-ag-2026-ai-report",
              "nist-ai-600-1-genai-profile"
            ]
          }
        ],
        "2": [
          {
            "text": "Keep the pre-AI manual workflow documented and staffed: sepsis screening criteria, dictation or typed notes, pharmacist-handled refill calls. Rehearse switching to it.",
            "source_ids": [
              "wong-2021-jama-netw-open-sepsis-covid-alerts",
              "wcax-2026-kinney-pullback"
            ]
          },
          {
            "text": "Offer a non-AI route for patients at all times (touch-tone or staffed line, human scribe or no scribe) rather than making the AI path the only way to get care.",
            "source_ids": [
              "vtdigger-2026-kinney-burt",
              "wcax-2026-kinney-pullback"
            ]
          },
          {
            "text": "When a model is decommissioned, review the decisions it influenced during the suspect period, such as notes, denials and missed alerts, and correct records where needed.",
            "source_ids": [
              "nist-ai-100-1-rmf",
              "finlayson-2021-nejm-dataset-shift"
            ]
          }
        ],
        "3": [
          {
            "text": "Keep paper or offline versions of the clinical criteria that models encode (sepsis screens, early-warning scores) so they can be applied when EHR-native models and the EHR are both unavailable.",
            "source_ids": []
          },
          {
            "text": "Include model and agent failure in downtime drills: practise a scenario where the EHR is up but a model is known to be wrong, not only one where everything is down.",
            "source_ids": []
          },
          {
            "text": "Report AI-related harm and near misses externally, through a Patient Safety Organization or FDA's reporting pathways for regulated devices, so other sites learn before they hit the same failure.",
            "source_ids": [
              "jc-chai-2025-ruaih-guidance"
            ]
          }
        ]
      },
      "standards": [
        {
          "name": "FDA draft guidance: AI-Enabled Device Software Functions, Lifecycle Management (Jan 2025, draft)",
          "what_it_requires": "Recommends (non-binding) that sponsors describe postmarket performance monitoring covering data drift, demographic shift and input-pipeline corruption, and how results reach users. Postmarket adverse-event reporting under 21 CFR 803 still applies.",
          "source_id": "fda-2025-ai-dsf-draft-guidance"
        },
        {
          "name": "FDA final guidance: Predetermined Change Control Plans for AI-Enabled Device Software Functions (Dec 2024, revised Aug 2025)",
          "what_it_requires": "A PCCP in the marketing submission sets out planned modifications, the protocol to develop, validate and implement them, and an impact assessment, so authorized changes can ship without a new submission.",
          "source_id": "fda-2024-pccp-final-guidance"
        },
        {
          "name": "NIST AI RMF 1.0 (AI 100-1) and Generative AI Profile (AI 600-1)",
          "what_it_requires": "Voluntary framework: monitor systems in production (MEASURE 2.4), keep mechanisms and named responsibility to supersede, disengage or deactivate AI (MANAGE 2.4), and plan post-deployment override, decommissioning and incident response (MANAGE 4.1). The GenAI profile adds confabulation and automation bias as named risks.",
          "source_id": "nist-ai-100-1-rmf"
        },
        {
          "name": "45 CFR 92.210 (Section 1557, patient care decision support tools)",
          "what_it_requires": "Covered entities must make ongoing reasonable efforts to identify decision-support tools that use race, color, national origin, sex, age or disability as inputs, and to mitigate discrimination risk (applies from May 2025).",
          "source_id": "hhs-45cfr92-210"
        },
        {
          "name": "45 CFR 170.315(b)(11) ONC certification, Decision Support Interventions",
          "what_it_requires": "Certified health IT must expose source attributes for predictive DSIs, including intended use, out-of-scope cautions, external validation, fairness and local-data validity monitoring. HTI-5 (proposed Dec 2025) would remove these model-card requirements.",
          "source_id": "onc-45cfr170-315-b11"
        },
        {
          "name": "CMS CY2024 MA rule FAQ on algorithms (42 CFR 422.101(c))",
          "what_it_requires": "Medicare Advantage plans may use algorithms to assist, but coverage decisions must rest on the individual patient's circumstances. A predicted length of stay alone cannot end post-acute care.",
          "source_id": "cms-2024-ma-faq-algorithms"
        },
        {
          "name": "Joint Commission + CHAI Responsible Use of AI in Healthcare guidance (Sept 2025)",
          "what_it_requires": "Voluntary guidance with seven elements, including ongoing local quality monitoring, risk and bias assessment, and voluntary blinded reporting of AI safety events.",
          "source_id": "jc-chai-2025-ruaih-guidance"
        }
      ],
      "elsewhere": {
        "text": "The EU AI Act (Regulation (EU) 2024/1689, in force 1 August 2024) treats AI safety components of regulated products as high-risk, with duties for risk management, data quality, logging, human oversight and accuracy. The 2026 Digital Omnibus, signed 8 July 2026, moved the high-risk dates to 2 December 2027 for stand-alone systems and 2 August 2028 for AI embedded in products such as medical devices, which remain under MDR/IVDR conformity assessment in the meantime. In the UK, the MHRA ran the AI Airlock regulatory sandbox for AI as a medical device from April 2024 to March 2025 and published its lessons in October 2025. Ontario's Auditor General found in May 2026 that all 20 provincially approved AI scribes produced inaccurate notes in procurement testing.",
        "source_ids": [
          "ec-ai-act-overview",
          "ep-2026-digital-omnibus-ai",
          "mhra-2025-ai-airlock-report",
          "ontario-ag-2026-ai-report"
        ]
      },
      "scoring_v01": {
        "likelihood": 4,
        "blast_radius": 3,
        "detectability": 5,
        "rationale": "Judgement: published external validations and audits keep finding deployed models and scribes that perform worse than claimed, so we score likelihood high. Blast radius is moderate per event, usually one model's users rather than a whole hospital, but vendor-wide models like the Epic sepsis score reach hundreds of sites at once. Detectability is the worst on the site because a wrong model raises no alarm and no tier change; it is found by audit, research or a clinician who refuses to comply."
      },
      "open_questions": [
        "How often do deployed clinical models degrade after go-live, and how long does it take to notice? There is no public registry of local performance or decommissioning decisions.",
        "What error rate do ambient scribes have in real use, as opposed to procurement tests and vendor-authored studies, and how many errors survive clinician sign-off?",
        "Does clinician review actually catch LLM omissions, or does fluent text lower scrutiny? Evidence on review effectiveness is thin.",
        "If HTI-5 removes the DSI model-card requirement, what will replace local transparency about validation and fairness for EHR-embedded models?",
        "How should agents that act on clinical records be validated and monitored, and who owns an agent's action under FDA, CMS and state pharmacy rules?"
      ],
      "watch_queries": [
        "\"sepsis model\" OR \"deterioration model\" external validation hospital",
        "\"dataset shift\" OR \"model drift\" clinical prediction deployment",
        "\"AI scribe\" OR \"ambient documentation\" error OR hallucination OR omission",
        "FDA \"AI-enabled device\" recall OR \"MAUDE\" artificial intelligence",
        "\"predictive algorithm\" denial Medicare Advantage OR prior authorization",
        "clinical \"AI agent\" OR \"agentic\" hospital incident OR error",
        "\"algorithmic bias\" clinical decision support race OR disability",
        "\"large language model\" clinical advice harm patient case report",
        "HTI-5 decision support interventions final rule",
        "Joint Commission \"AI safety event\" OR \"Responsible Use of AI\" certification"
      ],
      "last_reviewed": "2026-09-26",
      "reviewed_by": "opus-research",
      "last_audited": "2026-09-26",
      "audit_status": "verified"
    },
    {
      "id": "handoff",
      "name": "The human handoff",
      "one_liner": "The moment automation hands work back to a person who has stopped watching, stopped practising, or stopped hearing the alarms.",
      "definition": [
        "The human handoff layer is everything that has to be true of people for an automated hospital to fail safely: that clinicians notice when a machine is wrong, that someone owns each alarm and alert, that staff can take over when automation stops, and that they can still run the ward on paper. It covers automation bias (following or waiting for the machine), out-of-the-loop performance (taking over without situation awareness), skill atrophy and deskilling, alarm and alert fatigue, and downtime competence.",
        "The problem is old. Bainbridge's 1983 'Ironies of Automation' observed that automating a process leaves the operator with the tasks the designer could not automate, including taking over in abnormal conditions, while the manual skills needed for that takeover deteriorate when they are not used. Parasuraman and Riley (1997) described misuse (over-reliance) and disuse (ignoring automation, commonly after false alarms). Endsley and Kiris (1995) showed that operators of an automated system were slower to decide after it failed, and linked this to lower situation awareness from passive monitoring.",
        "Healthcare now has direct evidence of each mechanism. Experienced radiologists' accuracy collapsed when a purported AI suggested the wrong BI-RADS category (Dratsch 2023). Endoscopists' adenoma detection rate in non-AI colonoscopies fell from 28.4% to 22.4% after AI was introduced (Budzyń 2025). The Joint Commission logged 98 alarm-related sentinel events, 80 of them deaths, between January 2009 and June 2012. During EHR downtime, safety reports show downtime procedures missing or not followed in 46% of cases (Larsen 2018), and clinicians at Ascension in 2024 described medication errors and delayed labs after the switch to paper."
      ],
      "angle": "FailSystems' view: in an automated hospital the human is no longer the primary operator but the fallback, and a fallback that is never exercised decays. Every other layer on this site eventually fails into this one: when power, network, devices or models go, the plan is 'staff take over', and that plan silently assumes skills, attention and alarm ownership that automation has been eroding. The dangerous moment is not the outage itself but the handback, when a person with stale skills and little context is asked to act fast. We judge that rehearsal (unaided practice at tier 0, drills at tiers 2 and 3) is the only defense that addresses the cause rather than the symptom, and that it is the one hospitals most often cut.",
      "failure_modes": [
        {
          "id": "automation-bias",
          "name": "Automation bias: following the machine, or waiting for it",
          "mechanism": "Clinicians over-rely on decision support, making errors of commission (following incorrect advice) or omission (not acting because the system did not prompt). Reliance increases with workload, time pressure and task complexity, and when verifying the output is hard. Experience reduces but does not remove it: in a mammography experiment, very experienced readers' accuracy dropped from 82.3% to 45.5% when the AI suggestion was wrong.",
          "warning_signs": [
            "Clinician-AI agreement near 100% with no documented overrides",
            "Decisions made faster after AI deployment with no change in case mix",
            "Staff cannot explain the basis of a recommendation they acted on",
            "AI output placed where it is seen before the clinician's own assessment"
          ],
          "incident_ids": [
            "mammography-automation-bias-experiment-2023"
          ],
          "source_ids": [
            "goddard-2012-jamia-automation-bias",
            "lyell-coiera-2017-jamia-verification-complexity",
            "dratsch-2023-radiology-mammography-automation-bias",
            "fda-cds-guidance-2026",
            "parasuraman-riley-1997-human-factors",
            "sujan-2019-bmjhci-human-factors-ai"
          ]
        },
        {
          "id": "out-of-the-loop-handback",
          "name": "Out-of-the-loop handback",
          "mechanism": "When automation disconnects or fails, the person taking over has been monitoring passively and lacks current situation awareness. Endsley and Kiris found slower decisions after an expert system failed, attributed mainly to the shift from active to passive processing. Bainbridge noted that manual operators need 15 to 30 minutes to build a feel for a process before taking over, time an automated handback rarely allows.",
          "warning_signs": [
            "Automation disconnects or mode changes without a clear, explained display",
            "Takeover procedures exist on paper but are never practised against realistic failure signatures",
            "Supervisors monitor many automated streams at once with low event rates"
          ],
          "incident_ids": [
            "air-france-447-2009",
            "uber-tempe-2018"
          ],
          "source_ids": [
            "endsley-kiris-1995-human-factors",
            "bainbridge-1983-automatica-ironies",
            "bea-af447-final-report-2012",
            "ntsb-har-19-03-uber-tempe"
          ]
        },
        {
          "id": "skill-atrophy-deskilling",
          "name": "Skill atrophy and deskilling",
          "mechanism": "Skills that automation performs routinely are practised less and deteriorate, so a formerly experienced operator becomes an inexperienced one at the moment of takeover. In four Polish endoscopy centres, the adenoma detection rate of standard (non-AI) colonoscopy fell 6.0 percentage points in the three months after AI was introduced. Trainees who learn with automation from the start may never build the unaided skill at all (judgement; not yet measured).",
          "warning_signs": [
            "Unaided performance is not measured after AI rollout",
            "No protected unaided cases or simulator time",
            "New staff onboarded only on the automated workflow"
          ],
          "incident_ids": [
            "endoscopy-deskilling-study-2025"
          ],
          "source_ids": [
            "budzyn-2025-lancet-gh-deskilling",
            "bainbridge-1983-automatica-ironies"
          ]
        },
        {
          "id": "alarm-fatigue",
          "name": "Alarm fatigue and alarm disuse",
          "mechanism": "Most device alarm signals do not require clinical intervention (the Joint Commission cites estimates of 85 to 99 percent), so clinicians become desensitised and turn alarms down, off, or outside safe limits. Parasuraman and Riley described this as disuse: false alarms teach operators to ignore automation. Alarm fatigue was the most common contributing factor in alarm-related sentinel events reported to the Joint Commission.",
          "warning_signs": [
            "Hundreds of alarm signals per patient per day on a unit",
            "Default alarm limits never tailored to the patient",
            "Alarms found silenced or with widened limits on rounds",
            "ECG electrodes and sensors not changed on schedule"
          ],
          "incident_ids": [
            "alarm-response-death-massachusetts-2010",
            "mgh-monitor-alarm-off-2010"
          ],
          "source_ids": [
            "tjc-sea-50-alarms-2013",
            "sendelbach-funk-2013-aacn-alarm-fatigue",
            "cvach-2012-bit-alarm-fatigue",
            "parasuraman-riley-1997-human-factors"
          ]
        },
        {
          "id": "unowned-alarms",
          "name": "Unowned alarms and unrouted alerts",
          "mechanism": "An alarm only protects a patient if a specific person hears it and is responsible for acting. Joint Commission data list alarms not audible in all areas (25 events), alarms inappropriately turned off (36) and inadequate staffing to respond among contributors. The Joint Commission's alarm-safety goal (NPG.01.05.01 from January 2026, previously NPSG.06.01.01) requires hospitals to define who may set, change and turn off alarm parameters, but it generally excludes CPOE and other IT alerts, so AI and EHR alerts can fall outside any alarm-ownership policy.",
          "warning_signs": [
            "No written answer to 'who responds to this alert, and within how long?'",
            "AI or EHR alerts routed to a shared inbox or a role, not a person",
            "Central-station monitoring without a named watcher on every shift"
          ],
          "incident_ids": [
            "alarm-response-death-massachusetts-2010",
            "epic-sepsis-model-2021",
            "mgh-monitor-alarm-off-2010"
          ],
          "source_ids": [
            "tjc-sea-50-alarms-2013",
            "tjc-r3-issue-5-npsg-06-01-01",
            "tjc-npg-2026-hospital",
            "tjc-perspectives-2025-07",
            "tjc-npsg-2025-hospital"
          ]
        },
        {
          "id": "alert-override",
          "name": "Software alert fatigue and reflexive override",
          "mechanism": "Decision-support alerts with low specificity are overridden routinely: a review found drug safety alerts overridden in 49% to 96% of cases. A widely deployed sepsis model generated alerts for 18% of all hospitalised patients while missing 67% of sepsis cases at one academic centre, a burden the authors described as alert fatigue. Clinicians learn to click through, including on the rare correct alert.",
          "warning_signs": [
            "Override rates above 90% for an alert class",
            "Alert volume rises with each new model or rule, none retired",
            "No one reviews overridden alerts that preceded harm"
          ],
          "incident_ids": [
            "epic-sepsis-model-2021"
          ],
          "source_ids": [
            "vandersijs-2006-jamia-alert-override",
            "wong-2021-jama-im-epic-sepsis"
          ]
        },
        {
          "id": "downtime-incompetence",
          "name": "Lost downtime competence",
          "mechanism": "When the EHR or other systems go down, staff must run care on paper, but ONC notes that many organisations have employees who do not know how to work in a paper-based environment. Safety reports tied to downtime cluster in lab orders and results and medication, and patient identification and communication fail. Paper workflows are slower (lab results averaged 62% longer in one study) and the supporting infrastructure, such as forms and fax machines, may no longer exist.",
          "warning_signs": [
            "No downtime drill in the past 12 months",
            "Paper forms missing, outdated, or stored where no one can find them",
            "Read-only backup EHR credentials unknown to front-line staff",
            "Downtime events not followed by review"
          ],
          "incident_ids": [
            "ascension-2024"
          ],
          "source_ids": [
            "onc-safer-contingency-2025",
            "larsen-2018-jamia-ehr-downtime-events",
            "larsen-2019-aci-downtime-laboratory",
            "kff-2024-ascension-lapses",
            "ecri-top10-hazards-2026"
          ]
        }
      ],
      "detection": [
        {
          "text": "Measure unaided performance after any AI rollout, not just AI-assisted performance. In colonoscopy, the non-AI adenoma detection rate is the metric that exposed deskilling.",
          "source_ids": [
            "budzyn-2025-lancet-gh-deskilling"
          ]
        },
        {
          "text": "Track clinician-AI disagreement and override rates. Near-total agreement on a system with known error rates is a sign of automation bias, not accuracy.",
          "source_ids": [
            "goddard-2012-jamia-automation-bias",
            "dratsch-2023-radiology-mammography-automation-bias"
          ]
        },
        {
          "text": "Count alarm signals per bed per day and the fraction that were actionable; several hundred per patient per day, with most non-actionable, predicts alarm fatigue.",
          "source_ids": [
            "tjc-sea-50-alarms-2013"
          ]
        },
        {
          "text": "Track override rates by alert type. Rates approaching the 49 to 96 percent reported for drug alerts mean the alert is being ignored as a class.",
          "source_ids": [
            "vandersijs-2006-jamia-alert-override"
          ]
        },
        {
          "text": "Search incident reports for 'downtime' and code whether procedures were in place and followed; in one analysis 46% of downtime-related reports said they were not.",
          "source_ids": [
            "larsen-2018-jamia-ehr-downtime-events"
          ]
        },
        {
          "text": "Time key tasks during downtime drills (first paper medication order, first critical lab result delivered) and compare with normal operation.",
          "source_ids": [
            "larsen-2019-aci-downtime-laboratory",
            "onc-safer-contingency-2025"
          ]
        }
      ],
      "defenses_by_tier": {
        "0": [
          {
            "text": "Keep a set of unaided cases for every AI-assisted task and report unaided performance to the department quarterly.",
            "source_ids": [
              "budzyn-2025-lancet-gh-deskilling",
              "bainbridge-1983-automatica-ironies"
            ]
          },
          {
            "text": "Buy or build decision support that shows its inputs and reasoning so the clinician can check it, and treat time-critical uses as highest risk for automation bias.",
            "source_ids": [
              "fda-cds-guidance-2026",
              "goddard-2012-jamia-automation-bias"
            ]
          },
          {
            "text": "Inventory every alarm- and alert-generating system, AI and EHR alerts included, and name the role that responds to each and the response time expected.",
            "source_ids": [
              "tjc-r3-issue-5-npsg-06-01-01",
              "tjc-sea-50-alarms-2013"
            ]
          },
          {
            "text": "Retire or retune any alert class with an override rate above your threshold before adding a new one.",
            "source_ids": [
              "vandersijs-2006-jamia-alert-override",
              "wong-2021-jama-im-epic-sepsis"
            ]
          },
          {
            "text": "For each AI tool, write down which function it automates (information gathering, analysis, choosing an action, or carrying it out) and at what level, and keep high-consequence action selection at a level where a clinician must actively decide.",
            "source_ids": [
              "parasuraman-sheridan-wickens-2000-ieee"
            ]
          },
          {
            "text": "Train users that automation bias exists and emphasise their accountability for the final decision; both are among the few mitigators with evidence.",
            "source_ids": [
              "goddard-2012-jamia-automation-bias"
            ]
          }
        ],
        "1": [
          {
            "text": "Make every automation mode change visible and explained on screen: say what is off, since when, and what the clinician now owns.",
            "source_ids": [
              "bea-af447-final-report-2012",
              "endsley-kiris-1995-human-factors"
            ]
          },
          {
            "text": "Write and rehearse takeover procedures for each known failure signature (model offline, feed stale, sensor disagreement), not a generic 'use clinical judgement'.",
            "source_ids": [
              "bea-af447-final-report-2012",
              "bainbridge-1983-automatica-ironies"
            ]
          },
          {
            "text": "Where verification is hard, reduce the clinician's cognitive load before asking them to check the AI; automation bias tracks verification complexity.",
            "source_ids": [
              "lyell-coiera-2017-jamia-verification-complexity"
            ]
          },
          {
            "text": "Put a named human on every automated monitoring stream with a realistic watch load; do not rely on passive supervision of low-event streams.",
            "source_ids": [
              "ntsb-har-19-03-uber-tempe",
              "bainbridge-1983-automatica-ironies"
            ]
          }
        ],
        "2": [
          {
            "text": "Tailor alarm limits to the patient and change ECG electrodes and single-use sensors on the manufacturer's schedule to cut nuisance alarms.",
            "source_ids": [
              "tjc-sea-50-alarms-2013"
            ]
          },
          {
            "text": "Write down who may set, change and turn off alarm parameters, and audit silenced or widened alarms on rounds.",
            "source_ids": [
              "tjc-r3-issue-5-npsg-06-01-01"
            ]
          },
          {
            "text": "Test that critical alarm signals are audible in every area where the responder may be, including at night staffing levels.",
            "source_ids": [
              "tjc-sea-50-alarms-2013"
            ]
          },
          {
            "text": "Specify alarm priority and signal conventions to IEC 60601-1-8 in procurement so devices from different vendors are distinguishable by urgency.",
            "source_ids": [
              "iec-60601-1-8-handoff"
            ]
          }
        ],
        "3": [
          {
            "text": "Run an unannounced EHR downtime drill at least once a year on every clinical unit, and time the first paper order and first critical result.",
            "source_ids": [
              "onc-safer-contingency-2025",
              "cook-1998-how-complex-systems-fail"
            ]
          },
          {
            "text": "Train every clinician on paper ordering and charting and on activating the read-only backup EHR, and make sure they can find its login.",
            "source_ids": [
              "onc-safer-contingency-2025"
            ]
          },
          {
            "text": "Stock current paper forms for key EHR functions on each unit and write a patient-identification procedure for before, during and after downtime.",
            "source_ids": [
              "onc-safer-contingency-2025",
              "larsen-2018-jamia-ehr-downtime-events"
            ]
          },
          {
            "text": "Use one of the two exercises CMS already requires each year for a multi-day loss of the EHR and clinical systems, and revise the plan from what you learn.",
            "source_ids": [
              "ecfr-42cfr482-15",
              "ecri-top10-hazards-2026"
            ]
          },
          {
            "text": "Keep a downtime communication channel that does not depend on the EHR's computing infrastructure.",
            "source_ids": [
              "onc-safer-contingency-2025"
            ]
          }
        ]
      },
      "standards": [
        {
          "name": "Joint Commission NPG.01.05.01, clinical alarm safety (hospitals and critical access hospitals, from January 2026; previously NPSG.06.01.01, effective 2014)",
          "what_it_requires": "Leaders make alarm safety a priority, identify the most important alarm signals, and set policies on alarm settings, when alarms may be disabled or changed, who has authority to set, change or turn them off, and monitoring and response; staff must be educated on the alarm systems they are responsible for. It generally does not cover CPOE or other IT alerts.",
          "source_id": "tjc-npg-2026-hospital"
        },
        {
          "name": "42 CFR 482.15(d), CMS Conditions of Participation: emergency preparedness training and testing",
          "what_it_requires": "Hospitals must train all staff in emergency procedures initially and at least every 2 years, demonstrate staff knowledge, and run at least two exercises a year (one full-scale or functional, one additional such as a tabletop), analysing and documenting each.",
          "source_id": "ecfr-42cfr482-15"
        },
        {
          "name": "ONC SAFER Guide: Contingency Planning (2025 revision)",
          "what_it_requires": "Voluntary self-assessment. Recommends paper forms for key EHR functions, training and testing staff on downtime and recovery (including unannounced drills at least yearly and read-only backup EHR use), EHR-independent communication, and review of downtimes over 24 hours.",
          "source_id": "onc-safer-contingency-2025"
        },
        {
          "name": "FDA, Clinical Decision Support Software guidance (January 29, 2026)",
          "what_it_requires": "Non-binding. Defines automation bias and treats the level of automation and the time-critical nature of the decision as factors in whether a clinician can independently review the basis of a recommendation (criterion 4 for non-device CDS).",
          "source_id": "fda-cds-guidance-2026"
        },
        {
          "name": "IEC 60601-1-8:2006+A1:2012 (Ed. 2.1), alarm systems collateral standard",
          "what_it_requires": "Requirements and tests for alarm systems in medical electrical equipment: alarm categories by urgency, consistent alarm signals and control states, and their marking.",
          "source_id": "iec-60601-1-8-handoff"
        }
      ],
      "elsewhere": {
        "text": "The EU AI Act (Regulation (EU) 2024/1689) is the first law to name automation bias. Article 14(4)(b) requires high-risk AI systems to be provided so that overseers can remain aware of the tendency to over-rely on outputs, and 14(4)(d)-(e) require that they can disregard or override an output and stop the system safely. Article 26(2) requires deployers, such as hospitals, to assign human oversight to people with the necessary competence, training, authority and support. In England, DCB0160, mandated under section 250 of the Health and Social Care Act 2012, requires care organisations to apply clinical risk management to the deployment and use of health IT, which is where local handoff and downtime hazards are expected to be logged.",
        "source_ids": [
          "eu-ai-act-2024-1689",
          "aiact-explorer-articles-14-26",
          "nhs-dcb0160"
        ]
      },
      "scoring_v01": {
        "likelihood": 4,
        "blast_radius": 3,
        "detectability": 4,
        "rationale": "Judgement. Likelihood is high because alarm fatigue, alert override and weak downtime readiness are documented as common, not exceptional. Blast radius is usually one patient or one unit per event, but rises to system-wide when a large outage forces a whole network onto paper. Detectability is poor because skill loss and automation bias stay invisible until the handback, and routine metrics measure assisted rather than unaided performance."
      },
      "open_questions": [
        "How fast does unaided skill decay after AI adoption, does it plateau, and does it recover when AI is withdrawn? Budzyń compared only three months before and after, in one specialty.",
        "What drill frequency and format actually preserves paper-mode competence? SAFER recommends at least an annual unannounced drill, but outcome evidence for any frequency is thin.",
        "Does showing the basis of a recommendation reduce automation bias under real time pressure, or only in experiments?",
        "Who owns AI and EHR alerts that fall outside the Joint Commission's alarm-safety goal (NPG.01.05.01), and does any accreditor survey their response?",
        "Will clinicians trained from the start with AI assistance ever develop the unaided skill the fallback plan assumes?"
      ],
      "watch_queries": [
        "\"automation bias\" AND (clinical OR physician OR nurse OR radiolog*)",
        "deskilling AND \"artificial intelligence\" AND (endoscopy OR radiology OR pathology OR dermatology)",
        "\"alarm fatigue\" sentinel event OR death",
        "\"alert fatigue\" override rate clinical decision support",
        "\"EHR downtime\" OR \"electronic health record downtime\" patient safety",
        "hospital cyberattack nurses paper charting medication error",
        "\"out-of-the-loop\" OR \"takeover\" automation healthcare situation awareness",
        "\"human oversight\" AI Act Article 14 hospital deployer",
        "Joint Commission alarm OR \"NPG.01.05.01\" OR \"NPSG.06.01.01\" update",
        "\"never-skilling\" OR \"skill decay\" AI medical training"
      ],
      "last_reviewed": "2026-09-27",
      "reviewed_by": "opus-research",
      "last_audited": "2026-09-26",
      "audit_status": "verified"
    },
    {
      "id": "cascades",
      "name": "Cascades",
      "one_liner": "Not a layer: the way one failure crosses power, connectivity, devices, models and handoff, and spreads from one organisation to many.",
      "definition": [
        "Cascades is not a sixth layer. It is a propagation pattern that runs across the other five: power, connectivity, devices, models and handoff. A cascade starts as a failure in one layer, at one organisation, and turns into a failure in other layers or other organisations. Examples: a remote-access portal is breached and national claims processing stops. A security-software update crashes Windows machines and hospital services go offline. A pathology supplier is hit and three hospital trusts have to call for O-type blood donors.",
        "Safety science has argued for decades that serious accidents in complex systems do not come from one broken part. Perrow called accidents in systems that are both interactively complex and tightly coupled 'normal accidents': they are to be expected, not freak events. Reason's Swiss-cheese model describes harm reaching a patient only when latent conditions and active failures in several defensive layers line up. Leveson's STAMP and CAST treat accidents as a loss of control over the system, not a chain of failed parts. Cook and Rasmussen described how hospitals chasing efficiency 'go solid': the buffers that used to soak up a problem disappear, so that an event in one distant part of the hospital suddenly matters everywhere else. None of these models is settled: professionals disagree on what the parts of the Swiss-cheese model mean, and its critics call it too linear. We use them as lenses, not as laws.",
        "On this site a cascade is described by its path (the order in which it crossed layers, such as devices → connectivity → handoff) and its reach (one department, one hospital, a region, a country). Tracing the path matters because the controls that stop a cascade usually sit at the boundaries between layers and between organisations, and those boundaries are the parts nobody owns."
      ],
      "angle": "FailSystems' view: automation does not add many new ways for a single part to fail. What it adds is coupling. When a hospital runs on a shared EHR, a shared clearinghouse, a shared endpoint agent and a shared pathology network, the same fault reaches every place at once, and the paper workaround has usually withered from lack of use. In our judgement the incidents that do the most harm to patients in automated care will be cascades. Most of them will start outside the hospital, in a supplier the hospital does not control and may not know it depends on. Plan for common-mode failure, not for one component failing on its own.",
      "failure_modes": [
        {
          "id": "shared-supplier-concentration",
          "name": "Single shared supplier (concentration risk)",
          "mechanism": "Many hospitals, pharmacies or practices depend on one supplier for one function: claims clearing, pathology, identity, endpoint security. When that supplier fails, every customer loses the function at the same moment, and there is often no quick way to switch because contracts, interfaces and enrolments were built for one path. The customers also cannot see into the supplier, so they cannot judge when service will come back.",
          "warning_signs": [
            "One supplier handles a function for most of the organisations in your region or specialty",
            "No alternate supplier is enrolled or tested, or switching would take weeks of EDI or interface work",
            "Contract has no restoration-time commitment or incident-notification clause",
            "You cannot list which clinical workflows stop if that supplier is down for 7 days"
          ],
          "incident_ids": [
            "change-healthcare-2024",
            "synnovis-2024"
          ],
          "source_ids": [
            "witty-2024-senate-finance-testimony",
            "aha-2024-change-survey",
            "govuk-2026-csr-bill-critical-suppliers",
            "hscc-cms-2023-hospital-resiliency-landscape"
          ]
        },
        {
          "id": "common-mode-update",
          "name": "Common-mode update",
          "mechanism": "One software or content update is pushed to every installation at once and carries a latent defect that testing missed. Because every machine runs the same code, redundancy inside the hospital does not help: the primary and the backup crash together. Updates designed to ship fast, like security content, are the most exposed.",
          "warning_signs": [
            "Vendors can push updates to production endpoints with no staging ring under your control",
            "Primary and backup systems run the same agent, OS build or content version",
            "No inventory of which clinical workstations and servers run each kernel-level agent",
            "Recovery requires hands-on access to each machine"
          ],
          "incident_ids": [
            "crowdstrike-2024"
          ],
          "source_ids": [
            "crowdstrike-rca-channel-file-291",
            "tully-2025-jama-netw-open-crowdstrike",
            "microsoft-2024-crowdstrike-8-5m-devices"
          ]
        },
        {
          "id": "tight-coupling-lost-buffers",
          "name": "Tight coupling and lost buffers ('going solid')",
          "mechanism": "Efficiency work strips out slack: spare beds, spare stock, spare staff, manual steps. Activities then depend directly on events elsewhere in the system, so a delay in one place becomes a stoppage in another within hours. In a tightly coupled process there is no time to improvise before the next step needs the output of the failed one.",
          "warning_signs": [
            "Just-in-time supply with no local stock for time-critical items (blood, reagents, drugs)",
            "Bed occupancy routinely near 100%",
            "Paper downtime forms missing, out of date or never drilled",
            "Workflows where the next step cannot start without an electronic result"
          ],
          "incident_ids": [
            "synnovis-2024",
            "texas-winter-storm-2021"
          ],
          "source_ids": [
            "cook-rasmussen-2005-going-solid",
            "perrow-1999-normal-accidents",
            "nhsbt-2024-o-type-appeal"
          ]
        },
        {
          "id": "cross-infrastructure-interdependency",
          "name": "Cross-infrastructure interdependency",
          "mechanism": "Hospitals depend on utilities that depend on each other. Loss of electricity can stop water treatment and gas supply, which in turn stops hospital heating, sterilisation and toilets. Patients at home on powered equipment lose it at the same time and arrive at the emergency department. The hospital's generator covers its own electricity but not the water pressure or the patients' home equipment.",
          "warning_signs": [
            "Emergency plan assumes municipal water and gas stay up when the grid is down",
            "Boilers or chillers depend on mains water pressure",
            "No register of local patients on home oxygen, dialysis or other electricity-dependent equipment",
            "Regional plan assumes neighbours can take transfers during a region-wide event"
          ],
          "incident_ids": [
            "texas-winter-storm-2021"
          ],
          "source_ids": [
            "ferc-nerc-2022-feb2021-outages-tracking",
            "dshs-2021-winter-storm-deaths",
            "abc-2021-texas-hospitals-water",
            "circleofblue-2021-austin-hospitals-water"
          ]
        },
        {
          "id": "regional-spillover",
          "name": "Regional spillover to neighbouring hospitals",
          "mechanism": "A hospital that goes to downtime diverts ambulances and patients to its neighbours. The neighbours have their own systems intact but not the extra capacity, so waits, walk-outs and delays in time-critical care rise there too. One organisation's cyber incident becomes a capacity incident for the region.",
          "warning_signs": [
            "One health system holds a large share of regional inpatient capacity",
            "No regional agreement on diversion, transfers or shared downtime capacity",
            "EMS diversion hours climbing without a known cause",
            "Neighbouring hospitals are told about an outage by the news, not by the affected organisation"
          ],
          "incident_ids": [
            "san-diego-ransomware-spillover-2021",
            "wannacry-nhs-2017"
          ],
          "source_ids": [
            "dameff-2023-jama-netw-open-adjacent-eds",
            "nao-2017-wannacry",
            "ecfr-42cfr482-15"
          ]
        },
        {
          "id": "defensive-disconnection",
          "name": "Defensive disconnection",
          "mechanism": "Cutting network links to contain an attack is often the right call, but it causes a cascade of its own. Partners that relied on the link lose the service, and organisations that were never infected shut systems down as a precaution because they lack clear central advice. The outage from containment can be larger than the outage from the attack.",
          "warning_signs": [
            "No pre-agreed criteria for when to disconnect from a partner or supplier",
            "No plan for running the services that ride on that connection while it is cut",
            "Central incident guidance takes hours to reach local sites"
          ],
          "incident_ids": [
            "change-healthcare-2024",
            "wannacry-nhs-2017"
          ],
          "source_ids": [
            "witty-2024-senate-finance-testimony",
            "nao-2017-wannacry"
          ]
        },
        {
          "id": "latent-conditions-aligning",
          "name": "Latent conditions lining up",
          "mechanism": "Most cascades need several existing weaknesses at once: an unpatched system, a portal without multi-factor authentication, a missing bounds check, an untested backup. Each one is tolerable alone and may sit unnoticed for months. The cascade happens when a trigger finds a path through all of them. Use the Swiss-cheese picture with care. A survey of quality and safety professionals found they read its parts (holes, slices, arrow) in very different ways, and critics argue it is too static and linear. Its defenders still consider it useful because it is systemic.",
          "warning_signs": [
            "Known findings (unpatched hosts, missing MFA, failed audits) carried forward year after year",
            "Assessments that were done but gave no one the power to require fixes",
            "Incident reviews that stop at the first human or technical 'root cause'"
          ],
          "incident_ids": [
            "wannacry-nhs-2017",
            "change-healthcare-2024",
            "crowdstrike-2024"
          ],
          "source_ids": [
            "reason-2000-bmj-human-error",
            "nao-2017-wannacry",
            "witty-2024-senate-finance-testimony",
            "crowdstrike-rca-channel-file-291",
            "leveson-2019-cast-handbook",
            "perneger-2005-swiss-cheese-holes",
            "larouzee-lecoze-2020-swiss-cheese-critics"
          ]
        }
      ],
      "detection": [
        {
          "text": "Map your dependencies before the event. List every external service a clinical workflow needs (clearinghouse, lab, e-prescribing, identity, endpoint agents, cloud EHR) and mark any service shared with most of your region. If you cannot produce this map, you cannot see a cascade coming.",
          "source_ids": [
            "dutch-safety-board-2020-it-outages",
            "nist-sp800-161r1"
          ]
        },
        {
          "text": "Watch external availability, not only internal alarms. During the CrowdStrike outage, researchers detected disrupted hospital services from outside by scanning network ports and FHIR endpoints every few hours. Hospitals and regional coalitions can use the same kind of outside-in monitoring to spot outages at peers and suppliers.",
          "source_ids": [
            "tully-2025-jama-netw-open-crowdstrike"
          ]
        },
        {
          "text": "Track regional load signals: EMS diversion hours, ambulance arrivals, left-without-being-seen rates and stroke-code volume at your own ED. A sudden rise with no local cause may be a neighbour's outage reaching you.",
          "source_ids": [
            "dameff-2023-jama-netw-open-adjacent-eds"
          ]
        },
        {
          "text": "Require suppliers to tell you when they activate their contingency plan. HHS has proposed requiring business associates to report contingency-plan activation within 24 hours. Until that is final, put it in the contract.",
          "source_ids": [
            "hhs-hipaa-security-nprm-2025"
          ]
        },
        {
          "text": "Treat a rise in workaround use as a signal. Manual claims, phoned results, O-negative use above baseline and paper orders all show that a coupled system has degraded, often before anyone declares an incident.",
          "source_ids": [
            "nhsbt-2024-o-type-appeal",
            "aha-2024-change-survey"
          ]
        }
      ],
      "defenses_by_tier": {
        "0": [
          {
            "text": "Build and keep a dependency map that links each clinical service to the suppliers, networks, devices and utilities it relies on. Review it every year and after any major change.",
            "source_ids": [
              "dutch-safety-board-2020-it-outages",
              "nist-sp800-161r1"
            ]
          },
          {
            "text": "Write cybersecurity and restoration requirements into supplier contracts: multi-factor authentication on remote access, a restoration-time commitment, notice within 24 hours of contingency activation, and yearly written evidence of safeguards.",
            "source_ids": [
              "hhs-hipaa-security-nprm-2025",
              "witty-2024-senate-finance-testimony",
              "hscc-cms-2023-hospital-resiliency-landscape"
            ]
          },
          {
            "text": "Require staged rollout for any vendor update that runs with kernel or administrator rights on clinical endpoints. Get control over the rollout ring in writing, and keep a sample of clinical workstations on a delayed ring.",
            "source_ids": [
              "crowdstrike-rca-channel-file-291"
            ]
          },
          {
            "text": "Do not let your primary and backup depend on the same thing. Where a function is life-critical, make sure the backup uses a different supplier, network path or software stack.",
            "source_ids": [
              "perrow-1999-normal-accidents",
              "microsoft-2024-crowdstrike-8-5m-devices"
            ]
          },
          {
            "text": "Close known latent conditions on a deadline: unpatched internet-facing systems, remote-access portals without MFA, unsupported operating systems. Give someone the authority to enforce the deadline.",
            "source_ids": [
              "nao-2017-wannacry",
              "witty-2024-senate-finance-testimony"
            ]
          }
        ],
        "1": [
          {
            "text": "Enrol an alternate for each concentrated supplier before you need it (for example, a second clearinghouse EDI enrolment, or a reference-lab agreement), and test the switch once a year.",
            "source_ids": [
              "aha-2024-change-survey",
              "ecfr-42cfr482-15"
            ]
          },
          {
            "text": "Set pre-agreed disconnection criteria with key partners: who can cut the link, what runs while it is cut, and what evidence brings it back.",
            "source_ids": [
              "witty-2024-senate-finance-testimony",
              "nao-2017-wannacry"
            ]
          },
          {
            "text": "Keep a read-only copy of the EHR that can print, and test it regularly, so clinicians keep access to records while the primary system or its network is down.",
            "source_ids": [
              "onc-safer-contingency-2025"
            ]
          },
          {
            "text": "When your blood-matching or lab capacity drops, tell your blood supplier the same day, so that O-type stock can be managed nationally rather than drained locally.",
            "source_ids": [
              "nhsbt-2024-o-type-appeal"
            ]
          }
        ],
        "2": [
          {
            "text": "Keep enough paper downtime forms in every care area for at least 8 hours of ordering, medication administration, lab and radiology, and run unannounced downtime drills at least once a year.",
            "source_ids": [
              "onc-safer-contingency-2025"
            ]
          },
          {
            "text": "Tell neighbouring hospitals and EMS early when you go on diversion or downtime, and agree in advance how to share load. Regional capacity is a shared resource in a cascade.",
            "source_ids": [
              "dameff-2023-jama-netw-open-adjacent-eds",
              "ecfr-42cfr482-15"
            ]
          },
          {
            "text": "Rank services by clinical criticality and restore in that order. HHS has proposed a 72-hour restoration target for critical systems. Test whether you could meet it.",
            "source_ids": [
              "hhs-hipaa-security-nprm-2025"
            ]
          },
          {
            "text": "Prepare all staff, not only IT, to deliver care during an extended cyber downtime, and prioritise the services that must stay safe.",
            "source_ids": [
              "tjc-sea-67-cyberattack"
            ]
          }
        ],
        "3": [
          {
            "text": "Have signed arrangements with other hospitals to take your patients when your operations are limited or stopped, as the CMS emergency preparedness rule requires. Check that those hospitals do not depend on the same supplier, grid segment or water system as you.",
            "source_ids": [
              "ecfr-42cfr482-15"
            ]
          },
          {
            "text": "Plan for region-wide events where every neighbour is degraded at once. During the Texas freeze, one Austin hospital found no other hospital could take a large number of transfers.",
            "source_ids": [
              "circleofblue-2021-austin-hospitals-water",
              "ferc-nerc-2022-feb2021-outages-tracking"
            ]
          },
          {
            "text": "Plan for your community's electricity-dependent patients coming to the ED for oxygen and power when the grid fails. Keep space, outlets and oxygen supply for them.",
            "source_ids": [
              "abc-2021-texas-hospitals-water",
              "dshs-2021-winter-storm-deaths"
            ]
          },
          {
            "text": "After recovery, review the incident as a control problem (CAST or similar), not a hunt for a root cause. Ask which constraints, feedback loops and decision-makers failed to stop the spread, including at suppliers.",
            "source_ids": [
              "leveson-2019-cast-handbook",
              "leveson-2024-cast-healthcare-handbook"
            ]
          }
        ]
      },
      "standards": [
        {
          "name": "CMS Conditions of Participation, Emergency Preparedness, 42 CFR 482.15",
          "what_it_requires": "Hospitals must base their emergency plan on a facility-based and community-based all-hazards risk assessment, have arrangements with other hospitals to receive patients if operations are limited or stop, keep a communication plan with primary and alternate means, and run exercises at least twice a year.",
          "source_id": "ecfr-42cfr482-15"
        },
        {
          "name": "HIPAA Security Rule NPRM, 90 FR 898 (Jan 6, 2025), RIN 0945-AA22 (proposed, not final)",
          "what_it_requires": "Proposes written procedures to restore critical systems and data within 72 hours, yearly written verification of business associates' technical safeguards, and business-associate notice within 24 hours of activating a contingency plan. It cites the Change Healthcare attack.",
          "source_id": "hhs-hipaa-security-nprm-2025"
        },
        {
          "name": "NIST SP 800-161 Rev. 1 (May 2022, updated Nov 2024)",
          "what_it_requires": "Guidance for building cybersecurity supply chain risk management into strategy, policy and risk assessment for the products and services an organisation buys, across organisation, mission and system levels.",
          "source_id": "nist-sp800-161r1"
        },
        {
          "name": "ASTP/ONC SAFER Guide: Contingency Planning (2025)",
          "what_it_requires": "Self-assessment practices for EHR downtime, including paper forms for at least 8 hours, a tested read-only backup EHR, and unannounced downtime drills at least once a year. CMS requires hospitals to attest to the SAFER Guides annually.",
          "source_id": "onc-safer-contingency-2025"
        },
        {
          "name": "Joint Commission Sentinel Event Alert 67 (Aug 2023)",
          "what_it_requires": "Guidance, not a standard. It calls on organisations to prepare all staff to keep care safe through an extended cyberattack downtime.",
          "source_id": "tjc-sea-67-cyberattack"
        }
      ],
      "elsewhere": {
        "text": "EU: the NIS2 Directive (EU) 2022/2555 requires essential and important entities to manage supply-chain security, including the quality and resilience of suppliers' products and services and cybersecurity terms in contracts with direct suppliers (Art. 21). It also provides for coordinated EU risk assessments of critical supply chains (Art. 22). UK: after Synnovis, the Cyber Security and Resilience (Network and Information Systems) Bill would let regulators designate 'critical suppliers' to essential services, and the government's factsheet uses Synnovis as its case study. The bill was still before Parliament in mid-2026. Netherlands: the Dutch Safety Board found in 2020 that hospitals' awareness of IT-failure risk had not kept pace with their dependence on IT. It recommended that hospitals map IT-to-care dependencies, test and drill regularly, and analyse serious outages in depth.",
        "source_ids": [
          "eu-2022-nis2-directive",
          "govuk-2026-csr-bill-critical-suppliers",
          "dutch-safety-board-2020-it-outages"
        ]
      },
      "scoring_v01": {
        "likelihood": 4,
        "blast_radius": 5,
        "detectability": 4,
        "rationale": "Judgement, first draft. Likelihood 4: five large healthcare cascades in 2017–2024, three of them in 2024 alone, suggest the pattern recurs every year or two somewhere in the US or UK. Blast radius 5: cascades are by definition failures that spread beyond a single organisation; Change Healthcare and CrowdStrike reached national scale. Detectability 4: the triggering weakness is usually latent and often sits inside a supplier the hospital cannot see into, although once a cascade is running it is obvious."
      },
      "open_questions": [
        "How much patient harm do cascades cause beyond the first organisation? Only Dameff 2023 measured spillover at neighbours, and it covered one region and one event.",
        "Which healthcare functions are most concentrated in single suppliers nationally (clearinghouses, pathology, e-prescribing, identity, EHR hosting), and what share of hospitals share each one?",
        "Do contract clauses (restoration times, contingency notice, staged rollout) actually shorten outages, or only move liability?",
        "How long does downtime proficiency last after a drill, and how often must paper workflows be practised for a hospital to reach tier 3 safely?",
        "Is disconnecting early to contain an attack net-beneficial for patients once the downstream outage it causes is counted?"
      ],
      "watch_queries": [
        "healthcare third-party vendor outage hospitals disrupted",
        "clearinghouse OR pathology OR \"e-prescribing\" cyberattack hospitals multiple",
        "hospital ransomware ambulance diversion neighboring hospitals",
        "software update outage hospitals EHR downtime",
        "\"single point of failure\" health system cyberattack",
        "cascading failure hospital power water outage",
        "\"critical supplier\" NHS cyber attack",
        "(ransomware[tiab] OR cyberattack[tiab]) AND (hospital*[tiab]) AND (spillover OR adjacent OR regional)",
        "\"going solid\" OR \"tight coupling\" healthcare safety",
        "CAST OR STPA analysis healthcare information technology outage"
      ],
      "last_reviewed": "2026-09-26",
      "reviewed_by": "opus-research",
      "last_audited": "2026-09-26",
      "audit_status": "verified"
    }
  ]
}
