When Playbook Meets Adversary
When the Playbook Meets the Adversary: Red Team Testing of OT Incident Response
CISA red team operators ran for three months inside a mature critical infrastructure environment and triggered no defensive response. Red Team IR testing exposes the gap between a documented plan and a live capability.
When the Playbook Meets the Adversary
In AA23-059A, CISA documented a red team engagement that established persistence through spearphishing, moved laterally across multiple sites, reached systems near sensitive operations, and then attempted to trigger defensive response. The organization detected none of it. CISA identified 13 events that should have generated alerts and did not.
That is what an IR program looks like when adversary pressure is real. The failure is rarely the existence of a plan. It is the capability gap between what the plan states and what teams can execute under time pressure across the IT-OT boundary.
The Plan Most Organizations Have, and the Capability Most Have Not Validated
SANS 2025 reported 57% of organizations have dedicated OT IR plans, rising to 70% in regulated organizations with threat intelligence. Most testing is still paper-based. SANS found 68% rely on tabletop exercises and 39% test annually. Adversarial red or purple exercises are concentrated in the fully prepared cohort.
Sygnia's 2026 CISO Survey found 73% of organizations would not be fully ready for a serious attack tomorrow. IBM's 2025 breach report tied tested IR capability to average savings of $2.66M per breach, directly linking validation to financial outcome.
What Red Team IR Testing Actually Looks Like
OT Red Team IR testing is controlled adversarial presence under production safety constraints. Engagements simulate realistic tradecraft aligned with MITRE ATT&CK for ICS, moving from reconnaissance to IT-adjacent exploitation toward simulated OT consequence.
The objective is not disruption. It is measurement. Key outputs include detection rate, time-to-detection, time-to-response, time-to-containment, escalation accuracy, decision authority clarity, and IT-OT coordination quality at handoff.
The Five Gaps That Show Up Almost Every Time
Across CISA findings, field assessments, and practitioner reporting, five recurring failure patterns surface: detection that never triggers; alerts that create tickets but no action; decision authority confusion under pressure; IT-OT coordination breakdown during containment; and IT-centric runbook misapplication in OT environments.
These gaps are practical failures, not abstract maturity issues. They are what separate a compliance-grade plan from an adversary-grade response capability.
What Tabletops Validate, and What They Cannot
Tabletops validate role clarity and decision logic on paper. They do not validate performance under adversarial stress. Realistic simulation and Red Team pressure validate whether tools alert, escalation paths activate, and teams execute when fatigue and uncertainty are present.
SANS reported organizations practicing realistic OT IR are 1.7x more likely to report strong preparedness. The same distinction applies across the program: controls that exist are not controls that stop adversaries.
What 42 Days of Undetected Access Buys the Adversary
Dragos reported average OT ransomware dwell time at 42 days in 2026 reporting. Mandiant reported a 14-day global median and a 9-day ransomware median. OT remains materially slower to detect and disrupt.
Extended dwell gives attackers time to map control behavior, identify command pathways, and stage operational consequence. Colonial Pipeline remains a board-level example of uncertainty cost when containment and recovery confidence are insufficient.
From Response Validation to Recovery Validation
This is the validation layer between governance and operations. It tests whether funded plans and fusion-center workflows actually perform under adversary conditions. The next layer is recovery validation, where consequence has already occurred and the question is how quickly controlled operations can be restored.
Sources: CISA AA23-059A, SANS 2025 State of ICS/OT Cybersecurity Survey, Sygnia 2026 CISO Survey, IBM Cost of a Data Breach Report 2025, Dragos 2026 OT/ICS Cybersecurity Year in Review, Mandiant M-Trends 2026, NIST SP 800-82 Rev. 3, IEC 62443-2-1, MITRE ATT&CK for ICS, and cited industry reporting.
Plans on paper do not stop adversaries.
Validated execution does.
Red Team proves the difference.
Visibility gaps turn valid attack progression into low-confidence noise, delaying response until impact grows.
Programs fail operationally when triage systems create records but do not trigger timely containment decisions.
Without pre-authorized decisions, teams debate escalation while adversary dwell time compounds risk.
Operationally safe containment often depends on conduit and remote-access control, not endpoint shutdown reflexes.
OT response needs site-specific procedures at the cyber-safety interface where generic enterprise playbooks are insufficient.
Tabletops validate understanding. Red Team validates execution under uncertainty, stress, and adversarial adaptation.
Longer dwell enables mapping, staging, and persistence that increase both response complexity and recovery uncertainty.
Mature programs validate in sequence: authority, operating model, adversary pressure, and post-impact recovery readiness.
Plan quality is necessary but incomplete without measured execution under realistic pressure.
Adversary validation measures what teams detect, decide, and contain when uncertainty is high.
Recovery readiness determines whether the organization returns to controlled operations in hours or weeks.
IR plans are common. Adversary-validated IR capability is not.
The five recurring gaps are operational and measurable, not theoretical.
Red Team testing converts documentation confidence into real incident performance data.
