Business Continuity
Operational Resilience and Cybersecurity: Two Sides of the Same Coin in Critical Infrastructure
Operational resilience and cybersecurity are not separate programs: one without the other is incomplete. How to integrate ISO 22301 and IEC 62443 into a coherent approach for critical infrastructure.
Operational resilience: beyond service availability
Operational resilience is an organization's ability to keep delivering its critical functions in the face of disruptions: not just to survive the event, but to adapt, respond and recover in a controlled way.
For a long time, in industrial organizations, "operational resilience" meant mainly physical redundancy: backup systems, uninterruptible power supplies, failover procedures, stocks of critical components. The disruption people had in mind was the mechanical failure, the power blackout, the natural disaster.
The cyber threat has added a new category of disruption that is radically different from the previous ones. A mechanical failure is localized, predictable in its patterns, recoverable with known procedures. A cyber attack can be distributed, designed to maximize impact, and can hit redundant systems simultaneously if they share the same digital infrastructure.
An organization with excellent physical resilience but no cyber resilience is only half resilient.
ISO 22301 and IEC 62443: how they integrate
ISO 22301 is the international standard for business continuity management systems (BCMS). It defines how an organization identifies its critical functions, analyzes the risks that threaten them, and builds plans to maintain or restore them in the event of an interruption.
IEC 62443 is the standard for the cybersecurity of OT environments. It defines architectures, technical requirements and processes to protect industrial control systems.
These two standards appear to operate in separate domains (one focused on business continuity, the other on technical security), but they have more points of contact than is often recognized:
Identifying critical assets: both require identifying which systems are critical to the organization's essential functions. An ISO 22301 Business Impact Analysis (BIA) should include OT systems; an IEC 62443 risk assessment should take into account the operational importance of the systems.
Scenario analysis: ISO 22301 requires identifying the most likely and most impactful interruption scenarios. A cyber attack on OT environments must be an explicit scenario in the BIA, not an eventuality handled only by the IT team.
RTO and RPO for OT systems: the Recovery Time Objectives and Recovery Point Objectives (how long you can afford to be down and how much data you can afford to lose) must also be defined for industrial control systems. This has direct implications for OT configuration backup requirements and recovery plans.
Integrated testing: business continuity exercises should include cyber scenarios. A disaster recovery exercise that simulates a physical failure but does not consider the possibility that the backup system is also compromised does not test real resilience.
The resilience lifecycle: prevention, detection, response, recovery
Operational resilience in the face of cyber threats follows a four-phase lifecycle, each with specific requirements for OT environments:
Prevention: reducing the likelihood that an incident occurs. In OT, this means the IEC 62443 technical controls: network segmentation, access management, patch management, remote access security. But perfect prevention does not exist; it is a reduction of risk, not its elimination.
Detection: identifying an incident in the shortest possible time. The Time To Detect is critical: every hour of an attacker's undetected presence in the OT network is an hour in which they can explore, gather information, and position themselves for action. Continuous monitoring of the OT network is the fundamental technical requirement of this phase.
Response: containing the incident while limiting the damage. As discussed in the article on OT incident response, response in industrial environments requires a balance between containment effectiveness and operational continuity. Response plans must be pre-defined, not improvised.
Recovery: bringing systems back to normal operation in a controlled and verified way. In the OT context, this includes: restoring PLC configurations from verified backups, functional testing before restarting in production, and verification that the attack vector has been eliminated before recovery.
Worst-case scenarios: exercises and continuity plans for OT
The value of a continuity plan is measured by how much it has been tested before it is needed in a real situation. Untested plans contain assumptions that turn out to be wrong at the worst possible moment.
For OT environments, continuity exercises have specific characteristics:
Realistic and specific scenarios: "a ransomware attack has encrypted the SCADA workstations of production plant 2" is more useful than "a cyber incident has hit the OT". Specificity forces you to identify the real dependencies and the concrete procedures.
Involvement of operations: OT continuity exercises cannot be run by the IT or security team alone. The production manager, the shift supervisor, the process engineer must be part of the exercise, because they are the ones who make operational decisions during a real incident.
Testing recovery from configuration backups: how many organizations have verified backups of their PLC configurations? How many have tested the recovery process on a test system? The answer is rarely reassuring. The continuity exercise is the moment to find out in a controlled context.
Degraded operations scenarios: in some cases, full recovery takes time. Continuity plans must include procedures for operating in a degraded mode in the meantime: manual instead of automatic, increased supervision, reduced production rates.
The final convergence: security and operations as a single function
The most important trend in the evolution of risk management in critical infrastructure is the convergence of cybersecurity and operational resilience as an integrated function, not as two parallel programs that coordinate occasionally.
This convergence shows up at the organizational level: OT security teams reporting to functions that include operational continuity; KPIs shared between security and operations; budgets that fund resilience in an integrated way instead of artificially splitting between "IT security" and "operational continuity".
It also shows up at the technological level: OT monitoring platforms that integrate anomaly detection and availability alerts; incident response plans that incorporate both the technical cyber response and the operational failover procedures; exercises that simulate hybrid scenarios where the cyber incident manifests as an operational interruption.
The cyber threat has made the conceptual separation between protecting systems and ensuring service continuity obsolete. In critical infrastructure, they are the same thing.
The MON5 angle
In the prevention, detection, response and recovery cycle described in this article, detection is the phase where many industrial organizations are most exposed: without OT network monitoring, the Time To Detect is measured in months. MON5 covers this phase with continuous monitoring and ML anomaly detection on native OT protocols, passively and without any impact on production.
Upstream, the DISCOVER phase provides the asset inventory and network topology that a serious Business Impact Analysis requires: you cannot define RTOs for systems you do not know about. To integrate resilience and security starting from the facts, the first step is an OT assessment.
Related articles
Cybersecurity
Incident response in OT environments: the critical differences from IT
IT incident response playbooks do not work in OT. Isolating a compromised system can halt production; powering off a device can cause physical damage. How to build OT IR that actually works.

Cybersecurity
Anomaly detection in OT: building the baseline and managing false positives
OT networks are repetitive and predictable, in theory the ideal environment for anomaly detection. In practice, legitimate-but-anomalous behavior generates a false-positive noise that is the main cause of failure for industrial monitoring projects.

Cybersecurity
CVE and EPSS in OT Environments: Which Vulnerabilities to Fix When You Can't Patch Everything
Patching everything in an OT environment is impossible. CVSS alone is not enough to set priorities. EPSS adds the missing dimension: the probability that a vulnerability is being actively exploited today.
Do you have visibility into your OT network?
MON5 maps assets, vulnerabilities and anomalies in real time — without stopping production.