Different Threats. The Same Operational Consequence.

Several generating units disappear from the control-room display. Alarms arrive, protection systems operate, and the wider system must compensate for lost capacity. The operators may not yet know whether the cause was equipment failure, physical interference, a communications problem, an incorrect engineering change, or malicious activity inside the OT environment. They do know that the process has changed and must be stabilized before anyone attempts restoration.

This distinction matters because critical-infrastructure risk is often organized around the threat. Cybersecurity manages cyberattacks. Physical security manages intrusions and sabotage. Engineering manages equipment failures. Emergency management prepares for wider disruption.

The control room does not experience these risks as separate categories. It experiences their consequences.

A real event, and a wider lesson

On September 1, German authorities began investigating two separate suspected sabotage incidents involving electricity infrastructure.

In Brandenburg, projectiles carrying conductive material reportedly caused short circuits at an extra-high-voltage substation and triggered the automatic shutdown of a generating unit. Near Cologne, conductive cables were placed across overhead lines and five coal-fired units totaling approximately 4,200 MW went offline.

The general electricity supply was not interrupted. The wider grid absorbed the loss while the affected facilities dealt with the immediate operational consequences.

The important OT lesson is not about who may have been responsible. It is that a physical action created a condition that could also be produced through an OT cyberattack.

An attacker with sufficient access to a control or protection environment could create a similar outcome through an unauthorized breaker command, manipulated setting, disrupted communication, or loss of trustworthy visibility. The mechanism would differ, but the control room could face the same immediate reality: equipment has tripped, information is incomplete, and restoration must wait until the condition is understood.

Protection responds to the condition, not the motive

A protection system is designed to recognize an unsafe electrical condition and act quickly enough to protect equipment and the wider system. It does not determine whether that condition came from equipment failure, accidental contact, deliberate physical interference, or malicious manipulation.

That is not a limitation. It is the protection system doing its job.

The control room follows a similar principle during the first moments of an incident. Operators stabilize the process based on what the equipment is doing, not on an early theory about who caused it.

This is why the first response to many OT events is remarkably similar. Protect people, stabilize the process, determine what has been lost, confirm whether control information remains trustworthy, contain the problem, preserve evidence, and restore only when engineering and operations agree that it is safe.

The investigation will eventually branch. A physical incident may require site searches, equipment inspection, video review, and law-enforcement coordination. A cyber incident may require analysis of remote sessions, network traffic, credentials, engineering changes, controller logic, and system logs.

Those investigative paths are different. The responsibility to keep the operation safe and recover it responsibly is not.

The same consequence should bring the same teams together

Imagine that several units trip unexpectedly. Operators stabilize the process. Protection engineers examine relay operations. Field personnel inspect equipment. Physical security checks the site, while cybersecurity reviews remote access, recent changes, identities, and network communications. Leadership assesses the effect on service, regulators, and customers.

No single team has the complete picture. If each function follows its own procedure without a common operational lead, evidence may be lost and decisions may conflict. A cyber analyst may see no malicious traffic and dismiss a security dimension. Engineering may restore equipment before establishing whether an unauthorized change remains present.

The answer is not to make cybersecurity responsible for every disruption. It is to define how operations, engineering, cybersecurity, physical security, emergency management, and leadership work together when the cause is uncertain but the consequence is real.

That operating model must exist before the incident.

Start with the loss of function

Traditional risk discussions often begin with ransomware, sabotage, insider activity, equipment failure, flooding, or communications loss. For OT resilience, I prefer to begin closer to the operation.

What happens if this critical function becomes unavailable? For a generating unit, that question exposes dependencies on the substation, station service power, protection and control systems, fuel, cooling, telecommunications, operators, engineering support, physical access, vendors, and known-good configurations.

Some of those dependencies are digital. Others are physical or organizational. The process depends on all of them at the same time.

This changes the assessment. Instead of producing separate lists of risks, the organization can see where different events converge on the same weakness. A substation may be vulnerable to physical damage, protection failure, communications loss, or malicious control activity. A remote facility may depend on one telecommunications path regardless of how it is lost.

The common consequence helps leadership identify where redundancy, protection, monitoring, response coordination, and recovery investment will provide value across several scenarios.

A wider outage is not the only measure of consequence

The German grid reportedly continued supplying electricity despite the loss of significant generation. That is resilience at the wider system level, but it does not mean the affected facilities experienced no consequence. Units were unavailable, protection operated, and equipment required inspection and controlled restoration.

The system may absorb the loss of a facility while that facility still experiences serious operational and commercial impact. No customer outage does not mean no operational consequence.

Boards should therefore look at resilience on both levels. Can the enterprise or grid tolerate the loss of a site? Can the site itself reach a safe state, maintain essential visibility, preserve evidence, and recover trusted operation?

Both questions matter.

Exercise the operational reality

A useful exercise does not need a complicated attacker story.

Begin with a critical process area becoming unavailable, but do not immediately explain why. Let protection show a valid trip while the SOC reports no obvious malicious traffic. Delay field inspection, introduce a recent vendor session, and add pressure for a restoration estimate.

Then observe how the organization works.

Does the control room know whom to call? Can the teams preserve the right evidence, determine which information is trustworthy, control vendor access, and agree who can authorize restart? Can leadership explain what is known, what remains uncertain, and what the business consequence may be?

The objective is not to guess the cause correctly in the first few minutes. It is to make good operational decisions while the cause remains uncertain.

The boardroom and control room need one definition of resilience

The boardroom may organize cyber risk, physical risk, operational risk, and business continuity into separate governance structures and budgets. The control room experiences them together.

The disciplines are not the same. Each requires specialized knowledge, but they must converge around the safe and reliable operation of the process.

OT cybersecurity must protect the digital systems supporting the operation, determine whether they can be trusted, preserve cyber evidence, and help prevent unauthorized actions from creating physical consequences.

But the program becomes truly operational when it can work with protection engineering, operations, physical security, maintenance, emergency management, vendors, and executive leadership during the same event.

The initiating cause still matters. It will shape the investigation, corrective action, accountability, and prevention of recurrence.

The consequence determines what the operation must do first.

Different threats can bring a facility to the same condition. In many cases, the first responsibilities remain the same: protect people, stabilize the process, understand what has been lost, establish trust, coordinate decisions, and restore safely.

That is the reality OT resilience must be designed around.

Next
Next

OT Cybersecurity Is Simpler Than You Are Being Told. Here Is the Trick.