OT Security Metrics That Matter: Measuring What Counts
In OT cybersecurity, we often measure what is easy, not what matters.
We count assets. We count vulnerabilities. We count firewall rules. We count training completion. We count patching rates. We count the number of findings closed before the audit. All of that may be useful, but it does not always answer the question leadership really needs to ask.
Can this plant detect, respond to, contain, and recover from a cyber event without putting people, production, or trust at risk?
That is the real test. I have seen many environments where the dashboard looked active, the report looked professional, and the numbers looked better than the previous quarter. But behind those numbers, the operation still had weak remote access, poor segmentation, untested recovery procedures, operators unsure how to escalate cyber concerns, and vendors with access that nobody reviewed.
That is why OT security metrics need to change.
Metrics should not exist only to satisfy auditors. They should not exist only because a tool can export them. They should not exist only to make an executive dashboard look complete.
The right metrics should help leaders see operational truth.
And operational truth in OT is not about how many things we counted. It is about whether the organization is becoming more resilient.
The Comfort of Easy Metrics
Easy metrics are attractive because they are simple to collect.
It is easy to report how many assets were discovered, how many vulnerabilities were found, how many users completed training, how many firewall rules exist, and how many tickets were closed.
These numbers create movement. They show activity. They give leadership something to review. But activity is not the same as effectiveness.
A high training completion rate does not mean operators know what to do when an HMI behaves strangely. A large asset inventory does not mean the organization understands which assets are critical to safety and production. A reduction in vulnerabilities does not automatically mean the most dangerous exposures were addressed. A closed audit finding does not mean the control will hold under pressure.
This is where easy metrics can become misleading.
They are not wrong by themselves. They become dangerous when leaders mistake them for maturity.
Compliance Metrics Are Necessary, But Not Enough
Compliance metrics matter. You need evidence. You need traceability. You need to know whether policies exist, whether access reviews happened, whether training was completed, and whether required controls are in place. But compliance metrics usually tell you whether something was done. They do not always tell you whether it works.
That difference is critical in OT.
- A metric may show that remote access is documented, but it may not show whether every session is approved, monitored, and controlled by the asset owner.
- A metric may show that incident response was reviewed, but it may not show whether the operator, engineer, safety lead, vendor, IT team, and cybersecurity team know how to act together during a live operational event.
- A metric may show that segmentation exists, but it may not show whether that segmentation supports safe containment without breaking required process communication.
This is why OT leaders should not stop at compliance measurement.
Compliance tells part of the story. Resilience reveals whether the organization can withstand pressure.
Metrics Must Speak the Language of the Plant
One reason OT metrics fail is that they are often written in cybersecurity language for cybersecurity people. That may work inside the security team, but it does not always work with plant leadership.
A plant manager may not care about the number of medium vulnerabilities closed last month unless that number is connected to process risk. An operations leader may not be moved by a firewall rule cleanup report unless it explains how exposure was reduced without disrupting production. A board may not understand detection coverage unless it is tied to downtime, safety, recovery, and business impact.
In OT, security metrics must be translated into operational language.
Instead of only reporting vulnerabilities, leadership should understand which vulnerabilities affect systems supporting critical process functions. Instead of only reporting patching percentage, leadership should understand which systems cannot be patched, why they cannot be patched, what compensating controls exist, and what risk remains. Instead of only reporting incident count, leadership should understand whether the organization is improving its ability to detect, escalate, contain, and recover safely.
That is when metrics become useful. They stop being cybersecurity numbers and become operational intelligence.
The Most Important Metric Is Often Readiness
Many OT programs measure systems better than they measure people.
That is a mistake.
In OT, people are not outside the security program. They are part of the defense model.
The operator who notices abnormal behavior is part of the detection process. The engineer who understands system dependencies is part of containment. The maintenance technician who recognizes an unsafe workaround is part of prevention. The vendor who follows access procedures is part of control. The site leader who supports escalation without blame is part of response maturity.
So if people are part of the defense, readiness must be measured too.
Do operators know how to report a suspected cyber event? Do engineers know what to do if an engineering workstation cannot be trusted? Have site leaders participated in a tabletop exercise that includes cyber, safety, production, and recovery? Are vendors included in response expectations? Does the shift team understand when a cyber issue could become an operational issue?
These questions may not fit neatly into a dashboard, but they matter. During a real incident, the first few decisions may not come from the SOC. They may come from the plant.
Detection Metrics Must Be Connected to Action
Many organizations measure their ability to detect.
Fewer measure whether they can act.
That is a major gap.
In OT, detection without response readiness can create noise. Alerts may be generated, but nobody knows who owns the next decision. The SOC may escalate, but operations may not understand the process impact. The plant may exhibit abnormal behavior, but cybersecurity may not know how to interpret its operational implications. A useful detection metric should not only ask whether an alert was generated.
It should ask whether the right people understood the alert quickly enough to make a safe decision. That is a different level of maturity.
Mean time to detect matters. But in OT, the mean time to understand may matter even more.
How long does it take to confirm whether the affected asset supports a critical function? How quickly can the team determine whether containment is safe? How quickly can operations, engineering, safety, and cybersecurity align on the next action? Those are the metrics that show whether detection is actually useful.
Response Metrics Must Respect OT Reality
In IT, response metrics often focus on speed.
How fast did we contain? How fast did we close? How fast did we restore?
In OT, speed matters, but only when it is safe.
- A response that quickly isolates a system but results in a loss of view is not successful.
- A response that blocks traffic but interrupts process communication is not successful.
- A response that closes the incident before validating the plant state is not successful.
OT response metrics should measure controlled action, not only fast action.
The better question is not only how fast the team responded. The better question is whether the response reduced risk without creating a bigger operational consequence.
That means measuring whether response playbooks were followed, whether operations approved containment actions, whether safety was consulted when needed, whether vendor actions were governed, and whether recovery was validated before declaring success.
This is the kind of measurement OT leaders need. Not metrics that reward speed at the expense of discipline.
Recovery Metrics Must Go Beyond System Availability
Recovery is another area where traditional metrics can mislead leadership.
In IT, a system may be considered recovered when it is back online. In OT, that is not enough.
A system can be online but not trusted. A controller can be running but still require logic validation. An HMI can display data, but it still needs confirmation that the values are accurate. A workstation can be rebuilt, but still requires engineering review.
A process can restart, but only after safety, operations, and vendor checks are complete.
Recovery in OT is not only technical restoration. It is the restoration of operational confidence.
So the metric should not be limited to the mean time to restore service. It is time to restore the system to a validated, trusted, and operationally safe state.
That distinction matters.
Bringing something back quickly is not the same as bringing it back correctly.
Metrics Should Drive Decisions, Not Decoration
A metric that does not drive action is just decoration.
It may look good on a dashboard. It may help fill a monthly report. It may create the appearance of governance.
But if it does not help leadership make better decisions, prioritize risk, invest wisely, improve behavior, reduce exposure, or strengthen recovery, then it is not doing its job.
If remote access sessions are increasing, leadership should ask whether approval, monitoring, and vendor control are keeping up. If critical assets are still not classified, leadership should ask how risk is being prioritized. If tabletop exercises show confusion between operations and cybersecurity, leadership should invest in readiness, not only tools. If vulnerabilities cannot be patched, leadership should ask whether compensating controls are actually validated. If incident response times look good but recovery validation is weak, leadership should challenge the definition of success.
This is how metrics become part of management.
Not reporting for the sake of reporting.
Decision support.
The OT CISO View
OT security metrics are broken when they measure activity but not readiness.
- They are broken when they satisfy the audit but do not help the plant survive.
- They are broken when they count findings but do not explain the consequences.
- They are broken when they focus on tools but ignore people.
- They are broken when they reward speed but ignore safety.
And they are broken when leadership receives a dashboard that looks mature while the operation remains exposed.
- The purpose of OT metrics is not to make the cybersecurity program look good.
- The purpose is to improve the organization.
Better at understanding risk. Better at prioritizing investment. Better at detecting what matters. Better at responding safely. Better at recovering with discipline. Better at protecting people, production, and trust.
That is the leadership mindset.
Final Thought
In OT, the best metric is not always the easiest to collect.
It is the one that tells the truth. Can we see what matters? Can we understand the consequence? Can we act without making the plant unsafe? Can we recover with confidence? Can leadership make better decisions because of what we measure?
If the answer is yes, the metric has value.
If the answer is no, it is probably noise.
And OT has enough noise already.