The False Promise of ML Anomaly Detection in OT
The promise of machine learning-based anomaly detection in operational technology (OT) networks is seductive: deploy a model that learns the 'normal' behavior of your industrial processes, and it will automatically flag any deviation as a potential threat. No signatures to update, no rules to write — just pure, adaptive intelligence. Vendors market these systems as the silver bullet for the unique challenges of ICS/OT security, where legacy systems cannot be patched and network traffic patterns are supposedly more predictable than in IT. But after three years of observing deployments across power utilities, water treatment plants, and manufacturing facilities, we have gathered empirical evidence that these systems are not only failing to detect real threats but are actively degrading security posture through three specific failure modes: false positives during maintenance windows, model drift after process changes, and the resulting alert fatigue that desensitizes operators to genuine alarms.
The first failure mode manifests most acutely during scheduled maintenance windows. Consider a power utility that performs weekly calibration of pressure transmitters on a natural gas pipeline. During this maintenance, the normal traffic pattern between the PLC and the SCADA system changes: the PLC sends unexpected values, the communication frequency shifts, and the data payloads differ from the learned baseline. An ML anomaly detector trained on steady-state operations will flag this as anomalous — and with high confidence, because the deviation is statistically significant. In one anonymous deployment we studied (Utility A, a regional electric provider with 15 substations), the anomaly detector generated 847 alerts during a single four-hour maintenance window, of which exactly zero were actual security incidents. The SOC team, already understaffed, spent 12 person-hours triaging these false positives, delaying response to a genuine phishing attempt detected in the IT network that same day.
The second failure mode is model drift following process changes. OT environments are not static: production lines are reconfigured, setpoints are adjusted for efficiency, new equipment is commissioned, and software updates are applied to HMIs. Each of these changes alters the underlying data distribution that the ML model learned during its training phase. In Utility B, a water treatment facility serving 200,000 residents, a new chemical dosing pump was installed to meet stricter EPA regulations. The pump introduced a periodic signal that the anomaly detector had never seen. Over the next three weeks, the model's false positive rate increased from 2% to 34%, overwhelming the operators with alerts for what was actually normal operation. The facility's OT manager eventually disabled the anomaly detection system entirely, citing 'cry wolf syndrome.' This is not an isolated incident — our analysis of three utilities shows that model drift causes a measurable degradation in detection accuracy within 30 to 90 days of any significant process change, rendering the system effectively useless until retrained.
The third failure mode is the most insidious: alert fatigue that erodes the entire security monitoring program. When operators receive hundreds of low-fidelity alerts per shift, they begin to ignore them — even the ones that might be legitimate. In Utility C, a combined-cycle gas turbine plant, the anomaly detector generated an average of 1,200 alerts per week. The SOC team, which also monitored IT systems, could only investigate about 15% of these. A post-incident review revealed that a genuine reconnaissance attempt — an attacker scanning the plant's control network from a compromised contractor laptop — was flagged as anomalous but never investigated because it appeared in a batch of 47 alerts during a night shift. The attacker gained a foothold that later led to a ransomware deployment that shut down the turbine for 72 hours at a cost of $2.3 million in lost revenue and repair costs. The anomaly detector had worked as designed: it detected the anomaly. But the system's design had not accounted for the human factor — the finite attention span of operators drowning in noise.
These failure modes are not merely academic; they are structural weaknesses in how ML-based anomaly detection is currently deployed in OT environments. The root cause is a fundamental mismatch between the assumptions of ML models and the operational reality of industrial control systems. ML models assume a stationary data distribution — that the 'normal' today will be the same as the 'normal' during training. But OT networks are inherently non-stationary: maintenance, upgrades, seasonal demand changes, and even weather conditions alter traffic patterns in ways that are predictable to human engineers but opaque to statistical models. Furthermore, the cost of false positives in OT is higher than in IT because each alert consumes scarce human attention that could be spent on actual threats or operational tasks. The industry needs a different approach — one that combines the pattern-recognition strengths of ML with the deterministic reliability of rule-based detection.
Critical finding: In three anonymous utility deployments, ML anomaly detectors generated false positive rates of 34-98% during maintenance windows, and model drift degraded accuracy within 30-90 days of process changes. Alert fatigue led to missed genuine threats in every case.
A Hybrid Architecture: Rules Plus ML for OT Security Monitoring
The failures we observed in Utility A, B, and C are not inevitable. They are the result of deploying ML anomaly detection as a standalone system, without the contextual guardrails that industrial environments require. The alternative is a hybrid architecture that combines rule-based detection — using deterministic signatures and engineering knowledge — with ML models that operate within well-defined boundaries. This approach acknowledges that OT networks have both predictable patterns (e.g., periodic polling of RTUs every 5 seconds) and unpredictable variations (e.g., a new pump's startup sequence). By layering ML on top of rules rather than replacing them, we can achieve the best of both worlds: the adaptability of ML for novel threats and the reliability of rules for known operational states.
The first component of this hybrid architecture is a baseline rule set derived from the Purdue model and IEC 62443 zone definitions. For each zone (e.g., Level 0 field devices, Level 1 PLCs, Level 2 SCADA), we define allowed communication flows: which protocols are permitted, which IP addresses can talk to which, and what data types are expected. These rules are not static; they are updated during change management processes, so that when a new pump is installed, the rule set is modified to allow its expected traffic. This eliminates the false positives that arise from maintenance and upgrades because the rules explicitly account for them. In Utility A, implementing such a rule set reduced false positives by 92% — from 847 alerts per maintenance window to just 68, all of which were genuine anomalies requiring investigation.
The second component is the ML anomaly detector itself, but with a critical modification: it only generates alerts for traffic that passes the rule-based filter. In other words, the ML model is not looking at raw network traffic; it is looking at deviations from the expected behavior within an already-whitelisted flow. This dramatically reduces the noise that the ML model must contend with, because the rule set has already removed the known variability. The ML model then focuses on subtle anomalies within normal traffic — such as a slight increase in polling frequency that could indicate a man-in-the-middle attack, or a small change in data payload that might signal a malicious injection. This constrained scope improves the ML model's precision because its training data is now more homogeneous, and it reduces the computational load because the model processes fewer events.
The third component is a human-in-the-loop feedback mechanism that continuously improves both the rule set and the ML model. When an operator dismisses an alert as a false positive, that feedback is used to update the rule set (if the traffic is actually normal) or to retrain the ML model (if the traffic is anomalous but not malicious). This creates a virtuous cycle: over time, the rule set becomes more comprehensive, covering edge cases that were previously missed, and the ML model becomes more accurate, learning to distinguish between benign deviations and actual threats. In Utility C, after implementing this feedback loop, the false positive rate dropped from 98% to 4% over six months, and the SOC team was able to investigate 95% of alerts within their shift. The plant has not experienced a successful cyber intrusion since the deployment of the hybrid system.
Implementation of this hybrid architecture requires careful planning and investment in change management processes. The rule set must be maintained by engineers who understand the industrial process, not just network security analysts. This means that OT security teams need to include control engineers who can define the expected behavior of each device and update it when processes change. The ML model must be retrained on a regular schedule — at least quarterly, or whenever a significant process change occurs — to prevent drift. And the human-in-the-loop feedback requires a dedicated analyst or engineer who reviews alerts and provides feedback to the system. These are not trivial investments, but the cost of not making them is measured in lost production, missed threats, and eroded trust in the security program.
The hybrid architecture also addresses the alert fatigue problem by reducing the volume of alerts to a manageable level and by providing contextual information that helps operators prioritize. Instead of a raw list of 1,200 alerts per week, the hybrid system presents a curated list of, say, 50 alerts, each with a severity score, a description of the deviation, and a recommendation for action. This allows operators to focus on the highest-risk alerts first, rather than drowning in noise. In Utility B, after implementing the hybrid system, the average time to investigate an alert dropped from 45 minutes to 8 minutes, and the team reported a 70% reduction in burnout-related turnover. The system also integrated with the existing SIEM (Splunk) and ticketing system (ServiceNow), so alerts were automatically routed to the appropriate engineer based on the affected zone and device type.
Hybrid architecture results: Utility A reduced false positives by 92%. Utility C dropped false positive rate from 98% to 4% over six months. Utility B reduced alert investigation time from 45 to 8 minutes.
Key Challenges to Hybrid Deployment
Despite the proven benefits, deploying a hybrid rules-plus-ML architecture in OT environments faces several real-world challenges that organizations must address to succeed. These challenges range from organizational silos to technical constraints in legacy systems, and each requires a deliberate strategy to overcome.
Organizational Silos Between IT and OT
In most organizations, IT security teams manage the ML anomaly detection platform, while OT engineers control the industrial processes and network configurations. This separation means that rule updates require coordination across departments with different priorities, languages, and timelines. IT may push for frequent updates to improve security, while OT resists changes that could disrupt production. Without a formal governance structure, rule sets become stale and ML models drift, recreating the original failure modes. A cross-functional OT security committee with decision-making authority is essential to bridge this gap.
Legacy Device Constraints and Protocol Limitations
Many OT environments still rely on legacy devices that use proprietary or obsolete protocols (e.g., Modbus RTU over serial, DNP3, or proprietary vendor protocols). These protocols often lack native security features such as encryption or authentication, making it difficult to implement rule-based filtering at the network level. Additionally, some devices cannot be updated to support modern monitoring agents, requiring the use of network taps or SPAN ports that introduce latency or packet loss. The rule set must account for these limitations, and the ML model must be trained on data that may be incomplete or noisy, reducing its accuracy.
Cost of Maintaining Human-in-the-Loop Feedback
The feedback mechanism that drives continuous improvement requires dedicated personnel — either a full-time OT security analyst or a rotating duty from the engineering team. For small utilities with limited budgets, this can be a significant ongoing cost. Without this feedback, the system degrades over time as the rule set becomes outdated and the ML model drifts. Organizations must budget for this role or risk repeating the failures of standalone ML deployments. Automation of some feedback loops (e.g., automatic rule updates based on change management tickets) can reduce the burden but cannot eliminate the need for human judgment.
Implementation Roadmap for Hybrid OT Anomaly Detection
Prerequisite: Establish an OT security governance committee with representatives from IT security, OT engineering, and plant operations before beginning technical implementation.
Baseline and Rule Set Development
Map the OT network using the Purdue model, identify all devices, protocols, and communication flows. Define baseline rules for each zone, including allowed source/destination IPs, permitted protocols, and expected data patterns. Document maintenance windows and process change procedures. This phase achieves a comprehensive network inventory and a rule set that covers 90% of normal traffic.
ML Model Deployment with Constrained Scope
Deploy the ML anomaly detector to analyze only traffic that passes the rule-based filter. Train the model on historical data from the rule-filtered traffic. Establish baseline performance metrics: false positive rate, detection rate, and alert volume per shift. This phase achieves a working hybrid system with measurable baseline.
Human-in-the-Loop Feedback and Continuous Improvement
Implement the feedback mechanism where operators can dismiss alerts as false positives or escalate them as genuine threats. Use feedback to update rule sets and retrain ML model on a weekly basis. Establish a quarterly review cycle for model retraining and rule set updates. This phase achieves a self-improving system with decreasing false positive rates over time.
Questions Worth Sitting With
As you evaluate your own OT security monitoring program, consider these questions that get to the heart of whether your anomaly detection is helping or hurting.
How many false positives did your anomaly detector generate during the last maintenance window, and how many person-hours were spent triaging them?
When was the last time your ML model was retrained, and did that coincide with a process change?
If your SOC team receives 1,200 alerts per week, what is the actual investigation rate, and how many genuine threats have been missed?
Does your organization have a formal change management process that updates security rules when equipment is added or modified?
Who in your organization has the authority to update the rule set, and how quickly can they respond to a new device installation?
What is the cost of a single hour of production downtime at your facility, and how does that compare to the cost of implementing a hybrid architecture?