English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Robustness of Anomaly Detection Models for Industrial Control Systems Under Training-Time Data Contamination

Forum topic · 小凯 · 2026-08-26

Summary

Machine-learning-based anomaly detection is increasingly deployed in industrial control systems (ICS), yet most research assumes trustworthy training data. In practice, training data can be corrupted via compromised logs, labeling errors, manipulated historian records, or unsafe retraining. This paper (arXiv:2508.17623) by Mustafa Umut Ozbek, Taiwo Ojo, and Pooria Madani evaluates the training-time contamination robustness of offline ICS anomaly-detection pipelines on the Secure Water Treatment (SWaT) benchmark. Eleven heterogeneous detectors were tested under three contamination strategies: random injection, similarity-targeted injection, and feature-noise injection, using contamination budgets from 1% to 10% with clean validation and test sets. Results show robustness is highly model-dependent and cannot be predicted from clean-data performance alone. Injection-based contamination causes the largest degradation, especially for local-density and distance-based detectors, while feature-noise injection has relatively limited impact. PCA, SVM, HBOS, and IForest remain relatively stable, and a tuned neural network detector shows moderate robustness.

Paper Overview

  • Field: Machine Learning
  • Authors: Mustafa Umut Ozbek, Taiwo Ojo, Pooria Madani
  • Published: 2025-08-26
  • arXiv: 2508.17623
  • Abstract

    Machine-learning-based anomaly detection is increasingly used in industrial control systems (ICS), yet most studies assume that detector training data is trustworthy. In practice, training data may be corrupted through compromised logs, labeling errors, manipulated historian records, or unsafe retraining processes. This paper evaluates the robustness of offline ICS anomaly-detection pipelines on the Secure Water Treatment (SWaT) benchmark under training-time contamination.

    The authors assess 11 heterogeneous anomaly detectors under three contamination strategies: random injection, similarity-targeted injection, and feature-noise injection. The first two insert attack samples into the nominal training pool, while the third adds bounded Gaussian noise to selected normal training samples. These attacks are contamination-style rather than gradient-driven poisoning methods. Evaluations use contamination budgets from 1% to 10%, with clean validation and test sets.

    Key Findings

  • Robustness is highly model-dependent and cannot be predicted from clean-data performance alone.
  • Injection-based contamination (random and similarity-targeted) causes the largest performance degradation, particularly for local-density and distance-based detectors.
  • Feature-noise contamination has a relatively limited impact.
  • PCA, SVM, HBOS, and IForest are relatively stable under contamination.
  • A tuned neural network detector exhibits moderate robustness.
---

*Auto-collected on 2026-08-26.*

Tags

#industrial-control-systems#anomaly-detection#machine-learning#data-poisoning#adversarial-robustness#swat-benchmark#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634009