Paper Overview
Field: Machine Learning Authors: Xiaona Zhou, Muntasir Wahed, Tianjiao Yu, Constantin Brif, Ismini Lourentzou Published: 2026-05-28 arXiv: 2605.30344
Summary
Recent advances in vision-language models (VLMs) have achieved impressive performance across many tasks, but prior studies report that large language or multimodal models perform poorly when applied to anomaly pattern discovery in sequential data. Public anomaly detection benchmarks typically provide interval annotations without natural-language explanations, making it difficult to fine-tune VLMs to produce grounded, interpretable decisions.
To fill this gap, the authors construct VisAnomBench, a curated benchmark built from public time-series datasets. High-quality anomaly explanations are generated using multiple large VLMs and filtered through fine-grained, task-specific rewards. By fine-tuning on this benchmark, they develop VisAnomReasoner, a parameter-efficient VLM for time-series anomaly detection.
Results
- VisAnomBench: VisAnomReasoner achieves more accurate anomaly localization, with precision improvements of at least 21.23 percentage points and F1 improvements of 23.87 percentage points, consistently outperforming all baselines.
- TSB-AD-U benchmark: Additional experiments demonstrate strong cross-benchmark generalization, with precision gains of 9.57 percentage points and F1 gains of 13.39 percentage points.
*Auto-collected on 2026-06-01*