Title: Training-free Logical Anomaly Detection via VLM-based Automated Prompting and Patch-Text Alignment
-Journal/Conference: The 19th European Conference on Computer Vision - ECCV 2026
-Authors: Taeheon Jin, Dongwoo Kim, Jaehoon Lim, ChanMin An, MINJAE KWEUN, Jaeyo Chang, Jongpil Jeong
-DOI:
-Journal/Conference Link: https://eccv.ecva.net/
Abstract: While Vision-Language Models (VLMs) show promise in industrial anomaly detection, their scalability is hindered by reliance on manual prompt engineering and susceptibility to hallucinations. To address these limitations, we propose Automated Prompting and Patch-Text Alignment for Anomaly Detection (APTAD), a training-free multimodal framework that unifies structural and logical anomaly detection. At the core of our approach is an Automation-of-Thought (AoT) process, which autonomously generates precise inspection prompts and extracts structured text features without human intervention. By leveraging these extracted features, our framework naturally extends to few-shot anomaly detection. Furthermore, to ensure reliability, APTAD incorporates an Anomaly Reasoning module that performs deterministic comparisons. This mechanism provides explicit, explainable detection results while effectively eliminating VLM hallucinations. Extensive experiments demonstrate that APTAD achieves outstanding performance, yielding an image-level AUROC of 94.3% on the MVTec LOCO AD benchmark (full-data) and 98.3% on the VisA dataset (4-shot).
-Status: Submitted (2026/03/06)