MATCH: Multi-Agentic Evidence Grounding for Explainable Hate Video Detection
IEEE Transactions on Circuits and Systems for Video Technology, 2026
DOI: 10.1109/TCSVT.2026.3672052
IEEE document: 11424606
Version: Author-hosted manuscript
Abstract
The growing prevalence of hate videos promoting intolerance, bigotry, and discrimination presents significant psychosocial threats to both individuals and society. Current detection methods often rely on black-box models, which lack interpretability - a crucial factor for fostering more reliable content moderation and trustworthy AI. To bridge this gap, we propose MATCH, the first attempt to achieve interpretable hate video detection via multiple Large Multimodal Model (LMM) agent collaboration. Our method facilitates a novel diversely generate-then-verify paradigm, where LMM agents work in tandem to generate diverse clues and verify them to yield more faithful explanations. MATCH begins by proposing a new Dual-Perspective Proposing paradigm, where two LMM agents are regarded as Proposers to independently identify evidential clues from opposing angles - hate and non-hate. Leveraging these comprehensive clues, we introduce an innovative Spatiotemporal Evidence-Grounded Verification mechanism. In this mechanism, a third LMM agent acts as a Verifier, rigorously validating and reconciling the proposed clues against spatiotemporal evidences directly extracted from video content, yielding coherent and faithful explanations. Finally, these explanations are integrated with video features, enabling accurate identification of complex and ambiguous hateful content. Extensive experiments conducted on three benchmark datasets demonstrate that MATCH not only achieves state-of-the-art performance, but also provides reliable and trustworthy rationales for the predictions.
Keywords
Hate video detection; interpretability; large multimodal models; agent collaboration.