MATCH: Multi-Agentic Evidence Grounding for Explainable Hate Video Detection

Kaiju Li, Rongpei Hong, Jian Lang, Jin Wu, Fan Zhou, Jingkuan Song

IEEE Transactions on Circuits and Systems for Video Technology, 2026

DOI: 10.1109/TCSVT.2026.3672052

IEEE document: 11424606

Version: Author-hosted manuscript

Abstract

The growing prevalence of hate videos promoting intolerance, bigotry, and discrimination presents significant psychosocial threats to both individuals and society. Current detection methods often rely on black-box models, which lack interpretability - a crucial factor for fostering more reliable content moderation and trustworthy AI. To bridge this gap, we propose MATCH, the first attempt to achieve interpretable hate video detection via multiple Large Multimodal Model (LMM) agent collaboration. Our method facilitates a novel diversely generate-then-verify paradigm, where LMM agents work in tandem to generate diverse clues and verify them to yield more faithful explanations. MATCH begins by proposing a new Dual-Perspective Proposing paradigm, where two LMM agents are regarded as Proposers to independently identify evidential clues from opposing angles - hate and non-hate. Leveraging these comprehensive clues, we introduce an innovative Spatiotemporal Evidence-Grounded Verification mechanism. In this mechanism, a third LMM agent acts as a Verifier, rigorously validating and reconciling the proposed clues against spatiotemporal evidences directly extracted from video content, yielding coherent and faithful explanations. Finally, these explanations are integrated with video features, enabling accurate identification of complex and ambiguous hateful content. Extensive experiments conducted on three benchmark datasets demonstrate that MATCH not only achieves state-of-the-art performance, but also provides reliable and trustworthy rationales for the predictions.

Keywords

Hate video detection; interpretability; large multimodal models; agent collaboration.