VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding


Abstract

Synthetic video generation has gained significant attention for its realism and broad applications, but remains prone to violations of common sense and physical laws. This highlights the need for reliable abnormality detectors that understand such principles and are robust to hallucinations. To address this, we introduce VideoHallu, a benchmark of over 3,000 video QA pairs built from synthetic videos generated by models like Veo2, Sora, and Kling, paired with expert-crafted counterintuitive QA to evaluate the critical thinking abilities of Multi-modal Large Language Models (MLLMs) on abnormalities that are perceptually obvious to humans but often hallucinated due to language priors. VideoHallu evaluates MLLMs’ abnormality detection abilities with examples across alignment, consistency, commonsense, and physics. We benchmark SoTA MLLMs, including GPT-4o, Gemini-2.5-Pro, Qwen-2.5-VL, Video-R1, and VideoChat-R1. We observe that models perform well on many real-world benchmarks like MVBench and MovieChat, but still struggle with basic physics-based and commonsense reasoning in synthetic videos. Moreover, we post-train current SoTA MLLMs with Group Relative Policy Optimization (GRPO) using both real-world and synthetic commonsense/physics datasets, improving abnormality detection and critical thinking.

Proceedings
Advances in Neural Information Processing Systems (NeurIPS)
Date
BibTex