中文CaST-Bench: 面向视频问答的因果链引导时空推理基准测试
ENCaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering
arXiv新论文提出CaST-Bench基准,用于评估时空视频因果推理。它要求模型识别并定位多步因果链,填补了现有基准缺乏细粒度证据的空白,推动VLM从表面感知深入因果机制理解。
arXiv:2605.23216v1 Announce Type: new Abstract: Cause-and-effect reasoning in video is a significant challenge for Vision-Language Models (VLMs), as it requires going beyond surface-level perception to a deeper understanding of causal mechanisms. However, existing benchmarks rarely provide the fine-grained, grounded evidence needed to rigorously evaluate this capability. To address this gap, we introduce CaST-Bench, a benchmark for Causal Chain-Grounded Spatio-Temporal Video Reasoning. CaST-Bench presents complex causal questions that require models to identify and localize a chain of multiple