中文DRIVESPATIAL:面向自动驾驶视觉语言模型时空智能的基准
ENDRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving
该论文指出,现有自动驾驶视觉-语言基准集中于单视图、静态或单源问答,无法测试视觉语言模型(VLM)的时空智能。作者引入新基准Driv...,要求模型整合多视角观测、保持物体跨视角和时间连续性,并推理空间关系与未来动态,填补了动态场景推理的评估空白,推动VLM在自动驾驶中的应用。
arXiv:2605.23176v1 Announce Type: new Abstract: Spatiotemporal intelligence in autonomous driving (AD) requires an agent to integrate multi-view observations into a coherent scene representation, maintain object continuity across viewpoints and time, and reason about spatial relations, interactions, and future dynamics. However, existing AD vision-language benchmarks largely focus on single-view, static, ego-centric, or single-source question answering, leaving it unclear whether current Vision-Language Models (VLMs) can truly construct and reason over dynamic driving scenes. We introduce Driv