中文EgoInteract:合成第一人称视频用于交互理解与预测
ENEgoInteract: Synthetic Egocentric Videos Generation for Interaction Understanding and Anticipation
该研究提出EgoInteract,一个可控的第一人称视频生成模拟器,旨在解决真实数据集成本高、存在隐私和偏见等问题。它能够生成具有时间一致性的复杂人-物交互视频,有望降低数据采集成本、扩大交互模式覆盖范围,为第一人称感知任务提供灵活的合成数据来源。
arXiv:2605.18214v2 Announce Type: replace Abstract: Collecting large-scale egocentric video datasets with dense spatial and temporal annotations is costly, slow, and often constrained by environmental biases, privacy constraints, and limited coverage of interaction patterns. While synthetic data has shown strong potential in several vision domains, its use for egocentric perception remains relatively underexplored, especially for tasks requiring temporally coherent human-object interactions. In this work, we introduce EgoInteract, a controllable simulator for egocentric video generation design