中文SimInsert: 基于区域稀疏注意力融合的无缝视频对象插入
ENSimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion
SimInsert提出一种无需训练的视频物体插入范式,将任务解耦为单帧编辑与语义运动描述,利用图像到视频的生成先验。该方法无需显式运动工程或重新训练,提升了灵活性与泛化能力,可实现时空连贯且交互真实的插入效果。实际意义在于降低了视频编辑门槛,拓展了生成式AI在视频内容创作中的应用。
arXiv:2605.23245v1 Announce Type: new Abstract: Video object insertion requires ensuring spatio-temporal coherence and interactive realism, extending far beyond simple content placement. However, current approaches are often hindered by a reliance on explicit motion engineering or resource-intensive retraining, restricting their flexibility and generalization. To bridge this gap, we present \textit{SimInsert}, a training-free paradigm that efficiently decouples the task into intuitive single-frame editing and semantic motion description. By harnessing the robust generative priors of image-to-v