Hot eventLive
CrossEdit论文:跨模态训练实现音视频编辑
1 reports1 sources4 hr ago updated
Get the story
AI overview
研究者提出 CrossEdit,一个面向图像、音频与视频的统一全模态编辑模型,声称可零样本完成音视频电影场景编辑。其方法利用声学信号的加性特性,程序化生成大规模组合式音频编辑对,并结合自监督 AV 掩码重建与精选跨模态任务微调,使指令遵循能力迁移到未见的模态与指令组合。该成果发布于 arXiv 的 Audio and Speech 栏目,属一手研究报道。目前尚无第三方验证或后续进展披露,报道也未给出具体量化指标或开源信息。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
Oct 8, 2026
- arXiv · Audio and SpeechCrossEdit:跨模态训练实现丰富的音视频编辑
研究者提出 CrossEdit,一个面向图像、音频与视频的统一全模态编辑模型,可零样本完成音视频电影场景编辑。团队利用声学信号的加性特性程序化生成大规模组合式音频编辑对,并结合自监督 AV 掩码重建与精选跨模态任务微调,使指令遵循能力迁移到未见的模态与指令组合。
Heat trend
Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.