Hot eventLive
Soundwich:视频生成的分层可控音频框架
1 reports1 sources3 hr ago updated
Get the story
AI overview
Soundwich 是一种训练免框架,将冻结的联合音视频流匹配模型转为生成多个同步且可独立编辑的音频 stem,并与共享视频耦合。它通过共享场景表征保持跨 stem 的全局音视频一致性,同时保留源级分离,并将跨模态交互路由到对应视觉源,提升音画一致性。该框架于 2026-10-02 发布于 arXiv · Computer Vision。
Generated from reports · updated 2 hr ago
Timeline
Follow the coverage from different angles.
Oct 2, 2026
- arXiv · Computer Vision精选Soundwich:分层可控音频的视频生成
Soundwich 是一种训练免框架,将冻结的联合音视频流匹配模型转为生成多个同步且可独立编辑的音频 stem,并与共享视频耦合。它通过共享场景表征保持跨 stem 的全局音视频一致性,同时保留源级分离,并将跨模态交互路由到对应视觉源,提升音画一致性。
Heat trend
There is not enough continuous observation data to draw a trend yet.