Skip to content
Hot eventLive

Soundwich:视频生成的分层可控音频框架

1 reports1 sources3 hr ago updated

Get the story

AI overview

Soundwich 是一种训练免框架,将冻结的联合音视频流匹配模型转为生成多个同步且可独立编辑的音频 stem,并与共享视频耦合。它通过共享场景表征保持跨 stem 的全局音视频一致性,同时保留源级分离,并将跨模态交互路由到对应视觉源,提升音画一致性。该框架于 2026-10-02 发布于 arXiv · Computer Vision。

Generated from reports · updated 2 hr ago

Timeline

Follow the coverage from different angles.

Oct 2, 2026
  1. arXiv · Computer Vision精选
    Soundwich:分层可控音频的视频生成

    Soundwich 是一种训练免框架,将冻结的联合音视频流匹配模型转为生成多个同步且可独立编辑的音频 stem,并与共享视频耦合。它通过共享场景表征保持跨 stem 的全局音视频一致性,同时保留源级分离,并将跨模态交互路由到对应视觉源,提升音画一致性。

Heat trend

There is not enough continuous observation data to draw a trend yet.