Skip to content
arXiv · Audio and Speech· Haoyin Yan, Chengwei Liu, Zheng Xue, Xiaotao Liang, Jifa Cai, Zeyu Zhao, Jingjing Wang·· 4 hr agoAI score23

LIFT-SE:先做语言推理再做流变换的生成式语音增强

LIFT-SE: Linguistic Inference Followed by Flow Transformation for Generative Speech Enhancement

AI brief

LIFT-SE 是一个两阶段生成式语音增强框架,在 QRes-Codec 内把语言推理与声学合成解耦,先基于自监督前端蒸馏出的帧对齐特征自回归预测干净 codec token,再用条件流匹配把高斯噪声输运到由预测 token 条件的连续潜变量,最后由冻结解码器重建增强波形。

Source: arXiv · Audio and Speech · arxiv.org