Skip to content
Hot eventLive

Encoder-free Speech-LLM通过元数据监督预训练提升说话人辨识与副语言理解

1 reports1 sources3 hr ago updated

Get the story

AI overview

2026-10-02,arXiv Audio and Speech 频道发布一手论文,提出 metadata-supervised pretraining(MSP)方法,用于无编码器 Speech-LLM 的训练。该方法利用说话人身份、情感等语音元数据进行监督预训练,并引入 speaker-aware utterance composition(SAUC)与随机跨度掩码正则化,以提升模型在说话人辨识与副语言理解方面的能力。

Generated from reports · updated 2 hr ago

Timeline

Follow the coverage from different angles.

Oct 2, 2026
  1. arXiv · Audio and Speech
    教大语言模型识别说话人:Metadata-Supervised Pretraining 用于无编码器语音-LLM

    论文提出 metadata-supervised pretraining(MSP),利用说话人身份和情感等语音属性训练无编码器 Speech-LLM,并引入 speaker-aware utterance composition(SAUC)和随机跨度掩码正则化。

Heat trend

There is not enough continuous observation data to draw a trend yet.