CoLMbo-SV:可解释说话人验证的接地语言模型
Get the story
2026年10月5日,arXiv Audio and Speech 发布一手报道,介绍 CoLMbo-SV:一个面向可解释说话人验证的接地语言模型。该模型将预训练说话人编码器接入语言模型,并输入显式声学测量,从而在生成结构化声学比较报告的同时保持强说话人判别能力。研究还发布配套数据集 VoxReason,提供带实测声学属性与比较报告的配对录音。在 VoxCeleb1-O 上,CoLMbo-SV 取得 0.99% EER,相对最强音频语言基线降低约 80% 验证错误,数值接地得分 0.82。
Generated from reports · updated 15 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Audio and SpeechCoLMbo-SV:可解释说话人验证的接地语言模型
CoLMbo-SV 将预训练说话人编码器接入语言模型并输入显式声学测量,在生成结构化声学比较报告的同时保持强说话人判别能力。配套数据集 VoxReason 提供带实测声学属性与比较报告的配对录音;在 VoxCeleb1-O 上取得 0.99% EER,相对最强音频语言基线降低约 80% 验证错误,数值接地得分 0.82。
Heat trend
Current heat 7·Comparable peak 10(Oct 5)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.