Skip to content
Hot eventLive

多语言LLM音素到字素研究降低WER至7.66%

1 reports1 sources3 hr ago updated

Get the story

AI overview

2026年10月8日,arXiv(Audio and Speech,一手)发表关于多语言语音识别中基于LLM的音素到字素(P2G)映射的研究。研究在十语言CV-Lang10基准上展开,针对语言感知生成与跨语言数据失衡问题,评估了DANP与简化SKM(S-SKM)等考虑S2P不确定性的鲁棒训练策略;其中S-SKM以蒙特卡洛近似避免P2G训练中基于CTC的S2P概率加权。结合鲁棒训练与低资源过采样,平均WER从10.56%降至7.66%。目前未见后续报道或独立验证。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 8, 2026
  1. arXiv · Audio and Speech
    多语言语音识别中基于 LLM 的音素到字素映射研究

    研究在十语言 CV-Lang10 基准上探索多语言 LLM 音素到字素(P2G)映射,以应对语言感知生成与跨语言数据失衡。论文评估 DANP 与简化 SKM(S-SKM)等考虑 S2P 不确定性的鲁棒训练策略,S-SKM 以蒙特卡洛近似避免 P2G 训练中基于 CTC 的 S2P 概率加权。结合鲁棒训练与低资源过采样,平均 WER 从 10.56% 降至 7.66%。

Heat trend

Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –

02.557.510Oct8Oct8Oct8Oct8

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.