音素引导初始化方法用于LLM语音识别
Get the story
论文提出一种面向基于 LLM 的语音识别的音素引导初始化方法:在端到端框架内,先分别预训练音频编码器完成语音到音素(S2P)转换、LLM 完成音素到字素(P2G)转换,再将二者连接并端到端微调。实验覆盖日语 CSJ、中文 AISHELL-1,以及 Common Voice 25.0 中的 Tatar 和 Urdu 低资源语言,结果显示该方法匹配或优于级联 S2P-P2G 基线和无 P2G 初始化的端到端模型。该论文已被 IEEE SLT 2026 接收。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Audio and Speech面向基于 LLM 的语音识别的音素引导初始化方法
论文提出音素引导初始化方法,在端到端框架内先分别预训练音频编码器完成语音到音素(S2P)转换、LLM 完成音素到字素(P2G)转换,再连接并端到端微调。在日语 CSJ、中文 AISHELL-1 及 Common Voice 25.0 的 Tatar 和 Urdu 低资源语言上,该方法匹配或优于级联 S2P-P2G 基线和无 P2G 初始化的端到端模型。论文已被 IEEE SLT 2026 接收。
Heat trend
Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.