Skip to content
Hot eventLive

ASR非语言发声建模研究提出三种策略

1 reports1 sources4 hr ago updated

Get the story

AI overview

2026年10月8日,arXiv平台Audio and Speech栏目发布一手研究论文《Beyond Words:ASR中非语言发声的有效建模》。研究针对自动语音识别(ASR)中非语言发声(NVs,如笑声、呼吸、咳嗽、哭声)的建模问题,提出三种数据中心策略:一是两阶段课程学习,先将模型映射到通用token,再在目标类别上微调;二是token间迁移,从高资源事件(笑声、呼吸)向稀有事件(哭泣)迁移;三是带类别平衡的语音转换增强。论文目前仅公开上述方法框架,未披露实验数据或进一步验证结果。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 8, 2026
  1. arXiv · Audio and Speech
    Beyond Words:ASR 中非语言发声的有效建模

    研究提出三种数据中心策略提升 ASR 对非语言发声(NVs,如笑声、呼吸、咳嗽、哭声)的识别:两阶段课程学习(先映射到通用 token 再在目标类别上微调)、从高资源事件(笑声、呼吸)向稀有事件(哭泣)的 token 间迁移,以及带类别平衡的语音转换增强。

Heat trend

Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –

02.557.510Oct8Oct8Oct8Oct8

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.