DirectSpeech2LLM:缓解语音大模型提示过拟合的端到端框架
Get the story
DirectSpeech2LLM 提出端到端框架,用于缓解 Speech-LLM 仅用 ASR 指令训练后无法泛化到新指令的提示过拟合问题。该框架在冻结的 LLM 嵌入矩阵上计算基于距离的 CTC 损失,并用贪心 CTC 标签生成几何与时间对齐的语音嵌入。仅用 960 小时 LibriSpeech ASR 数据训练,即在 ASR 上超越级联系统,并零样本泛化到语音翻译与情感识别,接近级联系统上限。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Computation and LanguageDirectSpeech2LLM:缓解 Speech-LLM 提示过拟合的端到端框架
DirectSpeech2LLM 提出端到端框架,缓解 Speech-LLM 仅用 ASR 指令训练后无法泛化到新指令的提示过拟合。该框架在冻结的 LLM 嵌入矩阵上计算基于距离的 CTC 损失,并用贪心 CTC 标签生成几何与时间对齐的语音嵌入。仅用 960 小时 LibriSpeech ASR 数据训练,即在 ASR 上超越级联系统,并零样本泛化到语音翻译与情感识别,接近级联系统上限。
Heat trend
Current heat 9·Comparable peak 10(Oct 7)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.