Skip to content
Hot eventLive

Prosody-to-Text:用低通滤波语音预测文本

1 reports1 sources3 hr ago updated

Get the story

AI overview

2026年10月9日,arXiv 计算与语言栏目发布一项研究,提出 Prosody-to-Text 方法,仅用约 450Hz 截止的低通滤波语音(12 个最低 Mel 频段)从韵律模式恢复原文。研究者微调 Whisper 模型,报告 WER 为 36%,10% 的语句被完美恢复,40% 的语句 WER 不超过 25%;在给定正确前缀时,下一个 token 预测准确率达 79%。结果表明低频语音特征与词汇内容的关联强于此前认知,或可用于引导现代 LLM 的文本生成。该报道为一手来源,目前尚无后续验证或独立复现报道。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 9, 2026
  1. arXiv · Computation and Language
    Prosody-to-Text:基于低通滤波语音预测文本

    研究者微调 Whisper 模型,仅使用约 450Hz 截止的低通滤波(12 个最低 Mel 频段)从韵律模式恢复原文,WER 为 36%,10% 的语句被完美恢复,40% 的语句 WER 不超过 25%。在给定正确前缀时,下一个 token 的预测准确率达 79%。结果表明低频语音特征与词汇内容的关联强于此前认知,或可用于引导现代 LLM 的文本生成。

Heat trend

Current heat 9·Comparable peak 10(Oct 9)·Comparable change over 24 hours –

02.557.510Oct9Oct9Oct9Oct9

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.