Skip to content
Hot eventLive

JoyAI-Voice 2.0 全连续自回归语音生成模型发布

1 reports1 sources4 hr ago updated

Get the story

AI overview

JoyAI-Voice 2.0 提出端到端拟人语音生成模型,采用全连续双编码器架构,将原始语音编码为连续 latent 并分块,由语义-声学双编码器分解后融合,送入因果自回归 Transformer 进行规划,再由局部扩散 Transformer 渲染完整 latent,最终以 48 kHz 合成语音。该模型相关论文发布于 arXiv 的 Audio and Speech 栏目,属一手研究报道。目前公开信息仅涉及模型架构与合成流程,未披露发布时间、参数规模、数据集、评测指标或开源计划等细节,也无其他来源报道可供交叉验证。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 8, 2026
  1. arXiv · Audio and Speech
    JoyAI-Voice 2.0:基于语义-声学联合表征的全连续自回归语音生成模型

    JoyAI-Voice 2.0 提出端到端拟人语音生成模型,采用全连续双编码器架构,将原始语音编码为连续 latent 并分块,由语义-声学双编码器分解后融合送入因果自回归 Transformer 规划,局部扩散 Transformer 渲染完整 latent 以 48 kHz 合成。

Heat trend

Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –

02.557.510Oct8Oct8Oct8Oct8

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.