Skip to content
Hot eventLive

Sigma:首个大规模连续扩散语言模型

1 reports1 sources16 hr ago updated

Get the story

AI overview

2026年10月5日,arXiv 计算与语言栏目发布论文,提出 Sigma——首个基于可引导低维 ODE/SDE 潜轨迹的大规模(3B/8B)连续扩散语言模型。训练上采用分块似然优化,并以自回归模型预训练权重热启动。论文指出,推理阶段分类器无关引导与得分温度对高保真推理和编码至关重要。评测显示,Sigma 在 GSM8K、Minerva、HumanEval、MBPP 及 MATH-500、AIME 上与离散模型表现相当。目前公开信息仅含该论文内容,尚无后续独立验证或第三方复现报道。

Generated from reports · updated 16 hr ago

Timeline

Follow the coverage from different angles.

Oct 5, 2026
  1. arXiv · Computation and Language
    Sigma:大规模连续扩散语言模型

    论文提出 Sigma,首个基于可引导低维 ODE/SDE 潜轨迹的大规模(3B/8B)连续扩散语言模型,通过分块似然优化训练,并用自回归模型预训练权重热启动。推理中,分类器无关引导与得分温度对高保真推理和编码至关重要;在 GSM8K、Minerva、HumanEval、MBPP 及 MATH-500、AIME 上与离散模型表现相当。

Heat trend

Current heat 7·Comparable peak 10(Oct 5)·Comparable change over 24 hours –

02.557.510Oct5Oct5Oct5Oct6

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.