Skip to content
Hot eventWatching

STEMMA:面向大音频语言模型的歌曲到分轨多音频推理框架

1 reports1 sources1 days ago updated

Get the story

AI overview

据 arXiv(Audio and Speech)报道,STEMMA 是一个以制作溯源为核心的多音频音乐问答框架,用于判断片段是否来自同一曲目或段落、以及哪些 stems 属于哪些混音。它采用关系优先的构建策略:先指定目标关系,再检索满足条件的片段与难负样本,标签直接来自目录溯源而非语言模型生成。配套发布 STEMMA-Bench 与曲目不重叠的训练集 STEMMA-Instruct,在两个 LALM 上微调可提升多音频推理能力,同时保持单音频音乐理解。

Generated from reports · updated 1 days ago

Timeline

Follow the coverage from different angles.

Oct 9, 2026
  1. arXiv · Audio and Speech
    STEMMA:面向大音频语言模型的歌曲到分轨多音频推理框架

    STEMMA 是一个以制作溯源为核心的多音频音乐问答框架,判断片段是否来自同一曲目或段落、哪些 stems 属于哪些混音。它采用关系优先的构建策略,先指定目标关系再检索满足条件的片段与难负样本,标签直接来自目录溯源而非语言模型生成。配套 STEMMA-Bench 与曲目不重叠的训练集 STEMMA-Instruct,在两个 LALM 上微调可提升多音频推理能力,并保持单音频音乐理解。

Heat trend

Current heat 5·Comparable peak 10(Oct 9)·Comparable change over 24 hours -50%

02.557.510Oct9Oct9Oct10Oct10

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.