Skip to content
Hot eventLive

研究用原始视频训练Qwen3-1.7B

1 reports1 sources3 hr ago updated

Get the story

AI overview

2026年10月9日,arXiv计算机视觉方向发布一手论文,研究在原始视频中训练Qwen3-1.7B的可行性。研究将视频帧编码为连续视觉token,让模型预测下一个视觉token,并在YT-Temporal-1B原始片段上完成训练。结果显示,四个视频基准平均提高2.9分,十个图像基准平均提高5.1分;14个文本基准得分为48.9对48.0,文本表现保持。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 9, 2026
  1. arXiv · Computer Vision
    研究用原始视频中训练 Qwen3-1.7B

    论文研究中训练 Qwen3-1.7B 使用无标注原始视频的可行性,把帧编码为连续视觉 token 并让模型预测下一个视觉 token。在 YT-Temporal-1B 原始片段上中训练后,四个视频基准平均高 2.9 分、十个图像基准平均高 5.1 分,14 个文本基准平均 48.9 对 48.0,文本表现保持。

Heat trend

Current heat 9·Comparable peak 10(Oct 9)·Comparable change over 24 hours –

02.557.510Oct9Oct9Oct9Oct9

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.