Hot eventLive
研究用原始视频训练Qwen3-1.7B
1 reports1 sources3 hr ago updated
Get the story
AI overview
2026年10月9日,arXiv计算机视觉方向发布一手论文,研究在原始视频中训练Qwen3-1.7B的可行性。研究将视频帧编码为连续视觉token,让模型预测下一个视觉token,并在YT-Temporal-1B原始片段上完成训练。结果显示,四个视频基准平均提高2.9分,十个图像基准平均提高5.1分;14个文本基准得分为48.9对48.0,文本表现保持。
Generated from reports · updated 3 hr ago
LatestOct 9
论文报告在原始视频上训练Qwen3-1.7B,视频与图像基准提升,文本表现基本保持。Timeline
Follow the coverage from different angles.
Oct 9, 2026
- arXiv · Computer Vision研究用原始视频中训练 Qwen3-1.7B
论文研究中训练 Qwen3-1.7B 使用无标注原始视频的可行性,把帧编码为连续视觉 token 并让模型预测下一个视觉 token。在 YT-Temporal-1B 原始片段上中训练后,四个视频基准平均高 2.9 分、十个图像基准平均高 5.1 分,14 个文本基准平均 48.9 对 48.0,文本表现保持。
Heat trend
Current heat 9·Comparable peak 10(Oct 9)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.