Skip to content
Hot eventLive

Tiny-Scale中文BERT预训练策略对比研究

1 reports1 sources4 hr ago updated

Get the story

AI overview

2026年10月8日,arXiv Computation and Language 发表一项一手研究,在4层、256维、8.7M参数的tiny-scale中文BERT上,使用同一语料(1.29M句中文Wikipedia)和相同超参数,对MLM、WWM与MacBERT三种预训练策略进行对照比较。该研究旨在小规模模型条件下评估不同中文预训练策略的效果差异,目前报道仅披露实验设置,尚未公布具体结果数据。

Generated from reports · updated 3 hr ago

Timeline

Follow the coverage from different angles.

Oct 8, 2026
  1. arXiv · Computation and Language
    Tiny-Scale Chinese BERT 预训练:MLM、WWM 与 MacBERT 策略的对照比较

    该研究在 4 层、256 维、8.7M 参数的 tiny-scale Chinese BERT 上,用同一语料(1.29M 句中文 Wikipedia)和超参数对照 MLM、WWM 与 MacBERT 三种预训练策略。

Heat trend

Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –

02.557.510Oct8Oct8Oct8Oct8

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.