Hot eventLive
Tiny-Scale中文BERT预训练策略对比研究
1 reports1 sources4 hr ago updated
Get the story
AI overview
2026年10月8日,arXiv Computation and Language 发表一项一手研究,在4层、256维、8.7M参数的tiny-scale中文BERT上,使用同一语料(1.29M句中文Wikipedia)和相同超参数,对MLM、WWM与MacBERT三种预训练策略进行对照比较。该研究旨在小规模模型条件下评估不同中文预训练策略的效果差异,目前报道仅披露实验设置,尚未公布具体结果数据。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
Oct 8, 2026
- arXiv · Computation and LanguageTiny-Scale Chinese BERT 预训练:MLM、WWM 与 MacBERT 策略的对照比较
该研究在 4 层、256 维、8.7M 参数的 tiny-scale Chinese BERT 上,用同一语料(1.29M 句中文 Wikipedia)和超参数对照 MLM、WWM 与 MacBERT 三种预训练策略。
Heat trend
Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.