LLM蒸馏多标签主题分配:生成式与判别式学生模型对比
Get the story
该研究围绕基于 LLM 知识蒸馏的多标签主题分配任务,对比生成式与判别式学生模型。研究在 1B、4B、8B 参数规模下展开,判别式基线采用 DeBERTa-V3 与 ModernBERT。结果显示,在结构化商品评论上,判别式模型优于超轻量生成式模型;但在多轮对话数据上,1B 生成式模型表现反超。此外,在 112 个主题的大规模标签扩展与长尾分布条件下,生成式模型保持稳健,而判别式基线 Macro-F1 下降 35%。目前优化模型已在商品评论与对话场景完成生产部署,并满足严格延迟要求。
Generated from reports · updated 3 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Machine Learning Theory基于 LLM 知识蒸馏的多标签主题分配:生成式与判别式学生模型的对比分析
研究在 1B、4B、8B 参数规模下对比生成式与判别式(DeBERTa-V3、ModernBERT)学生模型,发现判别式在结构化商品评论上优于超轻量生成式,但 1B 生成式模型在多轮对话数据上反超。生成式模型在 112 个主题的大规模标签扩展与长尾分布下保持稳健,判别式基线 Macro-F1 下降 35%。优化模型已在商品评论与对话场景完成生产部署,满足严格延迟要求。
Heat trend
Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.