Hot eventLive
研究:说服性SFT提升LLM说服力,偏好优化无额外增益
1 reports1 sources4 hr ago updated
Get the story
AI overview
一项研究考察了三个连续训练阶段对大语言模型说服力的影响:在阴谋论数据上做监督微调(SFT)、在论辩数据上追加说服性 SFT,以及采用 Identity Preference Optimization(IPO)。研究结果显示,在阴谋论数据上进行 SFT 会降低模型的说服力,而在论辩数据上追加说服性 SFT 能提升模型的说服力;随后采用 IPO 进行偏好优化并未带来额外增益。该研究为连续训练阶段如何影响大语言模型说服力提供了实证观察。
Generated from reports · updated 3 hr ago
LatestOct 8
研究指出说服性SFT可提升LLM说服力,IPO偏好优化无额外增益。Timeline
Follow the coverage from different angles.
Oct 8, 2026
- arXiv · Computers and Society连续训练阶段如何影响大语言模型说服力:错位、监督微调与偏好优化的作用
一项研究考察了三个连续训练阶段对大语言模型说服力的影响:在阴谋论数据上做监督微调(SFT)、在论辩数据上追加说服性 SFT,以及采用 Identity Preference Optimization(IPO)。
Heat trend
Current heat 9·Comparable peak 10(Oct 8)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.