arXiv · Computers and Society· Antony Dalmiere (LAAS-TRUST), Pascal Marchand (LAAS-TRUST, INSA Toulouse), Guillaume Auriol (LAAS-TRUST, INSA Toulouse), Vincent Nicomette (LAAS-TSF, LAAS)·· 4 hr agoAI score36
连续训练阶段如何影响大语言模型说服力:错位、监督微调与偏好优化的作用
Successive Training Stages and Large Language Model Persuasion: Effects of Misalignment, Supervised Fine-Tuning, and Preference Optimization
AI brief
一项研究考察了三个连续训练阶段对大语言模型说服力的影响:在阴谋论数据上做监督微调(SFT)、在论辩数据上追加说服性 SFT,以及采用 Identity Preference Optimization(IPO)。
Source: arXiv · Computers and Society · arxiv.org