Skip to content
arXiv · Machine Learning Theory· Joss Armstrong·· 5 hr agoSelectedAI score62

Source Identification Is Not Fitness Testing: Measuring the Limits of Synthetic-Data Attribution

Source Identification Is Not Fitness Testing: Measuring the Limits of Synthetic-Data Attribution

AI brief

该研究测试了合成数据溯源在重复训练中的可靠性与筛选价值。以金融风险文本为素材,原始生成文本的生成器归属准确率达 98.7%,改写后降至 53.1%(释义)和 29.0%(风格改写),生成与人类检测仍接近完美。

Why it matters

原文用金融风险文本实验表明溯源准确率随改写大幅下降,且溯源与选择有用训练数据是两个不同问题。

Source: arXiv · Machine Learning Theory · arxiv.org