Ian Osband 提出 horizon loss 改进策略梯度训练
Get the story
2026年10月5日,Ian Osband 在 arXiv(Statistics Machine Learning)发表《Planning to Learn》,提出 horizon loss。该方法把精确策略梯度按剩余学习量截断,介于交叉熵与零时域极限之间。论文称,分类器期望奖励即期望准确率;精确策略梯度虽无探索、信用分配与动作采样噪声,却因短视而输给交叉熵。实验在 MNIST 与 ImageNet 上进行,ResNet-50、ResNet-101、ViT-S/16 在固定学习率下 top-1 准确率提升,标签噪声越大增益越明显。目前未见后续验证或争议报道。
Generated from reports · updated 16 hr ago
Timeline
Follow the coverage from different angles.
- arXiv · Statistics Machine LearningPlanning to Learn:把策略梯度按剩余学习量截断的 horizon loss
论文提出 horizon loss,把精确策略梯度按剩余学习量截断,介于交叉熵与零时域极限之间。分类器期望奖励即期望准确率,精确策略梯度虽无探索、信用分配与动作采样噪声,却因短视而输给交叉熵。MNIST 与 ImageNet 上 ResNet-50、ResNet-101、ViT-S/16 在固定学习率下 top-1 准确率提升,标签噪声越大增益越明显。
Heat trend
Current heat 7·Comparable peak 10(Oct 5)·Comparable change over 24 hours –
The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.