Skip to content
arXiv · Machine Learning Theory· Sreejeet Maity, Aritra Mitra·· 4 hr agoAI score22

从不可靠轨迹中学习:对抗鲁棒的联邦 Q-Learning

Learning from Unreliable Trajectories: Adversarially-Robust Federated Q-Learning

AI brief

论文研究多智能体在共享 MDP 中通过中央服务器协作学习最优状态-动作值函数,并提出 Robust Async-Fed-Q 算法,结合智能体端对方差缩减的 Bellman 最优算子估计与服务器端鲁棒聚合。

Source: arXiv · Machine Learning Theory · arxiv.org