Skip to content
Hot eventLive

Bellman-Certified Rounding for Sparse Policy Deployment in MDPs

1 reports1 sources5 hr ago updated

Get the story

AI overview

该论文研究有限 MDP 中连续策略混合舍入为稀疏二进制策略时如何保留折扣回报,方法通过 2d+2 次 Bellman 求解得到可复用包络,在舍入前提供均匀与候选特定保证。候选特定界将结构化测试集上的认证覆盖率从 48.2% 提升至 74.1%,在 γ=0.95 时局部积分把中位界损失比从 402.3 降至 2.08。论文证明当预算线性增长时线性维度依赖不可避免,且精确全局曲率阈值判定是 NP 难问题。

Generated from reports · updated 4 hr ago

Timeline

Follow the coverage from different angles.

Oct 2, 2026
  1. arXiv · Machine Learning Theory
    Bellman-Certified Rounding for Sparse Policy Deployment in MDPs

    论文研究有限 MDP 中连续策略混合舍入为稀疏二进制策略时如何保留折扣回报,通过 2d+2 次 Bellman 求解得到可复用包络,在舍入前提供均匀与候选特定保证。候选特定界将结构化测试集上的认证覆盖率从 48.2% 提升至 74.1%,在 γ=0.95 时局部积分把中位界损失比从 402.3 降至 2.08。论文证明当预算线性增长时线性维度依赖不可避免,且精确全局曲率阈值判定是 NP 难问题。

Heat trend

Current heat 9·Comparable peak 10(Oct 2)·Comparable change over 24 hours –

02.557.510Oct2Oct2Oct2Oct2

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.