Skip to content
arXiv · Artificial Intelligence· Jundong Hu, Shekar Ramachandran·· 3 hr agoSelectedAI score62

测量微任务资格缺口:现成小语言模型何时足以支撑 Agent Harness

Measuring the Microtask Eligibility Gap: When Is an Off-the-Shelf SLM Enough for an Agent Harness?

AI brief

论文构建了 4 个微任务基准,以低成本非 LLM 基线为阈值 τ,用 CI 感知规则检验 Qwen3 0.6/1.7/4/8B(FP16、贪心、无微调),结果 16 个配置均未通过资格缺口。

Why it matters

用可复现的 CI 置信区间规则衡量 SLM 微任务资格,帮助从业者在 harness 中更稳妥地分配 SLM 与基线。

Source: arXiv · Artificial Intelligence · arxiv.org