Skip to content
Hot eventLive

MintEval:自然语言到策略代码的行为等价性基准

1 reports1 sources16 hr ago updated

Get the story

AI overview

2026年10月5日,研究者在 arXiv(Software Engineering 栏目)发布一手报道,推出 MintEval 基准,用于检验 LLM 是否真正实现了用户要求的交易策略。该基准用可组合构建块程序化生成参考策略,并将其回译为交易员指令;评测时逐 bar 在相同市场数据与交易摩擦下比较动作,而非比较代码相似度或收益。目前报道仅介绍该基准的设计思路与评估方式,未披露具体实验结果、数据规模或参与模型清单。

Generated from reports · updated 16 hr ago

Timeline

Follow the coverage from different angles.

Oct 5, 2026
  1. arXiv · Software Engineering
    MintEval:LLM 是否实现了你要求的交易策略?自然语言到策略代码的行为等价性基准

    研究者推出 MintEval 基准,用可组合构建块程序化生成参考策略并回译为交易员指令,逐 bar 在相同市场数据与摩擦下比较动作而非代码相似度或收益。

Heat trend

Current heat 7·Comparable peak 10(Oct 5)·Comparable change over 24 hours –

02.557.510Oct5Oct5Oct5Oct6

The trend compares only the same participants observed continuously; its range may be smaller than the current heat count. Move or click on the chart to inspect hourly heat; use the left and right arrow keys to switch.