跳到正文
arXiv cs.AI· Chenglin Yang·· 10 小时前AI 评分24

评估智能体动作的堆叠防线而非单层:确定性规则与 LLM 门控会独立失效吗?

Evaluate the Stack, Not the Layer: Do Deterministic and LLM Gates for Agent Actions Fail Independently?

AI 导读

研究用 1,119 条已标注智能体动作测试运行时门控堆叠,发现任意两个 LLM 裁判组合仅相当于约 1.2 至 1.4 层乘法等效层,而规则层加一个裁判可达 1.86 至 2.09 层。

来源:arXiv cs.AI · arxiv.org