arXiv cs.AI· Chenglin Yang·· 10 小时前AI 评分24
评估智能体动作的堆叠防线而非单层:确定性规则与 LLM 门控会独立失效吗?
Evaluate the Stack, Not the Layer: Do Deterministic and LLM Gates for Agent Actions Fail Independently?
AI 导读
研究用 1,119 条已标注智能体动作测试运行时门控堆叠,发现任意两个 LLM 裁判组合仅相当于约 1.2 至 1.4 层乘法等效层,而规则层加一个裁判可达 1.86 至 2.09 层。
来源:arXiv cs.AI · arxiv.org