跳到正文
arXiv cs.LG· Amir Rafe, Subasish Das·· 6 小时前AI 评分31

自动化决策门基准测试:System One 决策模型对比训练分类器与大语言模型

Benchmarking System One decision models against trained classifiers and language models for automated decision gates

AI 导读

一项基准测试将六个家族的八个 System One 决策模型(含托管模型 Jev)与四个开源生成模型置于同一语义请求下评测。有标签时微调 DeBERTa-v3-large 在几乎所有基准上准确率最高;无标签时 Jev 在流程与意图任务上无显著对手,Gemma-4-31B 在 CLINC-150 上超过它但成本与延迟更高。

来源:arXiv cs.LG · arxiv.org