跳到正文
arXiv cs.AI· Laur\`ene Vaugrante, Thilo Hagendorff·· 9 小时前AI 评分31

LLM-as-a-Judge 的设计选择影响有多大?提示词、评分量表与模型的系统对比

How Much Do LLM-as-a-Judge Design Choices Matter? A Systematic Comparison of Prompt Designs, Rating Scales, and Models

AI 导读

研究系统评估了 10 个推理模型在 LLM-as-a-Judge 中的设计选择影响,涵盖句子情感与毒性评分(每类超 500 条)及问答对二元准确率分类(n=600)两项任务。

来源:arXiv cs.AI · arxiv.org