arXiv cs.AI· Laur\`ene Vaugrante, Thilo Hagendorff·· 9 小时前AI 评分31
LLM-as-a-Judge 的设计选择影响有多大?提示词、评分量表与模型的系统对比
How Much Do LLM-as-a-Judge Design Choices Matter? A Systematic Comparison of Prompt Designs, Rating Scales, and Models
AI 导读
研究系统评估了 10 个推理模型在 LLM-as-a-Judge 中的设计选择影响,涵盖句子情感与毒性评分(每类超 500 条)及问答对二元准确率分类(n=600)两项任务。
来源:arXiv cs.AI · arxiv.org