跳到正文
arXiv cs.CL· Longwei Cong, Sonja Hahn, Sebastian Gombert, Leon Camus, Fabian Zehner, Hendrik Drachsler, Ulf Kroehne·· 1 天前AI 评分22

超越分数对齐:LLM-as-a-Judge 残余评判难度的心理测量学分析

Evaluating LLM-as-a-Judge Beyond Score Alignment: A Psychometric Analysis of Residual Judging Difficulty

AI 导读

研究从心理测量学视角分析 LLM-as-a-Judge,对人工与 LLM 评分分别拟合 Many-Facet Rasch 模型,提出"残余难度"指标以衡量评判难度。在 SummEval 上测试 17 个开源权重 LLM 评委发现,潜在摘要质量的中等对齐并不意味着残余难度对齐,且这种不匹配强烈依赖评估维度:一致性维度偏向 LLM 更难,连贯性维度偏向人类更难。

来源:arXiv cs.CL · arxiv.org