跳到正文
arXiv cs.CL· Clayton Cohn, Joyce Fonteles, Kirk Vanacore, Gianni Mazza, Candida Crawford, Tom Hooper, Gautam Biswas, Rene Kizilcec·· 10 小时前AI 评分26

跨模型 LLM 共识不等于有效:K-12 数学辅导对话中学生失败模式诊断研究

Agreement Is Not Validity: Cross-Model LLM Consensus in Diagnosing Student Failure Modes in K-12 Math Tutoring Dialogue

AI 导读

一项针对 K-12 数学辅导对话的探索性研究发现,LLM 对五种学生失败模式(不确定性、错误归因、算子选择、概念缺口、程序失误)的分类中,人机一致性仅为中等(kappa = .524-.597),而跨模型一致性显著更高(kappa = .755-.781;alpha = .769)。这表明跨模型共识可能制造正确性的假象,模型间的一致不能替代对推断构念有效性的独立证据。

来源:arXiv cs.CL · arxiv.org