arXiv cs.CL· Tobias Hallmen, Fabian Deuser, Robin-Nico Kampa, Norbert Oswald, Elisabeth Andr\'e·· 9 小时前AI 评分31
细粒度情绪识别基准的失败在读取方式:VLM 感知能力被低估
The Failure Is in the Readout: Fine-Grained Emotion Recognition Benchmarks Measure Elicitation, Not Perception
AI 导读
研究发现 EmoNet-Face-HQ 细粒度情绪识别基准的失败源于答案读取方式而非模型感知能力:改用 logits 逐类二分类读取后,11 个开源 VLM 全部超过专家一致性锚点(κ_w=0.507-0.586),其中 3 个显著优于微调模型 EIF(Small κ_w=0.551,Large κ_w=0.534)。
来源:arXiv cs.CL · arxiv.org