跳到正文
arXiv cs.CV· Jiahao Meng, Yue Tan, Qi Xu, Haochen Wang, Zhongwei Ren, Weisong Liu, Yuhao Wang, Renrui Zhang, Xiangtai Li, Haodong Duan, Yunhai Tong, Ming-Hsuan Yang·· 4 小时前AI 评分30

VideoZeroBench:用时空证据验证探测视频 MLLM 的能力边界

VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification

AI 导读

研究团队推出 VideoZeroBench,一个覆盖 13 个视频领域、经两轮交叉验证的长视频问答基准,要求模型同时给出答案并定位时间区间与关键帧框。在 19 个受测模型中,Gemini-3.7-Flash 取得最高标准 QA 准确率 24.8%(Level-3),但答案正确且时空定位准确(Level-5)时无一模型超过 1.8%。代码与数据已公开。

来源:arXiv cs.CV · arxiv.org