跳到正文
arXiv cs.CV· Uddipan Basu Bir, Vincent Christlein, Andreas Maier, Mathias Zinnen·· 3 小时前AI 评分22

轻量级视觉语言模型如何实现文档 OCR 与结构化 JSON 提取

From Pixels to Structure: Lightweight Vision-Language Models for Document OCR and Structured JSON Extraction

AI 导读

一项对比研究测试了八款开源轻量级 VLM(最高 7B 参数)在三所大学遗产馆藏上的 OCR 到结构化提取表现,模型需从文档图像中提取文本并生成符合 schema 的 JSON。

来源:arXiv cs.CV · arxiv.org