arXiv cs.CV· Uddipan Basu Bir, Vincent Christlein, Andreas Maier, Mathias Zinnen·· 3 小时前AI 评分22
轻量级视觉语言模型如何实现文档 OCR 与结构化 JSON 提取
From Pixels to Structure: Lightweight Vision-Language Models for Document OCR and Structured JSON Extraction
AI 导读
一项对比研究测试了八款开源轻量级 VLM(最高 7B 参数)在三所大学遗产馆藏上的 OCR 到结构化提取表现,模型需从文档图像中提取文本并生成符合 schema 的 JSON。
来源:arXiv cs.CV · arxiv.org