面向场景文字理解的通用多模态大模型OCR提示检索增强方法
CSTR:
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:

国家自然科学基金(U23B2031); 校企联合项目(K160161028)


OCR Prompting Retrieval-augmented Method for General Multi-modal Large Models towards Scene Text Understanding
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    场景文字理解是真实世界视觉语言任务的核心能力,在视觉问答、图像描述等场景中具有重要应用价值.随着多模态大模型的快速发展,人工智能在场景文字理解领域取得了突破性的进展.然而,通用多模态大模型在该场景中依然普遍存在OCR信息利用不足、OCR识别错误和跨模态理解语义偏差等常见问题.因此,本文提出一种面向场景文字理解的通用多模态大模型OCR提示检索增强方法(OCR Prompting Retrieval-Augmented,OP-RA).该方法首先利用OCR系统提取图像中的文本信息,并分别从视觉模态与文本模态对OCR Token进行特征建模,关联跨模态的OCR视觉语言特征.在此基础上,结合任务语义信息,通过相似度计算实现任务感知的OCR检索增强,从候选中筛选出最相关的Top-K个关键OCR文本,并以结构化Prompt的形式输入大模型,继而提升模型的推理与理解能力.该方法无需额外训练,具备良好的即插即用特性.在TextCaps与TextVQA两个典型场景文字理解任务上的实验结果表明,所提出方法在BLIP-2,LION及InfiniteVL等多种大模型上均取得稳定性能提升,验证了其有效性与良好的泛化能力.

    Abstract:

    Scene text understanding is a core capability of real-world vision-language tasks and has important application value in scenarios such as visual question answering and image captioning. With the rapid development of multimodal large models, artificial intelligence has achieved breakthrough progress in the field of scene text understanding. However, general multimodal large models still commonly suffer from problems including insufficient utilization of OCR information, OCR recognition errors, and semantic deviations in cross-modal understanding. Therefore, this paper proposes an OCR Prompting Retrieval-Augmented (OP-RA) method for general multimodal large models toward scene text understanding. The method first extracts textual information from images using an OCR system, models the features of OCR tokens from both visual and textual modalities, and associates cross-modal visual-linguistic features of OCR. On this basis, combined with task semantic information, it achieves task-aware OCR retrieval augmentation via similarity calculation, selects the most relevant Top-K key OCR texts from candidates, and inputs them into large models in the form of structured prompts, thereby improving the reasoning and understanding capabilities of the models. The proposed method requires no additional training and exhibits favorable plug-and-play properties. Experimental results on two typical scene text understanding tasks, TextCaps and TextVQA, demonstrate that the proposed method achieves consistent performance improvements across various large models including BLIP-2, LION, and InfiniteVL, verifying its effectiveness and good generalization ability.

    参考文献
    相似文献
    引证文献
引用本文

宋子杰,葛乐乐,张耀,胡珍珍.面向场景文字理解的通用多模态大模型OCR提示检索增强方法.软件学报,2027,38(5):

复制
相关视频

分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-04-27
  • 最后修改日期:2026-06-18
  • 录用日期:
  • 在线发布日期: 2026-09-14
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62562563 传真:010-62562533 Email:jos@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号