Abstract:Scene text understanding is a core capability of real-world vision-language tasks and has important application value in scenarios such as visual question answering and image captioning. With the rapid development of multimodal large models, artificial intelligence has achieved breakthrough progress in the field of scene text understanding. However, general multimodal large models still commonly suffer from problems including insufficient utilization of OCR information, OCR recognition errors, and semantic deviations in cross-modal understanding. Therefore, this paper proposes an OCR Prompting Retrieval-Augmented (OP-RA) method for general multimodal large models toward scene text understanding. The method first extracts textual information from images using an OCR system, models the features of OCR tokens from both visual and textual modalities, and associates cross-modal visual-linguistic features of OCR. On this basis, combined with task semantic information, it achieves task-aware OCR retrieval augmentation via similarity calculation, selects the most relevant Top-K key OCR texts from candidates, and inputs them into large models in the form of structured prompts, thereby improving the reasoning and understanding capabilities of the models. The proposed method requires no additional training and exhibits favorable plug-and-play properties. Experimental results on two typical scene text understanding tasks, TextCaps and TextVQA, demonstrate that the proposed method achieves consistent performance improvements across various large models including BLIP-2, LION, and InfiniteVL, verifying its effectiveness and good generalization ability.