spot_img
HomeResearch & DevelopmentAutomated Dataset Creation for Text-Based Visual Question Answering

Automated Dataset Creation for Text-Based Visual Question Answering

TLDR: A new AI pipeline, ‘Text-VQA Aug,’ automates the synthesis of large-scale datasets for Visual Question Answering tasks that involve text in images. By integrating advanced models for text spotting, object grounding, captioning, and question generation, it efficiently creates and validates question-answer pairs. The pipeline has successfully generated around 72,000 QA pairs from 44,000 images, significantly reducing the need for manual annotation and providing crucial data for training specialist Text-VQA models.

Creating large datasets for Visual Question Answering (VQA) tasks, especially those involving text within images (known as Text-VQA), has traditionally been a challenging and time-consuming process. It often requires extensive human annotation, which is both tedious and costly. However, with the rapid advancements in foundation models that can handle both vision and language, and the maturity of Optical Character Recognition (OCR) systems, there’s a growing need for automated solutions.

Researchers have introduced a novel, unsupervised, and training-free pipeline called “Text-VQA Aug” designed to automatically synthesize high-quality Question-Answer (QA) pairs from images containing scene text. This innovative pipeline harnesses the power of multiple large multimodal models and algorithms to streamline the creation and validation of Text-VQA datasets, addressing the critical shortage of task-specific pre-training data for specialist models.

The Text-VQA Aug pipeline integrates several key components. It begins with **text spotting**, which involves both detecting and recognizing text in an image. For this, the GLASS module is utilized. Following this, **region of interest (ROI) detection** is performed to identify local contexts, associating text occurrences with foreground and background objects. Kosmos-2, a large grounding model known for its zero-shot object grounding capabilities, plays a crucial role here by generating object crops.

Once text is spotted and objects are cropped, the pipeline moves to **caption generation**. LLaVA-R, an enhanced version of LLaVA specifically tuned for text-rich images, generates descriptive captions for these object crops, often incorporating the spotted text. Since an image can have multiple object crops, a **caption aggregation** step concatenates these local captions to create a comprehensive global description for the entire image.

The next crucial phase is **OCR-based answer selection**. An algorithm identifies potential answers from the spotted OCR tokens by checking their presence and sequence within the aggregated image description. These potential answers are then fed into a **question generation** module. The Intel Neural Chat 7B model, a fine-tuned large language model, generates contextually appropriate questions based on the image description and the selected OCR-based answers.

Finally, the generated QA pairs undergo a **post-processing and validation** step. The same large language model is used to verify the correctness and completeness of the synthesized answers for their respective questions, filtering out any hallucinated or imprecise pairs. Additionally, questions with undesirable lengths are also filtered to maintain dataset quality.

This comprehensive pipeline successfully synthesized approximately 72,000 Text-VQA QA pairs from 44,581 unique images sourced from the OpenImages V5 dataset. This significantly larger dataset, compared to existing human-annotated ones, provides a valuable resource for pre-training specialist Text-VQA models. The pipeline demonstrates the ability to generate diverse questions, with about 38.5% of them not directly containing OCR tokens but still requiring an understanding of the image’s holistic context.

Also Read:

The implications of this work are far-reaching. Beyond addressing the data scarcity problem for Text-VQA models, similar pipelines could be developed for various applications. These include assistive tools for visually impaired individuals, visual search functionalities in retail and e-commerce, enhanced interactive learning materials in education, analysis of text on medical devices in healthcare, and even license plate recognition in security and surveillance systems. This research highlights a scalable approach to leveraging large multimodal models for automated data synthesis, paving the way for more robust and capable AI systems. You can read the full research paper here.

Nikhil Patel
Nikhil Patelhttps://blogs.edgentiq.com
Nikhil Patel is a tech analyst and AI news reporter who brings a practitioner's perspective to every article. With prior experience working at an AI startup, he decodes the business mechanics behind product innovations, funding trends, and partnerships in the GenAI space. Nikhil's insights are sharp, forward-looking, and trusted by insiders and newcomers alike. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -