spot_img
HomeResearch & DevelopmentBridging the Gap: How Visual-Language Models Perceive But Don't...

Bridging the Gap: How Visual-Language Models Perceive But Don’t Always Use Evidence

TLDR: This research paper investigates why Vision-Language Models (VLMs) sometimes provide incorrect answers even when the correct visual information is present in an image. The authors discovered a phenomenon called “seeing but not believing,” where deeper layers of VLMs correctly attend to relevant visual evidence, but this perception doesn’t translate into accurate final answers. They attribute this to textual bias and under-utilization of visual context. To address this, they propose Visual Evidence Augmentation (VEA), a training-free method that highlights these deep-layer evidence regions. VEA consistently improves VLM accuracy across various models and tasks by making the internally perceived visual cues explicit.

Vision-Language Models (VLMs) have made incredible strides in understanding and responding to queries that involve both images and text. From answering questions about a picture to describing complex scenes, these models often perform remarkably well. However, despite their advanced capabilities, VLMs sometimes stumble, providing incorrect answers even when the crucial visual information needed to answer correctly is clearly present in the image. This puzzling behavior has led researchers to ask a fundamental question: Do VLMs fail because they simply don’t ‘see’ the evidence, or because they ‘see’ it but don’t effectively ‘use’ it?

A recent research paper titled “SEEING BUT NOT BELIEVING: PROBING THE DISCONNECT BETWEEN VISUAL ATTENTION AND ANSWER CORRECTNESS IN VLMS” by Zhining Liu, Ziyi Chen, Hui Liu, Chen Luo, Xianfeng Tang, Suhang Wang, Joy Zeng, Zhenwei Dai, Zhan Shi, Tianxin Wei, Benoit Dumoulin, and Hanghang Tong delves deep into this very question. Their work systematically investigates the internal workings of VLMs to uncover the root cause of these failures.

The “Seeing But Not Believing” Phenomenon

The researchers found something quite surprising: VLMs often *do* perceive the correct visual evidence, even when their final answer is wrong. They term this phenomenon “seeing but not believing.” This means the problem isn’t always a lack of perception, but rather a failure to effectively leverage the perceived information during the reasoning and answer generation process.

Their detailed analysis of how attention shifts across different layers within a VLM revealed several key insights:

  • Layer-wise Transition: Shallow layers of the model tend to focus primarily on the text of the question. As the information moves deeper into the model, the focus gradually shifts towards the image. This mirrors how a human might first read a question and then look at a picture for an answer.
  • Deep-Layer Visual Grounding: The deeper layers of VLMs don’t just glance at the image; they act like a spotlight, concentrating their attention on very specific, localized regions that correspond to the key visual evidence needed to answer the question. They effectively filter out irrelevant visual clutter.
  • The Paradox: Most intriguingly, these deep layers often lock onto the correct visual evidence even when the model ultimately produces an incorrect answer. The model ‘sees’ the right information, but fails to ‘believe’ or integrate it into a factual response.

Why the Disconnect?

The paper suggests two main reasons for this “seeing but not believing” paradox:

  1. Textual Information Dominance: Many current VLMs are built with a large language model (LLM) as their backbone, paired with a comparatively smaller visual encoder. This architectural imbalance can lead to a strong bias towards linguistic signals. The model might place “blind faith in text,” sometimes even hallucinating answers based on common textual co-occurrences rather than grounding them in the visual input.
  2. Visual Context Under-utilization: Similar to issues observed in Retrieval-Augmented Generation (RAG) systems where models fail to fully exploit retrieved information, VLMs can under-utilize the visual context. If the image contains a lot of irrelevant information, the model might struggle to make full use of the salient evidence, even if it has identified it.

Introducing Visual Evidence Augmentation (VEA)

To address this critical gap, the researchers propose a novel, training-free method called Visual Evidence Augmentation (VEA). VEA works during the inference (answer generation) phase and doesn’t require any additional training of the VLM. Here’s how it functions:

VEA first identifies the specific “visual-grounding layers” within the VLM that are best at attributing attention to relevant evidence. It then extracts the attention patterns from these layers, denoises and smooths them to create a clear highlighting mask. This mask is then overlaid onto the original image, visually emphasizing the regions the model’s deep layers already deemed important. This augmented image is then fed back into the VLM, guiding it to better utilize the visual information.

Also Read:

Impactful Results

The experiments conducted across various VLM families, including LLaVA, Qwen, Gemma, and InternVL, and multiple visual question answering tasks, showed consistent and significant improvements in accuracy with VEA. The method proved particularly effective for smaller VLMs, helping to compensate for their weaker visual grounding capabilities, but also provided benefits for larger models. Furthermore, VEA demonstrated remarkable robustness against various types of visual noise and corruptions, maintaining accurate reasoning even when raw visual inputs were severely degraded.

In essence, the research highlights that VLMs already possess the internal capability to identify crucial visual evidence. The challenge lies in making these internal signals explicit and ensuring they are effectively carried forward into the reasoning and generation of answers. VEA offers a practical and effective way to bridge this gap between perception and reasoning, advancing our understanding and reliability of VLMs. You can read the full paper here: SEEING BUT NOT BELIEVING: PROBING THE DISCONNECT BETWEEN VISUAL ATTENTION AND ANSWER CORRECTNESS IN VLMS.

Nikhil Patel
Nikhil Patelhttps://blogs.edgentiq.com
Nikhil Patel is a tech analyst and AI news reporter who brings a practitioner's perspective to every article. With prior experience working at an AI startup, he decodes the business mechanics behind product innovations, funding trends, and partnerships in the GenAI space. Nikhil's insights are sharp, forward-looking, and trusted by insiders and newcomers alike. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -