TLDR: A study evaluated how well Vision Language Models (VLMs) like GPT-4o can simulate the visual perception of people with low vision. Researchers created a benchmark dataset from 40 low vision participants, collecting their vision details and image perception responses. They found that VLMs, when given minimal prompts, often infer beyond specified vision abilities, leading to low agreement with human responses. However, combining detailed vision information with a single example that includes both open-ended and multiple-choice responses significantly improved the simulation’s accuracy, reaching up to 70% agreement. The study concludes that while promising, VLMs are “Not There Yet” for fully accurate low vision simulation and should be used cautiously as a complementary tool.
Recent advancements in artificial intelligence, particularly in Vision Language Models (VLMs), have opened new avenues for simulating human behavior. These powerful models, like GPT-4o, are increasingly capable of complex reasoning and problem-solving. However, their application in the critical domain of accessibility, specifically in simulating the visual perception of individuals with low vision, has remained largely unexplored until now.
A new research paper titled “Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision” by Rosiana Natalie, Wenqian Xu, Ruei-Che Chang, Rada Mihalcea, and Anhong Guo from the University of Michigan delves into this very question. The study investigates the extent to which current VLMs can accurately interpret images from the perspective of someone with impaired vision.
The researchers compiled a unique benchmark dataset by conducting a survey with 40 participants who have low vision. This comprehensive survey collected detailed information about their vision, including visual acuity, field, and color perception, as well as how their vision has progressed. Crucially, participants also provided both open-ended descriptions and multiple-choice answers to up to 25 image perception and recognition tasks. This rich dataset allowed the researchers to understand the diverse range of visual abilities among individuals with low vision, with participants’ accuracy on multiple-choice questions ranging from 0% to 99%.
Using this collected data, the team constructed various prompts for VLMs (specifically GPT-4o) to create simulated agents for each participant. They experimented with different levels of information provided to the VLM, including minimal vision details, brief medical diagnoses, or comprehensive descriptions of the participant’s visual experience. They also varied the inclusion of example image responses, exploring single or multiple examples, and whether these examples included open-ended descriptions, multiple-choice answers, or both.
The evaluation focused on the agreement between the VLM-generated responses and the actual participants’ answers. The findings revealed some significant insights. When VLMs were given minimal or no specific vision information, they tended to infer beyond the specified visual limitations, often producing surprisingly detailed image descriptions that did not align with the actual low vision experience. This resulted in a low agreement score of 0.59 between the VLM agents and the participants.
Even when only vision information (diagnosis, brief, or detailed) was provided without examples, the agreement remained low at 0.59. This suggests that simply telling a VLM about a visual impairment isn’t enough to constrain its output to accurately reflect that impairment. However, a notable improvement was observed when example image responses were included in the prompts. The agreement scores increased significantly, reaching up to 0.67 with examples. The most effective approach was found to be a combination of both detailed vision information and example image responses, which boosted the agreement to 0.70.
Interestingly, the study found that a single example combining both open-ended descriptions and multiple-choice responses offered significant performance improvements over using either format alone. Providing additional examples beyond a single one yielded minimal further benefits, indicating the power of well-chosen, comprehensive examples.
While the 70% agreement score shows promising potential, the researchers acknowledge that VLMs are “Not There Yet” for standalone deployment in simulating low vision. They emphasize the ethical considerations of disability simulations, highlighting the risks of reinforcing stereotypes or misrepresenting lived experiences. The paper advocates for the responsible use of these simulations as complementary tools, not replacements for direct user research, and stresses the importance of human-in-the-loop validation.
Also Read:
- Assessing GPT-5’s Capabilities in Mammogram Interpretation
- Improving Video Quality Assessment with Integer-Only Loss Fine-tuning
Looking ahead, these simulated agents could have broader applications, such as personalizing assistive technologies like SeeingAI or Be My AI to better align with individual users’ unique vision needs. They could also be used for automated accessibility evaluations of visual media on platforms like social media, proactively identifying and addressing accessibility issues. The study also proposes a cost-effective and generalizable pipeline for collecting diverse vision data, which is crucial for developing more accurate and representative simulated agents. As AI models continue to advance, the potential for more reliable and effective low vision simulations grows, paving the way for more inclusive design processes and personalized user experiences. You can read the full paper here.


