spot_img
HomeResearch & DevelopmentUnpacking AI's View of the City: A Montreal Street...

Unpacking AI’s View of the City: A Montreal Street Scene Perception Study

TLDR: A new research paper introduces a benchmark to evaluate how Vision-Language Models (VLMs) align with human perception of urban scenes. Using 100 Montreal street images and 30 perception dimensions, the study found that VLMs perform better on objective, visually grounded properties (like vegetation) than on subjective appraisals (like overall impression). Model performance correlates with human agreement, and there’s a slight drop in accuracy on synthetic images. The findings highlight VLMs’ potential for factual urban audits but caution against relying on them for subjective interpretations, emphasizing the need for human-centered evaluation protocols and transparency.

Understanding how people perceive and interpret city scenes is crucial for effective urban design and planning. A new research paper delves into this fascinating area, specifically examining how well modern Vision-Language Models (VLMs) align with human judgments when looking at urban environments. The study introduces a novel benchmark designed to test these advanced AI models on urban perception.

The researchers, led by Rashid Mushkani from the UniversitĂ© de MontrĂ©al and Mila – Quebec AI Institute, curated a unique dataset of 100 street-level scenes from Montreal. This collection was evenly split between actual photographs and photorealistic synthetic images, providing a diverse and challenging set for evaluation. Twelve participants from various community groups contributed 230 annotation forms, assessing the images across 30 different dimensions. These dimensions covered a wide range of attributes, from objective physical characteristics like ‘Spatial Configuration’ and ‘Vegetation’ to more subjective impressions such as ‘Overall Impression’ and ‘Cultural Elements’. All French responses from participants were carefully normalized into English for consistency.

Seven prominent Vision-Language Models were put to the test in a zero-shot setup, meaning they received no specific training for this task. The models included claude-sonnet, openai-o4-mini, gpt-4.1, gemini-2.5-pro, grok-2-vision, qwen2.5-vl, and llama-4-maverick. A structured prompt and a deterministic parser were used to ensure a fair and reproducible evaluation. The performance of the models was measured using accuracy for single-choice items and Jaccard overlap for multi-label items, while human agreement was assessed using Krippendorff’s α and pairwise Jaccard.

Key Findings on AI’s Urban Perception

The study revealed several significant insights into how VLMs interpret urban scenes:

  • Objective vs. Subjective Alignment: Models showed stronger alignment with human judgments on visible, objective properties. Dimensions like ‘Spatial Configuration’, ‘Human Presence’, and ‘Vegetation’ were easier for the models to interpret, achieving higher scores. Conversely, subjective appraisals such as ‘Sustainability’, ‘Public Amenities’, and ‘Cultural Elements’ proved much more challenging for the AI.
  • Top Performers: Among the evaluated models, claude-sonnet emerged as the top system, achieving a macro score of 0.31 across all dimensions and a mean Jaccard overlap of 0.48 on multi-label items. openai-o4-mini followed closely.
  • Human Agreement Matters: A clear trend emerged: higher human agreement on a particular dimension correlated with better model scores. This suggests that models perform better where human judgments are more stable and consistent.
  • Synthetic Images: The use of photorealistic synthetic images resulted in a modest but consistent decrease in model scores compared to real photographs. This indicates that while synthetic data can be useful, there are still subtle differences that affect AI perception.
  • Distributional Mismatches: For subjective dimensions like ‘Overall Impression’, models didn’t just make random errors; they often exhibited different underlying priors than humans. For instance, models tended to overuse ‘Not applicable’ and under-produce labels like ‘Accessible’ or ‘Comfortable’ compared to human annotators.

Also Read:

Implications for Urban Planning and AI Research

The findings have important implications for both practical urban analysis and future AI research. For urban practitioners, current VLMs can be valuable tools for assisting with factual components of streetscape audits, such as identifying seating, vegetation, or space typology. They can pre-annotate images, streamlining the review process for human experts. However, the research cautions against relying on these models for subjective appraisals, where human nuance and context-dependent interpretations are paramount.

For researchers, the benchmark highlights the need for evaluation protocols that account for inter-annotator variability and report both accuracy and distributional fit, especially when dealing with subjective items. Future models could benefit from incorporating structured visual geometry for better spatial reasoning and learning to express calibrated uncertainty for subjective judgments. The study also emphasizes the importance of participatory co-design with local communities to align model outputs with situated values.

In conclusion, this benchmark provides a crucial step towards understanding the alignment between human and AI perception of urban scenes. It offers baseline measurements for contemporary VLMs and a transparent framework for future evaluations. The dataset, prompts, and evaluation harness are being released to support reproducible and uncertainty-aware research in participatory urban analysis. You can find the full research paper here: Do Vision–Language Models See Urban Scenes as People Do? An Urban Perception Benchmark.

Meera Iyer
Meera Iyerhttps://blogs.edgentiq.com
Meera Iyer is an AI news editor who blends journalistic rigor with storytelling elegance. Formerly a content strategist in a leading tech firm, Meera now tracks the pulse of India's Generative AI scene, from policy updates to academic breakthroughs. She's particularly focused on bringing nuanced, balanced perspectives to the fast-evolving world of AI-powered tools and media. You can reach her out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -