TLDR: A study by Dirk HR Spennemann investigated the prompt fidelity of ChatGPT4o / DALL-E3 text-to-image visualisations, finding that DALL-E3 incorrectly rendered 15.6% of explicitly specified attributes. Errors were lowest for paraphernalia, moderate for personal appearance (attire, glasses), and highest for personal attributes like age and hair. This research highlights reliability issues in AI’s ability to accurately interpret and visualize specific details from prompts, impacting bias detection and model evaluation.
In the rapidly evolving world of Artificial Intelligence, text-to-image models like ChatGPT4o and DALL-E3 have captivated imaginations by transforming written descriptions into visual art. However, a recent study delves into a crucial question: how faithfully do these AI models adhere to the specific details provided in a prompt? The findings reveal measurable gaps in what the AI is asked to create versus what it actually visualizes.
Authored by Dirk HR Spennemann, the research titled “Prompt fidelity of ChatGPT4o / Dall-E3 text-to-image visualisations” examines the accuracy with which DALL-E3 renders attributes explicitly specified in prompts generated by ChatGPT4o. This study is particularly relevant as AI-generated images become more prevalent, raising concerns about representational accuracy and potential biases.
The study utilized two public-domain datasets: 200 visualisations of women working in cultural and creative industries, and 230 visualisations of museum curators. For each image, the researchers assessed the accuracy of various attributes, including personal details like age and hair, appearance aspects such as attire and glasses, and paraphernalia like name tags and clipboards. The prompts for these images were initially generated by ChatGPT4o in response to general requests, ensuring a consistent starting point for the DALL-E3 image generation process.
The results indicate that while DALL-E3 generally performs well, it deviated from prompt specifications in 15.6% of all attributes examined (out of 710 instances). The types of errors varied significantly across different attribute categories. The AI demonstrated the highest fidelity for paraphernalia, such as name tags and clipboards, with the lowest error rates. Errors were moderate when it came to personal appearance attributes like attire and glasses. However, the most significant inaccuracies were observed in the depiction of the person themselves, particularly concerning age and hair, where the AI struggled most to render the specified details correctly.
These findings highlight a critical challenge for generative AI: even when prompts are explicitly detailed, the models do not always comply perfectly with every instruction. This selective adherence has important implications, especially for tasks requiring controlled representation. For instance, if an AI is asked to generate images of diverse age groups, and it consistently misrepresents age, it can perpetuate subtle stereotypes and distort demographic accuracy. This impacts efforts to audit AI models for bias and ensures transparency in their outputs.
Also Read:
- Assessing GPT-4o’s Ability to Detect Pneumonia from X-Rays
- Unveiling Bias: A Deep Dive into Race and Gender Representation in AI-Generated Occupational Personas
The research serves as a baseline for understanding the current state of prompt fidelity in text-to-image AI, especially given that the datasets were generated before the latest ChatGPT native image generation model was rolled out. It underscores the ongoing need for improvements in AI’s ability to interpret and accurately visualize complex human instructions, ensuring that the generated content truly reflects the user’s intent. For a deeper dive into the methodology and detailed results, you can read the full research paper here.


