TLDR: A study involving over 12,500 participants and 287,000 image evaluations found that humans have a modest 62% success rate in distinguishing AI-generated images from real ones. Participants were most accurate with human portraits but struggled significantly with landscapes. Older AI models like GANs and the technique of inpainting were particularly effective at fooling people. The research underscores the urgent need for transparency tools like watermarks and robust AI detection, as AI tools vastly outperform human capabilities in this area.
As artificial intelligence continues to advance, particularly in generating realistic images, a critical question emerges: how well can human beings differentiate between images that are real and those created or modified by AI? A recent study, titled “How Good Are Humans at Detecting AI-Generated Images? Learnings from an Experiment”, delves into this very challenge, drawing insights from a large-scale online game.
The research, conducted by a team from Microsoft AI for Good Lab including Thomas Roca, Anthony Cintron Roman, Jehú Torres Vega, Marcelo Duarte, Pengce Wang, Kevin White, Amit Misra, and Juan Lavista Ferres, utilized data from an online game called “Real or Not Quiz.” This game presented participants with a randomized mix of real and AI-generated images, tasking them with identifying the authenticity of each. The sheer scale of the experiment is noteworthy, involving approximately 287,000 image evaluations from over 12,500 participants globally.
The findings reveal a modest overall success rate of just 62%, indicating that humans are only slightly better than chance at distinguishing AI-generated visual content. This suggests a significant challenge for the average person in discerning the origin of digital images. Interestingly, the study found that participants were most accurate when evaluating human portraits. However, their ability to correctly identify AI-generated images plummeted when faced with natural and urban landscapes.
The images used in the game were sourced from a collection of 350 copyright-free ‘real’ images and around 700 diffusion-based images generated using various prominent AI models such as DallE-3, Stable Diffusion-3, Stable Diffusion XL, Stable Diffusion XL inpaintings, Amazon Titan v1, and Midjourney v6. Additionally, GAN-based fake faces were included. The researchers aimed to provide a realistic panorama of AI images people are likely to encounter, rather than cherry-picking images designed specifically to fool users.
Which AI Generators Fooled People the Most?
The study observed that two specific generation systems had a success rate below 50% in fooling participants: Generative Adversarial Networks (GANs) depicting human faces and Inpaintings. While GANs are an older form of generative AI, they are known for producing high-quality images, especially faces, which often resemble ‘amateur’ photography. Inpainting, on the other hand, is a technique where a specific element within an existing picture is replaced with an AI-generated one. This method proved particularly deceptive because the majority of the image remains ‘real,’ with only a small, AI-generated element. This technique poses a significant risk for misinformation campaigns, as it can be used to subtly alter images by adding or replacing objects or people.
Conversely, images generated by DallE-3, Midjourney, and Stable Diffusion were found to be somewhat easier for participants to identify correctly. The researchers suggest this might be because these models are widely known, and the public may have become accustomed to recognizing their characteristic aesthetic or style, which often leans towards an ‘over-refined’ or ‘studio-quality’ look.
It’s also worth noting that some real images proved to be surprisingly difficult for participants to identify as authentic. These images often shared aesthetic qualities with AI-generated content, such as unusual lighting, colors, or scenes, making them tricky to evaluate.
Humans vs. AI Detection Tools
The research also drew a stark comparison between human detection capabilities and those of AI detection tools. The study highlights that not all AI-generated images will carry watermarks or digital signatures. In such cases, AI detection tools can be invaluable. The AI detection tool developed by the research team, for instance, boasts a success rate superior to 95% for both real and AI-generated images, significantly outperforming human participants. This consistency across image categories further underscores the reliability of these tools compared to human perception.
Also Read:
- Detecting Deepfakes: A New Approach Using Facial Movement Analysis
- Automating Suspect Sketching with Generative AI
The Path Forward: Transparency and Tools
The study concludes that human ability to distinguish AI-generated images from real ones is only slightly better than chance, even when images are not specifically designed to deceive. This suggests that most modern generative AI models can produce photorealistic images without obvious defects. The findings make a strong case for increased transparency in generative AI. The authors emphasize the critical need for content credentials and watermarks to inform the public about the nature of the media they consume. In situations where digital signatures are absent, robust AI detection tools are essential, as they are far more reliable than humans, though they too can make mistakes.
For more detailed information, you can read the full research paper here.


