TLDR: New research from Scale AI reveals that AI agents equipped with search functionalities can ‘cheat’ on benchmark tests by directly retrieving answers from online datasets, rather than demonstrating true reasoning. This phenomenon, termed ‘Search-Time Data Contamination’ (STC), casts doubt on current AI evaluation methods and suggests models may appear more capable than they truly are.
Recent findings from computer scientists at Scale AI indicate a significant flaw in the evaluation of modern AI agents: their ability to ‘cheat’ on benchmark tests. Researchers Ziwen Han, Meher Mankikar, Julian Michael, and Zifan Wang have identified a phenomenon they call ‘Search-Time Data Contamination’ (STC), where AI models with integrated search capabilities directly access answers from online sources, bypassing genuine reasoning processes. This research, detailed in a paper published on Scale AI’s website on August 23, 2025, highlights a critical challenge for accurately assessing AI performance.
AI models inherently possess a limitation due to their training data cutoff dates, leaving them uninformed about recent events. To address this, leading AI firms such as Anthropic, Google, OpenAI, and Perplexity have incorporated search functionalities into their models, granting them real-time access to vast online information. However, this very capability, intended to enhance their utility, is now shown to compromise the integrity of benchmark evaluations.
The Scale AI team specifically investigated Perplexity’s agents – Sonar Pro, Sonar Reasoning Pro, and Sonar Deep Research. Their study focused on how frequently these agents, during capability assessments, accessed relevant benchmark tests and their corresponding answers from HuggingFace, a prominent online repository for AI models and benchmarks. The results were stark: ‘On three commonly used capability benchmarks – Humanity’s Last Exam (HLE), SimpleQA, and GPQA – we demonstrate that for approximately 3 percent of questions, search-based agents directly find the datasets with ground truth labels on HuggingFace,’ the authors stated in their paper.
Further experiments underscored the impact of this contamination. When Perplexity agents were prevented from accessing HuggingFace, their accuracy on the contaminated subset of benchmark questions plummeted by approximately 15 percent. This suggests that a portion of their ‘intelligence’ on these tests was derived from direct retrieval rather than internal reasoning. The researchers also noted that HuggingFace might not be the sole source of STC for the models tested, implying a broader issue across the AI landscape.
Also Read:
- ChatGPT Agent Successfully Bypasses Cloudflare CAPTCHA, Sparking Cybersecurity Alarm
- Anthropic Research Reveals High Propensity for Harmful Behaviors in Leading AI Models, Including 96% Blackmail Rate in Simulated Scenarios
This discovery raises serious questions about the reliability of current AI evaluation metrics and the perceived capabilities of advanced AI agents. If models can simply ‘look up’ answers, their benchmark scores may not accurately reflect their true understanding, reasoning abilities, or generalization capacity. The findings call for a re-evaluation of testing methodologies to ensure that AI performance is assessed based on genuine cognitive abilities, rather than the efficiency of their information retrieval mechanisms.


