TLDR: A new research paper introduces ACES, a sandbox to study how AI shopping agents make purchasing decisions. It reveals that while agents excel at instruction following, they can make irrational choices regarding price and ratings. Their selections are heavily influenced by product position and platform endorsements (like ‘Overall Pick’ tags), but negatively by ‘Sponsored’ tags. The study also shows that AI-optimized product descriptions can significantly boost sales, indicating a new frontier for seller strategies in an AI-mediated e-commerce landscape. The findings highlight the need for platforms, sellers, and regulators to adapt to these evolving AI buyer behaviors.
The world of online shopping is on the cusp of a major transformation, moving from human-driven browsing to autonomous AI agents making purchasing decisions on behalf of consumers. A new research paper titled “What Is Your AI Agent Buying? Evaluation, Implications, and Emerging Questions for Agentic E-Commerce” delves into this fundamental shift, exploring how these AI agents behave, what influences their choices, and the broader implications for e-commerce.
Authored by Amine Allouah, Omar Besbes, Josu´e D Figueroa, Yash Kanoria, and Akshit Kumar, the paper introduces a novel sandbox environment called ACES (Agentic e-CommercE Simulator). This innovative setup allows researchers to rigorously study AI shopping agents by pairing a vision-language model (VLM) agent with a fully customizable mock marketplace. This environment enables precise control over various factors like product positions, prices, ratings, reviews, sponsored tags, and platform endorsements, providing a clear view of how these elements causally influence AI agent purchasing decisions.
The research outlines a simplified three-step interaction for the AI agent: Veni (opening the browser and loading the mock-app), Vidi (searching for a product and capturing a screenshot of the catalog), and Emi (querying the VLM with the screenshot and user prompt to declare a selection). This streamlined approach, focusing on the critical choice step, allows for clean measurement of agent behavior without the complexities of a full end-to-end shopping journey.
One of the key findings relates to the basic rationality and instruction-following capabilities of these agents. While the latest models, such as Claude Sonnet 4, GPT-4.1, and Gemini 2.0/2.5 Flash, demonstrate near-perfect instruction following (e.g., choosing a product within a budget or of a specific color), they still exhibit surprising failures in basic economic rationality tests. For instance, some advanced models failed to select the lowest-priced or highest-rated product, even when all other attributes were identical. These failures were more common when differences were subtle, suggesting perceptual limitations or an inability to consistently prioritize optimal choices. Reasons for these failures included treating products as identical, making sub-optimal choices without justification, or dismissing optimal options by attributing differences to “display errors” or deeming them insignificant.
The study also sheds light on how AI agents react to common e-commerce levers. Product ranking and position on a page significantly influence selection, but the specific biases vary widely across different AI models. All models favor the top row, but their preferences for columns differ, undermining the idea of a universal “top” spot. For example, moving a product from the bottom right to a favorable top-row position could increase its selection rate five-fold for some models. Badge effects are also notable: sponsored tags tend to reduce selection likelihood, implying agents discount advertising, while platform endorsements like “Overall Pick” tags lead to a substantial increase in selection, indicating they are perceived as credible signals. In terms of product attributes, AI agents show human-like directional preferences, favoring cheaper, better-rated, and more-reviewed products, though their sensitivity to these factors varies significantly.
Perhaps one of the most intriguing aspects of the research is the exploration of seller-side AI agents. The study demonstrates that an AI agent can recommend minor tweaks to product descriptions that, in some cases, lead to dramatic increases in market share for the focal product. This suggests a future where sellers might use AI to optimize their listings specifically for AI buyer preferences, akin to search engine optimization (SEO) but for AI agents. These interventions, even single-pass ones, can materially shift selection shares, highlighting a new strategic dynamic in e-commerce.
The implications of these findings are far-reaching. For platforms, it means adapting layouts and ranking systems to account for AI-specific biases and exploring new monetization strategies beyond traditional advertising, such as offering listing optimization services to sellers. Brands and sellers will need to continuously monitor and adapt their product listings as AI models evolve, potentially leading to the emergence of new AI agent companies specializing in seller-side optimization. Consumers, while benefiting from reduced search friction, must be aware of the potential for irrational purchases by their AI assistants and the need for agents that better align with their preferences. Finally, the paper calls for AI shopping agent developers and regulators to establish standard evaluations and potentially mandate disclosures of model updates, given that even minor changes in an AI model can act as exogenous demand shocks, significantly altering market shares and biases.
Also Read:
- EU Digital Fairness Act: A Call for Regulation of Autonomous AI Agents
- Do AI Models Care About Threats or Rewards? A Deep Dive into Prompting Effectiveness
This research provides a crucial early look into the behaviors of AI shopping agents and the complex ecosystem they will create. For more in-depth information, you can read the full research paper here.


