AIToday
Large Language ModelsAI Safety & AlignmentTHE DECODERPublished: Aug 28, 2026, 04:00 JST2 min read

AI shopping agents are unreliable, study finds

AI shopping agents are unreliable, study finds

Key takeaway

  • AI shopping agents are unreliable, said a Wharton School study. Tested models changed picks by up to 99 percentage points.

  • Even source order or user memories altered choices.

  • No model stayed fully consistent.

3 Key Points

  1. What happened

    AI shopping agents tested by Wharton School researchers changed product picks by up to 99 percentage points when shown a single external source like Wirecutter. Even the order of sources or a user memory like "I love hiking!" shifted selections, sometimes toward pricier products.

  2. Why it matters

    The study suggests AI agents won't consistently buy the best or even optimal product for shoppers. Two users with the same query, or the same user on a different day, can get different recommendations with no visible reason. This unpredictability also means sellers will find it harder to optimize for AI shopping than traditional SEO, per the researchers.

  3. What to watch

    The most stable model, Gemini 3.5 Flash, still picked the objectively best product only 86% to 92% of the time. No model was fully consistent, so AI shopping agents are not yet ready to buy on your behalf.

Ask the AI about this article →

Context & Analysis

The study highlights how fragile AI shopping recommendations are. Even slight changes in context—like seeing a single external source or a short user memory—swung decisions widely. For example, a hiking memory pushed picks toward pricier watches despite a clearly superior $29.99 option, showing AI agents don't prioritize objective quality. Sellers face a new challenge: they can't predict which model a shopper uses, what it read, or how it processes information, making AI optimization harder than SEO. This research suggests that until models become more consistent, trusting them to buy on your behalf is risky.

FAQ

How were the AI shopping agents tested?
Researchers tested six models using the ACES simulator, showing each a screenshot of a product page with a fitness watch grid. They then varied external sources, their order, and user memory snippets to see if recommendations changed.
Which external source had the biggest influence on recommendations?
Wirecutter had the strongest pull, boosting the Fitbit Inspire 3 pick by up to 99 percentage points for Gemini 3.5 Flash and 90 for Claude Opus 4.8.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleChinese open-source AI gains ground with U.S. businesses