Back to blog
SEO & GEO

How ChatGPT and Google AI Mode Pick Products

We ran 100 real shopping questions through ChatGPT and Google AI Mode, then audited all 138 products they recommended. The two agreed only 38% of the time.

Share
Summarize with

Ask ChatGPT and Google AI Mode the same shopping question and you will often get two different answers. We know, because we asked both engines the same 100 real shopping questions and they named the exact same top product only 38% of the time. A different brand entirely, just as often.

Then we opened every product either one recommended, 138 in total, and audited each page field by field. This report is what we found: what the pick actually runs on, and what it costs a brand when its own data is the reason an answer goes wrong.

Why we looked now

In July 2026, Google shipped its first reporting for AI Overviews and AI Mode visibility. The detail that made us pay attention was where it lives: inside Google Merchant Center, the product feed, not Search Console. Queries are grouped into three stages:

DiscoveryEvaluationPurchase

That the pick is now real enough for Google to build reporting for it, and that the reporting sits in the feed rather than in Search, is the backdrop for this study. We set out to see how the picks actually work, from the shopper's side: both engines, on the public consumer surfaces real shoppers use, all 138 recommended products opened and audited page by page.

The study, in six numbers

Everything in this report traces back to these six numbers. The sections that follow are the work behind each one.

100
real buying questions, run through both engines
138
products they recommended, every one opened and audited
38%
of the time the two engines named the same exact product
17
recommended products had zero machine-readable data on the page
3
confirmed cases of a quoted price matching no real listing
60%
average data completeness on the pages that had any at all

A note on method: both engines, the public consumer surfaces shoppers actually use, not pinned API calls. The full method is in the appendix at the end.

It is tempting to draw AI shopping as a straight pipe: get shortlisted, get read, get told to the shopper. It does not work that way. Whether you show up at all is decided by three layers of leverage, each built on a different clock; and whether what gets said about you is true is a separate question sitting underneath all three.

Owned · fast
Your feed
The biggest lever for landing in the shopping cards the engines fan out to, and you can fix it this quarter.
Owned · fast
Your product page
A primary source the AI actually reads: its price drawer and most of the offers it cites come straight off the page. It is also the surface that turns an AI-sent visitor into a sale.
Earned · slow
Your brand presence
Being known and talked about. This is what wins the picks a small brand cannot buy into this month. It compounds, and no one wins it in a quarter.

Two of those are fast and fully yours, your feed and your page. The third, brand presence, is the long game, and it is the honest reason a challenger will not out-rank the category default overnight. None of it is out of reach; it is just on different clocks. And most brands have not even claimed the fast two.

The second axis
Showing up is one question. Whether what the AI says about you is true, right price, in stock, actually buyable, is a completely separate one. It is the job most brands are failing, and it is the one we could measure directly.

The four findings below measure both axes: how stable the pick is, and how often the story told about the picked product is broken.

Finding 1: ask two AIs the same question, you get a coin toss

We asked both engines the same 100 questions. They named the exact same top product only 38% of the time. Just as often, they landed on a different brand entirely.

Both engines, same question
Same product 38 Same brand only 22 Different brand 40

What does the split look like on real questions?

Same product · 38 of 100
Stand mixer: both said KitchenAid Artisan. Rice cooker: both said the Zojirushi NS-ZCC10. Car mount: both said the iOttie Easy One Touch 5.
Same brand, different model · 22 of 100
They agree on the brand, then split on the SKU: Instax Mini 13 vs Mini 12, Theragun Prime Plus vs Prime, ghd Chronos vs Max. Same brand, different page, different price.
Different brand entirely · 40 of 100
Smart lock: Yale vs Schlage. Bike lock: LITELOK vs ABUS. Garment steamer: Conair vs Rowenta.

A category with one obvious flagship pulls both engines onto the same pick; a crowded one pulls them apart, and even a shared brand often resolves to two different models. The pick is not guaranteed stable for a single engine either: in a smaller repeat check, the same question sometimes came back with a different top answer. So 38% is a snapshot, not a fixed number a brand can climb toward, which is why tracking where you actually show up matters more than chasing a single agreement score.

Finding 2: 17 products got the pick with nothing a machine could read

We checked all 138 recommended products for machine-readable product data, then scored every page we could fetch and parse against a 15-point rubric: price, availability, GTIN, brand, ratings, reviews, shipping, returns, and more.

17
of the 138 recommended products carried zero on-page structured data
still
picked, some by both engines on the same product

We checked the obvious alternative explanation first, that our own tool simply failed to see data rendered by JavaScript. Every blocked or suspicious page was re-checked in a real browser before being counted as zero. One retailer's listings looked zero-schema on two products, then a third product in a different category from the same retailer turned out fully populated, so we correct the record rather than round it off: this is template-specific, not company-wide.

Something other than the page got these 17 products into the running. We checked: all 17 were listed in Google Shopping, the feed layer these engines are known to pull from, 15 of them through the brand's own storefront. The page was invisible to machines; the feed was not. On-page structured data is not the gate to getting picked; feed presence is the far stronger signal. Field completeness across the schema-bearing pages averaged 60%, closer to 50% once you count the zero pages as zero.

So your feed gets you the shortlist. Telling the truth once someone looks is a separate job. The next two findings show what happens when nobody does it.

Finding 3: three prices matched no real listing at all

Separate from missing data, some products carried data that contradicted itself. A gaming chair's schema said $1,099 and out of stock while the live page sold it for $579, in stock. One page ran two blocks for the same SKU, one in stock, one sold out. One brand's data claimed its listings stay valid until the year 3025.

A stranger pattern showed up inside the AI's own price-comparison drawer: numbers that matched no listing we could find, checked against each product's official page and schema. The error was not the merchant's, and it ran in both directions.

$10
Oura Ring 5, real price $399, quoted twice in two sessions
+51%
a $430 recovery device, overstated to $650
−23%
an $867 dumbbell set, understated to $668

We opened the AI's own price drawer on 11 products and asked a simpler question: was the seller it fronted even the cheapest row in its own list? On the 9 with a real multi-seller comparison, it was not in 6 of them, off by up to 31%. The honest read: a shopper reading the AI's answer and a shopper reading your own page can be told two different things by your own systems, at the same time.

Finding 4: the AI kept selling products nobody could buy

We also found confident recommendations for products a shopper could not actually purchase. The recommendation layer moves faster than the truth underneath it.

  • A top pick recommended with zero reviews, no listed price, and no working way to purchase it.
  • A robot mower recommended as a top pick, discontinued by its own manufacturer months earlier.
  • A charger recommended to US shoppers that is no longer sold in the US at all.
  • A recommended store so geo-blocked the product page would not even open from the shopper's own region.

Why does this happen? Availability and lifecycle signals are the weakest fields in the entire data layer. When they go stale, the shopper's dead end lands on your brand, not the AI's.

The two engines don't just disagree, they behave differently

Along the way, it became clear the two engines shop in visibly different ways.

Google AI Mode
Educates first, then recommends
Splits the answer into named subsections
Showed a labeled ad inside about half the answers
ChatGPT
Leads with one pick plus one alternative
Ends by asking a clarifying question
Organic today, ad pipe already firing

Where were the recommendations sourced from? We logged what each answer leaned on:

Editorial buying guides
72%
The brand's own page
58%
Forums & user reviews
19%

Earned authority gets you cited; your own page gets quoted alongside it in most answers. But here is the trap in measuring any of this by citations: the brand that gets recommended and the source that gets cited are usually not the same one. Ask for daily training running shoes and the answer is the ASICS Novablast 5, while the citation chips point to RunRepeat, a review publisher, not ASICS. Score visibility by citations and you would conclude ASICS is nowhere, which is absurd; it is the exact product the shopper is being sent to buy. Across nearly every query the pattern held: citations skewed to publishers, not to the recommended brand's own page. Citations are the lagging shadow of earned presence, not a scoreboard for it.

Where this study sits

The strongest independent work on AI shopping points the same way. Here is how this study relates to it.

  • Semrush and Peec AI traced ChatGPT's carousel back to Google Shopping: the pick sat in Google Shopping's top results roughly 75 to 83% of the time. We measured the layer they did not: what the recommended pages themselves carry.
  • Profound, across a month-long snapshot of sampled product offers, finds the opposite half of the same story: about 88% of what ChatGPT actually cites is still read off the product page, not the feed. That is why this report treats feed and page as two separate jobs.
  • Ahrefs tracked 1,885 pages that added schema and found no measurable lift in AI citations afterward. Our own 17 schema-less products, recommended anyway, is the same conclusion from the opposite direction.

What is new here: a field-level data audit of the exact products two AI engines recommended, the same-question agreement rate between them, verified cases of AI-quoted prices matching no real listing, and a catalog of recommended-but-unbuyable products.

For the content side of the same question, why AI-written pages tend to backfire in search rather than build the authority that gets you cited, see our companion piece on why you shouldn't publish AI blog posts.

What this means for your brand, and what to do about it

No single trick makes an engine pick you, and the biggest picks turn on brand presence built over years, not a schema tag shipped this week. That is the long clock, and it is real work. But on the fast clock the wins are concrete and fully yours: your feed decides whether you are even eligible, and your product page decides what the AI reads, repeats, and sells on your behalf. Right now, most pages are telling machines a broken, half-empty, or self-contradicting story.

The useful part: every failure in this report is a fixable data problem.

  • No machine-readable data (Finding 2) needs complete product schema published on the page.
  • Conflicting or duplicate markup (Finding 3) needs collapsing to one authoritative source.
  • Wrong prices and currencies (Finding 3) need aligning to what the page actually sells.
  • Out-of-stock and discontinued items (Finding 4) get disapproved in Google Shopping and skipped by AI, so buyability has to be watched.
  • Missing GTINs are how Google Shopping and AI match your products at all: the shortlist layer itself.

These are the fixable ones, and they are exactly what BrandyBee is built to close. Point it at your store and it scans against these gaps, scores your AI-readiness, and fixes them: it publishes complete, conflict-free product schema, rewrites thin product descriptions into copy AI engines actually read, and diagnoses your Google Merchant Center feed health. The point is not to game the pick; it is to make AI get your products right, so what these engines say about you is what is true.

See what your store is missing →

Appendix: how the study was run

  • Surfaces. The public chatgpt.com web interface and Google AI Mode, in desktop Chrome. Models as served to real shoppers, not pinned API versions.
  • Sessions. All captures ran logged out, in anonymous sessions: no account, no history, no personalization. Key reproductions, including the mispricing cases, were confirmed again in a separate fresh session.
  • Prompts. 100 natural "best [product] for [need]" questions across dozens of categories.
  • Audit. A 15-field rubric over each recommended product page's structured data, every fetchable page scored field by field. Plain fetches first, blocked or JavaScript-rendered pages re-verified in a real browser.
  • Feed check. Each of the 17 zero-schema products was looked up in Google Shopping to confirm feed presence, a public view of the merchant-feed layer, checked independently of the recommendation regime.
  • Stability. Paired engine-vs-engine runs across all 100 questions. A smaller set of identical prompts was re-run to gauge how much the top pick moves between asks; a fuller repeat-run pass is still open.