How AI Models Choose Products: What We Measured
A new question, a borrowed playbook
AI assistants now answer shopping questions directly. When a customer asks "what is the best matcha powder for lattes," the model names three to five products. Everyone else is invisible. A new industry has formed to help brands win these recommendations, and its advice is borrowed from search engine optimization: improve your Google rank, add structured markup, get into listicles.
We tested whether that advice works, and what actually drives the choice. We registered the study publicly before collecting the predictor data, wrote our pass and fail rules in advance, and committed to reporting the result in either direction. The study covers 46 shopping queries, 1,679 real products, and 104,198 published texts, measured against the actual recommendations of GPT-5 and Gemini 2.5.
Two gates, in order
A model recommends a product in two stages. First, the product must exist in the model's memory. This is a hard gate: in our sample, products unknown to both models were recommended zero times out of sixteen opportunities. Second, among products the model knows, the text published about the product decides the margin. The standard SEO playbook (Google rank, markup, listicles) showed no measurable effect at any stage.
Our original working theory held that the quantity of well-placed text was the primary driver. Our own registered test rejected that. Whether the model knows the product at all is the stronger force, and we report this because we committed to it in advance.
Two very different failures look identical
Brands cannot see why AI assistants ignore them. Two very different failures look identical from the outside. In one, the model has never absorbed enough text about the product to know it exists. In the other, the model knows the product but has it filed under the wrong meaning, or the surrounding text is too thin to win. The first is a discoverability problem. The second is an identity and evidence problem. The fixes are different, the costs are different, and no existing tool tells a brand which problem it has.
Everyone measures outputs; no one measures inputs
Every funded tool in this space monitors outputs: it counts how often a brand gets mentioned by AI models. None of them measures the input side, meaning what the model knows and where the product sits in meaning. Our earlier descriptive study (49 queries, roughly 1,960 products) measured the industry's own prescriptions directly. Google rank showed a correlation near zero with being recommended. 86 percent of measured winning products carried no schema markup. Listicle presence and crawlability were null. The playbook is imported from search and it does not transfer.
What we measured
1. Discoverability: the model has to know you exist
We probe each model directly: describe this product. A substantive answer means the product exists in the model's memory. A blank or invented answer means it does not.
The result is close to absolute. Zero of 16 well-placed products unknown to both models were ever recommended. For Gemini alone, zero of 31 products it failed to recognize were ever picked by it. In our sample of established retail products, only 24 of 1,679 were unknown to both models, so this gate mostly affects small and new brands. Where it applies, nothing else matters.
well-placed products unknown to both models that were ever recommended.
products Gemini didn't recognize that Gemini ever picked.
products in our real-retail sample unknown to both models.
What builds existence is accumulated published text over time. Fame, for a language model, is text with compound interest. Years of coverage everywhere outweigh weeks of coverage in the right place. This is the registered result that overturned our own theory: across all well-placed products, whether the model knows the product correlates with being recommended at r = 0.18, versus r = 0.08 for the amount of supporting on-topic text, and r = negative 0.07 for raw mention volume.
Familiarity beat density.
Fame r = +0.18, density r = +0.08, volume r = −0.07. Our pre-registered falsification rule fired: density is not the primary mechanism. We report this because we committed to before collecting the data.
products carrying schema markup were LESS likely to exist in model memory (r = −0.17; 86.8% known with schema vs 97.9% without). Structured markup is a machine-readability standard; models do not read the web at recommendation time, they remember it.
2. Product identity: what the model thinks you are
Existing is necessary. Being remembered as the right thing comes next. We measure this by taking the model's own stored description of a product and placing it on the meaning map for the query.
Second, the models' remembered descriptions of picked products sit measurably closer to the query meaning than those of unpicked products. This was strong in our original word-level analysis and remains present for GPT-5 on our frozen instrument (p = .048), though absent for Gemini. Third, identity lives in sentences rather than labels. Having the category keyword in the product title showed a correlation of 0.01 with being picked. Title informativeness measured negative 0.04. Models absorb descriptions of what a product is and does. They do not read titles the way search engines do.
Category term in the title: r = +0.01. Title informativeness: r = −0.04. GPT-5 memory-position effect: p = .048; absent for Gemini.
3. Recommendation: what decides the pick among the known
Among products the model knows, placed inside the right meaning, three measured forces separate winners from losers.
Supporting text wins by a small, real margin. In 507 matched pairs of near-identical products (same query, same placement, one picked and one passed over), the picked product carried more supporting on-topic published text (Wilcoxon p = .002). The median edge is a single supporting text, and the denser product won 54 percent of pairs. This is a tiebreaker, and it is the only text-side lever that passed a registered test.
Matched twins test.
507 pairs, same query, same placement. Picked twin carried more supporting on-meaning text. Wilcoxon p = .002, one-tailed. Median edge: one snippet; denser winner in 54% of pairs.
An earlier estimate that a product's own website copy predicts being passed over (r ≈ −0.25) failed two independent reproduction attempts under stricter definitions. We withdraw it rather than defend it. No claim about own-site text direction is made in this study.
The margin varies sharply by category. Text placement's pull on picks measures r = 0.22 in food and beverage, 0.17 in footwear, 0.15 in beauty, and 0.13 in electronics and fitness, against roughly zero in apparel, baby, and home and kitchen. In contested categories, text visibly separates winners. In others, the game currently sits elsewhere, likely in brand memory alone.
small-brand products inside the meaning region that were ever recommended in the memory channel. Occupying an uncrowded gap did not help: pick rates were zero in crowded and uncrowded neighborhoods alike. In memory, there is no wedge. Whether the search channel opens one is the registered question of our next study.
The three rings have different residents. Close to the query's meaning, publications do the talking: the innermost ring is 55% editorial, and forums sit closest of all (median closeness 0.556). Far from the meaning, social takes over: the outer ring is nearly two-thirds Reddit and YouTube (37% + 26%). Community platforms are loud, but mostly off-topic.
Same matched-twins test, opposite direction: in the 507 pairs, the picked product's neighborhood was less negative by 3.9 percentage points (24.8% vs 29.0% negative; Wilcoxon p = .0002). The model does not only count the praise; the arguments against you weigh in.
The winners' text edge is spread thin and wide: a little more editorial (10.8 vs 10.0), a little more Reddit (2.7 vs 2.4), a little more YouTube (3.7 vs 3.4), a little more retail (2.5 vs 2.3), each p < .05, all at once. Nothing here supports "just get on X." The pattern is broad, on-topic, positive coverage, everywhere or nowhere.
Forums sit closest to what shoppers mean (median closeness 0.556): people in a car-seat forum talk in car-seat words. Retail 0.515, editorial 0.509. Reddit (0.437) and YouTube (0.457) sit farthest, plentiful, conversational, off-meaning.
One query, two neighborhoods
Query: "best convertible car seat for infants". Each dot is a real text snippet we retrieved about a product. Distance from center is the snippet's semantic closeness to the query. Amber dots mention the query's meaning vocabulary (safety, install, rear-facing, infant, newborn, comfort, budget, fit, insert, weight limit, pediatrician). Hover or click a dot to read it.
Study corpus, July 2026. Hover cross-highlights the same snippet in both the ring and the list.
4. Gaps and overlaps between the models
The two models agree on where winners come from and disagree on who wins. Their actual picks overlap by only about 5 percent, yet 92 percent of both models' picks fall inside the same pre-drawn meaning regions.
Their gates differ in strictness. Gemini's memory gate is airtight in our data: it never once picked a product it could not describe. GPT-5 is looser: it recommended 54 of 192 products that our describe-it probe scored as unknown to it, which suggests GPT-5 can recognize a product in context without being able to describe it cold. GPT-5 also responds more to text placement (r = 0.06 among known in-region products, versus no measurable signal for Gemini), and only GPT-5 shows the memory-position effect on our instrument. The practical consequence for a brand: existence and identity must be checked per model, because being visible to one says little about the other.
overlap between GPT-5's and Gemini's picks. They disagree on WHO wins. They agree on WHERE winners come from.
products GPT-5 recommended despite our probe scoring them as unknown to it. Gemini: never once.
5. What failed the measurement
Google rank: r near zero. Schema markup: absent for 86 percent of measured winners. Listicle presence and crawlability: null. Raw volume of mentions: slightly negative. Raw social result counts on Reddit and YouTube: null. Keyword in the title: null. Each of these is a measured result on real recommendation outcomes, and each is currently sold as an AI visibility tactic.
r ≈ −0.00 on recommendation.
86% of measured winners had none.
null.
r = −0.07. Slightly negative.
null.
category term in the title: r = +0.01. Title informativeness: r = −0.04.
median text age around a product: picked 377 days vs not picked 368 (p = .46; matched pairs p = .71). Newer text did not win.
mean result position 4.84 vs 4.86 (paired p = .36). We already knew your own rank was null; the rank of your coverage is null too.
organic vs video vs answer-box shares: no difference (~98% organic for everyone).
With these, every Google-presentation variable we measured is null on the memory channel. The model remembers the words, not the search page.
A two-stage process, and one wrong playbook
Recommendation by a model is a two-stage process. Discoverability comes first and is close to binary: the model either holds the product in memory or it does not, and no on-page tactic substitutes for the years of accumulated third-party text that builds that memory. Identity and evidence come second: among known products, being remembered as the right thing, in the right meaning neighborhood, with third-party text nearby, decides the margin, and the size of that margin depends on the category.
The industry's current playbook optimizes none of this. It optimizes search signals that our measurements show do not transfer.
For a brand, the actionable sequence is: measure which gate you are standing at, per model. If the models do not know you, the work is earning durable third-party coverage, and on-page tactics are premature. If the models know you but file you wrongly or thinly, the work is meaning-dense description and earned text near your target query. These are different investments, and mistaking one problem for the other wastes the budget.
What's nextFrom the raw search archive behind this study, we extracted 30,464 People-Also-Ask entries, 11,737 unique questions humans actually asked across our 46 categories ("Is Anker still a good brand?", "What brand has the widest toe box?"). This corpus seeds our next registered study: whether the search channel opens doors that memory keeps shut.
Both gates, one card
A short glossary
n, how many products or pairs a number is based on. Bigger n means a steadier number. We report n next to every statistic.
r (correlation), how strongly two things move together, from −1 to +1. Zero means no relationship. In this field, r above 0.25 is a strong signal, 0.10 to 0.25 is weak but real, below 0.10 is noise. Example: fame at +0.18 is a weak-but-real force; Google rank at −0.00 is nothing.
p (p-value), the chance of seeing a result this strong if there were actually no effect. Below .05 means the pattern is very unlikely to be luck. Our matched-twins result, p = .002, means about a 1-in-500 chance of luck.
Wilcoxon test, a comparison method for paired items that ranks differences instead of averaging them, so a few extreme products cannot distort the verdict. We used it for the 507 matched twins because text counts are lumpy.
Confirmatory vs exploratory, confirmatory results answer questions we registered publicly before collecting the data, with pass and fail rules written in advance. Exploratory results are honest patterns we found afterward; they guide the next registered study but carry no proof weight yet.
How the study was run
Pre-registered at OSF before predictor data collection (embargoed). Measurement instrument (embedding model, meaning regions, thresholds) frozen before any recommendation was captured. Outcomes: spontaneous recommendations from GPT-5 and Gemini 2.5, fixed prompt, temperature 0, no browsing, two runs per model. Predictors: 104,198 search-retrieved texts per fixed collection rules, scored on the frozen instrument, classified for polarity and name collision at temperature 0 with raw responses stored for audit. Confirmatory and exploratory results are labeled throughout. Commercial interest declared: Metrisque sells measurement based on this instrument. Adverse outcomes were registered in advance and are reported above. The full results addendum is published and citable: DOI 10.5281/zenodo.21417361.