Metrisque
Research · The study
RESEARCH

Understanding came before the recommendation.

Rankings and mentions explained little on their own. How a business was understood, named, categorized and matched to the question, explained most of what we saw. This page states the scope of that work and its limits.

What we asked

Buyer questions in the words buyers use, put to ChatGPT, Gemini and Claude repeatedly rather than once, so a single unstable answer cannot decide a reading.

What we recorded

Whether a brand was named at all, the category the assistants filed it under, the territory the answer described, the pages that matched the question, and the sources cited.

What stood out

Being absent from an answer usually traced back to something earlier: the brand was not known, it was filed under the wrong category, or no page answered the question directly.

Where the study stops

It describes readings taken on dated runs against a fixed question set. It does not measure buyer behaviour, traffic or revenue, and it does not predict future answers.

What we measured

92%

92% of AI recommendations landed inside the meaning territory drawn before asking. Pre-registered study, memory channel (models answering with nothing open). Re-measured with live search enabled: the territory held while the specific names changed between runs.

1,916

1,916 products measured across three assistants (OpenAI, Google, Anthropic models).

104,198

104,198 published texts read in the analysis.

Sampling frame, runs and method

Sampling frame

1,916 products drawn from consumer and software categories against a fixed set of buyer questions written in the words buyers use. Questions and products were fixed before any answer was captured. A matched-twins subset of 507 near-identical product pairs (same question, same placement, one recommended and one passed over) carried the confirmatory text test.

Runs and dates

Pre-registered July 2026, with outcome capture in July 2026 and the live-search re-measurement run afterwards. Memory channel: fixed prompt, temperature 0, nothing open, two runs per model per question. Search channel: the same questions re-asked with live search enabled.

Analysis method

The measurement instrument (embedding model, meaning regions, thresholds) was frozen before any recommendation was captured. Recommendations were then tested for whether they fell inside the pre-drawn territory, and 104,198 retrieved published texts were scored on that frozen instrument and classified for polarity and name collision, with raw responses stored for audit. Paired comparisons used the Wilcoxon signed-rank test. Confirmatory and exploratory results are labelled separately.

Limitations

  • Answers vary between identical runs, so a single run is never treated as a reading.
  • The 92% is the memory channel and is not a claim about search-enabled percentages.
  • Results are category-dependent, so a figure from one category does not transfer to another.
  • Every figure describes dated runs on a fixed question set. It does not measure buyer behaviour, traffic or revenue, and it does not predict future answers.

Citations

Pre-registration: OSF registry, registered before predictor data collection and held under embargo until release. Published record and results addendum: DOI 10.5281/zenodo.21417361. Commercial interest declared: Metrisque sells measurement built on this instrument.

The measured territory behind each reading is described in the Ring 1 methodology. To see the shape of the output before signing up, read the sample report.