Research/August 27, 2026

17,691 AI Answers: 20% of Brands Took Two Thirds of the Visibility

Robin Pautigny

Robin Pautigny

Co-founder, Refine

17,691 AI Answers: 20% of Brands Took Two Thirds of the Visibility

Summary

Between May 22 and August 26, 2026, we ran 17,691 prompt tests across 20 brands, 298 tracked prompts and five model versions. Aggregate mention rates look reasonable — Gemini 26.8%, Perplexity 21.1%, ChatGPT 16.9%. The per-brand medians tell a different story: 1.1% on Gemini, 0.7% on Perplexity, 4.2% on ChatGPT. On every engine, the top 20% of brands captured roughly two thirds of all visibility, and 36% of brand-engine pairs scored exactly zero. AI search is not a level playing field that rewards presence — it is a concentrated one that rewards being the established answer.

What We Measured

Most published numbers about AI search visibility come from one-off tests: someone asks ChatGPT twenty questions, counts the brands, and writes it up. That tells you what happened on a Tuesday afternoon. It does not tell you what a brand can expect over a quarter.

This is first-party measurement from our own product. Between May 22 and August 26, 2026, Refine ran tracked prompts daily against multiple answer engines and recorded, for each run, whether the customer's brand was named in the answer. Every number below comes from that dataset.

  • 17,691 prompt runs
  • 20 brands
  • 298 tracked prompts
  • 5 model versions across ChatGPT, Gemini and Perplexity
  • 97 days of continuous measurement

How mention rate is defined

Mention rate is the share of answers in which the brand is named at all — not its rank, not its sentiment, not whether it was recommended. It is the lowest bar a brand can clear: did the engine say your name. Everything that follows measures that single question.

The Averages Say One Thing. The Medians Say Another.

Aggregate every run together and the picture looks workable. Gemini named the tracked brand in 26.8% of answers, Perplexity in 21.1%, ChatGPT in 16.9%. Roughly one answer in five mentions you. A marketing team reading that would conclude they have a foothold.

Then compute the same figure per brand and take the median instead. On Gemini it drops from 26.8% to 1.1%. On Perplexity, from 21.1% to 0.7%. On ChatGPT, from 16.9% to 4.2%.

Aggregate vs median mention rate

Gemini: 26.8% aggregate, 1.1% median. Perplexity: 21.1% aggregate, 0.7% median. ChatGPT: 16.9% aggregate, 4.2% median. The gap is not noise — it is the shape of the distribution. A small number of brands are mentioned constantly, and they carry the average for everyone else.

This is the single most important thing in the dataset, and it is invisible in every benchmark that publishes averages alone. If you read "one in five answers mentions your brand" and your own tracking shows 1%, you are not underperforming a norm. You are the norm.

Two Thirds of Visibility Goes to the Top 20%

Ranking brands by mention rate within each engine and measuring what the leaders capture makes the concentration explicit. The top fifth of brands took 67% of all visibility on ChatGPT, 68% on Perplexity, 56% on Gemini, and 66% on the newer GPT-5 build. Across the whole dataset: 67%.

What is striking is not the concentration itself — most distributions have a long tail — but how consistent it is. Four engines, built by three companies, on different architectures and different retrieval stacks, converge on almost exactly the same split. Whatever makes a brand the default answer travels across models.

Share of total visibility captured by the top 20% of brands

Perplexity 68% · ChatGPT 67% · GPT-5 66% · Gemini 56% · All engines combined 67%.

The practical reading: these engines are not sampling the market. They are converging on a short list. And the short list is largely the same one across providers, which means a visibility problem on one engine is rarely just an engine problem.

A Third of Brand-Engine Pairs Score Exactly Zero

Of the 69 brand-engine pairs in the dataset with enough runs to measure, 25 returned a mention rate of exactly 0.0%. Not "low". Zero. Across an entire quarter of daily testing, on prompts the brand chose to track because they matter to its buyers, the engine never once said its name.

Widen the threshold slightly and 40% of pairs sit below 1%. At the other end, five pairs cleared 50%, topping out at 80.4%. The median across every pair is 2.4%.

A zero is more actionable than a low score, and worth separating in your own reporting. A brand at 3% is in the consideration set and losing on ranking. A brand at 0% is not in the retrieval set at all — the engine has nothing to cite. Those are different problems with different fixes, and averaging them together hides both.

Why Engines Disagree About the Same Brand

The per-brand data shows wide spreads across engines for the same company on the same prompts. One brand scored 78.2% on ChatGPT; another reached 80.4% on GPT-5 while sitting far lower elsewhere. Gemini ranged from 0% to 58.7% depending on the brand.

The mechanism is retrieval, not preference. These engines answer from what they can find and cite. Perplexity leans heavily on live search results. Gemini draws on Google's index. ChatGPT blends trained knowledge with browsing. A brand well covered in third-party listings and comparison pages surfaces on the search-heavy engines; a brand with strong owned content but no third-party coverage does better where trained knowledge dominates.

The diagnostic that follows from this

If your scores diverge sharply across engines, the gap is usually in your citation sources rather than your product or your positioning. Strong on ChatGPT and weak on Perplexity points at thin third-party coverage — you are not on the listicles and comparison pages that live search retrieves.

What This Means If You Are Not in the Top 20%

The concentration is the finding, so treat it as the strategy. Three things follow from the numbers.

  • Benchmark against the median, not the average. If a vendor tells you the industry norm is 20%, ask whether that is a mean. In this dataset the mean is roughly ten times the median on two engines out of three.
  • Separate zeros from low scores. A 0% is a retrieval problem: nothing citable exists for the engine to find. A 3% is a ranking problem: you are findable but not preferred. Only the second is fixed by better content on your own site.
  • Go where the citations already are. Because the same short list recurs across engines, the pages that produce it are largely shared — listings, comparison articles, community threads. Getting onto those pages moves several engines at once.

The uncomfortable implication of a 67% concentration is that publishing more content on your own domain, on its own, rarely moves a brand from the tail into the head. The brands in the top fifth are there because other sites talk about them.

Limits of This Data

Twenty brands is a real sample, not a large one. It is drawn from Refine customers, which skews toward B2B companies that already suspect they have an AI visibility problem — the median here is probably lower than the true market median, because brands that dominate their category rarely go looking for a tracking tool.

The prompts are customer-chosen, so they cluster on commercially valuable queries rather than sampling all of language. Run counts vary by engine: Gemini and Perplexity carry 5,898 runs each, ChatGPT 2,191, and one newer model only 132, which is why we have not broken that one out. And these are three months in one year of a fast-moving field; the same measurement in 2027 may look nothing like this.

All figures are aggregated across brands. No individual company, domain or prompt is identified anywhere in this article, and we have excluded any breakdown that would isolate a single customer.

What we would not hedge on is the shape. A 67% concentration that reproduces itself across four independently built engines is not a sampling artifact. It is how these systems behave.

Short on time? Have an assistant summarise this page for you.