Skip to content
AdsChatGPT now shows ads.Get early access

Tracking & Analytics/October 9, 2026

How Many Prompts Should You Track for GEO?A Practical Sizing Guide

Robin Pautigny

Robin Pautigny

Co-founder, Refine

How Many Prompts Should You Track for GEO? A Practical Sizing Guide

Summary

Most B2B brands get a trustworthy AI visibility signal from 30 to 100 well-chosen prompts, and larger catalogs or multi-market brands need 150 to 500. What matters more than raw count is coverage of buyer intent, repeated runs per prompt, and a set stable enough to compare month over month.

Quick answer

Track 30 to 100 prompts if you are a focused B2B or local brand, 100 to 250 if you sell several products or serve several segments, and up to 500 if you run multiple markets or a large catalog. Run each prompt several times per engine, and keep the core set fixed.

The Short Answer

There is no universal number, but there is a useful range. For most teams, 30 to 100 prompts is enough to see a real pattern in how AI engines talk about your brand. Below roughly 20 prompts, one lucky or unlucky answer can swing your headline metric by several points.

Above 500, most teams stop reading the data. The tracker becomes a spreadsheet nobody opens, and decisions slow down instead of speeding up.

  • Under 20 prompts: a sanity check, not a measurement.
  • 30 to 100 prompts: the sweet spot for a single product line or one market.
  • 100 to 250 prompts: multi-product SaaS, agencies tracking one client, brands with distinct segments.
  • 250 to 500 prompts: multi-country brands, marketplaces, large catalogs.

Why Prompt Count Changes Your Numbers

AI answers are probabilistic. The same question asked twice can name different brands, in a different order, with different sources. Your visibility score is an estimate, and like any estimate it gets tighter as the sample grows.

Here is the intuition. If you track 10 prompts and your brand appears in 4, a single answer changing moves your score by 10 points. With 100 prompts, the same single change moves it by 1 point. More prompts mean less noise per prompt, which means trends you can actually trust.

This is why a score that jumps from 40% to 52% on a tiny set is usually not a win. It is variance. On a larger, stable set, the same move is far more likely to be real.

How Refine handles this

Refine runs each tracked prompt repeatedly across ChatGPT, Gemini, Perplexity, Claude, Copilot and Mistral, then reports visibility as a rate rather than a single yes or no. You see the trend and the spread, so a one-off answer does not masquerade as progress.

A Sizing Framework by Brand Type

Start from how many distinct things a buyer could ask AI about you, not from a budget line. A useful rule: you want roughly 5 to 10 prompts per product, segment or use case you care about.

  • Early-stage startup, one product: 25 to 50 prompts. Focus on category, comparison and problem-led questions.
  • Growth-stage SaaS with 3 to 5 use cases: 80 to 150 prompts, grouped by use case.
  • Agency tracking a client: 50 to 120 per client, sharing a template across accounts.
  • Ecommerce or marketplace: 150 to 400, segmented by product category.
  • Multi-market brand: 40 to 80 per language, written natively, not machine-translated.

For multi-language tracking, resist the urge to translate your English list. People phrase questions differently in French, German or Spanish, and AI engines draw on different sources per language.

How to Allocate Prompts Across Intent

Raw count hides the real question: are you measuring the moments that influence revenue? Split your set by buyer intent, and weight it toward the bottom of the funnel.

  • Discovery (about 25%): broad category questions such as the best tools for a job.
  • Consideration (about 35%): comparisons, alternatives and shortlists.
  • Decision (about 25%): pricing, integrations, fit for a specific profile.
  • Brand and reputation (about 15%): what AI says when someone asks about you directly.

Branded prompts are cheap to win and easy to over-represent. Keep them to a minority, otherwise your average looks healthy while you are invisible on every unbranded question that brings new buyers.

Sampling Runs: The Other Half of Reliability

Prompt count is only one dimension. The second is how many times you run each prompt. A set of 50 prompts run 5 times each gives you 250 data points per engine per period, which is far stronger than 50 prompts run once.

If budget forces a tradeoff, a smaller prompt set with repeated runs usually beats a huge set sampled once. Repetition tames randomness, while breadth only adds coverage.

Rule of thumb

Aim for at least 3 runs per prompt per engine per reporting period, and judge movement only when it persists across two periods or exceeds the spread you normally see.

How to Grow Your Set Without Drowning

Start small and expand with evidence. A good growth path takes about a quarter.

  • Week 1: launch 30 to 50 prompts covering your top use cases and top 3 competitors.
  • Weeks 2 to 4: read the answers, note which sources AI cites, and spot missing question types.
  • Month 2: add 20 to 40 prompts where you were surprisingly absent or where competitors dominate.
  • Month 3: freeze a core set for trend lines, and keep a separate exploratory set that you rotate.

The frozen core is what lets you say visibility rose from one quarter to the next. The exploratory set is where you find new opportunities. Mixing the two destroys comparability.

Source real prompts wherever you can: sales calls, support tickets, search queries from Google Search Console, and the questions your customers literally type into chat tools.

Common Sizing Mistakes

  • Tracking only branded prompts, which flatters your score.
  • Rewriting prompts every month, which breaks the trend line.
  • Counting near-duplicates as separate prompts, which inflates the set without adding information.
  • Chasing 1,000 prompts before acting on the first 50.
  • Ignoring engine differences: a prompt that matters on ChatGPT may be irrelevant on Perplexity for your audience.

Finally, remember that tracking is a means, not the goal. The right number of prompts is the smallest set that reliably tells you what to fix next: which pages to improve, which sources to earn mentions on, and which competitors to watch.

Short on time? Have an assistant summarise this page for you.