Summary
A prompt universe is the fixed set of questions you track across AI engines to measure your brand visibility. Get it wrong and your GEO metrics are meaningless; get it right and you have a reliable read on how ChatGPT, Gemini, Perplexity and others describe you. This guide explains what a prompt universe is, the four prompt types every one needs, a step-by-step way to build yours, how many prompts you actually need, and how to keep the set current as your market shifts.
The short answer
A prompt universe is the curated list of questions you repeatedly ask AI engines to measure your visibility. Build it by starting from real buyer intent, covering four prompt types (category, comparison, use-case and branded), keeping it broad enough to be representative but focused enough to track consistently, and revising it on a schedule. The prompts you choose define what your GEO metrics actually mean, so treat this as the foundation of the whole program, not a throwaway first step.
The Short Answer
If you want to measure how visible your brand is in AI search, you first have to decide which questions count. That set of questions is your prompt universe. It is the single most important design decision in any GEO program, because every metric you report, mention rate, share of voice, sentiment, is calculated against these prompts. Choose prompts nobody actually asks and your dashboard will look impressive while telling you nothing about real demand.
The goal is a set of prompts that mirrors how real buyers in your category talk to ChatGPT, Gemini, Perplexity and other engines when they are researching a purchase. Not the questions you wish they asked, and not only the ones where you already win, but the honest spread of what your market types into an AI when it wants a recommendation. Get that representative sample right and everything downstream becomes trustworthy.
What a Prompt Universe Is and Why It Matters
In classic SEO you track keywords. In GEO you track prompts, and prompts are richer: they are full natural-language questions, often multi-part, and the same underlying intent can be phrased a dozen ways. A prompt universe is your deliberately chosen, stable collection of these questions, run repeatedly over time so you can measure trends rather than one-off answers.
It matters for one blunt reason: it is the denominator of your visibility. If your prompt universe is 20 questions and you appear in 8 of them, your mention rate is 40 percent. Add 20 softball branded questions you always win and the same brand suddenly reports 70 percent, with zero real improvement. The prompts define the score. That is why a vanity prompt set is worse than no tracking at all, because it manufactures false confidence. A good prompt universe is honest about the questions where you lose, because those are exactly the ones worth fixing.
The Four Types of Prompts Every Universe Needs
A representative universe is not just a pile of keywords with question marks. It should deliberately balance four intent types, because each reveals something different about your position in AI answers:
- Category prompts - open recommendation questions like "what is the best tool to track brand visibility in AI?" These are high-intent, highly contested, and the truest test of whether engines consider you a default option.
- Comparison prompts - head-to-head and shortlist questions like "Refine vs Profound" or "top alternatives to X". They reveal who you are bracketed with and whether you win direct match-ups.
- Use-case prompts - situation-driven questions like "how can a marketing team monitor AI mentions across ChatGPT and Gemini?" These surface whether engines connect your product to specific jobs to be done.
- Branded prompts - questions that name you directly, like "what is Refine and who is it for?" These check accuracy and sentiment: does the engine describe you correctly, and favourably?
Skipping a type blinds you to part of the picture. Only tracking branded prompts flatters you; only tracking category prompts hides whether engines even get your basics right. A healthy universe leans toward category and comparison prompts, because that is where recommendations are won or lost, while still covering use-case and branded questions to catch accuracy and positioning problems.
How to Build Your Prompt Universe Step by Step
You do not need a research team to build a solid first universe. A focused afternoon following a clear sequence gets you most of the way:
- Start from buyer intent, not keywords. List the real decisions your buyers make and the questions that precede them. Talk to sales, read support tickets, and skim the questions people actually type.
- Draft prompts in natural language. Write full questions the way a person speaks to an AI, not clipped keyword phrases. Include the messy, multi-part phrasings real users use.
- Cover the four types. Deliberately write category, comparison, use-case and branded prompts so no intent is missing.
- Add your competitors by name. Include comparison prompts against the specific rivals you care about, so you can benchmark share of voice against them.
- Vary the phrasing. Add two or three phrasings of your most important questions, since wording changes which brands surface and you want a stable read, not an artifact of one sentence.
- Prune the vanity. Cut prompts you only added because you win them. Keep the ones that reflect genuine demand, even the uncomfortable ones.
Once drafted, group the prompts by theme or funnel stage so you can slice your metrics later, seeing, for example, that you are strong on use-case questions but weak on open category recommendations. That grouping turns a flat list into a diagnostic tool.
Building your prompt universe in Refine
Refine is built around the prompt universe as the core object. You define the prompts that matter for your category, tag them by intent or funnel stage, and add the competitors you want to benchmark against. Refine then runs the whole set repeatedly across ChatGPT, Gemini, Perplexity, Claude, Copilot and Mistral, and reports mention rate, share of voice and sentiment per prompt and per engine over time. Because the set is fixed and re-run on a schedule, you get comparable trends instead of one-off screenshots, and you can see exactly which prompts and which engines are dragging your visibility down.
How Many Prompts Is Enough?
There is a real tension here. Too few prompts and your metrics are noisy and easy to skew; too many and the set becomes expensive to run and hard to keep meaningful. For most brands, a first universe of roughly 20 to 50 well-chosen prompts is the sweet spot: broad enough to be representative across the four types and your key competitors, small enough to run frequently and reason about.
Remember that each prompt should be run multiple times per engine, because AI answers are non-deterministic and a single response is closer to a coin flip than a measurement. So a 30-prompt universe across six engines with several runs each is already hundreds of data points per cycle. Scale the number of prompts to what you can run consistently, not to what looks comprehensive on a slide. It is far better to track 25 prompts reliably every week than 200 prompts once and never again.
Keeping Your Prompt Universe Alive
A prompt universe is not a set-and-forget artifact. Your category evolves, new competitors appear, buyers start asking new questions, and the way people phrase requests to AI shifts as the tools change. If you freeze your universe for a year, your metrics slowly stop reflecting reality. But you also cannot churn it constantly, because changing the prompts breaks the very trend comparison you built the universe to enable.
The balance is a stable core plus a reviewed edge. Keep the majority of your prompts fixed so your trend lines stay comparable quarter over quarter, and review the set on a schedule, say monthly, to add emerging questions and retire ones that no longer matter. When you do add or remove prompts, note the change so you can interpret any shift in your metrics as an editing effect rather than a real move. Treated this way, your prompt universe becomes a living map of how your market interrogates AI, and the reliable foundation every other GEO decision rests on.
Short on time? Have an assistant summarise this page for you.

