Playbooks/September 3, 2026

How to Win "X vs Y" Comparison Prompts in AI Search

Robin Pautigny

Robin Pautigny

Co-founder, Refine

How to Win "X vs Y" Comparison Prompts in AI Search

Summary

Comparison prompts sit at the bottom of the AI funnel: the person asking has already decided to buy something and is choosing between named options. Models answer them by stitching together comparison pages, review-site tables, forum threads and their own parametric memory of your category. This playbook explains how that assembly works, the three positions you can occupy in a versus answer, how to build comparison content that survives extraction, and how to track your position over time instead of guessing.

The direct answer

To win an "X vs Y" prompt, you need three things in place: a comparison page that states differences as plain, attributable facts rather than marketing claims; consistent third-party corroboration on review sites and forums the model retrieves from; and a category framing where you are one of the two or three names that belong in the conversation. Models do not judge which product is better. They reproduce the consensus they can find, then hedge. Your job is to make the consensus accurate and easy to lift.

The Short Answer

A comparison prompt is any question that names two or more options and asks the assistant to choose between them, or asks for alternatives to a specific product. "Notion vs Coda for a small agency." "What is a good alternative to Zendesk?" "Is Datadog worth it compared to Grafana Cloud?" These are the prompts where an AI answer directly replaces a shortlist a buyer would otherwise have built themselves.

When a model answers one, it is not running a benchmark. It is retrieving whatever comparison material it can reach, blending that with what it already absorbed about your category during training, and producing a balanced-sounding summary with a recommendation hedged behind conditions. That process rewards a very specific kind of content, and it punishes the comparison pages most companies actually publish.

The practical consequence: you cannot win a versus prompt by arguing that you are better. You win it by being the source that makes the comparison easy to write, and by making sure the independent sources the model checks say roughly the same thing you do.

Why Comparison Prompts Are Worth More Than Any Other Prompt

Not all AI visibility is equal. A mention inside an answer to "what is generative engine optimization" reaches someone who is still reading. A mention inside "best GEO tool for a five-person marketing team, compared" reaches someone with a budget and a deadline. The second is worth an order of magnitude more, and it is also far more contested, because your competitors are named in the same sentence.

Comparison prompts also behave differently from informational ones in three ways worth planning around:

  • They are unusually stable. Definitional answers get rewritten as models improve; versus answers change mainly when the underlying sources change, which means work you do here compounds.
  • They lean harder on retrieval. Because the model is asked for specifics such as pricing, limits and integrations, it reaches for live sources more often than it does on conceptual questions.
  • They are winner-concentrated. Most versus answers name two to four products. There is no page two of an AI answer, so the difference between fourth and fifth place is the difference between existing and not.

This is why comparison coverage deserves its own tracked prompt set rather than being folded into a general visibility score. Averaging a strong position on definitional prompts with an absence on commercial ones produces a number that looks fine and tells you nothing.

Where the Model Actually Gets Its Comparison Data

Pull the sources panel on a few versus answers across ChatGPT, Perplexity, Gemini and Copilot and a consistent pattern shows up. Four layers feed almost every comparison answer:

  • Vendor comparison pages, including yours and your competitors. These are read, but discounted: the model knows a page hosted by one of the two products is not neutral, and it tends to lift only the factual rows.
  • Review marketplaces such as G2, Capterra and TrustRadius, which are structurally ideal for models because they already contain normalised attributes, ratings and pros-and-cons lists.
  • Community threads on Reddit, Hacker News, Stack Overflow and niche Slack or Discord archives, which supply the caveats and the "we switched because" stories that make an answer feel grounded.
  • Independent roundups and blog comparisons from consultants, agencies and adjacent tools, which often carry more weight than any of the above because they read as third-party analysis.

Layer one is the only one you fully control, and it is the one models trust least. That asymmetry is the whole game. Teams that spend their entire GEO budget rewriting their own versus page and none of it on the other three layers usually see their position stall, no matter how good the page becomes.

The Three Positions You Can Hold in a Versus Answer

Before you optimise anything, work out which position you currently hold for each comparison prompt, because the fix is different in each case.

Named as a principal is the position where the prompt itself contains your name, or the model treats you as one of the two default options in the category. Here your risk is not absence but misrepresentation: outdated pricing, a missing feature, a limitation you removed two releases ago. The work is factual correction, not promotion.

Named as an alternative is the position where you appear in the second half of the answer, after the two principals, usually in a sentence beginning "you might also consider". This is the most common position for challengers and the most improvable one. The lever is category framing: the model needs a reason to believe you belong in the same set, which usually means being listed alongside the principals somewhere it trusts.

Absent is the position where you never appear, even though you compete directly. Nine times out of ten the cause is not a weak comparison page but a missing entity association. The model has no reliable evidence that your product does the thing being compared, because nobody outside your own domain has ever written it down in a retrievable place.

Diagnosing this at scale

Doing this by hand across a dozen competitors and six engines is where most teams give up, because a single manual check is one sample of a probabilistic system. Refine runs your comparison prompts repeatedly across ChatGPT, Gemini, Perplexity, Claude, Copilot and Mistral, records which brands were named in each response and which sources were cited, and shows the drift over time. That turns "I think we lost ground on the versus queries" into a number you can put in a report, and it surfaces which specific third-party page moved the needle.

Build the Comparison Page a Model Can Actually Use

Most vendor comparison pages are written to close a deal in the last five percent of the funnel. They are heavy on persuasion, light on specifics, and they omit anything unflattering. A model reads that page, recognises it as promotional, and extracts almost nothing. The version that gets extracted looks quite different.

Lead with a plain, one-paragraph verdict that includes the condition under which the competitor is the better choice. Counterintuitively, this is the single highest-leverage change you can make. A page that says "choose them if you need X, choose us if you need Y" reads as neutral analysis, gets quoted more often, and is far more likely to survive as the framing the model reuses. A page that says you win on every axis gets treated as an ad.

Then make the facts liftable. Attributes in a real table with one claim per cell. Pricing with the number and the date you last verified it. Feature statements written as complete sentences that make sense out of context, because that is exactly how they will be quoted: "Refine tracks six AI engines including Mistral" survives extraction, while a checkmark in a column labelled "multi-engine" does not.

  • One page per pairing, named for the pairing, rather than one mega-page comparing you to eight tools at once.
  • A short direct-answer block at the top, ideally under sixty words, that answers the versus question outright.
  • Dated, verifiable specifics: prices, limits, supported integrations, availability, with a last-updated date on the page.
  • An honest section on who your product is not for, which is the passage independent reviewers quote most often.
  • Structured data where it genuinely applies, plus internal links from the pages that already rank so crawlers reach it.

Finally, refresh on a schedule. Comparison facts rot faster than any other content type, and a page with 2025 pricing actively teaches models something false about you that then persists in answers for months.

Seed the Third-Party Layer You Do Not Control

The other three source layers are influenced rather than controlled, which makes the work slower but the payoff more durable. The goal is simple to state: for every comparison you care about, there should be at least one credible page you do not own where both names appear and yours is described accurately.

On review marketplaces, completeness matters more than star count. Models lift attribute tables and category placement, so an unfinished profile with an outstanding rating loses to a fully completed profile with a good one. Make sure your category assignment matches how buyers phrase the comparison, and keep a steady trickle of recent reviews rather than a single burst, since recency is visible in what gets retrieved.

In communities, the only sustainable approach is genuine participation, ideally from people who actually use the product. Threads that answer a real comparison question honestly, including where you are the wrong choice, get retrieved and quoted for years. Astroturfed recommendation threads are increasingly easy for both moderators and models to discount, and a visible pattern of them can suppress your brand entirely.

For independent roundups, the practical move is to make yourself easy to include: a current public pricing page, a clear one-line category description you use consistently everywhere, screenshots and logos available without a form, and a factual brief you can send when someone asks. Most writers building a comparison roundup will use whatever is easiest to verify.

Measure It: The Comparison Scorecard

Track comparison prompts as their own cohort, separate from your general visibility number, and report four things per pairing per engine.

  • Inclusion rate: across repeated runs of the same prompt, what share of answers name you at all.
  • Position class: principal, alternative, or absent, which tells you which lever to pull next.
  • Factual accuracy: whether the claims made about you in the answer are correct, tracked as a simple pass or fail with the specific error noted.
  • Source mix: which pages the engine cited, and how many of them you own versus how many are third-party.

The fourth metric is the leading indicator. When the share of third-party sources in your comparison answers rises, inclusion and position almost always follow within a few weeks. When every cited source is your own domain, you are one model update away from disappearing.

Run the cohort weekly rather than daily. Comparison answers are stable enough that daily sampling mostly measures noise, and weekly gives you enough runs per prompt to distinguish a genuine position change from the ordinary variance of a probabilistic system. Then pick the single pairing with the largest gap between commercial value and current inclusion, fix the layer that is weakest, and re-measure. Comparison visibility is not won in a quarter-long campaign. It is won one pairing at a time.

Short on time? Have an assistant summarise this page for you.