Tracking & Analytics/August 16, 2026

How to Set Up an AI Visibility Dashboard Your Team Will Actually Use

Robin Pautigny

Robin Pautigny

Co-founder, Refine

How to Set Up an AI Visibility Dashboard Your Team Will Actually Use

Summary

An AI visibility dashboard earns its place only if it answers three questions at a glance: are we present, are we ahead of named competitors, and what changed. This guide covers the prompt universe to build it on, the five metrics worth a tile, the sampling cadence that separates signal from noise, and the layout and weekly ritual that keep the thing in use past week three.

A useful AI visibility dashboard answers three questions in under ten seconds: are we present in the answers our buyers actually see, are we ahead of the competitors we actually lose deals to, and what changed since last week. If a view cannot do that at a glance, it is a report, and reports get opened once.

Most teams build the report. They start from whatever the tool can export, mention counts, engine breakdowns, a wall of prompt-level rows, and end up with something impressive in a launch meeting and abandoned by the third week. The dashboards that survive are narrower, slower to change, and built around a decision someone actually has to make on Monday morning.

The ten-second test

Before you add any tile, ask what decision it changes. If a number can go up or down without anyone doing anything differently, it belongs in an appendix, not on the dashboard. Applied honestly, this test removes about half of what most teams plan to build.

What an AI Visibility Dashboard Actually Measures

Rank tracking measures a stable, shared artifact: broadly the same results page for everyone at a given moment. AI answers have no equivalent. Ask the same question twice and you can get two different shortlists, in a different order, citing different sources. The unit of measurement is therefore not a position but a distribution: how often, across many runs, your brand appears, and how it is framed when it does.

That distinction changes what a dashboard should show. Three observable layers are worth tracking, and they matter in roughly this order.

  • Presence: whether your brand appears at all in the answer to a given prompt, on a given engine, in a given run. This is the base layer and the first thing that moves when you fix something.
  • Framing: how you are described when you do appear. A brand named as the recommended choice in the second sentence and a brand listed in a closing "other options include" line score identically on presence and mean entirely different things commercially.
  • Sources: which URLs the engine cited to build the answer. This is the only layer that tells you where to act, because it points at the specific pages, review profiles and community threads the model is reading about your category.

Anything you cannot map back to one of those three layers is probably a vanity number.

Start With the Prompt Universe, Not the Metrics

The biggest determinant of whether a dashboard gets trusted is not metric design. It is whether the prompts behind it sound like your buyers. A dashboard built on internal product vocabulary will report improvements nobody in sales recognises, and it will quietly lose credibility the first time it disagrees with what a rep heard on a call.

Aim for twenty-five to sixty prompts, written in the language customers use on discovery calls, and grouped into families so the dashboard can be read by intent rather than as one undifferentiated average.

  • Category discovery: "what tools track brand visibility in AI search", "how do I monitor ChatGPT mentions". High volume, low intent, but this is where models form their default shortlist.
  • Comparison and alternatives: "X vs Y", "alternatives to X". Lower volume, highest commercial value, and the family where absence is most expensive.
  • Use-case and fit: "best tool for a five-person marketing team", "AI visibility tracking for agencies". These decide whether you get recommended to the segment you actually serve.
  • Brand and reputation: "is X any good", "X pricing", "X reviews". This is where factual drift surfaces first, and where a wrong answer does the most damage.

Then freeze the list for at least a quarter. A prompt universe that changes every month produces a trend line that measures your editing habits rather than your visibility.

The Five Metrics That Deserve a Tile

Five numbers cover almost every decision a marketing team makes about AI search. A sixth usually costs more attention than it returns.

  • Presence rate: the percentage of runs, across all prompts and engines, where your brand appears. One headline figure, plus a breakdown by prompt family. This is the number you quote to the executive team.
  • Share of voice: your mentions as a proportion of all brand mentions in the same answers, measured against a fixed list of three to five named competitors. Presence tells you whether you exist; share of voice tells you whether you are winning.
  • Position in answer: where you land when you appear: first recommendation, middle of the list, or afterthought. Even a simple three-bucket split exposes a problem that presence rate hides completely.
  • Sentiment and factual accuracy: the share of appearances that describe you correctly and favourably. Track accuracy separately from sentiment. A model that is enthusiastic about a plan you discontinued is a different problem from one that is lukewarm about the right thing.
  • Source concentration: the domains cited most often across your prompt universe, and your standing on each. This is what converts the dashboard from a scoreboard into a work queue.

Three things are usually worth cutting. Raw mention counts, because they scale with how many prompts you added rather than with your visibility. Engine-by-engine tabs nobody compares, which are better served by a single filter. And any attempt to attribute revenue to AI answers before you have a stable presence baseline to attribute it against.

Sampling, Cadence, and Why a Single Run Proves Nothing

Generative engines are non-deterministic. The same prompt, on the same day, from the same account, can return different shortlists. A dashboard that samples each prompt once a week is not measuring visibility; it is measuring noise and presenting it with two decimal places.

A workable minimum is five runs per prompt, per engine, per weekly cycle. That is enough to separate a real shift from sampling variance for most prompt universes, and it gives you a threshold for reaction: at five samples, week-on-week moves under roughly ten points are usually noise. Print that threshold next to the chart so nobody escalates a two-point dip.

Collect on a schedule rather than on demand, from a clean logged-out context, and store every raw answer. Six months in, the stored answers matter more than the aggregates, because they are the only thing that lets you explain why a number moved.

The part worth automating first

Sampling is the piece teams underestimate. Fifty prompts across five engines at five runs each is 1,250 answers a week, every one of them needing to be parsed for brand mentions, position, sentiment and citations. Done by hand, that survives about two cycles. This is the layer Refine automates: it runs your prompt universe across ChatGPT, Gemini, Perplexity, Claude and Copilot on a schedule, scores presence and share of voice against the competitors you name, and keeps the cited sources, so the dashboard points at pages you can act on rather than numbers you can only watch.

Designing the View Your Team Will Actually Open

Layout does more for adoption than metric sophistication. Three zones, in this order, works reliably.

  • Zone 1: Four headline tiles: presence rate, share of voice, position mix, accuracy. Each with a week-on-week delta and the sample size behind it. Nothing else competes for the top of the screen.
  • Zone 2: A competitor grid: named competitors as rows, prompt families as columns, presence rate in the cells. This is the view that ends up screenshotted into board decks, and the one that starts the most useful arguments.
  • Zone 3: A changes feed: every prompt whose result moved beyond the noise threshold this week, linked to the raw answers and the sources cited. This is where the actual work comes from.

Two details separate dashboards that get used from ones that get bookmarked. The first is annotations: let whoever owns the dashboard write a line against any week. Shipped the comparison page, fixed the G2 profile, competitor launched. Six weeks later, those notes are the only way to interpret the trend. The second is a single named owner. A dashboard owned by a team is a dashboard owned by nobody.

Turning It Into a Weekly Ritual

The dashboard is not the deliverable. The twenty-minute meeting it makes possible is. Keep the agenda fixed and short.

  • Read the four headline tiles and their deltas. Anything inside the noise threshold gets no discussion at all.
  • Open the changes feed and take the two largest real moves, up or down. Read the actual answers, not the summary of them.
  • Look at what was cited in those answers. The fix is almost always a source problem, an outdated listing, a review profile nobody owns, a comparison article written by a competitor, rather than a content-volume problem.
  • Assign one owner and one date per fix, annotate the dashboard, and stop.

Expect the first four to six weeks to produce baseline rather than progress. Retrieval-heavy engines like Perplexity can reflect a new or corrected page within days, but the durable component, how the model itself describes your category, moves over months. A dashboard built to show weekly wins will disappoint. One built to show whether the trend is heading the right way, with the evidence to explain why, is what keeps a GEO program funded past its first quarter.

Short on time? Have an assistant summarise this page for you.