Summary
A preprint published on 5 October 2026 by Amine Aziz Alaoui of GetMint Research models brand visibility in ChatGPT using more than a million answers collected between January and September 2026. Its core finding: the daily visibility rate of a single prompt swings by about 14 points from one day to the next, mostly for reasons that fade within two to three days, while the underlying level drifts by around 4 to 4.5 points a day. No number of runs on a single day pins a prompt down to better than about ±28 points. Reliable numbers come from pooling prompts into topics and reading them over roughly two weeks. This article summarizes the study and what it changes for anyone tracking AI visibility.
The short answer
A single prompt’s ChatGPT visibility on a single day is a weak signal: according to the study, it can sit about ±28 points away from the prompt’s real level, however many times you run it that day. A topic of around 20 prompts read over 14 days is accurate to about ±7 to ±8 points in calm periods. Track topics, not prompts; read trends over two weeks, not days; and treat sudden moves across all your prompts as possible engine changes before crediting or blaming your own work.
The Study in One Paragraph
The paper is “A Measurement Model for Brand Visibility in ChatGPT Answers” by Amine Aziz Alaoui (GetMint Research, preprint v1.3, DOI 10.5281/zenodo.23156371). It uses eight months of production monitoring from GetMint, an AI visibility platform, on the ChatGPT consumer web interface in logged-out sessions: more than a million answers to tens of thousands of prompts, most observed once a day. The author models the visibility rate of a prompt (the share of answers that mention a brand) as a slowly drifting level plus a day effect that does not last, observed through random draws, with occasional shocks that hit every prompt at once.
We are not affiliated with the study. We cover it because it is one of the most rigorous public treatments of a question every AI visibility tool, including ours, has to answer: how much can you trust a visibility number, and for how long?
Finding 1: Visibility Moves 14 Points From One Day to the Next
On days where the same prompt was run several times, the author could separate sampling noise (the same question giving different answers) from real movement in ChatGPT’s behavior. For prompts with a mid-range visibility rate, the true daily rate moved by about 14 points between consecutive days, 20 points at two to three days, and 25 to 33 points at gaps of one to six weeks.
That movement splits into two parts:
- A day effect of about 14 to 16 points: what ChatGPT retrieves and how it writes on a given day. It fades within two to three days, and neither day is more “true” than the other.
- A level drift of about 4 to 4.5 points per day for one prompt, which accumulates over time: roughly 16 points after two weeks and 24 points after a month.
The paper also shows why many teams never see this. The common shortcut of one run per day, with variance estimated over the whole window, mathematically subtracts the very movement it should measure. It produces a misleading “stable for two weeks” reading whatever the engine does. The author notes that the first version of his own work made this mistake, which is a good reason to take the correction seriously.
Why this matters
If your dashboard says a prompt went from 40% to 55% visibility since yesterday, the study suggests that is well within ordinary day-to-day movement. Reacting to it, or reporting it as a win, is reading noise.
Finding 2: More Runs on One Day Will Not Fix It
The intuitive fix for noisy AI answers is to run each prompt more often. The study shows where that stops working. Running a prompt more times on the same day only tightens your estimate of that day’s rate, not of the prompt’s underlying level, because the day effect is shared by every run on that day.
In numbers from the paper: for a single prompt, 96 runs in one day give about ±29 points on the level, 300 runs about ±28, and no number of same-day runs goes below about ±28. For a topic of 20 prompts, 30 runs each on one day (600 answers) gives about ±7, not the ±4 that sampling math alone promises, and going to 100 runs each does not improve it. Beyond about 30 runs per prompt, a same-day burst buys nothing more.
Finding 3: Topics Over Two Weeks Are the Reliable Unit
Precision comes from pooling days and prompts, not runs. On calm days, the day-to-day moves of different prompts are almost uncorrelated (around 0.004 to 0.009 in one client’s portfolio, even for the same prompt translated into two countries), so they cancel out when you average a topic.
The study’s margins for a 20-prompt topic at one run per prompt per day:
- One day: about ±23 points.
- 7 days: about ±9 points.
- 14 days: about ±7 to ±8 points.
- 30 days: about ±8 points, because the level itself moves during the month.
- A single prompt never gets below about ±32 points, whatever the window.
The practical conclusion is clear: the unit of AI visibility monitoring is the topic, not the prompt, and the useful reading window is about two weeks. In simulation, a 14-day moving average shown with a band of about ±7 behaved close to the author’s more sophisticated filter. A 30-day average, shown with the narrow band that sampling math gives, covered the true level only about half the time.
Finding 4: ChatGPT Changes Without Announcing It
Some days, every prompt moves together. The author built a “sentinel” that flags days when tracked prompts shift in the same direction far more than chance would allow. Between May and July 2026 it dated seven such common shocks on ChatGPT web, six with no public engine event within two days. It also caught an unannounced model version switch on 8 August 2026, matched to the version label shown in the answers.
The reverse also happened: the public launch of GPT-5.6 on 9 July 2026 showed no detectable effect on brand mentions. As the paper puts it, announcements and impact are distinct objects. You cannot rely on release notes to explain your numbers. You need to detect changes from the data itself.
The study also found that some segments move more than others: retail prompts and mid-size organizations diverged from the average curve. The author treats these segment results as exploratory.
What This Means for Your AI Visibility Tracking
Whatever tool you use, these are the habits the evidence supports:
- Group prompts into topics of 15 to 20 or more and report the topic, not individual prompts. A single prompt is a diagnostic, not a KPI.
- Read trends over about two weeks. Daily charts are for spotting events, not for judging performance.
- Show uncertainty. A topic score without a margin of around ±7 to ±8 points invites over-reading.
- Watch for days when all prompts move together. Before crediting a content change, check whether competitors’ and control prompts moved the same way.
- Measure your own actions against control prompts. The paper describes a difference-in-differences approach: compare the prompts you targeted with comparable prompts you did not touch.
- Do not pay for huge same-day run counts. Past roughly 30 runs per prompt per day, the extra precision is an illusion.
For a broader view of why AI answers shift, see our guide on how often AI answers change, and for running proper experiments, how to A/B test GEO changes.
How to apply this with Refine
Refine runs your prompts every day across ChatGPT, Gemini, Perplexity, Claude, Copilot and Mistral. Tag your prompts by topic, read share of voice and mention rate at the tag level over two weeks rather than prompt by prompt, and compare your topics against competitors on the same prompts to separate engine-wide moves from your own gains. Start with a free [AI visibility audit](/audit).
Limits of the Study
The paper is a preprint and lists its own limits clearly. It covers one surface, ChatGPT web in logged-out sessions, and one platform’s client prompts, not the whole web. The day effect and drift are estimated on days with repeated runs, which are a minority concentrated in June and July 2026. Whether the speed of movement changes over time is left open, and the margins assume calm periods. Results for Gemini, Perplexity, Claude or Copilot may differ, although the author notes the method transfers to any engine.
Even with these caveats, the main lesson is solid and fits what practitioners see every day: an AI visibility number is a measurement with error, not a fact. Treat it like one, and your decisions about content, PR and budget will rest on signal instead of daily noise.
Short on time? Have an assistant summarise this page for you.

