Tracking & Analytics/August 12, 2026

How to Audit Your Brand's AI Sentiment Across ChatGPT, Gemini and Perplexity

Robin Pautigny

Robin Pautigny

Co-founder, Refine

How to Audit Your Brand's AI Sentiment Across ChatGPT, Gemini and Perplexity

Summary

An AI sentiment audit answers a question a visibility score cannot: when ChatGPT, Gemini or Perplexity mention your brand, what do they actually say about it? This guide covers the prompt set to build, how many samples you need before a reading is trustworthy, a coarse scoring scale that survives contact with noise, and how to trace a negative framing back to the source that caused it.

The short answer

An AI sentiment audit measures how favorably ChatGPT, Gemini, Perplexity and other assistants describe your brand when they mention it. You run one by freezing a prompt set, sampling each engine several times per prompt, scoring every answer on presence, position, framing, accuracy and attribution, then tracing the negative framings back to the sources that produced them. Sentiment is a source problem far more often than a wording problem.

What an AI Sentiment Audit Actually Measures

Most AI visibility programs start and stop at presence: did the assistant name us, yes or no. That is a useful first metric and a shallow one. A brand can appear in nine answers out of ten and still lose every deal those answers touch, because the mention reads as "a cheaper option for small teams" rather than "the tool most teams in this category standardize on."

A sentiment audit separates the mention from its framing. On every answer you collect, you are scoring five distinct things.

  • Presence — is the brand named at all, and does it appear on its own or only after the user names it?
  • Position — first recommendation, buried in a list of eight, or an also-ran in the closing sentence.
  • Framing — the adjectives and hedges attached to the mention. "Powerful but complex," "solid for basic needs" and "the default choice for serious teams" are three very different outcomes for the same brand.
  • Accuracy — does the description match reality? Wrong pricing, a discontinued feature or an outdated positioning is a sentiment problem even when the tone is warm.
  • Attribution — which sources the assistant cites or paraphrases. This is the field that turns a diagnosis into a fix.

Framing is the dimension teams underestimate. Large language models rarely say anything overtly negative about a brand. Negative sentiment in AI answers is usually damning with faint praise: a qualifier, a caveat, a "best for" that quietly excludes your actual buyer.

Why Sentiment Splits Across ChatGPT, Gemini and Perplexity

Running the same prompt on three assistants often produces three different opinions of your company. That is not a flaw in your measurement. It reflects how differently these systems assemble an answer.

  • ChatGPT leans heavily on what the model absorbed during training, blended with web retrieval when the question looks time-sensitive. Sentiment here is sticky: it reflects how your brand was described across the web months ago, not last week.
  • Perplexity is retrieval-first and cites as it goes. Sentiment tracks whatever five to ten pages it pulled, which means one well-ranked critical review can dominate the tone of an entire answer.
  • Gemini sits between the two and leans on Google index signals, so your Search visibility and your AI sentiment are more tightly coupled there than anywhere else.
  • Claude and Copilot add their own retrieval mixes. Copilot inherits Bing indexation in particular, and a page Bing never indexed cannot influence the answer no matter how good it is.

The practical consequence: never average sentiment into a single company-wide number before you have looked at it engine by engine. A brand can be described warmly by ChatGPT and dismissively by Perplexity, and the fixes for those two problems have almost nothing in common.

One score hides the problem

A blended AI sentiment figure across all engines is comfortable to report and nearly useless to act on. Refine keeps every raw answer, per engine and per run, so you can read the sentence Perplexity actually produced and see which cited page pushed it there. The number tells you something moved. The answer text tells you what to change.

Building a Prompt Set That Exposes Sentiment

Presence tracking works fine with category prompts. Sentiment does not. If you only ask "what are the best tools for X," you will collect polite one-line descriptions and learn very little. Sentiment surfaces when the question gives the model permission to be critical.

A workable audit set is 25 to 60 prompts spread across five families.

  • Category prompts: "best X tools in 2026," "top alternatives to Y." These establish baseline presence and position.
  • Comparison prompts: "Brand A vs Brand B," "how does A compare to B for mid-market teams." Comparisons force the model to state trade-offs, which is where framing becomes explicit.
  • Objection prompts: "is Brand A worth the price," "what are the downsides of Brand A," "Brand A complaints." This family produces the most actionable data and is the one most teams skip.
  • Fit prompts: "is Brand A a good fit for a 15-person agency," "should an enterprise use Brand A." These reveal whether the model has miscast your ideal customer.
  • Reputation prompts: "is Brand A reliable," "is Brand A still active," "who is behind Brand A." These surface stale or wrong facts that quietly cap every other answer.

Write the prompts in the words your buyers actually use. The fastest source is your own sales calls and support inbox, not a keyword tool. A prompt nobody asks is a prompt whose sentiment does not matter.

The Five-Step Audit

The method is simple. The discipline is in doing it the same way every time.

  • Freeze the prompt set. Rewriting prompts between runs makes the trend line meaningless. Add prompts over time if you must, but never quietly retire the ones that made you look bad.
  • Sample, do not check. LLM answers are probabilistic. Run each prompt at least five times per engine, in a clean session with memory and personalization turned off, and read the results as a distribution rather than a fact.
  • Capture the raw text. Store the full answer and the cited URLs, not just a score. Almost everything useful in an audit comes out of rereading answers later.
  • Score each answer on the five dimensions above, with the same scale every time. Consistency matters more than sophistication.
  • Trace the outliers. For every clearly negative answer, open the sources. Nine times out of ten the framing came from one specific page: an old review, a competitor comparison article, a Reddit thread, or your own outdated pricing page.

Turning Answers Into a Number You Can Track

You need a score to see movement over quarters, but keep the scale coarse. Fine-grained sentiment scoring on a probabilistic system creates false precision and invites arguments about noise.

A simple, defensible scale:

  • +2 — recommended without qualification, or named as the leading option.
  • +1 — mentioned positively but with a limiting qualifier, such as "good for small teams."
  • 0 — mentioned neutrally, one entry in a list, no evaluation attached.
  • -1 — mentioned with a caveat that would deter your target buyer, or described inaccurately.
  • -2 — actively steered away from, or a competitor recommended instead in a direct comparison.

Report two numbers per engine: the mean score across all mentions, and the share of answers where you appear at all. The pair matters more than either figure alone. Rising sentiment with falling presence usually means you are being cited by a narrower, more specialized set of sources, which is a very different situation from broad but lukewarm coverage.

Fixing Negative Sentiment at the Source

Once you have traced framings back to sources, the fixes are unglamorous and mostly not about your own homepage.

  • Correct the factual layer first. Wrong pricing, an old feature list or a stale founding story on your site, G2, Crunchbase or Wikipedia propagates everywhere. This is the cheapest win in the entire exercise.
  • Answer the objection prompts publicly. If "downsides of Brand A" returns a competitor page, you have ceded that question. An honest page naming real trade-offs tends to get cited precisely because it reads as balanced.
  • Fix the "best for" line. If every assistant says you are best for small teams and you sell to mid-market, that framing came from somewhere, usually your own positioning copy from two years ago, still indexed.
  • Earn third-party mentions with specifics. A review saying "fast setup, took us under an hour" gives the model something quotable. A review saying "great product" gives it nothing.
  • Be patient with parametric memory. Retrieval-driven engines can shift within days of a new page ranking well. Sentiment baked into training data moves on the timescale of model releases.

How Often to Re-Run the Audit

Monthly is the right cadence for the full scored audit. Weekly sampling is useful as a monitoring layer, catching a sudden drop or a new competitor entering answers, but weekly sentiment swings are almost always noise around a trend rather than the trend itself.

Judge the program over quarters. Set a baseline in month one, fix the factual layer in month two, and expect the first readable movement somewhere in month three or four, earlier on retrieval-heavy engines and later on ChatGPT. Teams that abandon the effort usually do so around week six, which is exactly when the data starts to become useful.

Short on time? Have an assistant summarise this page for you.