← The ClarAI blog · MEASUREMENT
MeasurementHonesty

Your AI Visibility Score Is Lying to You (Ours Might Be Too)

AI answers are nondeterministic, samples are small, and citation distributions are power-law. What honest AI visibility measurement looks like, and exactly how ClarAI implements it.

TThe ClarAI team
AI search visibility, ClarAI
AUGUST 5, 2026 · 7 MIN READ
A single confident visibility score shown next to its wide error bars, illustrating false precision.

An AI visibility score built from a handful of prompt runs is an estimate with error bars, not a reading. If a tool hands you a single confident number with no sample size, no coverage disclosure, and no raw answers behind it, you should treat that number as marketing. That includes ours: ClarAI publishes a visibility score too, and every limitation described below applies to us unless we specifically engineer around it. This post covers where the numbers wobble, what honest measurement looks like, how we implement each piece, and where our own score can still mislead you.

Why every score in this category wobbles

Start with the physics of the thing being measured. Large language model outputs are nondeterministic: the same prompt, sent to the same engine, returns different answers on different runs, and batch-level variance in inference means even nominally deterministic settings drift. A brand that appears in three of five runs today may appear in one of five tomorrow with nothing changed on its website.

Then add how tools sample. No vendor queries every plausible buyer prompt on every engine continuously; everyone samples a subset, because every query costs real money. Small samples swing hard. A mention rate computed from four prompts moves 25 points when a single answer changes. Layer on the shape of citations: a small set of domains dominates AI answers in most categories while the long tail is unstable run to run, which means position changes deep in the tail are mostly noise. Put together, many of the leaderboard movements this category sells as insight happen entirely within the noise floor.

This is not a fringe complaint. The sharpest 2026 criticism of the category calls the false precision out directly: Canonry published a piece titled "Every AI Visibility Tool Is Lying to You", Brainlabs argued in "Your AI Visibility Data Is Wrong (And That's Okay)" that AI visibility numbers are probabilistic estimates that fluctuate run to run with no reliable baseline, and Petra Labs, in "How Accurate Are AI Visibility Tools?", found a single brand's score swinging 32 points across three ChatGPT access surfaces on the same day. The emerging buyer checklist from that criticism is the right one: does the tool sample multiple prompts across multiple engines, does it disclose sample sizes and variance, and does it show the raw evidence.

“No third-party tool has access to our internal ranking or AI systems.”

Google Search Central, "Optimizing your website for generative AI features on Google Search" (2026)

Google's warning is aimed at tools claiming inside knowledge, and it settles the question of what any honest tool can be: an observer. Everything real in this category is measurement of the observable layer, the answers engines actually give and the pages they actually cite. Anything sold as more than that is now arguing with Google, not with us.

What honest measurement looks like

  • Sample sizes on the number, not in a methodology PDF. Every score, rate, and rank should carry how many prompts and runs produced it.
  • Coverage disclosure. If 4 of 72 tracked prompts were measured in the latest run, the interface should say so where the score is shown.
  • Multiple engines, reported separately. An average across engines hides more than it reveals; per-engine sample counts keep it honest.
  • Raw receipts. The actual answer text and the actual cited URLs should be stored and shown, so any number can be audited down to its evidence.
  • Unmeasured means unmeasured. Prompts that have not been tested should be labelled as such, not silently scored as zero or, worse, as maximum opportunity.
  • Change treated as a distribution. Small run-over-run movements should be presented as within normal variance, not narrated as wins and losses.

How ClarAI implements each of these

Coverage disclosure: ClarAI's dashboard and visibility pages carry a coverage badge that states how many of your tracked prompts the latest run actually measured, in the "4 of 72 prompts measured" form, so a thin sample can never quietly impersonate a full scan. Sample counts: mention-rate tiles state the number of prompts tested behind the percentage, and per-engine breakdowns report their own counts, so a 100 percent mention rate on a three-prompt sample reads as exactly what it is.

Raw receipts: every monitored prompt stores the full answer text per engine plus the citations that came back, and both are shown in the app, which means every score is auditable down to the sentences that produced it. Unmeasured honesty: our prompt scoring marks each prompt as measured or unmeasured, and unmeasured prompts receive a neutral opportunity value rather than the flattering maximum, with the interface disclosing when a degraded formula is in use. And there is no seeded fiction: a fresh project shows honest empty states until a real analysis has run, never placeholder numbers.

Where our own score can still mislead you

Symmetry demands this section. ClarAI's overall score still moves between runs for the same reasons every tool's does: the engines are nondeterministic and we are sampling. We run a call-budgeted sample of your highest-priority prompts per engine rather than the full prompt set on every run, which is a cost decision, and it is exactly why the coverage badge exists. We render a variance band beside the score once there are enough runs to compute one: the honest run-to-run spread, read from our own variance estimate rather than asserted as a single precise number. A small score movement that sits inside that band deserves the same skepticism this post recommends everywhere else, and we present it as within normal variance rather than narrating it as a win or a loss. Nondeterminism cannot be engineered away by anyone. It can only be disclosed, sampled against, and priced in.

Questions to ask any vendor, including us

  • How many prompts and runs produced this score, and where does the interface say so?
  • Can I read the actual AI answers and citations behind any number on this dashboard?
  • What happens to prompts you have not measured: are they excluded, zeroed, or honestly labelled?
  • How much does this score move between two runs where nothing changed on my site?
  • Do you claim any access to how engines rank internally? If yes, reconcile that with Google's public statement that nobody has it.

Get a score that shows its work

Run a free check on your domain. You get real answers, real citations, and the sample size printed next to every number.

Sources

Google Search Central, Optimizing your website for generative AI features on Google Searchdevelopers.google.com/search/docs/fundamentals/ai-optimization-guide
T
The ClarAI team

We build ClarAI, a platform for measuring and improving how brands appear in AI answers across ChatGPT, Perplexity, Gemini, Claude, and Google AI Overviews. When we appear in our own comparisons, we say so.