AI share of voice is directional, not precise. Anyone claiming otherwise is selling
Last reviewed
Short answer
The same prompt, re-run, returns substantially different sources: the published academic measurement found day-to-day source similarity of 0.34 to 0.42 and brand similarity of 0.45 to 0.59, and an industry study of 693,509 answers found repeat ChatGPT responses sharing only about 21% of cited domains. AI visibility is measurable as a trend across many runs. As a precise score, it is fiction.
The variance, with its actual sources
Two independent measurements, published four months apart. The academic one, an April 2026 arXiv paper from Aurora Intelligence, ran repeated prompts and computed similarity: cited sources on consecutive days overlapped at Jaccard 0.34 to 0.42, brands at 0.45 to 0.59, and same-day repeats were barely steadier, meaning the wobble is the model's own stochasticity, not the news cycle.
The industry one, Parse's July 2026 study across 693,509 answers, found two repeat ChatGPT answers to the same prompt sharing 21.2% of cited domains on average, with commerce and shopping among the most volatile categories at roughly 81% churn. Different methods, same conclusion, and the shopping-specific number is the one a merchant should remember.
What this does to a score out of one hundred
A visibility score is a sample statistic wearing a lab coat. Sampled from a system this variable, a single week's 34 against last week's 31 is indistinguishable from noise unless the sampling is heavy: the arXiv authors recommend at least seven runs per prompt per day and rolling aggregation over two to four weeks before trusting movement.
So the test for any AI visibility tool, ours included, is whether it is honest about this. Sampling cadence stated, trends over point values, no celebration of two-point moves. A dashboard that reports your AI rank to a decimal is describing its query budget, not your visibility.
Why we lead with the caveat instead of burying it
Entitled's tracker samples weekly and reports trend direction, and the app's own framing is watch the trend, not one run. That is not modesty as decoration, it is what the published measurement supports, and it is why this page cites the studies rather than asking you to trust us.
The commercial reason to say it out loud: merchants in this category have been burned by confident dashboards before. A tool that tells you what its numbers cannot mean earns the right to be believed about what they can.
How to consume visibility numbers, in four rules
Treat these as the contract with any measurement, manual or paid.
- Trends over levels: direction across weeks means something, a single score does not
- Windows over days: aggregate two to four weeks before calling a change real
- Same prompts over time: a changed prompt set resets the baseline, silently
- Mentions and traffic together: visibility sampling says you appear, analytics says it mattered, and only the pair tells the story
Frequently asked questions
If it's this noisy, why measure at all?
Because trends survive noise. A store moving from never mentioned to sometimes mentioned across a month of sampling is a real signal with business value, and it is invisible without measurement. Noise argues for statistical honesty, not for ignorance.
My tool showed a big drop this week. Panic?
Not on one week. Check whether the prompt set changed, then wait for the second week or re-sample. The published day-to-day churn produces scary-looking drops from nothing, and commerce prompts churn hardest of all measured categories.
Do the expensive enterprise tools escape this?
They buy narrower error bars with bigger query budgets, which is legitimate, and the variance floor remains because the engines are stochastic by design. The difference between tools is sampling honesty and volume, not access to a stable truth.
Is Google AI Mode as unstable as ChatGPT?
Less, per Parse's measurement: Google AI Overviews repeats shared about 31.5% of sources against ChatGPT's 21.2%. Steadier, still far from stable, and the same trend-first reading rules apply.
Sources
Keep reading
Score your own catalog
Entitled grades every product against the rules on this page, explains each failure, and fixes them with AI you approve. The score is free forever.
