An AI visibility score is the percentage of a fixed set of prompts where an AI engine mentions or cites your brand. Run 100 prompts, get named in 22 answers, and your visibility score is 22%.
That is the whole formula. Weighting, sentiment, position, citation depth: every added sophistication is a variation on counting appearances across a set of questions you picked.
Which means the score is only ever as meaningful as the prompt set behind it. A brand can move 20 points without touching its website, purely by changing the questions.
What does the score measure?
Presence across a question set, expressed as a percentage. It answers "how often does the engine bring us up". Nothing beyond that.
What does it not measure?
It is not traffic, and it is not revenue. A visibility score can climb while referral sessions stay flat, because most AI answers resolve without a click.
It is also not accuracy. An engine can name your brand in every answer, get your pricing wrong in all of them, and the score still reads 100%.
How is the score calculated?
Three inputs: a prompt set, a set of engines, and a rule for what counts as an appearance. The rule is where tools quietly disagree.
What is the base formula?
Visibility score = (answers containing your brand) divided by (total answers), expressed as a percentage.
Total answers means prompts multiplied by engines. Twenty prompts across five engines is 100 answers, not 20. Average per engine instead of pooling and the same data gives you a different number.
What counts, a mention or a citation?
Most tools count a mention: your brand name appears in the answer text. Stricter scores count only a citation, where the engine credits and links one of your pages, which is the behaviour Google documents for its own AI features.
The two produce very different numbers for the same brand. Well-known companies score high on mentions and low on citations. Niche publishers usually do the reverse.
A worked calculation
Twenty prompts, five engines, 100 total answers. Your brand is named in 22 and one of your pages is cited in 9.
| Metric | Count | Score |
|---|---|---|
| Total answers | 100 | Denominator |
| Answers naming the brand | 22 | Mention score 22% |
| Answers citing the domain | 9 | Citation score 9% |
| Answers naming a competitor | 61 | Competitor presence 61% |
The 13 point gap between mention and citation is the diagnostic figure. Engines know the brand, then reach for someone else's page when they need a source.
Why do two tools give different scores?
Same website, same week, two different numbers. Expected, and not usually a bug in either tool.
Different prompt sets
If the prompt lists differ, the scores are not comparable in any way. Largest source of disagreement by far, and the easiest one to check.
Ask any provider for the full prompt list before you compare their number to another. A score without a documented prompt set cannot be verified or reproduced.
Different counting rules
One tool counts brand mentions, another counts linked citations, a third weights by position in the answer. All three are defensible. All three give different percentages.
Weighting adds more divergence. A score that treats a first-paragraph mention as worth more than a closing aside will not match a flat count.
Different sampling
Answer engines are not deterministic, so the same prompt returns different text on different runs. Query once per prompt and the score is noisy. Sample repeatedly and average, and it settles.
This is why small changes between two runs rarely mean anything. Treat anything under a few points as noise unless it holds across two measurements.
Why does your score differ across engines?
The same brand routinely scores 30% on one engine and 5% on another. The engines retrieve differently, so that spread is signal, not error.
Some engines lean heavily on what already ranks organically, since Google documents that its AI features run on the same core ranking systems as organic search. Others retrieve live and cite what they find.
At least one may answer certain questions with no web search at all. In our July 2026 scan of one audit-related prompt, ChatGPT fired zero search queries and cited nothing. Perplexity cited 13 sources on the identical question.
That last case is worth isolating. If an engine is not retrieving for your prompts, publishing more content will not move your score there. The fix runs through third-party presence instead.
What counts as a good score?
There is no absolute benchmark. Any provider quoting one is quoting a number from a different prompt set, which makes the comparison the only real answer.
How do you score against competitors?
Run the identical prompt set for three to five rivals, then express the result as share of voice so the numbers sit on one scale. Twenty percent is strong if the category leader sits at 25% and weak if they sit at 70%.
It is the only reading that survives a change of prompt set, because a harder question list lowers everyone's number together.
How do you score against yourself?
The second useful reading is your own trend on a frozen prompt set. Freeze the questions, run every 8 to 12 weeks, watch direction rather than level.
Change the prompts and the trend resets. Most teams that lose faith in their score have quietly edited the question list somewhere between runs.
What does the score not tell you?
It cannot tell you why. A visibility score is a symptom reading, and every fix has to come from the layer underneath it: which prompts you lost, who won them, and what shape the winning page actually took.
It also cannot promise that a tactic will move it. The Princeton and Georgia Tech GEO study (KDD 2024) reported gains of up to 40% from optimization tactics, but that figure measured attribution share inside an already-retrieved document set, not visibility. C-SEO Bench, a NeurIPS 2025 benchmark, could not reproduce the effect.
Treat a falling score as a reason to go look, not as a finding in itself. Moving it is answer engine optimization work, and most of it happens at the level of how a single page is structured.
Where Unveilr fits
Unveilr treats the score as the entry point rather than the deliverable. Agents scan a frozen prompt set, detect which sources won the answers you lost, update the content that failed, then re-scan the same prompts to confirm the number moved for a reason.
Freezing the prompt set is what makes the trend trustworthy. Without it, a rising score and an easier question list look identical.
In one D2C case study, the brand moved from the 9th most-cited domain in its category to number 1, with ChatGPT visibility rising from 3.3% to 44.7%.

