Unveilr Book a demo
AI share of voice measure AI visibility brand mentions in ChatGPT AI visibility metrics AI search visibility tracking

AI Share of Voice: How to Measure Your Brand's Visibility in AI Search

Sanditya SrivastavaSanditya SrivastavaJun 30, 202615 min read
Unveilr: how to measure your brand's AI share of voice across ChatGPT and Perplexity.

AI share of voice is the percentage of brand mentions in AI answers that belong to you, measured against every competitor named in the same category prompts. You calculate it by running a fixed set of queries across ChatGPT, Perplexity, and other engines, counting how often each brand appears, and dividing your mentions by the total.

That one number tells you whether AI engines treat you as a default answer or skip you entirely. Below is what the metric means, how to compute it, the supporting metrics that make it useful, and how to set up tracking you can trust week over week.

What AI Share of Voice Means

Traditional share of voice measured your slice of paid impressions or organic rankings against rivals. AI share of voice (AI SOV) moves that idea into the answer box.

When someone asks ChatGPT "what's the best project management tool for agencies," the model names a handful of brands. AI SOV asks a simple question: across all the prompts your buyers actually type, what fraction of those named brands are you?

It matters because the answer box is often the only impression you get. Roughly 93% of Google AI Mode sessions end without a click, so if the model does not mention you, the user never sees your site at all. Being absent from the answer is closer to being invisible than being ranked tenth on a results page.

AI SOV is a competitive metric by design. A brand mention rate of 30% sounds healthy until you learn the category leader sits at 70%. Share of voice puts your number in context by dividing the total pool of mentions among everyone who shows up.

If you want the foundational concepts behind getting mentioned at all, start with what answer engine optimization is.

Share of voice, brand visibility, and share of model

Three terms circulate for overlapping ideas and they are not interchangeable. Mixing them up is the most common way an AI visibility report misleads the person who wrote it.

Brand visibility is absolute. It is the share of your prompt set where you appear at all, regardless of who else is named. It answers "am I present."

Share of voice is relative. It is your mentions divided by all brand mentions in the category. It answers "am I winning." You can hold visibility perfectly steady while share of voice falls, simply because a competitor started appearing more often in the same answers. Nothing about your performance changed; the denominator did.

Share of model is per-engine. The newer term usually means share of voice measured inside one engine's answers rather than across a blended set. It exists because the per-engine spread is so wide that a blended figure hides the only detail you can act on.

Use visibility to answer whether you are present, share of voice to answer whether you are winning, and share of model to answer where. Reporting one without the others is how a flat number gets mistaken for a stable position.

The Formula

The core calculation is straightforward:

AI SOV = (your brand mentions / total brand mentions in category) × 100

Say you run 30 category prompts across ChatGPT and Perplexity. The models name brands 200 times in total across those answers.

Your brand appears 50 times. Your AI share of voice is 50 / 200 = 25%.

Two details change the result more than people expect:

  • Define the category boundary carefully. "Total brand mentions" means the competitors you decide belong in your set. Include three rivals and your share looks strong. Include fifteen and it drops. Pick a competitor set that matches how buyers actually compare, and keep it stable so the number stays comparable over time.
  • Decide what counts as a mention. A passing reference ("tools like X, Y, and Z") is weaker than a recommendation ("I'd suggest X for this"). Many teams track both: raw mention share and recommendation share. The gap between them is where a lot of the real signal lives.

Why the same formula gives you a different answer each run

The arithmetic is trivial. The input is not, because the same prompt does not return the same answer twice.

Repeat overlap on commercial engines sits at roughly a Jaccard of 0.34 to 0.42, which means two runs of an identical prompt agree on well under half their sources. Compute share of voice from a single pass and you have measured one sample of a distribution, then reported it as a fact.

The fix is unglamorous. Run each prompt five to ten times per engine and average before you divide. If you intend to put the number in a board deck, the honest floor is higher still, because at small sample sizes a two-point move is indistinguishable from noise.

This is also why the trend matters more than the level. A single scan tells you very little. The same scan repeated the same way for eight weeks tells you almost everything.

The Core Metrics to Track

AI SOV is one number, and on its own it hides too much. These seven metrics give you the full picture. Track each one per platform and over time, never from a single answer, because results shift with every model update.

Metric What it measures How to read it
Brand mention rate Share of answers where your brand appears at least once The baseline. Industry average sits near 17.2% per the 2026 benchmark below
Recommendation rate Share of prompts where you are actively recommended, not just named Often a stronger pipeline signal than raw mentions. 100 prompts with 35 mentions but 12 recommendations = 12%
Prompt coverage Share of your defined prompt library where you show up at all Tells you how broad your presence is across topics
Share of voice Your mentions divided by total brand mentions in the category The competitive view. Under 15% suggests a citation gap, 25-40% is competitive, above 40% is strong
Model-specific visibility The same metrics broken out by engine Reveals where you win and where you are missing entirely
Visibility volatility How much your numbers swing between scans High volatility means your position is fragile and needs reinforcing
Sentiment The tone the model uses when describing you A brand can be mentioned often but framed poorly, flagging missing features or stale information

Why sentiment needs a closer look

The sentiment metric deserves attention. An engine might cite you in half its answers and still describe you as "more expensive than alternatives" or "limited on integrations." Frequent and unfavorable is its own kind of problem, and you only catch it by reading the actual language, not just the mention count.

Why Your Score Differs Across Platforms

You will not get one AI share of voice. You will get a different one on every engine, and the spread is wide.

You might sit at 40% inside ChatGPT and 15% inside Perplexity for the same prompt set. That is normal, and it reflects how differently these systems build answers.

The numbers back this up. An analysis of roughly 680 million AI citations in early 2026 found that only about 11% of cited domains overlapped between ChatGPT and Perplexity. Citation volume for the same brand varied by as much as 615x across platforms.

These are not the same channel with cosmetic differences. They are separate surfaces with separate rules.

How each engine builds answers

A few reasons for the divergence:

  • ChatGPT leans on parametric knowledge. It builds a lot of its answers from patterns baked into training rather than live web results. It cites sources about 87% of the time but names a specific brand in only around 20.7% of answers. Playbooks that move the needle on citation-heavy engines often do little here. See how to get cited by ChatGPT for the specifics.
  • Perplexity rewards freshness and structure. New editorial pieces with clean schema can appear in Perplexity citations within days. Our guide on how to rank in Perplexity covers what it favors.
  • Google AI Overviews track classical ranking. AI Overviews mention brands in about 61% of answers and cite sources roughly 85% of the time, and those citations correlate heavily with where you already rank in regular Google search. Without the underlying organic work, optimizing the overview alone goes nowhere. See how to rank in Google AI Overviews.

The practical takeaway: a single blended score across all engines averages away the detail you need. Track each platform on its own, then decide where the opportunity is largest.

Benchmarks

Benchmarks help you judge whether a number is good. A few from current 2026 data:

  • Average brand mention rate is about 17.2%. The State of AI Search 2026 report put the average rate at which a brand appears across AI answers at 17.2%. If you are above that, you are ahead of the typical brand. Top performers reach far higher.
  • Mention behavior swings hard by engine. ChatGPT names brands in roughly 20.7% of answers, Gemini in about 83.7%, AI Overviews in around 61%, and Google's AI Mode in about 37.6%. The same brand can look dominant on one engine and absent on another.
  • Share of voice bands. As a rough read on the competitive metric: under 15% points to a real citation gap, 25-40% is competitive in most categories, and above 40% signals strong visibility.

Treat these as orientation, not targets. A niche B2B category with three real players behaves nothing like consumer software with fifty. Your own trend line, measured the same way each week, is more useful than any industry average.

How to Set Up Tracking

You do not need a platform to start. You need a consistent method. The discipline is in keeping the inputs fixed so the outputs stay comparable.

  • [ ] Define a fixed prompt set of 20 to 30 queries. Mix three clusters: brand prompts ("is [your brand] any good"), category prompts ("best [category] tool for [use case]"), and comparison prompts ("[you] vs [competitor]"). Write them the way buyers actually ask.
  • [ ] Lock your competitor set. Decide which rivals count toward the "total mentions" denominator and keep that list stable. Changing it mid-stream breaks your trend line.
  • [ ] Run every prompt across each engine. Cover ChatGPT, Perplexity, Google AI Overviews, and Claude at minimum. Run each prompt five to ten times per engine and average, since a single pass samples a distribution rather than measuring it.
  • [ ] Record four things per answer. For each prompt and engine, log whether you were mentioned, your position relative to competitors, which sources were cited, and the sentiment of the description.
  • [ ] Scan on a fixed cadence. Run a full audit weekly or at least monthly, and spot-check your top prompts more often. Citation distributions shift within weeks as models update and indexes refresh.
  • [ ] Compute the metrics and watch the trend. Calculate mention rate, recommendation rate, prompt coverage, and share of voice per platform. The single most useful output is the direction of travel, not any one snapshot.

Doing this by hand across 30 prompts, five engines, and several runs each is a few hundred queries per scan. That is fine for a one-time baseline.

It gets heavy fast as a weekly habit, which is where tooling earns its place. See the best AEO and AI visibility tools for a comparison.

From measuring to moving the number

Measuring share of voice is only half the job. A dashboard that reports the number every week tells you where you stand, but it does not move you up the list.

The brands that grow share of voice fastest close the loop: they read what the data shows, which prompts they lose, which sources win the citation, and how each engine frames them, then update content to match, then re-scan to confirm the number moved.

This matters more with every model launch, because a new ChatGPT or Gemini release can reshuffle citation patterns within days, and content that won last month may not win this one. Unveilr is built around running that loop on a schedule, which is worth knowing about mainly at the point where the run counts above stop being survivable by hand.

Frequently Asked Questions

What is a good AI share of voice?

It depends on your category size, but as a general read: under 15% suggests a citation gap worth closing, 25-40% is competitive, and above 40% is strong. The more useful benchmark is your own trend. A score climbing from 18% to 28% over a quarter beats a flat 30%.

How is AI share of voice different from brand mention rate?

Mention rate is absolute: the share of answers where you appear. Share of voice is relative: your mentions as a fraction of all brand mentions in the category. You can have a healthy mention rate and still lose on share of voice if a competitor is named far more often in the same answers.

What is the difference between share of voice and brand visibility?

Brand visibility asks whether you are present at all across your prompt set. Share of voice asks how much of the total brand conversation is yours. Visibility can stay flat while share of voice falls, because a competitor appearing more often changes your denominator without changing anything you did. Report both, or a stable-looking number will hide a losing position.

What is share of model?

Share of model usually means share of voice measured inside a single engine rather than blended across all of them. The term exists because per-engine spread is extreme: the same brand can sit at 40% on ChatGPT and 15% on Perplexity for an identical prompt set. A blended average hides exactly the detail you would act on, so most serious reporting is per-model.

Why is my score so different on ChatGPT versus Perplexity?

Because they build answers differently. ChatGPT relies more on training-data knowledge and names brands in fewer answers, while Perplexity leans on fresh, well-structured web content. An analysis of 680 million citations found only about 11% domain overlap between the two. Track each engine separately rather than blending them.

How many prompts do I need to track?

Start with 20 to 30 queries spread across brand, category, and comparison intent. That is enough to produce a stable read without becoming unmanageable. Run each prompt five to ten times per engine, since repeat overlap sits around a Jaccard of 0.34 to 0.42 and a single answer is a sample rather than a measurement.

How often should I measure?

Run a full scan weekly or monthly, and spot-check your most important prompts more frequently. AI answers shift within weeks because of model updates and index changes, so a quarterly check misses too much movement to act on. Monthly is the practical floor for most teams.

Does sentiment matter or just mention count?

Both. An engine can mention you often while framing you as expensive or feature-poor. Frequent negative mentions can hurt more than absence, and you only catch them by reading the actual wording, not the count. Track sentiment alongside frequency so you see both at once.

Sources

About the Author

Sanditya Srivastava is the founder of Unveilr, an answer engine optimization (AEO) service that helps brands get cited and recommended across AI search platforms like ChatGPT, Perplexity, Google AI Overviews, Gemini, and Claude. He writes about how AI search is reshaping brand discovery.