A usable AI visibility audit report has seven sections: a headline score, the method, presence and citation results, a competitive benchmark, technical findings, content gaps, and a ranked action list. Anything missing one of those seven is a snapshot, not an audit.
This matters because the report is the deliverable. You can run flawless checks and still hand over something nobody acts on. That is the usual failure mode: a dashboard screenshot with no recommendation attached.
Below is what each section holds, a worked example of the numbers, and the questions worth asking before you accept a report from anyone.
What is the report actually for?
It answers three questions: where do we stand, why, and what do we change first. Answer only the first and you have built a scoreboard.
That third question is the expensive one, because every answer to it turns into a piece of answer engine optimization work that somebody has to scope, schedule and eventually pay for.
How do you test a finished report?
Hand it to someone who did not run the audit and ask what they would do on Monday. A specific page, a specific fix, a reason. If they cannot name all three, the report failed, however good the data underneath it was.
That test rules out most automated outputs. A visibility percentage with no cause attached is not actionable, and it mostly produces arguments about the number rather than work.
What are the seven sections?
1. Headline score and what it is measured against
One number, plus the comparison that gives it meaning. A score on its own is unreadable, because visibility percentages depend entirely on which prompts someone chose.
Comparison is the content here: your citation rate against three to five named competitors on the identical prompt set, because being cited in 18% of answers is only strong or weak when you set it next to somebody else's number.
2. Method and prompt set
The full prompt list, the engines tested, the dates, and whether sessions were fresh. This section exists so the audit can be repeated and compared later.
Skip it and every future report becomes incomparable. If the prompt set quietly shifts between runs, an improving score might just be an easier question list. Nobody will be able to tell which.
3. Presence and citation results
Two separate numbers per engine, never merged into one. Presence is whether your brand is named. Citation is whether your page is credited as a source.
The gap between them is the most diagnostic figure in the report. High presence with low citation means engines know who you are but will not source you, which points at content. The reverse points at brand recognition.
| Engine | Presence | Citation | Reading |
|---|---|---|---|
| ChatGPT | 30% | 5% | Known, rarely sourced |
| Perplexity | 20% | 18% | Cites what it finds |
| Google AI Overviews | 15% | 12% | Tracks organic rank |
| Claude | 22% | 15% | Cites what it finds |
| Gemini | 10% | 4% | Other formats winning |
Splits this wide are normal. In our July 2026 scan of one audit-related prompt, Perplexity returned 13 cited sources and Claude 11. ChatGPT returned none at all, answering from training data without ever running a search.
4. Competitive benchmark
The same two numbers for three to five rivals, plus the list of domains winning the slots you lose. That second list is the part clients actually use.
If the same handful of domains take every citation in your category, the report should say what format they use. The engine has already picked a shape it trusts for that question. You are competing against the shape, not just the domain.
5. Technical access findings
Crawler permissions, rendering, and whether key pages return content to a bot at all. Short when it passes. The whole report when it does not.
A single wildcard rule in robots.txt can zero out every other finding, which is why this section runs before anyone reads the content gaps. Check the file against OpenAI's crawler list and Google's, not from memory.
6. Content gaps
The prompts where competitors are cited and you are absent, mapped to the page that should have won. This is where an audit turns into a content plan.
Each gap needs a named target: an existing page to restructure, or a missing page to write. Without that mapping it is just a list of complaints.
7. Ranked action list
Every finding sorted by expected impact against effort, with an owner and a timeline. Access fixes go first. They are cheap, and they gate everything else.
Order matters more than completeness here. A 40-item list with no sequence gets read once and filed. Five ordered items get done.
A worked example
The table below is an illustrative summary, not a specific client's data. It shows the shape a finished summary takes, and the arithmetic sitting behind the headline number.
| Metric | You | Competitor A | Competitor B | Category best |
|---|---|---|---|---|
| Prompts tested | 24 | 24 | 24 | 24 |
| Answers mentioning brand | 5 | 11 | 8 | 11 |
| Presence rate | 21% | 46% | 33% | 46% |
| Answers citing domain | 3 | 9 | 4 | 9 |
| Citation rate | 13% | 38% | 17% | 38% |
| Share of voice | 17% | 38% | 28% | 38% |
Read left to right, this says the brand is neither unknown nor trusted. It turns up in a fifth of answers, gets credited in an eighth, and trails the leader by roughly three times on both.
The action that follows is content, not access: if the technical section came back clean, then the gap between 13% and 38% is entirely about which pages exist and how they happen to be written.
The bottom row is share of voice, the brand's slice of every mention in the set. The rows above it are the only ones you can cross-check against your own analytics.
What does a weak report look like?
Four warning signs, all common:
- A single visibility score with no competitor comparison
- No prompt list, so the result cannot be reproduced
- Presence and citation merged into one metric
- Recommendations that apply to any website ("add schema", "improve content")
That last one is the giveaway. Generic advice means the audit found the checklist but never actually found your site, and what you are holding is a template with your logo dropped on top of it.
What changes the cost?
Pricing varies too much to quote a useful range. Four things drive it, and knowing them is what lets you compare two quotes honestly.
- Prompt count. Testing 100 prompts costs more than 20 and rarely changes the first decision.
- Engine coverage. Each added engine multiplies the testing, not adds to it.
- Re-scans. A one-off report is cheaper than a tracked loop, and far less useful.
- Whether fixes are included. Diagnosis and remediation are separate scopes.
Ask which of the four a quote covers before you compare two numbers. Most disagreements about audit pricing turn out to be scope mismatches, not price ones.
Where Unveilr fits
Unveilr delivers this report as a repeating cycle rather than a one-time PDF. Agents scan the prompt set, detect which sources won each answer, update the losing content, then re-scan to verify the change held.
Re-scanning is what turns section 7 from a recommendation into evidence, and without it nobody ever finds out whether the ranked action list was right in the first place.
In one D2C case study, the brand moved from the 9th most-cited domain in its category to number 1, with ChatGPT visibility rising from 3.3% to 44.7%.

