Choosing an AEO agency comes down to five questions, and a provider who answers all five specifically is already in the top tier of a market full of repackaged SEO retainers. The label is new enough that anyone can wear it.
This is a selection guide, not a vendor comparison. For named options and how they stack up, the agency roundup covers that ground; here the job is knowing a good one when you see it, whoever they are.
The B2B SaaS angle matters because your category has specific failure surfaces, review listings, comparison prompts, documentation access, and a generic content agency will miss them. All of them sit inside answer engine optimization proper, and a provider should speak that language unprompted.
What are the five questions?
Each maps to a way engagements actually fail. Ask them in order and take notes on specificity, because vagueness on any of these predicts the retainer.
Ask every shortlisted provider the same five, in the same words. The comparison only works when the questions are held constant, exactly like a prompt set.
Can I see the prompt set you would track for us?
The prompt set is the unit of work, so a provider should build a draft one during the sales process, from your buyers' questions rather than keyword exports. A prompt set sourced from real buyer language is the difference between tracking your market and tracking a template.
Red flag: a fixed-size package ("50 prompts included") before anyone has asked what your buyers ask.
Which engines do you scan, and how often?
The answer should name engines and a cadence, and it should include re-scans after model updates, because answers reshuffle when models ship. Quarterly-only scanning in a market that moves monthly is a diagnosis subscription.
Red flag: "we optimize for AI" with no scanning infrastructure at all. That is a content agency with a new deck.
What does the monthly report show?
Per-prompt presence and citation, split by engine, including the losses. An aggregate visibility score is not accountability; the report anatomy you should expect is the same one you would demand from an audit.
Red flag: wins-only reporting, or a proprietary score with no per-prompt detail underneath it.
Who writes the content, and who fixes what it finds?
Diagnosis and remediation are different scopes, and the gap between them is where retainers quietly hollow out. If content is included, ask for volume and who holds the quality bar; if excluded, price your own production into the comparison.
Red flag: an audit deliverable with no owner for the fixes. You are buying a PDF.
What will you tell us we cannot win?
The honest answer includes prompts answered from training data with no retrieval, where publishing earns nothing, and surfaces like video-heavy answers where articles do not compete. A provider who names your unwinnable prompts upfront is pricing work rather than hope.
Red flag: guarantees. Nobody controls a model update.
What do the answers look like side by side?
| Question | Strong answer | Weak answer |
|---|---|---|
| Prompt set | Draft built from your buyer language | Package tier with a prompt count |
| Engines and cadence | Named engines, re-scan on model updates | "All major AI platforms", no cadence |
| Reporting | Per-prompt presence and citation, losses included | Proprietary aggregate score |
| Content and fixes | Volume, owner and quality bar stated | Audit-only, fixes unowned |
| Honest ceiling | Names your unwinnable prompts | Guarantees results |
Score the meeting against the table. Three strong answers is a real shortlist candidate; five is rare and worth paying for.
One more test costs nothing: ask what they would do in month one. The right answer is a baseline, and providers who start with deliverables before measurement are selling activity.
Does B2B SaaS need category-specific questions?
Two additions. Ask how they handle review-site listings, since those contest your recommendation prompts, and whether they audit documentation access, since login-walled docs are the most common self-inflicted invisibility in the category.
A generalist who has never thought about either will learn on your retainer. That is fine at a discount and expensive at rack rate.
Agency, tools, or in-house?
The agency question hides a prior question, which is whether an agency is the right delivery model at all. The decision is mostly about whose hours do the work.
Tools-only fits teams with spare content capacity: the subscription is cheap and every fix is your labour. An agency fits when there is no spare capacity, which is the common B2B SaaS case.
In-house wins only at prompt-set sizes that keep a hire permanently busy. Below a few hundred prompts, the salary out-costs any retainer on the market.
The hybrid is legitimate and increasingly common: tools plus a fractional internal owner, with an agency for production bursts. Nothing about the models is exclusive.
Whichever model wins, keep the prompt set yours. It encodes your buyer language and your baselines, and it should survive any vendor change intact.
What should the engagement look like after signing?
The first month is a baseline, not results. Month one should hand you five artifacts:
- The finalised prompt set, built from your buyer language and signed off
- A full scan of every prompt across the agreed engines
- Presence and citation recorded per prompt, per engine, losses included
- The full audit run against your site and listings
- The unwinnable list, prompts where no retrieval happens and content cannot help
Expect the baseline to be uncomfortable reading. A provider whose first report flatters you has already started curating, and month one is the easiest month to catch it.
Then the loop: fixes shipped against specific lost answers, re-scans confirming which moved, the report reading movement against the baseline. Judge the program on 8 to 12 week windows, because crawl and re-index cycles set the physics, and anything judged at week three will look like failure regardless of quality.
And keep the measurement portable. The prompt set, the scan history and the baselines should be yours to take, because vendor lock-in through unexportable measurement is this market's oldest trick, inherited directly from SEO.
Where Unveilr fits
Unveilr is one of the options this guide helps you interrogate, and the five questions are the ones we would rather be asked. The model is a managed loop: prompt set scoped from your buyer language, agents scanning across engines, content produced against losses, re-scans verifying each fix, all reported per prompt.
Ask us the fifth question especially. The unwinnable-prompt list is the most useful thing a first conversation produces, and it is the part most providers leave out of a pitch because it is the part that shrinks the scope.
For evidence rather than description, the Care Dale case study documents the loop on a shower-filter brand that moved from the ninth most-cited domain in its category to the most-cited, with ChatGPT visibility rising from 3.3% to 44.7%. Read it the way this guide tells you to read anyone's proof, by checking whether per-prompt detail sits underneath the headline number.

