Quick Answer: Perplexity works in four stages: it interprets your question, searches the web in real time, summarises the pages it found, and cites each source. The citations are numbered and link to the original pages. On Pro Search you choose the model, including GPT-5 and Claude 4.6 Sonnet, and Research runs dozens of searches in 2 to 4 minutes.
Perplexity describes itself as an answer engine rather than a search engine, and the difference is mechanical, not cosmetic. A search engine ranks pages and hands you the list; Perplexity retrieves pages, reads them, writes one answer and shows its working as numbered citations.
This page explains that mechanism as Perplexity documents it: what happens to a question, which bots touch your site, how sources are chosen and shown, and how the search modes differ. The playbook for getting cited on Perplexity is a separate page; this one is about how the machine runs.
What happens when you ask Perplexity a question
Perplexity interprets the question, searches the web, compiles the most relevant material into one answer, and cites every source it used. Its help centre breaks the process into exactly those four parts, per How does Perplexity work.
The interpretation step is what separates it from a keyword search. Perplexity says it uses "cutting-edge language models like GPT-5 and Claude 4.6 Sonnet to understand the context and nuances of your query" before any retrieval happens, so a badly phrased question is rewritten into something searchable.
At a glance
| Stage | What Perplexity says happens | What it means for a page |
|---|---|---|
| Understanding | A language model interprets the question and its context | The query that hits the index is not the user's literal words |
| Searching | It searches the internet for "authoritative sources like articles, websites, and journals" | Your page must be findable and fetchable at that moment |
| Summarising | It compiles the most relevant insights into one answer | Only the extractable parts of your page are used |
| Citing | Each answer carries numbered citations to the original sources | The citation is the visibility; there is no ranked list |
Follow-ups and memory
Perplexity keeps the context of the conversation, so a follow-up question is interpreted against the earlier ones. Its help centre calls this contextual memory. For a brand, it means the first answer in a thread frames every later question the user asks.
How Perplexity finds and fetches pages
Perplexity reaches your site through two user agents: PerplexityBot, which indexes pages for its search results, and Perplexity-User, which fetches a page at question time. Both are documented, with published IP lists, per Perplexity's crawler doc.
PerplexityBot builds the index
PerplexityBot "is designed to surface and link websites in search results on Perplexity" and, in Perplexity's own words, "is not used to crawl content for AI foundation models". Allow it in robots.txt and permit its published IP ranges if you want to appear in results. Perplexity says robots.txt changes can take up to 24 hours to take effect.
Perplexity-User fetches at question time
Perplexity-User visits a page "to help provide an accurate answer and include a link to the page in its response". Because a person triggered the request, Perplexity says this fetcher "generally ignores robots.txt rules". It controls which sites user requests can reach, and it is not used for crawling or training either.
Why firewalls matter more than robots.txt
Perplexity publishes WAF guidance for Cloudflare and AWS because a bot-blocking rule is the most common reason a page is never fetched. Its recommended rule matches the user-agent string and the IP source address together, using the JSON lists at perplexity.com/perplexitybot.json and perplexity.com/perplexity-user.json. The robots.txt rules for AI bots are only half the access question; the firewall is the other half.
How Perplexity picks and shows its sources
Perplexity does not publish a ranking formula; what it publishes is the kind of source it looks for and how the result is shown. The What is Perplexity page says responses are "supported by citations from reputable news organizations, academic publications, and established content sources", per Perplexity's own description.
What a citation looks like
Every answer carries numbered citations that link to the original page, and Perplexity calls this "source transparency". The numbers let a reader trace which claim came from which page. There is no separate ranked list of results underneath.
What the worked example shows
Perplexity's own example is the question "What are the health benefits of yoga?". It says it will scan health and fitness websites, medical journals and yoga blogs, compile the key facts into an overview, and return a concise answer with citations. Three source types, one answer, and a citation per fact is the pattern for almost every query.
The honesty clause
Perplexity's answer-engine page ends with "we encourage you to double-check sources for added confidence". The engine can retrieve a wrong or stale page and summarise it faithfully, so a citation is evidence the page was used, not that it was right. For a brand, that cuts both ways: an outdated page of yours can be cited as confidently as a current one.
Quick Search vs Pro Search vs Research
Perplexity offers a standard Search, a Pro Search with model choice and deeper retrieval, and a Research mode that runs an autonomous multi-search report. The differences are in how much retrieval happens and which model writes the answer, not in the four-stage mechanism.
| Mode | What Perplexity says it does | Who gets it |
|---|---|---|
| Search | The standard interpret, search, summarise and cite flow | Everyone |
| Pro Search | "Thorough research across multiple sources" with a model selector: Best (default), GPT-5, Claude 4.6 Sonnet and more | Free users get a limited number per day; Pro subscribers get extended access |
| Research | "Performs dozens of searches automatically, reads hundreds of sources" and delivers a report in "2-4 minutes" | Described under the Research feature |
Why the mode changes your odds of being cited
A standard search fetches a handful of pages; Research reads hundreds. A page that is the fourth-best source on a topic may never appear in a quick answer and may be cited routinely in a Research report. When you measure share of voice across AI engines, record which Perplexity mode produced each reading, because the two are different populations.
Model choice and the answer
On Pro Search the summarising model is the user's choice, and Perplexity also lists a Model Council feature and file and app creation on its help centre home. The retrieval is the same; the prose that comes out, and which fetched fact gets emphasised, can differ by model. That is one reason the same prompt can cite the same page and describe it differently.
The Comet browser
Perplexity also ships Comet, a browser for Mac, Windows, iOS and Android with the assistant built in, per the Comet page. Its product page describes tasks such as comparing coverage across news outlets and drafting email replies; it does not publish a separate retrieval mechanism, so treat it as the same engine reached from a browser.
How Perplexity differs from a search engine
Perplexity delivers one synthesised answer with citations instead of a list of links, and it sources the web in real time as you ask. Follow-ups stay in the same thread. Its help centre frames the contrast directly: traditional engines "present you with lots of links to sift through", while Perplexity is "delivering the precise knowledge you need without the extra steps and clicks".
What that changes for a brand
There is no position two. A page is either among the sources the answer used, and named in a citation, or it is absent. Ranking well in a conventional index still matters because the retrieval step has to find you, but it is the extractable sentence on the page that earns the citation, which is the whole premise of answer engine optimization.
Real-time retrieval and recency
Because Perplexity fetches at question time, the version of your page it reads is whatever is live now. That makes it the engine where a refreshed page shows up fastest and where a stale one is punished fastest. The comparison of Perplexity and ChatGPT for brand visibility goes further into that difference.
What the mechanism means for a page that wants to be cited
A page passes three checks in order: the bots can reach it, retrieval finds it relevant, and the model can lift a clear sentence from it. Failing any one produces the same result, no citation, which is why diagnosis has to be done in that order.
Access
Check your logs for both user agents over 30 days and verify the IPs against Perplexity's JSON lists. A run of 403s or challenges from a CDN means the firewall is refusing them, and Perplexity's own WAF guidance is the fix.
Relevance and extraction
If the page is fetched and never cited, the problem is on the page. Perplexity looks for articles, sites and journals it treats as authoritative, then lifts the parts that answer the rewritten question. Pages that state the answer first in a structure a model can quote are the ones that survive the summarising step; pages that make the reader work for it do not.
Where Perplexity sits among the engines
Perplexity is one of several AI search engines with live retrieval, and its citation-first design makes it the clearest window into whether your pages are being read. If Perplexity fetches your page and cites it, the content is extractable; if it fetches and ignores it, you have found the page to rewrite first.

