Unveilr Book a demo
How Does Perplexity Work Retrieval to Citations

How Does Perplexity Work in 2026

Unveilr banner: How does Perplexity work?

Quick Answer: Perplexity works in four stages: it interprets your question, searches the web in real time, summarises the pages it found, and cites each source. The citations are numbered and link to the original pages. On Pro Search you choose the model, including GPT-5 and Claude 4.6 Sonnet, and Research runs dozens of searches in 2 to 4 minutes.

Perplexity describes itself as an answer engine rather than a search engine, and the difference is mechanical, not cosmetic. A search engine ranks pages and hands you the list; Perplexity retrieves pages, reads them, writes one answer and shows its working as numbered citations.

This page explains that mechanism as Perplexity documents it: what happens to a question, which bots touch your site, how sources are chosen and shown, and how the search modes differ. The playbook for getting cited on Perplexity is a separate page; this one is about how the machine runs.

What happens when you ask Perplexity a question

Perplexity interprets the question, searches the web, compiles the most relevant material into one answer, and cites every source it used. Its help centre breaks the process into exactly those four parts, per How does Perplexity work.

The interpretation step is what separates it from a keyword search. Perplexity says it uses "cutting-edge language models like GPT-5 and Claude 4.6 Sonnet to understand the context and nuances of your query" before any retrieval happens, so a badly phrased question is rewritten into something searchable.

At a glance

Stage What Perplexity says happens What it means for a page
Understanding A language model interprets the question and its context The query that hits the index is not the user's literal words
Searching It searches the internet for "authoritative sources like articles, websites, and journals" Your page must be findable and fetchable at that moment
Summarising It compiles the most relevant insights into one answer Only the extractable parts of your page are used
Citing Each answer carries numbered citations to the original sources The citation is the visibility; there is no ranked list

Follow-ups and memory

Perplexity keeps the context of the conversation, so a follow-up question is interpreted against the earlier ones. Its help centre calls this contextual memory. For a brand, it means the first answer in a thread frames every later question the user asks.

How Perplexity finds and fetches pages

Perplexity reaches your site through two user agents: PerplexityBot, which indexes pages for its search results, and Perplexity-User, which fetches a page at question time. Both are documented, with published IP lists, per Perplexity's crawler doc.

PerplexityBot builds the index

PerplexityBot "is designed to surface and link websites in search results on Perplexity" and, in Perplexity's own words, "is not used to crawl content for AI foundation models". Allow it in robots.txt and permit its published IP ranges if you want to appear in results. Perplexity says robots.txt changes can take up to 24 hours to take effect.

Perplexity-User fetches at question time

Perplexity-User visits a page "to help provide an accurate answer and include a link to the page in its response". Because a person triggered the request, Perplexity says this fetcher "generally ignores robots.txt rules". It controls which sites user requests can reach, and it is not used for crawling or training either.

Why firewalls matter more than robots.txt

Perplexity publishes WAF guidance for Cloudflare and AWS because a bot-blocking rule is the most common reason a page is never fetched. Its recommended rule matches the user-agent string and the IP source address together, using the JSON lists at perplexity.com/perplexitybot.json and perplexity.com/perplexity-user.json. The robots.txt rules for AI bots are only half the access question; the firewall is the other half.

How Perplexity picks and shows its sources

Perplexity does not publish a ranking formula; what it publishes is the kind of source it looks for and how the result is shown. The What is Perplexity page says responses are "supported by citations from reputable news organizations, academic publications, and established content sources", per Perplexity's own description.

What a citation looks like

Every answer carries numbered citations that link to the original page, and Perplexity calls this "source transparency". The numbers let a reader trace which claim came from which page. There is no separate ranked list of results underneath.

What the worked example shows

Perplexity's own example is the question "What are the health benefits of yoga?". It says it will scan health and fitness websites, medical journals and yoga blogs, compile the key facts into an overview, and return a concise answer with citations. Three source types, one answer, and a citation per fact is the pattern for almost every query.

The honesty clause

Perplexity's answer-engine page ends with "we encourage you to double-check sources for added confidence". The engine can retrieve a wrong or stale page and summarise it faithfully, so a citation is evidence the page was used, not that it was right. For a brand, that cuts both ways: an outdated page of yours can be cited as confidently as a current one.

Quick Search vs Pro Search vs Research

Perplexity offers a standard Search, a Pro Search with model choice and deeper retrieval, and a Research mode that runs an autonomous multi-search report. The differences are in how much retrieval happens and which model writes the answer, not in the four-stage mechanism.

Mode What Perplexity says it does Who gets it
Search The standard interpret, search, summarise and cite flow Everyone
Pro Search "Thorough research across multiple sources" with a model selector: Best (default), GPT-5, Claude 4.6 Sonnet and more Free users get a limited number per day; Pro subscribers get extended access
Research "Performs dozens of searches automatically, reads hundreds of sources" and delivers a report in "2-4 minutes" Described under the Research feature

Why the mode changes your odds of being cited

A standard search fetches a handful of pages; Research reads hundreds. A page that is the fourth-best source on a topic may never appear in a quick answer and may be cited routinely in a Research report. When you measure share of voice across AI engines, record which Perplexity mode produced each reading, because the two are different populations.

Model choice and the answer

On Pro Search the summarising model is the user's choice, and Perplexity also lists a Model Council feature and file and app creation on its help centre home. The retrieval is the same; the prose that comes out, and which fetched fact gets emphasised, can differ by model. That is one reason the same prompt can cite the same page and describe it differently.

The Comet browser

Perplexity also ships Comet, a browser for Mac, Windows, iOS and Android with the assistant built in, per the Comet page. Its product page describes tasks such as comparing coverage across news outlets and drafting email replies; it does not publish a separate retrieval mechanism, so treat it as the same engine reached from a browser.

How Perplexity differs from a search engine

Perplexity delivers one synthesised answer with citations instead of a list of links, and it sources the web in real time as you ask. Follow-ups stay in the same thread. Its help centre frames the contrast directly: traditional engines "present you with lots of links to sift through", while Perplexity is "delivering the precise knowledge you need without the extra steps and clicks".

What that changes for a brand

There is no position two. A page is either among the sources the answer used, and named in a citation, or it is absent. Ranking well in a conventional index still matters because the retrieval step has to find you, but it is the extractable sentence on the page that earns the citation, which is the whole premise of answer engine optimization.

Real-time retrieval and recency

Because Perplexity fetches at question time, the version of your page it reads is whatever is live now. That makes it the engine where a refreshed page shows up fastest and where a stale one is punished fastest. The comparison of Perplexity and ChatGPT for brand visibility goes further into that difference.

What the mechanism means for a page that wants to be cited

A page passes three checks in order: the bots can reach it, retrieval finds it relevant, and the model can lift a clear sentence from it. Failing any one produces the same result, no citation, which is why diagnosis has to be done in that order.

Access

Check your logs for both user agents over 30 days and verify the IPs against Perplexity's JSON lists. A run of 403s or challenges from a CDN means the firewall is refusing them, and Perplexity's own WAF guidance is the fix.

Relevance and extraction

If the page is fetched and never cited, the problem is on the page. Perplexity looks for articles, sites and journals it treats as authoritative, then lifts the parts that answer the rewritten question. Pages that state the answer first in a structure a model can quote are the ones that survive the summarising step; pages that make the reader work for it do not.

Where Perplexity sits among the engines

Perplexity is one of several AI search engines with live retrieval, and its citation-first design makes it the clearest window into whether your pages are being read. If Perplexity fetches your page and cites it, the content is extractable; if it fetches and ignores it, you have found the page to rewrite first.

Frequently Asked Questions

Does Perplexity use its own search index or Google's?
Perplexity documents its own crawler, PerplexityBot, whose stated job is to surface and link websites in Perplexity's search results, and it publishes the bot's IP ranges. It does not describe using another company's index in its help centre or crawler documentation, so treat PerplexityBot's index as the documented retrieval source.
Which AI models does Perplexity use to write answers?
Perplexity's help centre names OpenAI's GPT-5 and Anthropic's Claude 4.6 Sonnet, plus a default Best option that picks the model for the question, with the selector available on Pro Search. The model writes the summary; retrieval happens before the model choice applies, so citations come from the same fetch regardless.
Does Perplexity train AI models on my website?
Perplexity states that PerplexityBot is not used to crawl content for AI foundation models and that Perplexity-User is not used for training either. Both bots are documented as serving search results and user requests only. That is a narrower purpose than some other vendors' training crawlers, which have separate user agents.
Why does Perplexity ignore my robots.txt?
Only Perplexity-User does, and only because a person asked for the page. Perplexity says user-triggered fetches "generally ignore robots.txt rules", while PerplexityBot, the indexing crawler, respects them. To stop user fetches you would need a firewall rule, which also removes you from answers where a user wanted your page.
How many sources does Perplexity cite per answer?
Perplexity does not publish a fixed number. A standard search compiles "the most relevant insights" from the pages it fetched, while Research mode reads hundreds of sources and can cite far more. The count you see depends on the mode, the question and how many fetched pages contributed a distinct fact.
Is Pro Search free?
Partly. Perplexity says free users can run a limited number of Pro Searches per day, and Pro subscribers get extended access along with the model selector. Standard Search has no such limit noted on the help page, so most everyday questions are answered through the free four-stage flow.
Can I see which sources Perplexity used?
Yes. Every answer carries numbered citations that link to the original pages, and Perplexity describes verifying them as part of the intended use. Selecting a number opens the source. Because the engine encourages readers to double-check, a cited page may receive a click from someone verifying the claim.

About the Author

Sanditya Srivastava is the founder of Unveilr, an answer engine optimization (AEO) service that helps brands get cited and recommended across AI search platforms like ChatGPT, Perplexity, Google AI Overviews, Gemini, and Claude. He writes about how AI search is reshaping brand discovery.