Analysis

Perplexity GEO

The engine whose citations track Google rankings most closely, and the one with the most contested crawling record.

Bar chart of where Perplexity's citations concentrate: YouTube 31.2%, Reddit 13.9%, Wikipedia 7.2% of mention share.

Perplexity is the AI answer engine whose citations track Google rankings most closely, which makes it the easiest of the four to influence with ordinary SEO. It’s also the one with the most aggressive crawling record.

Both facts are documented, and both change what you should do. This page covers how Perplexity retrieves, what it demonstrably cites, and where the GEO sales pitch runs ahead of the evidence.

How Perplexity retrieves

Perplexity runs its own crawler and its own index rather than leaning entirely on someone else’s search API. When it launched a Search API in September 2025 it described an index of more than 200 billion unique URLs, retrieved at passage level so individual snippets are scored rather than whole pages.

Two user agents do the work, and they behave differently.

AgentJobrobots.txt
PerplexityBotBuilds the index that surfaces and links sitesRespects disallow rules
Perplexity-UserFetches a page live for one user’s requestGenerally ignored, per Perplexity

Perplexity states plainly that Perplexity-User generally ignores robots.txt, on the reasoning that a person asked for that specific page. It also publishes machine-readable IP ranges at perplexity.com/perplexitybot.json, which is more than most operators do and lets you verify a fetch from your own logs.

PerplexityBot isn’t used for model training, which makes the blocking decision simpler here than for OpenAI or Anthropic. Disallowing it removes you from answers without protecting anything from a training set.

The crawling dispute

On 4 August 2025 Cloudflare published research alleging Perplexity used undeclared crawlers that presented themselves as an ordinary Chrome browser to fetch pages from domains which had blocked PerplexityBot in robots.txt. Cloudflare estimated the undeclared traffic at 3 to 6 million requests a day, against 20 to 25 million from the declared agent, and removed Perplexity from its verified bots programme.

Perplexity denied it the same week. A spokesperson called the report a sales pitch and said the bot Cloudflare named wasn’t Perplexity’s, later attributing the traffic to a third-party service it uses occasionally.

Neither side has moved since, so treat it as contested. The practical consequence is that a robots.txt disallow may not be doing what you assume, and your server logs are the only place you’ll find out.

Eric Wilkinson

Not sure which of this matters for your site?

I am Eric. Give me 30 minutes and I will look at your site and tell you which of these things would move anything for you, and which you can skip.

Book a free call →

What Perplexity cites

Ahrefs analysed 3.1 million US queries in July 2026 and found Perplexity’s citations concentrated hard on a handful of platforms: YouTube at roughly 31.2% mention share, Reddit at 13.9%, Wikipedia at 7.2%.

That distribution says something uncomfortable for the standard content plan. A video and a well-received forum thread reach further into Perplexity answers than another blog post on your own domain will.

Perplexity also overlaps with Google more than its competitors do. In an Ahrefs test across 15,000 long-tail queries, 28.6% of Perplexity’s cited links appeared in Google’s top 10, against 12% for ChatGPT, Gemini and Copilot combined. Ranking in Google is a more direct lever here than anywhere else.

Numbers you’ll see quoted that don’t survive checking include the average citations per Perplexity answer, given variously as 4.7, 8.2 and 21.87 across different agency blogs, none of which discloses a method. Where the “experts” cite a precise figure with no sample size attached, it usually came from another blog post.

What the evidence doesn’t support

Search Atlas tested schema coverage against citation visibility across OpenAI, Gemini and Perplexity in December 2025 and found no reliable relationship. Ahrefs’ controlled difference-in-differences test on 1,885 pages reached the same conclusion for the other engines, at effect sizes indistinguishable from zero.

A claim circulating widely holds that Princeton and Georgia Tech research proved schema markup lifts citation odds by 30 to 40%. The paper being pointed at is the original GEO study by Aggarwal and colleagues, accepted at KDD 2024, and it never tested schema markup. What it tested was content: adding citations, quotations & statistics raised visibility by up to about 40%, while keyword stuffing did not help and sometimes hurt.

That misattribution is worth remembering, because it’s the single most repeated claim in GEO marketing and the source says something different.

What to do about it

Rank in Google first. With 28.6% overlap, ordinary organic work moves Perplexity visibility more reliably than anything Perplexity-specific you could buy.

Then go where the citations concentrate. A YouTube presence & a real reputation on the forums your customers read do more here than another page on your site, which is a genuine departure from a classic SEO plan.

Check your logs before assuming a block worked, given the Cloudflare finding. And weigh PerplexityBot separately from the training crawlers, since blocking it costs you answers without protecting a training set. The Common Crawl bot page walks through that split, and the ChatGPT GEO, Gemini GEO and Claude GEO pages cover where the other engines diverge.

Eric Wilkinson

Eric Wilkinson

I run WilkiLeads, a one-person SEO and web consultancy. I work on search visibility, site rebuilds & the technical side of getting pages indexed, cited and found.

Book a free 30 minute call →
Start here · €99

Find out what’s holding your site back.

I review your website personally and write a prioritised action plan for your exact site: what to fix first, what to ignore, and what will move rankings.