Category: SEO

  • Claude GEO

    Claude GEO

    Claude has the thinnest evidence base of any major AI assistant, and that shapes what honest advice looks like. Anthropic has never publicly named the search index behind Claude’s web search, and no credible large-sample study of what Claude cites exists.

    Anyone selling you a Claude GEO package is therefore working from inference. Some of that inference is reasonable, and some of the numbers being quoted at you were invented by a content marketer.

    What follows is what Anthropic has documented, what independent researchers have established, and where the honest answer is that nobody knows.

    How Claude gets web pages

    Web search reached Claude.ai as a research preview for paid US accounts on 20 March 2025, then opened to all plans worldwide on 27 May 2025. Developers got a web_search tool on the Messages API on 7 May 2025, billed at $10 per 1,000 searches on top of tokens.

    Anthropic runs three crawlers with separate jobs, and blocking one leaves the other two working.

    Agent What Anthropic says it does
    ClaudeBot Collects web content for training
    Claude-User Fetches a page when a user’s query needs it
    Claude-SearchBot Crawls to improve search result quality

    All three honour robots.txt, including Crawl-delay, and each has to be disallowed by name, per subdomain. Anthropic doesn’t publish crawler IP ranges and says IP blocking isn’t a reliable substitute, since its bots run from public cloud address space.

    Citations are always on. In the API, every web search result carries the source URL, the title & up to 150 characters of the passage Claude drew from, and those citation fields don’t count toward your token bill.

    The search index nobody has confirmed

    Anthropic has never stated which index backs Claude’s web search, in documentation, a blog post or a press statement. That gap has been filled by circumstantial evidence rather than a company answer.

    On 19 March 2025, one day before web search launched, Anthropic’s Trust Center subprocessor list added Brave Search under the Web Search category. The independent researcher Simon Willison then matched Claude’s returned results against Brave’s own output and found a BraveSearchParams field inside Claude’s tool-call schema.

    That evidence is good. It’s still inference, and treating it as confirmed is how bad advice gets built. If Brave is the retrieval layer, then Brave’s index coverage of your site is the constraint, which is a different optimisation target from Google’s.

    [[WL_CTA]]

    Why there is no Claude citation data

    ChatGPT and Perplexity have been measured at scale by several data vendors. Claude has not, and the reason is mundane: it holds a smaller share of consumer search behaviour, so the tools that scrape AI answers point at bigger targets.

    Figures do circulate. Claude averages 5.67 citations per response. It overlaps 86.7% with Brave’s top results. Structured databases account for 68% of its source mix. Every one of those numbers traces back to an agency blog with no disclosed sample, no methodology and no date.

    The “experts” quoting them to you have not run a study. They have read another agency’s blog post, which read another one.

    What has been measured across Claude is stability, and the result should make anyone selling AI rank tracking uncomfortable. SparkToro and Gumshoe ran roughly 3,000 identical prompts across Claude, ChatGPT and Google’s AI with about 600 volunteers, and the odds of the same brand list appearing twice came in under 1 in 100.

    What carries over from ordinary SEO

    Retrieval systems that read the open web reward the same things, whichever index sits underneath. A page has to exist, be crawlable, answer the question in plain language and be linked to from places that matter.

    The measured correlations from the larger engines are the best available proxy. Across 75,000 brands, Ahrefs found branded web mentions correlated at 0.664 with AI citation, against 0.217 to 0.254 for raw backlink count, so being talked about beats being linked to on its own.

    The tactics with no evidence behind them have none here either. A controlled test on 1,885 pages found schema markup moved AI citations by an amount indistinguishable from zero, and Ahrefs found 97% of llms.txt files got zero requests of any kind in May 2026. Nothing about Claude changes that arithmetic.

    What to do about it

    Decide the three Claude bots deliberately rather than in one blanket rule. Blocking ClaudeBot keeps your pages out of future training; blocking Claude-User and Claude-SearchBot removes you from answers people are actively asking for. The same split logic is walked through on the Common Crawl bot page.

    There’s a live reason to think about training separately from search. Anthropic settled a US authors’ copyright class action for $1.5 billion, covering an estimated 482,000 works, with final court approval on 20 July 2026. Publishers weighing a licensing position have something concrete to point at.

    Beyond that, do the work you’d do anyway: pages that answer the question, links from real publications & a brand that gets mentioned where your category is discussed. For where the engines genuinely diverge, the ChatGPT GEO, Gemini GEO and Perplexity GEO pages have the specifics.

  • Perplexity GEO

    Perplexity GEO

    Perplexity is the AI answer engine whose citations track Google rankings most closely, which makes it the easiest of the four to influence with ordinary SEO. It’s also the one with the most aggressive crawling record.

    Both facts are documented, and both change what you should do. This page covers how Perplexity retrieves, what it demonstrably cites, and where the GEO sales pitch runs ahead of the evidence.

    How Perplexity retrieves

    Perplexity runs its own crawler and its own index rather than leaning entirely on someone else’s search API. When it launched a Search API in September 2025 it described an index of more than 200 billion unique URLs, retrieved at passage level so individual snippets are scored rather than whole pages.

    Two user agents do the work, and they behave differently.

    Agent Job robots.txt
    PerplexityBot Builds the index that surfaces and links sites Respects disallow rules
    Perplexity-User Fetches a page live for one user’s request Generally ignored, per Perplexity

    Perplexity states plainly that Perplexity-User generally ignores robots.txt, on the reasoning that a person asked for that specific page. It also publishes machine-readable IP ranges at perplexity.com/perplexitybot.json, which is more than most operators do and lets you verify a fetch from your own logs.

    PerplexityBot isn’t used for model training, which makes the blocking decision simpler here than for OpenAI or Anthropic. Disallowing it removes you from answers without protecting anything from a training set.

    The crawling dispute

    On 4 August 2025 Cloudflare published research alleging Perplexity used undeclared crawlers that presented themselves as an ordinary Chrome browser to fetch pages from domains which had blocked PerplexityBot in robots.txt. Cloudflare estimated the undeclared traffic at 3 to 6 million requests a day, against 20 to 25 million from the declared agent, and removed Perplexity from its verified bots programme.

    Perplexity denied it the same week. A spokesperson called the report a sales pitch and said the bot Cloudflare named wasn’t Perplexity’s, later attributing the traffic to a third-party service it uses occasionally.

    Neither side has moved since, so treat it as contested. The practical consequence is that a robots.txt disallow may not be doing what you assume, and your server logs are the only place you’ll find out.

    [[WL_CTA]]

    What Perplexity cites

    Ahrefs analysed 3.1 million US queries in July 2026 and found Perplexity’s citations concentrated hard on a handful of platforms: YouTube at roughly 31.2% mention share, Reddit at 13.9%, Wikipedia at 7.2%.

    That distribution says something uncomfortable for the standard content plan. A video and a well-received forum thread reach further into Perplexity answers than another blog post on your own domain will.

    Perplexity also overlaps with Google more than its competitors do. In an Ahrefs test across 15,000 long-tail queries, 28.6% of Perplexity’s cited links appeared in Google’s top 10, against 12% for ChatGPT, Gemini and Copilot combined. Ranking in Google is a more direct lever here than anywhere else.

    Numbers you’ll see quoted that don’t survive checking include the average citations per Perplexity answer, given variously as 4.7, 8.2 and 21.87 across different agency blogs, none of which discloses a method. Where the “experts” cite a precise figure with no sample size attached, it usually came from another blog post.

    What the evidence doesn’t support

    Search Atlas tested schema coverage against citation visibility across OpenAI, Gemini and Perplexity in December 2025 and found no reliable relationship. Ahrefs’ controlled difference-in-differences test on 1,885 pages reached the same conclusion for the other engines, at effect sizes indistinguishable from zero.

    A claim circulating widely holds that Princeton and Georgia Tech research proved schema markup lifts citation odds by 30 to 40%. The paper being pointed at is the original GEO study by Aggarwal and colleagues, accepted at KDD 2024, and it never tested schema markup. What it tested was content: adding citations, quotations & statistics raised visibility by up to about 40%, while keyword stuffing did not help and sometimes hurt.

    That misattribution is worth remembering, because it’s the single most repeated claim in GEO marketing and the source says something different.

    What to do about it

    Rank in Google first. With 28.6% overlap, ordinary organic work moves Perplexity visibility more reliably than anything Perplexity-specific you could buy.

    Then go where the citations concentrate. A YouTube presence & a real reputation on the forums your customers read do more here than another page on your site, which is a genuine departure from a classic SEO plan.

    Check your logs before assuming a block worked, given the Cloudflare finding. And weigh PerplexityBot separately from the training crawlers, since blocking it costs you answers without protecting a training set. The Common Crawl bot page walks through that split, and the ChatGPT GEO, Gemini GEO and Claude GEO pages cover where the other engines diverge.

  • Gemini GEO

    Gemini GEO

    Google is the one engine that has told you, in writing, exactly what to do about AI answers. The instruction is to keep doing SEO, and it’s published on Google’s own developer site.

    That leaves an awkward gap between what the documentation says and what a GEO retainer is being sold for. This page works through the mechanism Google has described, the citation data, and the traffic numbers behind the claim that search is dying.

    Query fan-out, in Google’s own words

    AI Mode doesn’t run your question as one search. Google describes the technique on its product blog as breaking a question into sub-topics and “issuing a multitude of queries simultaneously on your behalf,” across the web, the Knowledge Graph and the Shopping Graph.

    Robby Stein, Google’s VP of Product for Search, put it more plainly in April 2025.

    For any question, it makes a plan, breaks it down into related subtopics, and runs multiple Google searches to find the most helpful and reliable info.

    Robby Stein, VP of Product, Google Search, 16 April 2025

    The Gemini API documents the same behaviour for developers: the model decides whether a search helps, then “automatically generates one or multiple search queries and executes them.” Google bills each of those as a separate grounded use, which is a useful confirmation that the fan-out is real rather than a metaphor.

    This is the single most important mechanical difference from classic search. You aren’t competing for one query any more. You’re competing across a spray of sub-queries you never see, which is why pages ranking for adjacent long-tail terms get pulled into answers about something broader.

    What Google tells site owners to do

    Google’s guidance page on AI features is unusually direct, and it closes off three of the things being sold as GEO.

    There are no additional requirements to appear in AI Overviews or AI Mode. You don’t need to create new machine readable files, AI text files, or markup to appear in these features. There’s also no special schema.org structured data that you need to add.

    Google Search Central, AI features and your website, updated 10 December 2025

    The controls it does offer are the ones you already have: nosnippet, data-nosnippet, max-snippet and noindex, the same directives that governed classic snippets.

    Danny Sullivan, Google’s Search Liaison, made the same point on the Search Off the Record podcast in December 2025, saying structured data “still matters” while adding that “it’s not structured data and you win AI.” John Mueller, asked on Reddit whether extensive schema helps language models, answered “the short answer is yes, no, and it depends,” and flagged it as his personal view rather than official guidance.

    The controlled evidence agrees with the documentation. Ahrefs tested 1,885 pages that added JSON-LD against roughly 4,000 matched controls and found AI Mode citations moved 2.4%, indistinguishable from zero, while AI Overview citations fell 4.6%.

    [[WL_CTA]]

    How much AI citation tracks organic ranking

    Every large study finds a real relationship and none of them agree on its size. That disagreement is the most useful thing in this section, because it tells you how much confidence any single number deserves.

    Study Sample Cited pages also ranking organically
    Ahrefs, July 2025 ~1.9m AI Overview citations 76% in the top 10
    seoClarity, October 2025 362,000 queries ~56% in the top 20
    BrightEdge, September 2025 9 industries, 16 months 54.5% in the top 10
    Ahrefs, March 2026 863,000 keywords, ~4m URLs 38% in the top 10

    Ahrefs’ own figure halved in eight months on the same style of measurement, and nobody has explained why. In its March 2026 data the citations that weren’t in the top 10 split almost evenly between positions 11 to 100 and pages ranking beyond 100.

    Two conclusions survive the spread. Ranking well makes citation substantially more likely, and it guarantees nothing, because a page nobody can find in the SERP can still be pulled into an answer.

    The claim that SEO is dying

    Per-query, AI Overviews suppress clicks. Pew Research tracked 68,879 real Google searches from 900 US adults in March 2025 and found users clicked a traditional result on 8% of visits where an AI summary appeared, against 15% where it didn’t. They clicked a link inside the summary on 1% of visits, and ended the session entirely 26% of the time versus 16%.

    In aggregate, the picture is much calmer. Graphite analysed Similarweb data across the top 40,000 US websites and found organic search traffic down 2.5% year over year, with the largest sites up about 1.6%. Google stated in August 2025 that total organic click volume from Search was relatively stable year over year.

    Both are true at once. AI Overviews appear on a minority of queries, weighted toward informational ones, so a real per-query click collapse coexists with roughly flat totals for large sites, and the pain lands on mid-sized publishers.

    Meanwhile AI Mode itself is still tiny. SparkToro’s analysis of Similarweb panel data for January to April 2026 found 0.34% of Google searches moved into AI Mode. Organic search remains the channel most websites live on, and treating it as finished is a good way to lose the traffic you still have.

    Google-Extended and what it doesn’t do

    Google-Extended is a robots.txt token rather than a crawler. Google’s crawler documentation states it “doesn’t have a separate HTTP request user agent string” and that “crawling is done with existing Google user agent strings,” so it’s a permission flag on content Google already fetched.

    What it controls is narrow: whether your content “may be used for training future generations of Gemini models.” Google is explicit about the rest.

    Google-Extended does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.

    Google Search Central, Google crawlers documentation

    Agencies conflate those two things constantly, usually while quoting a fear about disappearing from Google. Disallowing Google-Extended is a decision about training data, and the related split across the other crawlers is worked through on the Common Crawl bot page.

    What to do about it

    Build for the fan-out. Cover the sub-questions around your main topic properly, on pages that can each answer one thing completely, because the fan-out is what surfaces them.

    Rank. Every study above says organic position raises citation odds, so the ordinary work of content quality & earned links is the work. Then get mentioned on the third-party pages that already rank for your category, since the fan-out reaches those too.

    Leave the schema rebuild and the llms.txt file alone until someone shows you a controlled test. Google has now told you twice, in its documentation and through its Search Liaison, that neither is required. The other engines differ in the details, covered on the ChatGPT GEO, Perplexity GEO and Claude GEO pages.

  • ChatGPT GEO

    ChatGPT GEO

    Most of what gets sold as GEO is SEO with a new invoice attached. The things that get a page cited in ChatGPT are, with a handful of exceptions, the things that got it ranking in Google in the first place.

    The exceptions are real and worth knowing. This page separates them from the sales pitch using published data, and it names the studies so you can check them yourself.

    Money is arriving regardless of what the mechanics turn out to be. Profound, one of the tools built to measure AI visibility, raised a $96 million Series C at a $1 billion valuation in February 2026. Gartner’s 2026 CMO Spend Survey puts 15.3% of marketing budgets against AI. Budgets are moving, so the category is becoming real whether or not the tactics are.

    How ChatGPT reads the web

    OpenAI runs four crawlers, and they do different jobs. Confusing them is the most expensive mistake on this page, because two of the four decide whether you can be cited at all.

    Agent What it does
    GPTBot Collects content that may be used to train foundation models
    OAI-SearchBot Surfaces websites in ChatGPT’s search features
    ChatGPT-User Fetches a page live when a user’s question needs it
    OAI-AdsBot Checks the safety of pages submitted as ads

    Block GPTBot and your content stays out of training. Block OAI-SearchBot and, in OpenAI’s own words, your pages “will not be shown in ChatGPT search answers, though can still appear as navigational links.” Plenty of sites blocked both in one line and then wondered why they went quiet.

    ChatGPT-User is the odd one. OpenAI documents that because the fetch is triggered by a person rather than a schedule, robots.txt rules may not apply to it.

    Retrieval doesn’t happen on every question. A Semrush analysis of 80 million ChatGPT clickstream records found roughly 46% of queries triggered a live web search, with the rest answered from what the model already held. For those, no amount of on-page work reaches the answer.

    What the citation studies show

    The largest study of what predicts a ChatGPT citation points at links. SE Ranking analysed 129,000 domains across 216,524 pages in 20 niches and found referring domains was the single strongest correlate: sites with up to 2,500 referring domains averaged 1.6 to 1.8 citations, while sites above 350,000 averaged 8.4.

    Ahrefs ran a different cut across 75,000 brands in December 2025 and found branded web mentions correlated at 0.664 with ChatGPT citation, against 0.217 to 0.254 for raw backlink count. Mentions of a brand on YouTube scored highest of anything measured, at roughly 0.737.

    Read those two together and the picture isn’t mysterious. Being widely linked and widely talked about is what predicts citation, which is the same thing that has predicted rankings for twenty years.

    One finding cuts the other way and deserves airtime. Ahrefs tested 15,000 long-tail queries in July 2025 and found only 12% of links cited by ChatGPT, Gemini and Copilot appeared in Google’s top 10 for the same prompt. Ranking first isn’t a ticket to being quoted, and a page buried at position 40 can still be pulled into an answer.

    Every study named here was run by a company that sells SEO software. The methodologies are disclosed and the samples are large, which is more than most of this field offers, and the incentive is still worth holding in mind.

    [[WL_CTA]]

    The tactics with no evidence behind them

    Schema markup is the one you’ll be sold hardest. Ahrefs ran the only properly controlled test I know of: 1,885 pages that added JSON-LD between August 2025 and March 2026, measured against roughly 4,000 matched control pages using difference in differences. ChatGPT citations moved 2.2%, statistically indistinguishable from zero.

    Google says the same thing in its own documentation, which is unusually blunt for Google.

    There are no additional requirements to appear in AI Overviews or AI Mode. You don’t need to create new machine readable files, AI text files, or markup to appear in these features. There’s also no special schema.org structured data that you need to add.

    Google Search Central, AI features and your website, updated 10 December 2025

    Then there’s llms.txt. Ahrefs checked server logs for all 137,210 domains in its analytics panel and found 97% of published llms.txt files received zero requests in May 2026, from anything. Of the requests that did arrive, about 1% came from an identifiable AI bot; the rest were SEO audit tools.

    Asked directly whether Google endorsed llms.txt, John Mueller answered on Bluesky in January 2026: “I’m tempted to say something snarky since this has come up so often, but to be direct, no.”

    E-E-A-T gets sold as a lever too. Google’s own documentation says that “while E-E-A-T itself isn’t a specific ranking factor,” a mix of signals that identify content with good E-E-A-T is useful, and that “rater data is not used directly in our ranking algorithms”. It’s a description of what good looks like, not a setting.

    The “experts” selling AI rank tracking have a harder problem still. SparkToro and Gumshoe ran roughly 3,000 identical prompts across ChatGPT, Claude and Google’s AI with about 600 volunteers, and found the odds of getting the same brand list twice were under 1 in 100. Same list in the same order came in near 1 in 1,000.

    What is genuinely different

    Three things behave differently enough to change what you do, and none of them are on-page.

    Third-party pages carry more weight than they do in classic search. A study of 30 million cited sources found Reddit the most-cited domain across every major AI answer engine, ahead of YouTube, LinkedIn and Wikipedia. Being the subject of a good thread outranks having a good page about yourself.

    Unlinked brand presence matters more than link count, per the Ahrefs correlation work above. Getting written about, reviewed & listed is closer to the job than building links to a target URL.

    Directories and aggregators absorb a disproportionate share of local answers. Whitespark studied 540 queries across three US cities and six service industries in May 2025 and found that for Houston plumbers, 60% of AI citations went to third-party publishers such as Yelp, Thumbtack and Reddit, against 40% to individual businesses.

    What to do about it

    Check your robots.txt first, because it’s the only thing here that can silently remove you. Decide GPTBot and OAI-SearchBot separately, and see the Common Crawl bot page for how the same split works across the other crawlers.

    Then do the SEO. Publish pages that answer the question completely, earn links from places people read, and get your brand mentioned on the third-party pages that already rank for your category. That list isn’t new, which is the point.

    Skip the schema rebuild, the llms.txt file & the E-E-A-T audit until someone shows you a controlled test. The engines differ in the details, so it’s worth reading the Gemini GEO, Perplexity GEO and Claude GEO pages for where they diverge.

  • Common Crawl Alternatives

    Common Crawl Alternatives

    Nothing replaces Common Crawl on its own terms. No other public archive gives away billions of raw pages a month with no account, no contract and no charge beyond bandwidth. The useful question is which substitute fits the job you were using it for.

    Four routes cover almost every case.

    Archives with a different history

    The Internet Archive holds far deeper history than Common Crawl and exposes it through its own CDX API at web.archive.org/cdx/search/cdx, with the same URL, timestamp, status and digest fields. It’s the right source for tracking how one page changed over years. Bulk access is a conversation with the organisation rather than an open bucket.

    The Open Web Index is the closest thing to a public sibling. Built by OpenWebSearch.eu, a Horizon Europe project that started in September 2022 with 14 European research institutions including CERN, it released a federated pan-European index for research and development use in August 2026 after the 42 month programme ended.

    Pre-cleaned corpora built from the same crawls

    Most people reaching for Common Crawl want text for a model, and the filtering is the expensive part. FineWeb, released by Hugging Face in May 2024, is around 15 trillion tokens processed from 96 Common Crawl snapshots between summer 2013 and April 2024, with a 1.3 trillion token educational subset called FineWeb-Edu. Google’s C4 came out of a single 2019 snapshot for training T5.

    These aren’t independent sources. They’re Common Crawl with the cleaning already done, so they inherit its coverage gaps along with its convenience.

    Running your own crawler

    Crawling yourself is the only option that gives you exactly the pages you want, at the freshness you want, under your own robots.txt compliance. It’s also the one with real running costs.

    Apache Nutch is the crawler Common Crawl’s main crawl is built on. StormCrawler, which runs on Apache Storm, is what the foundation uses for its daily CC-NEWS crawl. Scrapy suits smaller, targeted jobs. Budget for storage, politeness delays, IP reputation and the ongoing maintenance of a system that breaks quietly.

    Commercial indexes and datasets

    Ahrefs, Majestic and DataForSEO run their own crawlers & sell API access to link graphs and page data. Coverage is narrower than Common Crawl in raw page count, and considerably better on the freshness of links to a given domain. Vendors such as Bright Data and Zyte sell managed collection instead of an archive.

    If you need Use
    Historical versions of specific pages Internet Archive CDX
    An EU-governed open index Open Web Index
    Cleaned pretraining text FineWeb or FineWeb-Edu
    Exact pages, current, on your terms Nutch, StormCrawler or Scrapy
    Fresh link data for SEO Ahrefs, Majestic or DataForSEO
    Managed collection with support Bright Data or Zyte

    Before switching, confirm the original constraint. A lot of teams look for an alternative because the CDX API keeps returning 503, which the Common Crawl rate limit page solves without leaving the project, usually by moving the same query to the columnar index described on the Common Crawl index page.

    For what the archive holds and how it’s licensed, see Common Crawl.

  • Common Crawl Rate Limit

    Common Crawl Rate Limit

    Common Crawl doesn’t publish a request-per-second number for its URL index. It publishes a behaviour instead: send too many requests in a short window and the CDX server at index.commoncrawl.org starts returning HTTP 503, and your IP can be blocked outright.

    The foundation’s own FAQ describes that endpoint as frequently abused and therefore heavily rate limited. Treat 503 as a stop signal rather than an error to retry through.

    If you’re already blocked

    Common Crawl asks you to wait 24 hours before trying again. Retrying from the same IP during the block, or rotating through proxies to get around it, is the behaviour that caused the limit in the first place.

    What triggers it

    Four things account for most blocks. Parallel workers hitting the index with no delay between calls. Loops that re-query the same URL pattern instead of caching the answer. Plain HTTP requests, which aren’t supported and fail in ways that look like a network problem. Proxy rotation, which the foundation explicitly asks people not to use.

    None of it is exotic. A single unthrottled script that walks a list of 50,000 domains will find the limit within a few minutes.

    Staying under it

    Use HTTPS, always. The index server doesn’t serve plain HTTP, so http://index.commoncrawl.org produces client errors that get misdiagnosed as rate limiting.

    Sleep between calls and stay single-threaded. A short pause per request costs you minutes across a small job and saves you a day of being locked out.

    curl -s "https://index.commoncrawl.org/CC-MAIN-2026-30-index?url=example.com/*&output=json"
    sleep 1

    Cache every response to disk, keyed by URL and crawl id. Most scripts that hit the limit are asking the same question repeatedly across a rerun.

    Back off on 503 rather than retrying immediately. Exponential backoff with a ceiling, then stop the job and look at the logs.

    The fix that scales

    The CDX API is built for interactive lookups, and no amount of politeness turns it into a bulk tool. For anything bigger, the project points people at the columnar index, a Parquet copy of the same data at s3://commoncrawl/cc-index/table/cc-main/warc/.

    Athena, Spark and DuckDB read it directly, there’s no shared server to overload, and the same query that would take 50,000 API calls becomes one SQL statement over a partitioned table. Common Crawl’s own index server page recommends it for bulk filtering & aggregation.

    The Common Crawl index page covers both interfaces and when to reach for each. For background on how the archive is built and what a crawl contains, see Common Crawl.

  • Common Crawl Index

    Common Crawl Index

    The Common Crawl index maps a URL to the exact file, byte offset and length where its archived copy sits. Without it, finding one domain inside the 84.69 TiB of WARC data in a single crawl would mean reading all of it.

    Two indexes exist over the same data, and they suit opposite jobs.

    The CDX API

    The URL index server runs pywb at index.commoncrawl.org and answers queries over HTTPS. Each crawl has its own endpoint named after the crawl id, and the full list of collections is published as machine-readable JSON at /collinfo.json.

    curl -s "https://index.commoncrawl.org/CC-MAIN-2026-30-index?url=example.com/*&output=json"

    Each result line carries the URL, the fetch timestamp, the HTTP status, the MIME type, a content digest, and the three fields that matter for retrieval: filename, offset and length. Those three let you pull one page out of a multi-gigabyte archive with a ranged HTTP request against https://data.commoncrawl.org/.

    The server is heavily rate limited, and going too fast returns 503 and can block your IP for a day. That’s covered on the Common Crawl rate limit page.

    The columnar index

    The columnar index holds the same records in Apache Parquet at s3://commoncrawl/cc-index/table/cc-main/warc/, partitioned by crawl and subset. It carries 26 columns covering URL components, host and registered domain, TLD, MIME type, detected languages, content digest, fetch status & the WARC filename with record offset and length.

    Columnar storage is what makes it fast. A query touching three columns reads only those three, so a per-domain count over a whole crawl can scan a few megabytes rather than terabytes. Common Crawl’s own worked example on Athena scanned 2.12 MB and cost under a cent.

    SELECT url, fetch_status, warc_filename
    FROM ccindex
    WHERE crawl = 'CC-MAIN-2026-30'
      AND subset = 'warc'
      AND url_host_registered_domain = 'example.com'

    Athena, Apache Spark and DuckDB all read it in place. The data sits in the public commoncrawl bucket in us-east-1 and needs no AWS credentials, so --no-sign-request is enough for the S3 CLI.

    Which one to use

    Job Index
    Check whether one page was crawled CDX API
    List a single domain’s crawled URLs CDX API
    Every URL on a TLD Columnar
    Counts, joins or aggregation Columnar
    Anything scripted across thousands of hosts Columnar
    Anything you will rerun Columnar

    The dividing line is simple enough to apply without thinking about it. If a person is waiting for the answer, query the API. If a machine is, query Parquet.

    For what the index points at, and how a crawl is packaged into WARC, WAT and WET files, see Common Crawl.

  • Common Crawl Bot

    Common Crawl Bot

    CCBot is the crawler Common Crawl uses to build its public archive. It reads robots.txt before fetching, respects Crawl-delay, honours nofollow on links, and identifies itself with a user agent that names the project.

    CCBot/2.0 (https://commoncrawl.org/faq/)

    Older log lines may still show CCBot/1.0 (+https://commoncrawl.org/bot.html). Both belong to the same foundation.

    Verifying CCBot in your logs

    A user agent string is free to forge, so scrapers routinely pretend to be CCBot to inherit whatever allowances a site grants it. Common Crawl publishes the IP ranges the real crawler fetches from, which turns verification into a subnet check against your access logs.

    As of 11 August 2026 the published ranges are one IPv6 block, 2600:1f28:365:8000::/56, and four IPv4 blocks: 3.41.188.32/29, 18.97.9.168/29, 18.97.14.80/29 and 18.97.14.88/30. Anything claiming to be CCBot from outside those ranges isn’t.

    Blocking it

    Two lines in robots.txt stop future crawls.

    # robots.txt
    User-agent: CCBot
    Disallow: /

    Slowing it down instead of blocking it takes one more directive, useful on a small server where the crawl itself is the problem.

    User-agent: CCBot
    Crawl-delay: 10

    Check the file for conflicting rules before you add anything. A wildcard User-agent: * block further down doesn’t override a named CCBot group, because crawlers follow the most specific matching group and ignore the rest.

    What blocking does and does not do

    A robots.txt block keeps CCBot out of crawls that haven’t run yet. It reaches nothing that’s already published.

    Every crawl since 2008 stays online in the public S3 bucket, and copies live in derived corpora such as C4 and FineWeb, on academic mirrors, and on the disks of anyone who downloaded them. The foundation’s Opt-Out Ledger, published on 17 September 2025, records legal removal requests from the BBC, the Guardian, the Financial Times, Reuters, the Associated Press & others, plus more than 900 sites entered by the News/Media Alliance. Even those requests apply going forward.

    Blocking also costs you presence in the datasets that non-AI projects run on: link graph research, academic search tools and several small independent search engines all read Common Crawl because it’s the only web-scale corpus they can afford.

    Making the call

    Block CCBot if your business is selling access to your own text, if you’re a publisher acting alongside a licensing position, or if crawler load is measurably hurting a small server.

    Leave it open if you want your pages present in open research data and in the corpora that feed retrieval systems, and if the crawl traffic costs you nothing you’d notice.

    Decide once, write it down, and audit the file quarterly against the other crawlers you care about. For how the archive is built and what ends up in a crawl, see Common Crawl. For other sources of web-scale data, see Common Crawl alternatives.

  • What Is Common Crawl

    What Is Common Crawl

    Common Crawl is a free archive of the public web, published by a US nonprofit that has been collecting snapshots since 2008. The July 2026 crawl alone holds 2.14 billion pages drawn from 33.2 million registered domains. If your site is crawlable, a copy of it probably sits in there right now.

    That archive is the reason the name keeps coming up in conversations about AI. Filtered Common Crawl made up 60% of the training mix for GPT-3, at 410 billion tokens, according to Table 2.2 of the paper OpenAI published in 2020. Almost every open pretraining corpus since then has been distilled from the same source.

    This page covers what the archive contains, how it’s packaged, how to query it without getting blocked, and what a site owner can control.

    What Common Crawl is

    Common Crawl Foundation is a 501(c)(3) nonprofit founded in 2007 by Gil Elbaz, the engineer who built Applied Semantics into the product Google bought & turned into AdSense. Rich Skrenta runs it day to day as executive director. Peter Norvig and Joi Ito have served on its advisory board.

    The organisation does one thing: it crawls the web at scale and gives the results away. Funding came almost entirely from the Elbaz Family Foundation Trust until 2023, when Anthropic and OpenAI each donated $250,000.

    Its own homepage describes the corpus as over 300 billion pages spanning 15 years, cited in more than 10,000 research papers, growing by 3 to 5 billion new pages a month. Nothing sits behind a login. You can download a terabyte of it this afternoon for the cost of the bandwidth.

    Crawl size and schedule

    Crawls ship roughly monthly and carry an identifier in the form CC-MAIN-year-week. CC-MAIN-2026-30 ran from 7 to 25 July 2026 and is a fair sample of what a modern crawl looks like.

    Measure CC-MAIN-2026-30
    Pages fetched 2.14 billion
    Hosts 40.5 million
    Registered domains 33.2 million
    URLs not seen in earlier crawls 603 million
    Uncompressed total 364.01 TiB
    WARC, compressed 84.69 TiB
    WAT, compressed 14.09 TiB
    WET, compressed 5.89 TiB
    URL index 0.22 TiB

    The 603 million figure is the one worth sitting with. Roughly a quarter of every crawl is URLs the project had never fetched before, which is why a page published last month can appear in a corpus a research team downloads next year.

    Crawls aren’t a complete picture of the web and were never meant to be. Common Crawl samples: it follows links, respects robots.txt, and stops. Two consecutive crawls will disagree about which of your pages exist.

    WARC, WAT and WET file formats

    Each crawl ships as three parallel sets of files, all gzipped, all sharded into roughly 100,000 pieces. They hold the same fetches at three levels of detail.

    WARC Raw HTTP request and response WAT Links, headers, metadata WET Extracted plain text Index URL to file offset s3://commoncrawl us-east-1, anonymous read, no charge
    One crawl, four artefacts. WAT, WET and the URL index are all computed from the WARC files, which is why the raw set is fifteen times the size of the plain text.

    WARC files hold the complete fetch: the request Common Crawl sent, the response headers your server returned, and the raw body. This is the format the Internet Archive standardised, and it’s what you want if you care about markup, status codes or HTTP headers.

    WAT files hold computed metadata as JSON. Every outbound link, every title, every header, without the page body. Link graph work almost always starts here rather than in the WARCs.

    WET files hold extracted plain text and nothing else. They’re the smallest of the three at 5.89 TiB compressed for the July 2026 crawl, and they’re what most language model pipelines read first.

    The URL index and the columnar index

    Downloading 84 TiB to find one domain would be absurd, so the project publishes an index that maps URLs to a file, an offset & a length. Two versions exist, and picking the wrong one is the most common mistake people make.

    The CDX server at index.commoncrawl.org answers interactive lookups over HTTPS, and plain HTTP isn’t supported. It runs on pywb, takes a URL pattern plus a crawl identifier, and returns JSON lines. Both interfaces are compared on the Common Crawl index page, and the 503 responses it returns under load on the Common Crawl rate limit page.

    curl -s "https://index.commoncrawl.org/CC-MAIN-2026-30-index?url=example.com/*&output=json"

    The columnar index is the one for real work. It lives at s3://commoncrawl/cc-index/table/cc-main/warc/ in Apache Parquet, carries 26 columns covering URL parts, MIME type, language, fetch status, WARC filename and record offset, and it’s partitioned by crawl. Athena, Spark and DuckDB all read it directly.

    Rule of thumb

    Fewer than a few thousand lookups, interactive, one domain at a time: use the CDX API. Anything larger, anything scripted across many hosts, anything you plan to run twice: use the columnar index. The project says the same thing on its own index server page.

    CCBot, robots.txt and IP ranges

    The crawler identifies itself as CCBot. Its current user agent string is CCBot/2.0 (https://commoncrawl.org/faq/), and it reads robots.txt before fetching, honours Crawl-delay, and follows nofollow on links.

    Blocking it takes two lines.

    # robots.txt
    User-agent: CCBot
    Disallow: /

    Common Crawl also publishes the IP ranges CCBot fetches from, which lets you verify a request in your logs rather than trusting a user agent string anyone can forge. As of 11 August 2026 the list covers one IPv6 block, 2600:1f28:365:8000::/56, and four IPv4 blocks starting at 3.41.188.32/29. The decision itself, and what it costs you, is worked through on the Common Crawl bot page.

    How the data reaches language models

    Common Crawl is raw material, not a training set. Every serious model built on it runs its own filtering pass first, and those passes discard most of what they read.

    GPT-3 is the best documented case. The OpenAI paper lists filtered Common Crawl at 410 billion tokens with a 60% weight in the training mix, sampled at 0.44 epochs across 300 billion training tokens, deliberately undersampled relative to curated sources such as Wikipedia and the two books corpora.

    Google built C4, the Colossal Clean Crawled Corpus, from a single Common Crawl snapshot in 2019 to train T5. Hugging Face released FineWeb in May 2024, around 15 trillion tokens processed from 96 crawls covering summer 2013 to April 2024, along with a 1.3 trillion token educational subset called FineWeb-Edu.

    The practical consequence for a website owner is a lag. A model trained on FineWeb is reading a filtered copy of crawls that ran months or years before the model shipped, so the version of your site an assistant describes may be several rewrites out of date.

    The CC-NEWS dataset

    News moves faster than a monthly crawl, so the project runs a separate one. CC-NEWS was announced on 4 October 2016, built on StormCrawler rather than the Nutch-based main crawler, and it publishes WARC files daily instead of monthly.

    It sits at s3://commoncrawl/crawl-data/CC-NEWS/yyyy/mm/ with a warc.paths.gz manifest per month. Anyone building a news corpus, a media monitoring tool or a dated text collection starts here rather than in CC-MAIN.

    Licensing and the opt-out ledger

    The terms of use grant a limited, non-transferable, non-exclusive licence to use the service, and they push copyright responsibility onto you. Common Crawl states that crawled content may carry its own separate terms from whoever owns it, and that users must respect third-party rights.

    The archive isn’t public domain, and the foundation doesn’t claim otherwise. It reserves the right to remove any crawled content at any time, and it names a copyright agent for infringement claims.

    On 17 September 2025 the foundation published its Opt-Out Ledger, a public list of every legal removal request it has received. Site owners are pointed at robots.txt as the routine control, with legal requests handled by mail to the foundation.

    The 2025 publisher dispute

    On 4 November 2025 The Atlantic published an investigation by Alex Reisner into the archive’s contents. It reported that the corpus held large numbers of paywalled articles from outlets including The Economist, the Los Angeles Times, the Wall Street Journal, the New York Times and The Atlantic itself.

    Reisner reported that the foundation had told publishers it respected paywalls while continuing to distribute the material, and that removal requests were not being carried out. Skrenta, quoted throughout the piece, argued that models should be free to read what people can read.

    US news publishers followed with a cease and desist letter demanding that the foundation stop scraping and delete the archived material. The Opt-Out Ledger now lists the BBC, the Guardian, the Financial Times, the Washington Post, News Corp, DMG Media, Advance Publications, the Associated Press, Le Monde, Reuters and Hearst Newspapers, plus more than 900 news sites entered by the News/Media Alliance.

    One limitation runs through all of it. An opt-out stops future crawls, and it doesn’t reach into corpora already downloaded, mirrored and trained on by third parties around the world.

    What to check on your own site

    Start with your robots.txt, because it’s the only lever that works without a lawyer. Search it for CCBot, GPTBot, ClaudeBot, Google-Extended and PerplexityBot, and confirm the rules say what you think they say.

    Then look up your own domain in the index. A handful of CDX queries will tell you which of your URLs were fetched, what status codes they returned, and whether the crawler is seeing your canonical pages or a mess of parameter variants. Sites frequently discover that most of their crawled URLs are faceted duplicates.

    Weigh the block itself on evidence rather than instinct. Blocking CCBot keeps future crawls out of open corpora, and it also removes you from the datasets that academic search tools, link graph research & several small search projects are built on. There are alternatives to Common Crawl if you need web-scale data yourself, and none of them are free in the same way.

    If you’d rather not run any of this by hand, that’s a normal thing to hand off. I do it as part of a technical review.