Free tool · powered by the Common Crawl index

Is your site in LLM training data?

Common Crawl is the largest public web corpus and a primary training source for ChatGPT, Claude and Gemini. Check which of your domain's URLs it has captured across the latest crawls.

How it works

01

Enter your domain

Type any domain — no login, no crawl of your site.

02

We query Common Crawl

We read the latest monthly Common Crawl indexes for your domain in seconds.

03

See coverage & trend

Get captured URLs, section breakdown, freshness and a coverage trend you can export.

Common Crawl & LLM training data — FAQ

What is Common Crawl?
Common Crawl is a free, open repository of web crawl data — petabytes of pages collected roughly monthly since 2008. It is one of the largest publicly available web corpora and a primary training source for large language models such as ChatGPT, Claude and Gemini. Each monthly crawl is published as a dated index (for example CC-MAIN-2026-30) that anyone can query.
Does being in Common Crawl mean my site is in ChatGPT's training data?
Not with certainty, but it is the strongest public signal available. Common Crawl is a major training source for most large language models, so pages captured in it have had a real chance to enter model training data. It is not a guarantee: not every model ingests every crawl, and vendors filter and deduplicate what they use. Being absent from Common Crawl, however, strongly reduces the chance your content was seen.
How do I get my pages into Common Crawl?
Make sure your pages are publicly reachable and that your robots.txt does not block CCBot, Common Crawl's crawler. Keep an up-to-date XML sitemap, ensure pages return a 200 status, avoid heavy client-side rendering that hides content, and earn inbound links so the crawler discovers you. New or changed pages can take one or more monthly crawls to appear.
Why are none of my pages in Common Crawl?
The most common reasons are: your robots.txt blocks CCBot, the site is new or has few inbound links, pages return errors or redirects, or the content is rendered only in JavaScript that the crawler does not execute. This checker shows which URLs are captured across the latest crawls so you can spot the gap.
How often is Common Crawl updated?
Common Crawl publishes a new crawl roughly once a month, each covering a two-to-three-week window. Because indexes are separate, coverage of a single site can rise or fall from one month to the next — this tool plots that trend across the latest crawls.
Is Common Crawl coverage the same as SEO?
No. SEO optimizes how you rank in traditional search engines like Google. Common Crawl coverage is about whether your content exists in the corpus that trains AI models — the foundation of Generative Engine Optimization (GEO) and Answer Engine Optimization (AEO). A page can rank well in Google yet be missing from Common Crawl, or vice versa.
AI görünürlüğünü takip eden 100'den fazla markaya katılın

AI'ın markanız hakkında
ne düşündüğünü öğrenin

Markanız için ücretsiz AI görünürlük denetimi alın ve ChatGPT, Claude ve Gemini'nin sizi ne sıklıkta önerdiğini ve hangi rakiplerin önde olduğunu görün. LLM Görünürlük Skorunuzu keşfedin ve AI arama varlığınızı geliştirmek için uygulanabilir GEO önerileri alın.

3 gün ücretsiz deneme
2 dakikada sonuç
Ücretsiz plan mevcut
İstediğiniz zaman iptal