Free tool · powered by the Common Crawl index

Is your site in LLM training data?

Common Crawl is the largest public web corpus and a primary training source for ChatGPT, Claude and Gemini. Check which of your domain's URLs it has captured across the latest crawls.

How it works

01

Enter your domain

Type any domain — no login, no crawl of your site.

02

We query Common Crawl

We read the latest monthly Common Crawl indexes for your domain in seconds.

03

See coverage & trend

Get captured URLs, section breakdown, freshness and a coverage trend you can export.

Common Crawl & LLM training data — FAQ

What is Common Crawl?
Common Crawl is a free, open repository of web crawl data — petabytes of pages collected roughly monthly since 2008. It is one of the largest publicly available web corpora and a primary training source for large language models such as ChatGPT, Claude and Gemini. Each monthly crawl is published as a dated index (for example CC-MAIN-2026-30) that anyone can query.
Does being in Common Crawl mean my site is in ChatGPT's training data?
Not with certainty, but it is the strongest public signal available. Common Crawl is a major training source for most large language models, so pages captured in it have had a real chance to enter model training data. It is not a guarantee: not every model ingests every crawl, and vendors filter and deduplicate what they use. Being absent from Common Crawl, however, strongly reduces the chance your content was seen.
How do I get my pages into Common Crawl?
Make sure your pages are publicly reachable and that your robots.txt does not block CCBot, Common Crawl's crawler. Keep an up-to-date XML sitemap, ensure pages return a 200 status, avoid heavy client-side rendering that hides content, and earn inbound links so the crawler discovers you. New or changed pages can take one or more monthly crawls to appear.
Why are none of my pages in Common Crawl?
The most common reasons are: your robots.txt blocks CCBot, the site is new or has few inbound links, pages return errors or redirects, or the content is rendered only in JavaScript that the crawler does not execute. This checker shows which URLs are captured across the latest crawls so you can spot the gap.
How often is Common Crawl updated?
Common Crawl publishes a new crawl roughly once a month, each covering a two-to-three-week window. Because indexes are separate, coverage of a single site can rise or fall from one month to the next — this tool plots that trend across the latest crawls.
Is Common Crawl coverage the same as SEO?
No. SEO optimizes how you rank in traditional search engines like Google. Common Crawl coverage is about whether your content exists in the corpus that trains AI models — the foundation of Generative Engine Optimization (GEO) and Answer Engine Optimization (AEO). A page can rank well in Google yet be missing from Common Crawl, or vice versa.
Join 100+ brands already tracking AI visibility

Find out what AI
thinks about your brand

Get a free AI visibility audit for your brand and see how often ChatGPT, Claude, and Gemini recommend you – and which competitors are ahead. Discover your LLM Visibility Score and get actionable GEO recommendations to improve your AI search presence.

3 days free trial
Results in 1 minutes
Free tier available
Cancel anytime