Qortora · Search · Indexed page

commoncrawl.orgFetched 2026-08-17T12:17:15Z

Common Crawl - Blog

Explore Common Crawl's latest updates, insights, and stories. Stay informed on web data trends and our community's impact.

Open original source · Full cached text

Common Crawl - Blog The Data OverviewCDXJ IndexURL IndexWeb GraphsLatest CrawlCrawl StatsGraph StatsErrata Resources Get StartedAI AgentBlogExamplesCCBotInfra StatusOpt-Out LedgerFAQ Community Research PapersMailing List ArchiveHugging FaceDiscordCollaborators About AboutTeamJobsPrivacy PolicyTerms of Use Search AI Agent Contact Us Blog The latest news, interviews, technologies, and resources. News Announcing the First Stable Release of CC-Downloader Over a year ago we released cc-downloader, an experimental tool to politely download Common Crawl data. Today we're releasing its first stable version, with a Rust library and Python bindings. Pedro Ortiz Suarez Pedro is a Principal Research Scientist at the Common Crawl Foundation. Filter by Category or Search by Title See All Crawl Release Web Graphs News Analysis Thank you! Your submission has been received! Oops! Something went wrong while submitting the form. News Announcing the First Stable Release of CC-Downloader Over a year ago we released cc-downloader, an experimental tool to politely download Common Crawl data. Today we're releasing its first stable version, with a Rust library and Python bindings. Pedro Ortiz Suarez Pedro is a Principal Research Scientist at the Common Crawl Foundation. News Notes from HTTP Workshop Basel and IETF 126 Vienna Two weeks in Basel and Vienna, at the HTTP Workshop and IETF 126. Protocol adoption measured across the whole web, and an attempt to define what "machine readable" actually means. Thom Vaughan Thom is Principal Engineer at the Common Crawl Foundation. Web Graphs Host- and Domain-Level Web Graphs May, June, and July 2026 We are pleased to announce a new release of host-level and domain-level web graphs based on the crawls of May, June, and July 2026, consisting of 240.4 million nodes and 3.7 billion edges at the host level, and 118.0 million nodes and 2.8 billion edges at the domain level. Sebastian Nagel Sebastian is a Distinguished Engineer at the Common Crawl Foundation. Crawl Release July 2026 Crawl Archive Now Available The crawl archive for July 2026 is now available. The data was crawled between July 7th and July 25th, and contains 2.14 billion web pages (or 364.01 TiB of uncompressed content). We also announce some improvements and changes. Sebastian Nagel Sebastian is a Distinguished Engineer at the Common Crawl Foundation. News Common Crawl Joins Project Tapestry Common Crawl has joined Project Tapestry, a global initiative led by the AI Alliance to advance open, sovereign AI. We will contribute our expertise in responsible web data, multilingual coverage and culturally informed AI development. Common Crawl Foundation Common Crawl builds and maintains an open repository of web crawl data that can be accessed and analyzed by anyone. News Measuring Crawled Coverage of a Website in Common Crawl How can we measure how many pages we’ve crawled from a particular website? The answer is a lot more complicated than you might think. Greg Lindahl Greg is Chief Technology Officer at the Common Crawl Foundation. Analysis Is one vantage point enough? IPv6 across the top million web hosts We probed the top 1,000,000 web hosts for IPv6 from five vantage points on three continents. 31.7% work from everywhere, one vantage point turns out to be enough for the headline rate, and 5,530 hosts reveal why it isn't enough for the rest of the story. Thom Vaughan Thom is Principal Engineer at the Common Crawl Foundation. Analysis Turning 30,000 Arabic Domains Into a Better Crawl How we filtered, geolocated and categorised a donation of Arabic seed domains Laurie Burchell Laurie is a Principal Research Engineer at the Common Crawl Foundation. News Common Crawl Foundation at LREC 2026 The Common Crawl team attended the 16th International Conference on Language Resources and Evaluation in Palma, Mallorca, co-organizing a tutorial, presenting recent published work, and strengthening links with the research community. Malte Ostendorff Malte is a Senior Research Engineer at Common Crawl. News 13th Web-as-Corpus Workshop @ EMNLP 2026 The WaC-13 workshop invites research submissions on web data, corpus building, and linguistic analysis. Laurie Burchell Laurie is a Principal Research Engineer at the Common Crawl Foundation. Web Graphs Host- and Domain-Level Web Graphs April, May, and June 2026 We are pleased to announce a new release of host-level and domain-level web graphs based on the crawls of April, May, and June 2026. The graphs consist of 247.3 million nodes and 6.3 billion edges at the host level, and 121.1 million nodes and 3.9 billion edges at the domain level. Luca Foppiano Luca Foppiano is a Senior Engineer at the Common Crawl Foundation. Crawl Release June 2026 Crawl Archive Now Available We are happy to announce the release of the June 2026 crawl archive, consisting of 2.10 billion web pages, or 354.59 TiB of uncompressed content. Luca Foppiano Luca Foppiano is a Senior Engineer at the Common Crawl Foundation. News CommonLID Update: New Tools, Growing Impact CommonLID, a community-built language ID benchmark, has a new website and interactive leaderboard. Its paper was accepted to ACL 2026, with a poster session on 7 July. Source code, a PyPI package, and the dataset are now available. Laurie Burchell Laurie is a Principal Research Engineer at the Common Crawl Foundation. News Common Crawl Foundation at IIPC-WAC 2026 Common Crawl was well represented with contributions at the 2026 IIPC Web Archiving Conference and General Assembly. Common Crawl Foundation Common Crawl builds and maintains an open repository of web crawl data that can be accessed and analyzed by anyone. News The Columnar Index Is Now the URL Index We have renamed the Columnar Index to the URL Index, to be clearer about its purpose and to pave the way for more datasets in a columnar format. Common Crawl Foundation Common Crawl builds and maintains an open repository of web crawl data that can be accessed and analyzed by anyone. Analysis Introducing the AI Visibility Audit A free guide for SEOs and GEOs on how to check whether AI systems can actually reach a site, and how to stay visible in the crawl that trains them. Stephen Burns Stephen Burns is Web Intelligence Lead at the Common Crawl Foundation. Web Graphs Host- and Domain-Level Web Graphs March, April, and May 2026 We are pleased to announce a new release of host-level and domain-level web graphs based on the crawls of March, April, and May 2026. The graphs consist of 262.4 million nodes and 8.1 billion edges at the host level, and 118.8 million nodes and 4.3 billion edges at the domain level. Michael Paris Michael is a Senior Research Engineer at the Common Crawl Foundation. Crawl Release May 2026 Crawl Archive Now Available We are happy to announce the release of the May 2026 crawl archive, consisting of 2.16 billion web pages, or 365.56 TiB of uncompressed content. Michael Paris Michael is a Senior Research Engineer at the Common Crawl Foundation. News April 2026 Crawl Archive Now Available in a Hugging Face Storage Bucket As an early experiment in distributing Common Crawl data through another channel, the April 2026 crawl archive is now available in a Hugging Face Storage Bucket, alongside its existing home on AWS S3. Malte Ostendorff Malte is a Senior Research Engineer at Common Crawl. News You can now build directly on Common Crawl from the browser Browsers can now fetch Common Crawl data directly, no backend needed. Build SQL explorers, snapshot viewers and diff tools as static pages. Thom Vaughan Thom is Principal Engineer at the Common Crawl Foundation. Web Graphs Host- and Domain-Level Web Graphs February, March, and April 2026 We are pleased to announce a new release of host-level and domain-level web graphs based on the crawls of February, March, and April 2026. The graphs consist of 269.0 million nodes and 9.4 billion edges at the host level, and 124.6 million nodes and 4.8 billion edges at the domain level. Luca Foppiano Luca Foppiano is a Senior Engineer at the Common Crawl Foundation. Crawl Release April 2026 Crawl Archive Now Available We are pleased to announce that the crawl archive for April 2026 is now available, containing 2.19 billion web pages or 379.2 TiB of uncompressed content. Luca Foppiano Luca Foppiano is a Senior Engineer at the Common Crawl Foundation. News April 2026 Common Crawl Newsletter Check out our newsletter for April 2026, with updates on what we've been up to. Jen English Jen English is a seasoned professional with a core competency in web content curation, web crawling, taxonomies, and ontology creation. News Announcing a Change to Common Crawl Dataset Size Reporting Common Crawl is switching to reporting dataset sizes in nibbles. As an organisation dedicated to data preservation, we feel it would be remiss to allow this underrepresented unit to fall out of use. Our latest crawl now exceeds 689 tebibbles. Common Crawl Foundation Common Crawl builds and maintains an open repository of web crawl data that can be accessed and analyzed by anyone. Web Graphs Host- and Domain-Level Web Graphs January, February, and March 2026 We are pleased to announce a new release of host-level and domain-level web graphs based on the crawls of January, February, and March 2026. The graphs consist of 270.2 million nodes and 9 billion edges at the host level, and 120 million nodes and 4.4 billion edges at the domain level. Luca Foppiano Luca Foppiano is a Senior Engineer at the Common Crawl Foundation. Crawl Release March 2026 Crawl Archive Now Available We are pleased to announce the release of the March 2026 crawl, containing 1.97 billion web pages, or 344.64 TiB of uncompressed content. We also observed a dramatic increase in fetches over IPv6, explained by the enabling of Happy Eyeballs in the OkHttp library. Luca Foppiano Luca Foppiano is a Senior Engineer at the Common Crawl Foundation. Analysis IPv6 Adoption Across the Top 100K Web Hosts We probed the 100,000 most-linked web hosts for IPv6 support using the Common Crawl Web Graph. Only 36.9% are fully reachable over IPv6, with adoption ranging from 71% among the top 100 to 32% in the long tail. Thom Vaughan Thom is Principal Engineer at the Common Crawl Foundation. News Web Graph Statistics Gets a Proper Upgrade Our Web Graph Statistics site has been updated with interactive charts, a domain lookup tool for tracking harmonic centrality and PageRank over time, mobile improvements, unified rank tables with OR filtering, and merged degree plots. Thom Vaughan Thom is Principal Engineer at the Common Crawl Foundation. Analysis Measuring Web Accessibility from Crawl Archives A WCAG colour contrast audit of 240 top domains using Common Crawl's February 2026 archive finds four in ten colour pairings fall short of accessibility thresholds. Only one in five sites are fully compliant. Thom Vaughan Thom is Principal Engineer at the Common Crawl Foundation. News Announcing the Whirlwind Tour of Common Crawl's Datasets Using Java Introducing the second installment in our Whirlwind Tour series, covering crawl structure, index access, and content extraction, giving developers a practical foundation for building Java-based data workflows. Luca Foppiano Luca Foppiano is a Senior Engineer at the Common Crawl Foundation. Web Graphs Host- and Domain-Level Web Graphs December 2025 and January/February 2026 We're happy to announce the release of the Web Graphs for December 2025 and January/February 2026, consisting of 288.6 million nodes and 12.4 billion edges at the host level, and 134.2 million nodes and 5.4 billion edges at the domain level. Luca Foppiano Luca Foppiano is a Senior Engineer at the Common Crawl Foundation. News Introducing the New Examples & Resources Browser We've replaced our old Examples and Use Cases pages with a single searchable, filterable browser. 119 resources from 115 contributors, all in one place. Search, filter by type or language, sort, and share links. We welcome commu…