Qortora · Search · Indexed page

commoncrawl.orgFetched 2026-09-15T01:54:04Z

Common Crawl - Open Repository of Web Crawl Data

We build and maintain an open repository of web crawl data that can be accessed and analyzed by anyone.

Open original source · Full cached text

Common Crawl - Open Repository of Web Crawl Data The Data OverviewCDXJ IndexURL IndexWeb GraphsLatest CrawlCrawl StatsGraph StatsErrata Resources Get StartedAI AgentBlogExamplesCCBotInfra StatusOpt-Out LedgerFAQ Community Research PapersMailing List ArchiveHugging FaceDiscordCollaborators About AboutTeamJobsPrivacy PolicyTerms of Use Search AI Agent Contact Us Common Crawl maintains a free, open repository of web crawl data that can be used by anyone. Common Crawl is a 501(c)(3) non–profit founded in 2007. ‍ We make wholesale extraction, transformation, and analysis of open web data accessible to researchers. Overview Over 300 billion pages spanning 15 years. Free and open corpus since 2007. Cited in over 10,000 research papers. 3–5 billion new pages added each month. Featured Papers Open English–French datasets show neural quality filters amplify benchmark contamination Nathan Godey, et al. Gaperon: A Peppered English-French Generative Language Model Suite The largest validated African text-and-speech dataset to date Sheriff Issaka, et al. The African Languages Lab: A Collaborative Approach to Advancing Low-Resource African NLP How LLMs erase Northeast Indian languages Badal Nyalang Stereotyped by Silence: How LLMs Erase Northeast Indian Languages Through Omission and Orthographic Corruption Evaluating AI agent’s capacity to solve real-world CAPTCHA Xiangyu Wu, et al. MirrorCAPTCHA: Wild CAPTCHA, Wild Distribution, Wild Web-based Platform Meet Multimodal LLM Agents What the crawler keeps, and what it loses Michael Paris, Hande Celikkanat, Luca Foppiano Measuring What the Crawler Sees: Discovery Curves, Core Persistence, and Shell Dynamics in Longitudinal Web Crawls Tracking AI adoption across European firms using their websites Julio Garbers, Terry Gregory The Diffusion of Artificial Intelligence Across Firms: Evidence from Europe A new benchmark for web language identification Pedro Ortiz Suarez, et al. CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data Computation and Language Asier Gutiérrez-Fandiño, et al. esCorpius: A Massive Spanish Crawling Corpus More on Google ScholarCurated BibTeX Dataset Latest Blog Post News Web Graph Embeddings: An Experimental Dataset Release We are publishing an experimental dataset of graph embeddings built from the Common Crawl Web Graph: a 128-dimensional vector for each of 52.9 million web hosts, learned from hyperlinks alone, along with two interactive Hugging Face Spaces. Malte Ostendorff Malte is a Senior Research Engineer at Common Crawl. The Data Overview CDXJ Index URL Index Web Graphs Latest Crawl Crawl Stats Graph Stats Errata Resources Get Started AI Agent Blog Examples CCBot Infra Status Opt-Out Ledger FAQ Community Research Papers Mailing List Archive Hugging Face Discord Collaborators About About Team Jobs Privacy Policy Terms of Use © 2026 Common Crawl