Welcome.AIWelcome.AI
    Skip to content
    Common Crawl Foundation logo

    Common Crawl Foundation

    Open Repository of Web Crawl Data

    Technology

    How Common Crawl Foundation’s technology works — architecture, AI models, and technical capabilities.

    What type of AI does Common Crawl Foundation use?

    Common Crawl Foundation does not directly use AI but provides a significant dataset that has been a primary training source for many large language models. Studies indicate that 64% of large language models published between 2019 and 2023 were trained on filtered versions of its data.

    How does Common Crawl Foundation work?

    Common Crawl Foundation operates by crawling the web and maintaining an open repository of web data. It processes approximately three billion pages per cycle, storing over 10 petabytes of data collected since 2008. This data is freely accessible for analysis and research.

    What are Common Crawl Foundation's main features?

    Key features of Common Crawl include a vast dataset of over 300 billion pages, monthly crawls adding 3-5 billion new pages, and the availability of data in various formats like WARC and CDXJ. It also offers tools for querying and analyzing the data.

    What technology powers Common Crawl Foundation?

    Common Crawl Foundation utilizes CCBot, a web crawler based on Apache Nutch, to collect data. The data is stored on Amazon Web Services under the Open Data Sponsorship Program, ensuring efficient access and management of the vast dataset.

    Where AI Leaders Stay Informed

    The latest AI intelligence, case studies, and research — delivered to your inbox every week.

    Free to read. Unsubscribe anytime.

    Is this your company?

    Claim this profile to manage information and unlock premium features.