HN Hall of Fame Weekly email

Common Crawl

commoncrawl.org Research & data Datasets Web & internet Class of 2021-03 Hall of Fame
Screenshot of commoncrawl.org captured 2026-07-20
Page preview · captured 2026-07-20

Resurfaced independently across 9 calendar years, with breakout response in 2 of them.

submissions
13
submitters
13
observed span
2012–2025
peak thread · 61 comments
397 pts
latest 20+ return · 2025-03-04
27 pts

Submission timeline

2007–2026

One slot for every year since HN launched. Height is that year's peak points; orange marks a 100+ point or 50+ comment breakout. Select a bar to open its strongest thread.

First comments on top threads

HN comment order

I don't think that this is the answer to "only google can crawl the web". This is a huge archive suitable for making a web search engine maybe. What if you want to make a simple link previewer? An abstract crawler for scientific articles? Most websites are behind cloudflare which will block/captcha you, but happily whitelist only google & major social sites. Tha answer is measures that bring the web back to basics, not this over-SEOed bot infested ecosystem. FANGS…

If the implantation is semantically rich and complete enough, this might really help those who want to tackle the first of pg's "ambitious startup" ideas. If I have an idea for a search product, competing with Google isn't really the first roadblock my brain puts up. It's more like "sure brain, sounds swell; now, how do you propose to populate this engine of yours?"

Can a website be both for the open web and at the same time modify native scroll behavior?

I haven't seen this resource before. You can also search the index [1] I downloaded and ran the example code [2] to lookup a URL and fetch its content and the response was instantaneous! [1] http://index.commoncrawl.org/CC-MAIN-2025-08-index?url=ycombinator.com&output=json [2] https://commoncrawl.org/get-started

The first top-level comment from each of the four biggest threads, in HN’s own order. Excerpts are shortened; open a comment for full context.

Breakout years
2

100+ points or 50+ comments

Total points
638

reference only — not used in Hall rules or ranking

Total comments
74

reference only — not used in Hall rules or ranking

Every submission

DateTitle as submittedByPointsComments
2012-03-11Common CrawlFirst breakoutnamin1255
2013-02-01Open access repository of web crawl databpolania10
2015-08-26Common Crawl – An Open Repository of Web Crawl Datasinak10
2018-10-12An open repository of web crawl data that can be accessed and analyzedNicoJuicy10
2020-03-11Common Crawl – open repository of web crawl datar_singh30
2020-12-03Common Crawlgraderjs20
2021-03-26Common CrawlHall induction · Best threadAissen39761
2022-12-01Common Crawlaka87820
2023-01-02Common Crawlstefankuehnel20
2023-02-05Common Crawlturrini20
2023-03-05Common Crawlwildpeaks70
2023-04-17Common Crawlnotmysql_687
2025-03-04Common Crawl maintains a free, open repository of web crawl dataLatest 20+ point returndoener271