Submission timeline
2007–2026One slot for every year since HN launched. Height is that year's peak points; orange marks a 100+ point or 50+ comment breakout. Select a bar to open its strongest thread.
First comments on top threads
HN comment orderI don't think that this is the answer to "only google can crawl the web". This is a huge archive suitable for making a web search engine maybe. What if you want to make a simple link previewer? An abstract crawler for scientific articles? Most websites are behind cloudflare which will block/captcha you, but happily whitelist only google & major social sites. Tha answer is measures that bring the web back to basics, not this over-SEOed bot infested ecosystem. FANGS…
If the implantation is semantically rich and complete enough, this might really help those who want to tackle the first of pg's "ambitious startup" ideas. If I have an idea for a search product, competing with Google isn't really the first roadblock my brain puts up. It's more like "sure brain, sounds swell; now, how do you propose to populate this engine of yours?"
Can a website be both for the open web and at the same time modify native scroll behavior?
I haven't seen this resource before. You can also search the index [1] I downloaded and ran the example code [2] to lookup a URL and fetch its content and the response was instantaneous! [1] http://index.commoncrawl.org/CC-MAIN-2025-08-index?url=ycombinator.com&output=json [2] https://commoncrawl.org/get-started
The first top-level comment from each of the four biggest threads, in HN’s own order. Excerpts are shortened; open a comment for full context.
- Breakout years
- 2
- Total points
- 638
- Total comments
- 74
100+ points or 50+ comments
reference only — not used in Hall rules or ranking
reference only — not used in Hall rules or ranking
Every submission
| Date | Title as submitted | By | Points | Comments |
|---|---|---|---|---|
| 2012-03-11 | Common CrawlFirst breakout | namin | 125 | 5 |
| 2013-02-01 | Open access repository of web crawl data | bpolania | 1 | 0 |
| 2015-08-26 | Common Crawl – An Open Repository of Web Crawl Data | sinak | 1 | 0 |
| 2018-10-12 | An open repository of web crawl data that can be accessed and analyzed | NicoJuicy | 1 | 0 |
| 2020-03-11 | Common Crawl – open repository of web crawl data | r_singh | 3 | 0 |
| 2020-12-03 | Common Crawl | graderjs | 2 | 0 |
| 2021-03-26 | Common CrawlHall induction · Best thread | Aissen | 397 | 61 |
| 2022-12-01 | Common Crawl | aka878 | 2 | 0 |
| 2023-01-02 | Common Crawl | stefankuehnel | 2 | 0 |
| 2023-02-05 | Common Crawl | turrini | 2 | 0 |
| 2023-03-05 | Common Crawl | wildpeaks | 7 | 0 |
| 2023-04-17 | Common Crawl | notmysql_ | 68 | 7 |
| 2025-03-04 | Common Crawl maintains a free, open repository of web crawl dataLatest 20+ point return | doener | 27 | 1 |
