Submission timeline
2007–2026One slot for every year since HN launched. Height is that year's peak points; orange marks a 100+ point or 50+ comment breakout. Select a bar to open its strongest thread.
First comments on top threads
HN comment orderIn my experience, one of the hardest parts of writing a web crawler is URL selection: After crawling your list of seed URLs, where do you go next? How do you make sure you don't crawl the same content multiple times because it has a slightly different URL? How to avoid getting stuck on unimportant spam sites with autogenerated content? Because the author only crawled domains from a limited set and only for a short time, he did not need…
Originally I intended to make the crawler code available under an open source license at GitHub. However, as I better understood the cost that crawlers impose on websites, I began to have reservations. My crawler is designed to be polite and impose relatively little burden on any single website, but could (like many crawlers) easily be modified by thoughtless or malicious people to impose a heavy burden on sites. Because of this I’ve decided to postpone (possibly indefinitely) releasing the…
This is a good overview, but being from 2012 it's missing comment on one now common area, using a real browser for crawling/scraping. You will often now hear recommendations to use a real browser, in headless mode, for crawling. There are two reasons for this: - SPAs with front end only rendering are hard to scrape with a traditional http library. - Anti bot/scraping technology fingerprints browsers, looks at request patterns, and browser behaviour to try and detect and block…
This was written in 2012, Its even easier these days by using SQS and Cloud Formation. 250 Million is a small number you are better of first going through Common Crawl and then use data from crawls to build a better seed list. Common Crawl now contains repeated crawls conducted every few months and also urls donated by blekko. https://groups.google.com/forum/m/#!msg/common-crawl/zexccXg...
The first top-level comment from each of the four biggest threads, in HN’s own order. Excerpts are shortened; open a comment for full context.
- Breakout years
- 4
- Total points
- 967
- Total comments
- 215
100+ points or 50+ comments
reference only — not used in Hall rules or ranking
reference only — not used in Hall rules or ranking
Every submission
| Date | Title as submitted | By | Points | Comments |
|---|---|---|---|---|
| 2012-08-10 | How to crawl a quarter billion webpages in 40 hoursFirst breakout · Best thread | cing | 323 | 67 |
| 2016-01-08 | How to crawl a quarter billion webpages in 40 hours (2012) | _ao789 | 136 | 23 |
| 2018-07-05 | How to crawl a quarter billion webpages in 40 hours (2012) | allenleein | 296 | 61 |
| 2023-06-15 | How to crawl a quarter billion webpages in 40 hours (2012)Latest 20+ point return | swyx | 208 | 63 |
| 2026-04-15 | How to crawl a quarter billion webpages in 40 hours (2012)Hall induction | downbad_ | 4 | 1 |
