HN Hall of Fame Weekly email

How to crawl a quarter billion webpages in 40 hours

www.michaelnielsen.org Essays & writing Essays & articles Web & internet Class of 2026-04 Hall of Fame
Screenshot of www.michaelnielsen.org captured 2026-07-20
Page preview · captured 2026-07-20

Resurfaced independently across 5 calendar years, with breakout response in 4 of them.

submissions
5
submitters
5
observed span
2012–2026
peak thread · 67 comments
323 pts
latest 20+ return · 2023-06-15
208 pts

Submission timeline

2007–2026

One slot for every year since HN launched. Height is that year's peak points; orange marks a 100+ point or 50+ comment breakout. Select a bar to open its strongest thread.

First comments on top threads

HN comment order

In my experience, one of the hardest parts of writing a web crawler is URL selection: After crawling your list of seed URLs, where do you go next? How do you make sure you don't crawl the same content multiple times because it has a slightly different URL? How to avoid getting stuck on unimportant spam sites with autogenerated content? Because the author only crawled domains from a limited set and only for a short time, he did not need…

soult·323-point thread·

Originally I intended to make the crawler code available under an open source license at GitHub. However, as I better understood the cost that crawlers impose on websites, I began to have reservations. My crawler is designed to be polite and impose relatively little burden on any single website, but could (like many crawlers) easily be modified by thoughtless or malicious people to impose a heavy burden on sites. Because of this I’ve decided to postpone (possibly indefinitely) releasing the…

This is a good overview, but being from 2012 it's missing comment on one now common area, using a real browser for crawling/scraping. You will often now hear recommendations to use a real browser, in headless mode, for crawling. There are two reasons for this: - SPAs with front end only rendering are hard to scrape with a traditional http library. - Anti bot/scraping technology fingerprints browsers, looks at request patterns, and browser behaviour to try and detect and block…

This was written in 2012, Its even easier these days by using SQS and Cloud Formation. 250 Million is a small number you are better of first going through Common Crawl and then use data from crawls to build a better seed list. Common Crawl now contains repeated crawls conducted every few months and also urls donated by blekko. https://groups.google.com/forum/m/#!msg/common-crawl/zexccXg...

The first top-level comment from each of the four biggest threads, in HN’s own order. Excerpts are shortened; open a comment for full context.

Breakout years
4

100+ points or 50+ comments

Total points
967

reference only — not used in Hall rules or ranking

Total comments
215

reference only — not used in Hall rules or ranking

Every submission

DateTitle as submittedByPointsComments
2012-08-10How to crawl a quarter billion webpages in 40 hoursFirst breakout · Best threadcing32367
2016-01-08How to crawl a quarter billion webpages in 40 hours (2012)_ao78913623
2018-07-05How to crawl a quarter billion webpages in 40 hours (2012)allenleein29661
2023-06-15How to crawl a quarter billion webpages in 40 hours (2012)Latest 20+ point returnswyx20863
2026-04-15How to crawl a quarter billion webpages in 40 hours (2012)Hall inductiondownbad_41