HN Hall of Fame Weekly email

On Chomsky and the Two Cultures of Statistical Learning

norvig.com Essays & writing Essays & articles AI & data Class of 2016-06 Hall of Fame
Screenshot of norvig.com captured 2026-07-20
Page preview · captured 2026-07-20

The originally submitted URL now redirects to the address above. Updating the destination does not change the item’s HN history, Hall membership, or rank.

Resurfaced independently across 9 calendar years, with breakout response in 4 of them.

submissions
13
submitters
13
observed span
2011–2025
peak thread · 135 comments
313 pts
latest 20+ return · 2025-12-16
101 pts

Submission timeline

2007–2026

One slot for every year since HN launched. Height is that year's peak points; orange marks a 100+ point or 50+ comment breakout. Select a bar to open its strongest thread.

First comments on top threads

HN comment order

Chomsky's one paragraph quote at the beginning of this article is more clear and thoughtful than the rest of this. I feel the author's missing the point. In the case of language, observing and reporting statistical probabilities in written/spoken language output does very little to explain the cognitive systems used in acquiring and using language. Even one statistical anomaly serves to show that statistical learning is NOT the entire picture when it comes to language development. There was another article…

brockf·313-point thread·

I think that Norvig hits the nail on the head near the beginning of his piece: "I believe that Chomsky has no objection to this kind of statistical model [the Newtonian model of gravitational attraction]. Rather, he seems to reserve his criticism for statistical models like Shannon's that have quadrillions of parameters, not just one or two." This is no more than an objection to problems of fitting your chosen model to data. If you only have a small number…

mrow84·152-point thread·

> Chomsky has focused on the generative side of language The answers to "why" that Chomsky pushes so hard for are very valuable to adult language learners. There are basic syntactic rules to generating broadly correct language. Having these rules discovered and explained in the simplest possible form is irreplaceable by statistical models. Neural networks, much like native speakers can say "well this just sounds right," but adult learners need a mathematical theory of how and why they can generate…

I used to think the same way. But after spending the last few years getting frustrated trying to create more complex linguistic systems, e.g., dialog systems or scientific articles understanding, I am coming to the conclusion that the current statistical approach is a dead end. It’s actually impeding the field because it’s working so well for certain tasks that when people are trying to build systems with real understanding, they can’t match the performance obtained by gigantic language models. But…

pesenti·85-point thread·

The first top-level comment from each of the four biggest threads, in HN’s own order. Excerpts are shortened; open a comment for full context.

Breakout years
4

100+ points or 50+ comments

Total points
758

reference only — not used in Hall rules or ranking

Total comments
447

reference only — not used in Hall rules or ranking

Every submission