Submission timeline
2007–2026One slot for every year since HN launched. Height is that year's peak points; orange marks a 100+ point or 50+ comment breakout. Select a bar to open its strongest thread.
First comments on top threads
HN comment orderI'm confused by findings like this one, and I'm hoping someone here can educated me. There are many known universal approximations. Deep networks are one. SVMs are one. Heck, cubic splines are one, and they've been in use for nearly a hundred years IIRC. The problem has never been one of finding a sufficiently powerful approximator. It has been training that approximator. My understanding of the significant advancement made by deep learning is that we finally figured out how to…
This connects to a suspicion I have that a lot of the most impressive results, that have caught public attention, for LLMs and Sora are due to memorization. This is not to say that they can't generalize or mix learned patterns - I think they can do that too, which is quite impressive to a researcher, but these are demonstrated on tasks that the public might not care as much about, and the performance there isn't always great as well…
interview with Pedro: https://news.cs.washington.edu/2020/12/02/uncovering-secrets...
The first top-level comment from each of the four biggest threads, in HN’s own order. Excerpts are shortened; open a comment for full context.
- Breakout years
- 2
- Total points
- 594
- Total comments
- 244
100+ points or 50+ comments
reference only — not used in Hall rules or ranking
reference only — not used in Hall rules or ranking
Every submission
| Date | Title as submitted | By | Points | Comments |
|---|---|---|---|---|
| 2020-12-03 | Every Model Learned by Gradient Descent Is Approximately a Kernel Machine | breck | 4 | 0 |
| 2020-12-04 | Deep learning networks are approximately kernel machines | deathflute | 2 | 1 |
| 2020-12-05 | Every Model Learned by Gradient Descent Is Approximately a Kernel MachineFirst breakout · Best thread | scottlocklin | 406 | 107 |
| 2023-01-10 | Every Model Learned By Gradient Descent Is Approximately A Kernel Machine (2020) | optimalsolver | 2 | 0 |
| 2024-02-25 | Every model learned by gradient descent is approximately a kernel machine (2020)Hall induction · Latest 20+ point return | Anon84 | 176 | 136 |
| 2025-07-27 | Every Model Learned by Gradient Descent Is Approximately a Kernel Machine | LordNibbler | 4 | 0 |
