Submission timeline
2007–2026One slot for every year since HN launched. Height is that year's peak points; orange marks a 100+ point or 50+ comment breakout. Select a bar to open its strongest thread.
First comments on top threads
HN comment orderI read this article back when I was learning the basics of transformers; the visualizations were really helpful. Although in retrospect knowing how a transformer works wasn't very useful at all in my day job applying LLMs, except as a sort of deep background for reassurance that I had some idea of how the big black box producing the tokens was put together, and to give me the mathematical basis for things like context size limitations etc. I would strongly…
Illustrated Transformer is amazing as a way of understanding the original transformer architecture step-by-step, but if you want to truly visualize how information flows through a decoder-only architecture - from nanoGPT all the way up to a fully represented GPT-3 - nothing beats this: https://bbycroft.net/llm
that's a great arxiv translation, a model for ML elucidation re: Transformers, see also: https://towardsdatascience.com/the-fall-of-rnn-lstm-2d1594c7... which suggests that Transformers have been supplanted by simple conv2d networks that span both the input and the out; also mentioned are "hierarchical neural attention encoders", but no links; q.v. https://www.cs.cmu.edu/~hovy/papers/16HLT-hierarchical-atten...
The self-attention mechanism is explained very well in this blog post. Because of this it is very much worth a read for anybody interested in the state of the art of deep learning models for machine translation. Other parts of the Transformer model are glossed over more, though.
The first top-level comment from each of the four biggest threads, in HN’s own order. Excerpts are shortened; open a comment for full context.
- Breakout years
- 2
- Total points
- 734
- Total comments
- 105
100+ points or 50+ comments
reference only — not used in Hall rules or ranking
reference only — not used in Hall rules or ranking
Every submission
| Date | Title as submitted | By | Points | Comments |
|---|---|---|---|---|
| 2018-06-27 | The Illustrated Transformer – A Visual Explanation for Attention Is All You Need | jalammar | 3 | 1 |
| 2018-10-17 | The Illustrated Transformer | rerx | 17 | 2 |
| 2018-11-01 | The Illustrated Transformer | ghosthamlet | 65 | 4 |
| 2020-08-05 | The Illustrated Transformer – Jay Alammar – Visualizing ML One Concept at a Time | ai_ja_nai | 1 | 0 |
| 2020-10-29 | The Illustrated Transformer Visualizing machine learning one concept at a time | mrfusion | 2 | 0 |
| 2023-04-02 | The Illustrated Transformer | marcodiego | 4 | 0 |
| 2023-07-29 | The Illustrated Transformer | averylamp | 2 | 0 |
| 2024-02-17 | Easy Guide to Understanding Transformers | seshagiric | 3 | 0 |
| 2024-07-02 | The Illustrated Transformer (2018)First breakout | debdut | 162 | 11 |
| 2025-12-22 | The Illustrated TransformerBest thread · Latest 20+ point return | auraham | 475 | 87 |
