HN Hall of Fame Weekly email

The Illustrated Transformer

jalammar.github.io Books & learning Tutorials & guides AI & data Candidate
Screenshot of jalammar.github.io captured 2026-07-20
Page preview · captured 2026-07-20

Resurfaced independently across 5 calendar years, with breakout response in 2 of them.

submissions
10
submitters
10
observed span
2018–2025
peak thread · 87 comments
475 pts
latest 20+ return · 2025-12-22
475 pts

Submission timeline

2007–2026

One slot for every year since HN launched. Height is that year's peak points; orange marks a 100+ point or 50+ comment breakout. Select a bar to open its strongest thread.

First comments on top threads

HN comment order

I read this article back when I was learning the basics of transformers; the visualizations were really helpful. Although in retrospect knowing how a transformer works wasn't very useful at all in my day job applying LLMs, except as a sort of deep background for reassurance that I had some idea of how the big black box producing the tokens was put together, and to give me the mathematical basis for things like context size limitations etc. I would strongly…

Illustrated Transformer is amazing as a way of understanding the original transformer architecture step-by-step, but if you want to truly visualize how information flows through a decoder-only architecture - from nanoGPT all the way up to a fully represented GPT-3 - nothing beats this: https://bbycroft.net/llm

that's a great arxiv translation, a model for ML elucidation re: Transformers, see also: https://towardsdatascience.com/the-fall-of-rnn-lstm-2d1594c7... which suggests that Transformers have been supplanted by simple conv2d networks that span both the input and the out; also mentioned are "hierarchical neural attention encoders", but no links; q.v. https://www.cs.cmu.edu/~hovy/papers/16HLT-hierarchical-atten...

angel_j·65-point thread·

The self-attention mechanism is explained very well in this blog post. Because of this it is very much worth a read for anybody interested in the state of the art of deep learning models for machine translation. Other parts of the Transformer model are glossed over more, though.

rerx·17-point thread·

The first top-level comment from each of the four biggest threads, in HN’s own order. Excerpts are shortened; open a comment for full context.

Breakout years
2

100+ points or 50+ comments

Total points
734

reference only — not used in Hall rules or ranking

Total comments
105

reference only — not used in Hall rules or ranking

Every submission