Attention repaints every word using the words around it, taking most of its paint from the words that matter most.
The one job
Underneath every chatbot, assistant and clever demo, there is a single operation:
Given the words so far, guess what comes next.
That’s it. Everything else is that trick, repeated: write a word, add it to the sentence, guess again.
And it doesn’t give one answer. It scores every word it knows — tens of thousands of them. I gave a model “The patient has”:
| Next word | Confidence |
|---|---|
| been | 20.1% |
| a | 8.3% |
| not | 4.6% |
| no | 3.9% |
| had | 2.2% |
Look carefully: there’s no medicine in that list. No fever, no diabetes, no chest pain. The model knows *“The patient has ___“* needs a verb. It has no idea it’s standing in a hospital.
To make a good guess, it needs one more thing: a way to work out which of the earlier words actually matter right now. That’s attention.
Mixing paint
Take this half-finished sentence:
Patient has severe chest ___
Which earlier words matter for the next word?
| Earlier word | How much it matters |
|---|---|
| severe | 85% |
| patient | 10% |
| has | 2% |
Every word is a colour. The word chest starts out plain grey — on its own it could be a chest of drawers, a treasure chest, or a chest X-ray. The word severe is red: urgent, intense. Attention mixes the paint, stirring 85% of severe’s red into chest’s grey. So chest comes out reddish.
Drag the slider and watch the word change colour — and change meaning:
In “chest of drawers”, drawers is brown, so chest comes out brownish instead. Same word in, different colour out, depending on its neighbours.
And here’s the part people miss: when the model makes its guess, it doesn’t look at the word. It looks at the colour the word ended up as. Reddish chest leads to pain. Brownish chest leads to drawers.
Queries, keys and values
Each word produces three different vectors, by multiplying its embedding with three learned weight matrices:
| Question it answers | In the analogy | |
|---|---|---|
| Query (Q) | “What am I looking for?” | chest asking: which word tells me what kind of chest I am? |
| Key (K) | “What do I offer?” | severe advertising: I’m an intensity word |
| Value (V) | “What do I pass on?” | the actual red paint that gets mixed in |
The mechanism is three steps:
- Score. Compare every query against every key with a dot product. High score = this word is relevant to me.
- Normalise. Divide by \(\sqrt{d_k}\), then softmax. That turns raw scores into the percentages you saw: 85%, 10%, 2%. They always sum to 100%.
- Mix. Take a weighted average of the values using those percentages. The result replaces the word’s representation.
\[ \text{Attention}(Q,K,V) = \operatorname{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}\right)V \]
That’s the whole formula, and you already know every piece of it. The paint being mixed is V. The percentages come from Q·K.
TipWhy divide by \(\sqrt{d_k}\)?
Dot products of long vectors grow large. Feed a large number into softmax and it saturates: one word gets 99.9% and the rest get nothing, so gradients vanish and learning stalls. Dividing by \(\sqrt{d_k}\) keeps the scores in a sane range. It’s a numerical fix, not a deep idea — but interviewers love asking about it.
It cannot look forward
While generating text, a word may only attend to words before it. Future positions are set to \(-\infty\) before the softmax, so they get exactly 0% of the paint.
Without that mask, predicting the next word would be trivial: the model would have already seen the answer. This is the causal mask, and it’s the difference between GPT-style decoders and BERT-style encoders (which do see both directions, which is why BERT is good at understanding text but can’t generate it well).
Read a row: it’s one word choosing where to get its paint. The grey cells are the future, blocked out. Notice the last row — the word being predicted takes 85% of its paint from severe, which is exactly the table from earlier.
(Weights here are illustrative, to show the shape of a real attention matrix.)
Many heads, not one
One attention head mixes paint one way. Real models run dozens in parallel — each with its own Q, K and V matrices, each free to learn a different kind of relationship. One head might track which adjective modifies which noun; another, who she refers to; another, which bracket closes which.
Each head works in a smaller space (\(d_{\text{model}}/h\)), their outputs are concatenated, and a final matrix mixes them back together. Same total compute, many more relationships captured.
Attention has no sense of order
Look at the formula again: it’s a weighted sum. Shuffle the words and you’d get the same answer. Attention on its own has no idea about word order — “dog bites man” and “man bites dog” would be identical.
Order is added separately, as positional information baked into the vectors (modern models mostly use RoPE, which rotates Q and K by an angle that depends on position). That’s its own chapter.
Why long context gets expensive
Every word attends to every other word. \(n\) words means \(n^2\) scores.
| Tokens | Attention scores |
|---|---|
| 1,000 | 1 million |
| 10,000 | 100 million |
| 100,000 | 10 billion |
Double the text, quadruple the work. This one line explains most of what long-context pricing looks like, and most of the research into making attention cheaper (FlashAttention makes it memory-efficient without changing the maths; sliding-window and sparse attention change the maths to look at less).
Where it breaks
- Attention weights are not explanations. A high weight shows where the paint came from, not why the model answered as it did. Treating attention maps as interpretability is a known trap — Jain & Wallace (2019) made this case sharply.
- The middle gets lost. With long inputs, models reliably use the beginning and end more than the middle. Put what matters at the edges of the prompt.
- More context isn’t more understanding. Attention dilutes: with 100k tokens competing, each relevant one gets a thinner slice of the paint.
- The paint analogy has a limit. Mixing colours is lossy and commutative; attention operates in hundreds of dimensions where “mixtures” stay separable. It’s a picture of the weighting, not of the geometry.
Whiteboard check
Marker in hand, someone watching. Could you get through these without notes?
NoteExplain self-attention to someone non-technical, in under a minute.
Every word starts with a rough meaning. Attention lets each word look at the others and ask “which of you matters to me?”, then blends in a bit of each one in proportion to the answer. “Chest” next to “severe” ends up meaning something different from “chest” next to “drawers”. The model then predicts from the blended meaning, not the original word.
NoteWhy scale by the square root of d_k?
Dot products grow with dimension; large inputs push softmax into saturation, where one position takes almost all the weight and gradients vanish. Scaling by \(\sqrt{d_k}\) keeps the variance of the scores roughly constant so training stays stable.
NoteWhat breaks if you remove the causal mask during training of a decoder model?
The model can see the token it’s meant to predict, so loss drops to near zero and it learns nothing useful. At inference the future doesn’t exist, so the distribution it faces is completely different and generation falls apart.
NoteWhy multiple heads instead of one big head?
A single softmax-weighted average forces one relationship per position. Multiple heads in lower-dimensional subspaces let the model attend to several kinds of relationship at once — syntactic, coreference, positional — at the same total cost.
NoteWhere does the quadratic cost come from, and how do people avoid it?
Every query is scored against every key: \(O(n^2 d)\) time and, naively, \(O(n^2)\) memory. FlashAttention keeps the exact result while avoiding materialising the full matrix; sliding window, sparse and linear attention approximate by restricting or reformulating who can attend to whom.
TL;DR
- Attention = weighted averaging. Queries and keys set the weights, values are what gets mixed.
- Softmax makes weights sum to 1; \(\sqrt{d_k}\) keeps them numerically sane.
- The causal mask stops the model seeing the future — that’s what makes generation possible.
- Many heads capture many kinds of relationship at once.
- Cost is quadratic in sequence length. That’s the whole long-context story.