Attention: how a model decides what matters

Every word gets repainted using the words around it. That single idea is what made modern AI work — and what makes long context expensive.
attention
transformers
Author

Ashish Pandey

Published

September 20, 2026

Attention repaints every word using the words around it, taking most of its paint from the words that matter most.

The one job

Underneath every chatbot, assistant and clever demo, there is a single operation:

Given the words so far, guess what comes next.

That’s it. Everything else is that trick, repeated: write a word, add it to the sentence, guess again.

And it doesn’t give one answer. It scores every word it knows — tens of thousands of them. I gave a model “The patient has”:

Next word Confidence
been 20.1%
a 8.3%
not 4.6%
no 3.9%
had 2.2%

Look carefully: there’s no medicine in that list. No fever, no diabetes, no chest pain. The model knows *“The patient has ___“* needs a verb. It has no idea it’s standing in a hospital.

To make a good guess, it needs one more thing: a way to work out which of the earlier words actually matter right now. That’s attention.

Mixing paint

Take this half-finished sentence:

Patient has severe chest ___

Which earlier words matter for the next word?

Earlier word How much it matters
severe 85%
patient 10%
has 2%

Every word is a colour. The word chest starts out plain grey — on its own it could be a chest of drawers, a treasure chest, or a chest X-ray. The word severe is red: urgent, intense. Attention mixes the paint, stirring 85% of severe’s red into chest’s grey. So chest comes out reddish.

Drag the slider and watch the word change colour — and change meaning:

In “chest of drawers”, drawers is brown, so chest comes out brownish instead. Same word in, different colour out, depending on its neighbours.

And here’s the part people miss: when the model makes its guess, it doesn’t look at the word. It looks at the colour the word ended up as. Reddish chest leads to pain. Brownish chest leads to drawers.

Queries, keys and values

Each word produces three different vectors, by multiplying its embedding with three learned weight matrices:

Question it answers In the analogy
Query (Q) “What am I looking for?” chest asking: which word tells me what kind of chest I am?
Key (K) “What do I offer?” severe advertising: I’m an intensity word
Value (V) “What do I pass on?” the actual red paint that gets mixed in

The mechanism is three steps:

  1. Score. Compare every query against every key with a dot product. High score = this word is relevant to me.
  2. Normalise. Divide by \(\sqrt{d_k}\), then softmax. That turns raw scores into the percentages you saw: 85%, 10%, 2%. They always sum to 100%.
  3. Mix. Take a weighted average of the values using those percentages. The result replaces the word’s representation.

\[ \text{Attention}(Q,K,V) = \operatorname{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}\right)V \]

That’s the whole formula, and you already know every piece of it. The paint being mixed is V. The percentages come from Q·K.

TipWhy divide by \(\sqrt{d_k}\)?

Dot products of long vectors grow large. Feed a large number into softmax and it saturates: one word gets 99.9% and the rest get nothing, so gradients vanish and learning stalls. Dividing by \(\sqrt{d_k}\) keeps the scores in a sane range. It’s a numerical fix, not a deep idea — but interviewers love asking about it.

It cannot look forward

While generating text, a word may only attend to words before it. Future positions are set to \(-\infty\) before the softmax, so they get exactly 0% of the paint.

Without that mask, predicting the next word would be trivial: the model would have already seen the answer. This is the causal mask, and it’s the difference between GPT-style decoders and BERT-style encoders (which do see both directions, which is why BERT is good at understanding text but can’t generate it well).

Read a row: it’s one word choosing where to get its paint. The grey cells are the future, blocked out. Notice the last row — the word being predicted takes 85% of its paint from severe, which is exactly the table from earlier.

(Weights here are illustrative, to show the shape of a real attention matrix.)

Many heads, not one

One attention head mixes paint one way. Real models run dozens in parallel — each with its own Q, K and V matrices, each free to learn a different kind of relationship. One head might track which adjective modifies which noun; another, who she refers to; another, which bracket closes which.

Each head works in a smaller space (\(d_{\text{model}}/h\)), their outputs are concatenated, and a final matrix mixes them back together. Same total compute, many more relationships captured.

Attention has no sense of order

Look at the formula again: it’s a weighted sum. Shuffle the words and you’d get the same answer. Attention on its own has no idea about word order — “dog bites man” and “man bites dog” would be identical.

Order is added separately, as positional information baked into the vectors (modern models mostly use RoPE, which rotates Q and K by an angle that depends on position). That’s its own chapter.

Why long context gets expensive

Every word attends to every other word. \(n\) words means \(n^2\) scores.

Tokens Attention scores
1,000 1 million
10,000 100 million
100,000 10 billion

Double the text, quadruple the work. This one line explains most of what long-context pricing looks like, and most of the research into making attention cheaper (FlashAttention makes it memory-efficient without changing the maths; sliding-window and sparse attention change the maths to look at less).

Where it breaks

  • Attention weights are not explanations. A high weight shows where the paint came from, not why the model answered as it did. Treating attention maps as interpretability is a known trap — Jain & Wallace (2019) made this case sharply.
  • The middle gets lost. With long inputs, models reliably use the beginning and end more than the middle. Put what matters at the edges of the prompt.
  • More context isn’t more understanding. Attention dilutes: with 100k tokens competing, each relevant one gets a thinner slice of the paint.
  • The paint analogy has a limit. Mixing colours is lossy and commutative; attention operates in hundreds of dimensions where “mixtures” stay separable. It’s a picture of the weighting, not of the geometry.

Whiteboard check

Marker in hand, someone watching. Could you get through these without notes?

Every word starts with a rough meaning. Attention lets each word look at the others and ask “which of you matters to me?”, then blends in a bit of each one in proportion to the answer. “Chest” next to “severe” ends up meaning something different from “chest” next to “drawers”. The model then predicts from the blended meaning, not the original word.

Dot products grow with dimension; large inputs push softmax into saturation, where one position takes almost all the weight and gradients vanish. Scaling by \(\sqrt{d_k}\) keeps the variance of the scores roughly constant so training stays stable.

The model can see the token it’s meant to predict, so loss drops to near zero and it learns nothing useful. At inference the future doesn’t exist, so the distribution it faces is completely different and generation falls apart.

A single softmax-weighted average forces one relationship per position. Multiple heads in lower-dimensional subspaces let the model attend to several kinds of relationship at once — syntactic, coreference, positional — at the same total cost.

Every query is scored against every key: \(O(n^2 d)\) time and, naively, \(O(n^2)\) memory. FlashAttention keeps the exact result while avoiding materialising the full matrix; sliding window, sparse and linear attention approximate by restricting or reformulating who can attend to whom.

TL;DR

  • Attention = weighted averaging. Queries and keys set the weights, values are what gets mixed.
  • Softmax makes weights sum to 1; \(\sqrt{d_k}\) keeps them numerically sane.
  • The causal mask stops the model seeing the future — that’s what makes generation possible.
  • Many heads capture many kinds of relationship at once.
  • Cost is quadratic in sequence length. That’s the whole long-context story.