LARGE LANGUAGE MODELS / INTERMEDIATE

Inside The Transformer: Attention Explained For Non-Specialists

The transformer underpins well-known language models such as OpenAI's GPT series. Here is what attention does, why it was a breakthrough and how the pieces of a transformer fit together.

Checked against primary sources and independently reviewed on . Sources are listed at the end.

The transformer is a type of neural network introduced in 2017, and it underpins well-known large language models such as OpenAI’s GPT series. OpenAI’s 2023 technical report, for example, describes GPT-4 as a transformer-based model trained to predict the next token.1 The name appears throughout product documentation and research papers, yet the idea at its centre, attention, is rarely explained in plain terms.

This article explains what attention does with a worked example, why it replaced earlier designs, and what the main parts of a transformer are for. It builds on the ideas of tokens and embeddings from How An LLM Reads And Writes.

The Problem Attention Solves

The meaning of a word depends on the words around it. Take the sentence “The auditor rejected the invoice because it had no purchase order.” A person knows instantly that “it” means the invoice, not the auditor. Change the ending to “because she found no purchase order” and the focus shifts to the auditor.

Before transformers, the leading language models were mostly recurrent networks, which read text one word at a time and carried a running summary forward. Information from early words had to survive many steps to influence later ones, and because each step waited for the one before it, training was hard to spread across many processors at once.2

In 2017, Vaswani and colleagues, most of them at Google, proposed an architecture “based solely on attention mechanisms”, dropping recurrence entirely.2 They demonstrated it on machine translation, where it beat the previous best results on standard English to German and English to French benchmarks while taking less time to train.

What Attention Does

Attention lets a token look at other tokens in the input and decide how much each one matters to its own meaning. In the original design’s encoder, every token could look at every other token. In the text-generating models described later, a token can look only at itself and the tokens before it. For the word “it” in the example, the model works out a set of weights over the earlier words and then builds a new representation of “it” that mixes in information from the words with the highest weights.

Theauditorrejectedtheinvoicebecauseithad no purchase orderstrongweak
An illustrative view of attention from the token 'it', which can look only at itself and earlier words. Thicker, darker lines mean higher weight. In a trained model, one attention head might link 'it' most strongly to 'invoice'.

The mechanism works with three learned transformations of each token’s vector, usually called the query, key and value. A useful analogy is a library search. The query is what a token is looking for, each key describes what another token offers, and comparing the query against every key produces the attention weights. The values are the information actually passed along, blended according to those weights.

Because each token is compared with all the tokens it is allowed to see in a single step, the model can connect words that are far apart as easily as neighbours. During training, the calculation for a whole passage can run in parallel on modern chips. That parallelism is a large part of why transformers could be trained on far more data than earlier designs.

Several Heads At Once

A single attention calculation captures one kind of relationship. The transformer runs several in parallel, called attention heads, and combines their outputs before moving on. Vaswani and colleagues called this multi-head attention.2 The largest GPT-3 model, for comparison, used 96 heads in each layer.5

In trained models, different heads often settle on different jobs. A 2019 study of Google’s BERT, a transformer built for understanding text rather than generating it, found heads that mostly looked at the next or previous word, heads that linked verbs to their direct objects, and heads that connected later mentions of something to earlier ones.6

Keeping Track Of Word Order

Attention on its own does not know the order of the tokens: “the auditor rejected the invoice” and “the invoice rejected the auditor” would look alike. The original paper fixed this by adding fixed mathematical patterns, called positional encodings, to each token’s embedding before the first layer.2 Later work proposed other approaches, such as rotary position embeddings in 2021, which build position into the attention calculation itself.7

How A Transformer Is Put Together

  1. Output: Next-Token ProbabilitiesA final layer turns the representation at the last position into a score for every possible next token.
  2. Feed-Forward NetworkProcesses each token on its own, adding learned knowledge and transformations.
  3. Masked Multi-Head Self-AttentionEach token gathers information from itself and the tokens before it.
  4. Positional InformationTells the model where each token sits. Added here in the original design; some later designs apply it inside attention instead.
  5. Token EmbeddingsEach input token is turned into a vector.
A simplified decoder-style transformer of the kind used in the first GPT model. The middle block is repeated many times, with each copy refining the representation of every token.

The highlighted attention and feed-forward pair forms one transformer block, and large models stack many of these blocks; the largest GPT-3 model had 96 of them.5 The original 2017 design had two halves: an encoder that read the source sentence and a decoder that wrote the translation.2 The first GPT model, published by OpenAI in 2018, used only a decoder-style stack.8 In that kind of stack, each token may attend only to itself and earlier tokens, a restriction called masking. It is what makes next-token prediction work, because the model cannot peek at the answer it is trying to predict.

Why This Matters In Practice

Two practical consequences follow from the design. First, because each token is compared with every token it can see, the work done by self-attention grows with the square of the input length: double the input and that part of the work roughly quadruples.2 This is one reason models have a maximum context length and why long prompts cost more to process.

Second, the model has no separate store of facts. What it learned is spread across the parameters of its attention and feed-forward layers, and when Lewis and colleagues proposed retrieval-augmented generation in 2020, they treated updating what a model had absorbed into its parameters as an unsolved research problem.9 That is one reason techniques such as retrieval-augmented generation, covered in Using LLMs Well, supply facts in the prompt instead.

Footnotes

  1. OpenAI, “GPT-4 Technical Report”, arXiv:2303.08774, March 2023. arxiv.org ↩

  2. A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser and I. Polosukhin, “Attention Is All You Need”, arXiv:1706.03762, June 2017. arxiv.org ↩ ↩2 ↩3 ↩4 ↩5 ↩6

  3. S. Jain and B. C. Wallace, “Attention is not Explanation”, NAACL 2019, arXiv:1902.10186. arxiv.org ↩

  4. S. Wiegreffe and Y. Pinter, “Attention is not not Explanation”, EMNLP 2019, arXiv:1908.04626. arxiv.org ↩

  5. T. B. Brown et al., “Language Models are Few-Shot Learners”, arXiv:2005.14165, May 2020, Table 2.1. arxiv.org ↩ ↩2

  6. K. Clark, U. Khandelwal, O. Levy and C. D. Manning, “What Does BERT Look At? An Analysis of BERT’s Attention”, arXiv:1906.04341, June 2019. arxiv.org ↩

  7. J. Su, Y. Lu, S. Pan, A. Murtadha, B. Wen and Y. Liu, “RoFormer: Enhanced Transformer with Rotary Position Embedding”, arXiv:2104.09864, April 2021. arxiv.org ↩

  8. A. Radford, K. Narasimhan, T. Salimans and I. Sutskever (OpenAI), “Improving Language Understanding by Generative Pre-Training”, 2018. cdn.openai.com ↩

  9. P. Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”, NeurIPS 2020, arXiv:2005.11401. arxiv.org ↩

Knowledge Hub content is general information. It is not legal advice, a compliance certification, a guarantee of security or a substitute for an assessment of your own systems. Standards and rules change; check the sources for the latest position.