Transformer Architecture
The transformer, introduced in the 2017 paper Attention Is All You Need, is the architecture behind essentially every modern large language model. Its key idea is self-attention: for each token, the model computes how relevant every other token in the context is, and mixes their information accordingly. Unlike earlier recurrent networks, which processed text one step at a time, attention looks at the whole sequence at once, which makes training massively parallel on GPUs.
You don't need the full math to use LLMs well, but understanding the architecture explains a lot: why long contexts get expensive, why generation is sequential, what the "KV cache" is, why mixture-of-experts models are cheap to run for their size, and where the limits of context windows come from.
TL;DR
- Text becomes token IDs (tokenization), then embedding vectors.
- A stack of transformer blocks refines those vectors. Each block has multi-head self-attention and a feed-forward network, with residual connections and normalization.
- Attention computes, per token, a weighted mix of other tokens' values, with weights from query·key similarity.
- LLMs are decoder-only transformers with causal masking: each token attends only to earlier tokens, and the model predicts the next token.
- Positional information comes from schemes like RoPE; attention itself is order-agnostic.
- Inference uses a KV cache to avoid recomputation; mixture-of-experts (MoE) activates only part of the network per token.
Quick Example
Scaled dot-product self-attention with a causal mask, the core operation, in PyTorch:
Real models run many such heads in parallel, stack dozens of layers, and use fused kernels (FlashAttention) instead of materializing the full score matrix.
Core Concepts
From Tokens to Vectors
Each token ID indexes an embedding table, producing a vector of size d_model (thousands of dimensions in large models). These vectors are the model's working representation, and each layer updates them. At the end, a final projection (often sharing weights with the embedding table) maps each vector to a score for every vocabulary token: the logits used to sample the next token.
Self-Attention
For each token, the model computes three vectors by multiplying by learned matrices:
- Query (Q): what this token is looking for.
- Key (K): what this token offers to be matched on.
- Value (V): the information it passes along if attended to.
Attention weights are softmax(Q·Kᵀ / √d), and the output is the weighted sum of values. This lets a pronoun pull in information from the noun it refers to, a closing bracket find its opening one, or an answer reference a fact from thousands of tokens earlier.
Attention over n tokens costs O(n²) in compute (and in memory, naively), which is the fundamental reason long contexts are expensive.
Multi-Head Attention
Instead of one attention operation, a layer runs many heads in parallel, each with its own Q/K/V projections, then concatenates their results. Different heads learn different relationships: syntax, coreference, position, copying. Efficiency variants share keys and values across heads, as in multi-query and grouped-query attention (GQA), or compress them, as in multi-head latent attention, to shrink the KV cache.
The Transformer Block
- The feed-forward network (MLP, typically with SwiGLU activations) processes each token independently and holds much of the model's stored knowledge.
- Residual connections add each sublayer's output to its input, which keeps gradients flowing through very deep stacks.
- Normalization (usually RMSNorm, applied pre-sublayer) stabilizes training.
Large models stack dozens to over a hundred of these blocks.
Positional Information
Attention is permutation-invariant; without extra signals, "dog bites man" and "man bites dog" look the same. Transformers inject position via:
- Sinusoidal or learned absolute embeddings (the original transformer, GPT-2).
- Rotary Position Embeddings (RoPE), which rotate Q and K vectors by position-dependent angles so attention depends on relative distance. It's the standard in modern LLMs, and it's extended to longer contexts with scaling techniques (YaRN, NTK scaling).
- ALiBi, a distance-based penalty on attention scores.
Encoder, Decoder, Decoder-Only
Decoder-only models are trained on one objective, predicting the next token, at enormous scale, and that turned out to generalize to almost every language task.
Inference Mechanics
Autoregressive Generation
Generation is a loop: run the model on the context, sample one token, append it, and repeat. The prefill phase processes the whole prompt in parallel (compute-bound); the decode phase produces tokens one at a time (memory-bandwidth-bound). That's why time-to-first-token and tokens-per-second are separate metrics.
The KV Cache
Keys and values for earlier tokens don't change when new tokens are added, so they're cached. Each decode step computes Q, K, and V only for the newest token and attends over the cached keys and values. The KV cache grows linearly with context length and batch size and often dominates GPU memory. That's the motivation for GQA, cache quantization, paged attention (vLLM), and prompt caching. See LLM inference.
Mixture-of-Experts
In an MoE layer, the single feed-forward network is replaced by many "expert" networks plus a router that sends each token to a few of them (for example 2 of 64). The model has a huge total parameter count but a much smaller active parameter count per token, giving more capacity at a lower per-token compute cost. Many frontier and open models (Mixtral, DeepSeek-V3, Qwen MoE variants) use MoE.
Best Practices
Reason About Cost via Context Length
Longer prompts mean quadratic attention compute in prefill and a larger KV cache in decode. Trimming irrelevant context, retrieving selectively (RAG), and caching shared prefixes reduce cost and latency more than most prompt tweaks. See context engineering.
Choose Architectures by Task
Use decoder-only LLMs for generation and reasoning, encoder models or embedding models for classification, search, and similarity at scale, and small language models where latency and cost dominate.
Learn by Building a Small One
Implementing a tiny GPT (a few million parameters) in PyTorch, following Karpathy's nanoGPT, for example, demystifies attention, masking, and training loops faster than any diagram.
Common Mistakes
"The Model Looks Up Answers in a Database"
Transformers store knowledge implicitly in their weights, mostly in the feed-forward layers. Facts are reconstructed by computation, not retrieved from records, which is one root cause of hallucinations.
"Attention Weights Explain the Model's Reasoning"
Attention maps show where information flowed in one layer and head, not why the model produced an answer. Interpretability research uses more careful tools (probing, attribution, sparse autoencoders) to explain behavior.
"Bigger Context Means the Model Uses All of It Equally"
Models can attend to every token in the window, but in practice they may underweight information buried in the middle of very long contexts. Put critical instructions and data where the model uses them best, and evaluate with your own long-context tests.
FAQ
Why did transformers replace RNNs and LSTMs?
Recurrent networks process tokens sequentially, which limits training parallelism and makes long-range dependencies hard to learn. Transformers process all positions in parallel during training and connect any two tokens directly through attention, so they scale efficiently to huge datasets and models.
What does "parameters" mean in a model's size?
The number of learned weights: embedding tables, attention projections, feed-forward matrices, and normalization parameters. A 70B model has about 70 billion of them. For MoE models, distinguish total parameters from active parameters per token.
What is FlashAttention?
An exact attention algorithm that computes attention in tiles held in fast on-chip GPU memory, avoiding materializing the full n×n score matrix in slower memory. It gives the same results with much less memory traffic, making long contexts far faster and cheaper.
Are there alternatives to transformers?
State-space models (Mamba), linear-attention variants, and hybrids that mix attention with recurrent layers aim to scale linearly with sequence length. Several production models use hybrid designs, but full-attention transformers remain the dominant architecture.
Related Topics
- Large Language Models — The models built on transformers
- Tokenization — The input to the transformer
- LLM Sampling — Turning output logits into text
- Context Windows — Limits that attention imposes
- LLM Inference — Serving transformers efficiently
- PyTorch — The framework most transformers are built in