What Is a Transformer in an LLM and How Does It Work?

A Transformer is the neural network architecture behind most modern large language models. It reads a sequence of tokens, lets every token exchange information with every other token through self-attention, and repeats that process across many layers to build a context-aware representation that predicts the next token. You can see this happen step by step in interactive visualizations such as AnimatedLLM, which is designed to show how large language models work under the hood.

The core idea: tokens talking to each other

Older sequence models processed text one step at a time, so information from early words had to survive a long chain of steps. A Transformer instead looks at the whole sequence at once. Each token produces a query, a key, and a value. Attention compares a token's query against every token's key to decide how much to pull from each value.

The result is that the meaning of a word is not fixed at input time. It is assembled from the surrounding context, layer by layer.

Self-attention in plain terms

For each token position:

  1. Query — what this token is looking for.
  2. Key — what each token offers as a match.
  3. Value — the information actually passed along when a match is strong.

Scores are scaled, passed through a softmax so they sum to one, and used to weight the values. A token that attends strongly to "bank" and "river" ends up with a representation shaped by both.

Multi-head attention, feed-forward layers, and the supporting parts

A single attention pattern is limiting, so the model runs several attention heads in parallel. Each head can specialize: one may track syntax, another may link pronouns to their referents. Their outputs are concatenated and projected back to the model dimension.

After attention, each position passes through a position-wise feed-forward network. This is where most parameters live, and it transforms the mixed context into a richer feature representation.

Two structural pieces keep deep stacks trainable:

Component Role
Residual connection Adds the layer input back to its output, giving gradients a short path
Layer normalization Stabilizes the scale of activations across the stack

Without these, stacking dozens of layers tends to be unstable.

The full data flow

  1. Tokenization — text is split into tokens and mapped to integer IDs.
  2. Embedding — each ID becomes a vector.
  3. Positional encoding — since attention has no built-in order, position information is added so the model knows token order.
  4. Transformer blocks — attention plus feed-forward, wrapped in residual and normalization steps, repeated many times.
  5. Output projection — the final hidden state is turned into scores over the vocabulary.
  6. Prediction — the next token is selected, then fed back in for the next step.

Encoder, decoder, and why decoder-only won

  • Encoder-only models (BERT-style) read the full sequence bidirectionally and are strong for classification and embedding tasks.
  • Encoder-decoder models (original translation-style Transformers) encode a source and generate a target.
  • Decoder-only models use causal masking so each position can only attend to earlier positions. This matches next-token prediction directly, scales well, and is the basis of most current LLMs.

Common misconceptions

  • Attention weights are not a full explanation. They show one signal among many; feed-forward layers and residual streams also carry information.
  • More parameters does not mean more understanding. Capability depends on data, training, and architecture choices together.
  • Positional encoding is not optional. Remove it and the model loses word order.

Seeing it for yourself

Interactive demos let you watch attention matrices and hidden states change as you move through layers. For example, you can trace how a pronoun's representation shifts toward its antecedent in deeper layers, which makes the abstract mechanics concrete. AnimatedLLM is one such resource for exploring these internals visually.

animatedllm.github.io
Understand how large language models work under the hood.
qdrant.tech
Qdrant is an Open-Source Vector Search Engine written in Rust. It provides fast and scalable vector similarity search service with convenient API.