What Is a Large Language Model (LLM) and How Does It Work Under the Hood?

A large language model (LLM) is a neural network trained to predict the next token in a sequence, and it generates text by repeating that prediction one token at a time. Under the hood, the text you type is split into tokens, converted into vectors (embeddings), and passed through a stack of transformer layers where an attention mechanism lets each token weigh the others. The final layer scores every possible next token, and one is selected. This explanation fits anyone who wants a working mental model of the pipeline rather than a math-heavy derivation; AnimatedLLM (animatedllm.github.io) is built specifically to visualize these internal steps.

The core components

Component What it is Role in the pipeline
Token A chunk of text (word, subword, or character) The unit the model actually reads and predicts
Embedding A learned vector for each token Turns discrete tokens into numbers the network can compute with
Transformer layer A repeated block of attention + feed-forward sublayers Mixes information across tokens and transforms it
Attention A weighting scheme over other tokens Lets each position pull in relevant context
Output scores (logits) A score per vocabulary token Ranked to pick the next token

The vocabulary is fixed, so every input and output is expressed in terms of tokens the model already knows.

From text to prediction, step by step

  1. Tokenize. Input text is split into tokens. A word may become one token or several subword pieces.
  2. Embed. Each token maps to a vector. Position information is added so the model knows order.
  3. Pass through transformer layers. Each layer applies attention (tokens exchange information) and a feed-forward network (each position is transformed).
  4. Produce logits. The final representation is projected onto the vocabulary, giving a score for every possible next token.
  5. Select a token. The highest score wins under greedy decoding; sampling methods can pick lower-ranked tokens for variety.
  6. Append and repeat. The chosen token is added to the sequence, and the loop runs again to produce the next one.

The expected result at each loop is a single new token; a full response is just this loop repeated.

How attention works, in plain terms

Attention answers: "for this token, which other tokens matter right now?" Each token produces a query, and every token produces a key and a value. The query is compared against all keys to get weights, and the values are combined using those weights. So a pronoun can gather information from the noun it refers to, and a verb can look back at its subject.

This is why context length matters: attention can only weigh tokens inside the window it is given. Multi-head attention runs several of these weightings in parallel so different relationships can be tracked at once.

Training vs. inference

These are the same architecture used in two different modes:

  • Training: the model sees text with the next token known, predicts it, measures the error, and adjusts weights across many examples. This is where knowledge and language patterns are learned.
  • Inference: weights are frozen. The model only predicts forward, token by token, with no learning. This is what happens when you use a chat model.

A useful intuition: training is studying with an answer key; inference is answering without one.

Building intuition with interactive visualizations

Reading the steps is not the same as seeing them. AnimatedLLM is an interactive resource whose stated purpose is to help you understand how large language models work under the hood. Use it to watch tokens become vectors, follow attention weights between positions, and trace how a prediction is produced. For example, if you want to see why a model completes "The cat sat on the ___" with "mat," step through the attention view to observe which earlier tokens the final position is weighting most heavily.

Common sticking points

  • Tokens are not words. A single word can be multiple tokens, which is why models sometimes miscount letters.
  • Embeddings are not meanings you can read off. They are learned coordinates; similarity is geometric, not definitional.
  • Attention is not memory. It operates within the current context window, not a persistent store.
  • One token at a time. Fluency comes from repetition, not from generating a whole sentence in one pass.

If you want to go from "I've heard of transformers" to "I can trace the pipeline," work through the tokenize → embed → attend → score → select loop once, then use an interactive demo to watch each stage on real text.

animatedllm.github.io
Understand how large language models work under the hood.
firecrawl.dev
Firecrawl is the web data API to search, scrape, and interact with the web at scale. Turn any source into clean Markdown or structured data your agen…
humanloop.com
Humanloop is joining Anthropic to accelerate the adoption of AI, safely.
jan.ai
Jan is an open-source alternative to ChatGPT. Run open-source AI models locally or connect to cloud models like GPT, Claude and others.
llmeval.com
LLMEval is a research series dedicated to building comprehensive, fair, and robust evaluation frameworks for large language models.
pygpt.net
One desktop app for cloud and local AI. Chat, build agents, connect your knowledge with RAG, and get things done with skills, tools and plugins. Free…