How Large Language Models Work: From Your Prompt to the Answer
You type a question, and a few seconds later ChatGPT, Claude or Gemini writes a fluent, often impressively useful answer. It can feel like the machine understands you. Under the hood, though, a large language model does one surprisingly simple thing, over and over: it predicts the next piece of text.
This post follows your prompt through a large language model (LLM), step by step, from raw text to the finished answer. It builds on how neural networks learn, but you can follow it without reading that first. At the end, you'll build a tiny language model in about 30 lines of Python.
The one-sentence version
An LLM is a very large neural network trained to predict the next token of text, and it writes answers by making that prediction again and again. Everything else, from grammar to facts to reasoning, emerges from getting very good at that single task across an enormous amount of text.
From prompt to answer
Here is the whole journey at a glance. We'll zoom into each box below.
An LLM writes one token at a time. Each new token is added to the text and the whole process repeats.
Notice the loop. An LLM doesn't compose its whole answer at once. It produces one token, adds it to the text, and runs the entire process again to get the next one. A 300-word answer means hundreds of trips around that loop.
Step 1: Text becomes tokens
Computers work with numbers, not letters, so the first job is to chop your text into tokens. A token is usually a common word or a piece of a word. Frequent words like "the" get a single token, while rarer words are split into parts. Each token has an ID number from the model's vocabulary, which typically contains tens of thousands of tokens.
Tokens are often pieces of words. Each token maps to a number the model can work with.
This is why LLM limits and prices are counted in tokens rather than words. In English, a token averages about three-quarters of a word, so 1,000 tokens is roughly 750 words.
Step 2: Tokens become meaning, as numbers
A token ID on its own is just a label. The next step turns each token into an embedding: a long list of numbers, often thousands long, that captures its meaning. You can think of it as coordinates on an enormous map of meaning. Words with similar meanings, such as "king" and "queen" or "Paris" and "London", end up close together on that map.
Nobody writes these numbers by hand. Like the weights in any neural network, they are learned during training because they help the model make better predictions.
Step 3: The Transformer pays attention
The heart of every modern LLM is the Transformer, the architecture introduced in 2017 in the paper "Attention Is All You Need". Its key idea is attention: for every token, the model works out which other tokens in the text matter most for understanding it.
Take the sentence "The animal didn't cross the street because it was tired." What does "it" refer to? You know instantly that it's the animal. Attention is how the model figures that out:
Attention lets each word look at every other word. Here “it” pays most attention to “animal”, which is what it refers to.
A Transformer stacks dozens of these attention layers. Early layers pick up grammar and word relationships. Later layers combine them into more abstract patterns, such as who did what to whom, the tone of the text, or the steps of an argument. Large models spread hundreds of billions of learned numbers, called parameters, across these layers.
Step 4: Predicting the next token
After the final layer, the model produces a score for every token in its vocabulary: how likely is each one to come next? Those scores become probabilities.
The model scores every possible next token. Usually a likely one is picked, with a little randomness.
Then one token is chosen. Always picking the top token makes text repetitive and dull, so models usually sample with a little randomness, controlled by a setting called temperature. Low temperature means safe, predictable choices, good for facts and code. Higher temperature means more varied and creative text, but also more mistakes.
Step 5: Repeat until done
The chosen token is appended to the text, and the whole process runs again with the slightly longer text as input. This continues until the model produces a special "end" token or hits a length limit.
The amount of text a model can consider at once is its context window. It includes your prompt, any documents you've pasted in, the conversation so far, and the answer being written. Anything outside the window, the model simply cannot see.
How an LLM is trained
A model that only predicts the next word isn't automatically a helpful assistant. Chatbots are built in stages:
A chatbot is built in stages: first it learns language, then it learns to follow instructions, then it learns what good answers look like.
- Pretraining. The model reads a vast collection of text, including books, websites, articles and code, and learns to predict the next token. This is where it absorbs grammar, facts and patterns of reasoning. It is by far the most expensive stage, running for weeks or months on thousands of specialized chips.
- Instruction tuning. The model is trained on examples of questions paired with good answers, so it learns to respond to instructions instead of just continuing text.
- Preference tuning. People compare pairs of answers and pick the better one. The model is then adjusted to produce answers people prefer: more helpful, more honest, and less harmful. A common version of this is called reinforcement learning from human feedback (RLHF).
Build a tiny language model in Python
Real LLMs are huge, but the core idea fits in a few lines. The program below "trains" on a handful of sentences by counting which word follows which. It then generates text the same way an LLM does: predict probabilities for the next token, pick one, append it, repeat. Each numbered comment matches a step above.
import random
from collections import Counter, defaultdict
text = """the cat sat on the mat . the dog sat on the rug .
the cat chased the dog . the dog chased the ball .
the cat slept on the mat . the dog slept on the rug ."""
# 1. Tokenize: split the text into tokens (here, simply words)
tokens = text.split()
# 2. "Train": count which token follows which
next_counts = defaultdict(Counter)
for current, following in zip(tokens, tokens[1:]):
next_counts[current][following] += 1
# 3. Predict: turn counts into probabilities for the next token
def next_token_probs(word):
counts = next_counts[word]
total = sum(counts.values())
return {w: round(c / total, 2) for w, c in counts.most_common()}
print("After 'the':", next_token_probs("the"))
# 4. Generate: repeatedly sample the next token and append it
def generate(start, length=8, temperature=1.0, seed=7):
rng = random.Random(seed)
out = [start]
for _ in range(length):
probs = next_token_probs(out[-1])
words = list(probs)
# temperature < 1 sharpens the choice, > 1 flattens it
weights = [p ** (1 / temperature) for p in probs.values()]
out.append(rng.choices(words, weights=weights)[0])
return " ".join(out)
print(generate("the", temperature=0.2))
print(generate("the", temperature=1.5))
Running it prints:
After 'the': {'dog': 0.33, 'cat': 0.25, 'mat': 0.17, 'rug': 0.17, 'ball': 0.08}
the dog sat on the dog . the dog
the cat sat on the mat . the cat
Notice the first generated line: "the dog sat on the dog". Our toy model only ever looks at the previous word, so it has no idea the sentence already mentioned a dog. That is exactly the problem attention solves. A real LLM looks back across thousands of tokens at once, which is why its text stays coherent across whole pages.
Why LLMs make things up
Because an LLM is trained to produce likely text, not verified text, it can generate statements that sound right but are false. This is called hallucination. It is most likely with obscure facts, exact numbers, quotes and citations. Other important limits:
- Knowledge cutoff: a model knows only what was in its training data, unless it is connected to search or your own documents.
- Context limits: it can only consider what fits in its context window.
- Bias: it can reproduce biases present in its training text.
A popular fix for the first problem is retrieval-augmented generation (RAG): fetch relevant documents first, put them in the context window, and ask the model to answer from them. That will be covered in detail in the RAG posts.
Key terms, decoded
- Large language model (LLM)
- A very large neural network trained to predict the next token of text.
- Token
- A chunk of text, often a word or part of a word, that the model reads and writes.
- Embedding
- A list of numbers representing a token's meaning, learned during training.
- Transformer
- The neural network architecture behind modern LLMs, built around attention.
- Attention
- The mechanism that lets each token weigh how relevant every other token is.
- Parameters
- The learned numbers inside the model. Large models have billions of them.
- Temperature
- A setting that controls how random the choice of the next token is.
- Context window
- The maximum amount of text the model can consider at once.
- Hallucination
- Confident output that is false or made up.
Strip away the mystique and an LLM is a next-token predictor: tokens become numbers, attention connects them, probabilities pick the next token, and the loop repeats. The surprising part is how much capability emerges when that simple recipe is scaled up. Next, it's time to get hands-on and train a first machine learning model in Python.
Which step would you like explored in more depth? Let us know in the comments.

Comments
Post a Comment