What is a large language model?
The engine inside ChatGPT, Claude and Gemini is, at heart, a very well-read autocomplete. Here’s how it works, and why that explains its quirks.
Video transcript
ChatGPT, Claude and Gemini all run on something called a large language model. It sounds technical, but the idea behind it is surprisingly simple.
You know how your phone suggests the next word as you type? A large language model does the same thing, after learning from an enormous amount of text.
It splits your message into tokens, small chunks of text. Then it predicts a likely next token, adds it, and predicts again, over and over, until the answer is written.
First it’s trained on huge amounts of text, from books to web pages. Then people rate its answers, which fine-tunes it to be helpful, polite and safe.
Knowing how it works explains its quirks. Very long chats can lose track of the start. The same question can get two different answers. And it may miss recent news unless it searches the web.
Here’s a handy trick for long chats: ask for a summary of what you’ve agreed, then paste it into a fresh chat and carry on.
Most of all, remember that it predicts what a good answer sounds like. Predicting isn’t checking, so it can sound sure and still be wrong.
That’s a large language model, in plain English. Read the full guide below to get more from the assistant you use.
In 30 seconds
- A large language model writes one token at a time: autocomplete on a vast scale.
- It learns from huge amounts of text, then fine-tuning with human feedback makes it helpful.
- It can lose track of very long chats, and may not know recent news unless it searches.
Start a text with “Running a bit” and your phone suggests “late”. That’s autocomplete: a guess at the next word, based on what usually comes next. The AI inside ChatGPT, Claude and Gemini works on the same idea, at a scale that’s hard to picture.
Autocomplete on a vast scale
That AI is a Jargon busterLarge language model: An AI model trained on vast amounts of text to predict the next word. It powers assistants like ChatGPT, Claude and Gemini., or LLM. Like your phone’s keyboard, its job is to predict what comes next. The difference is scale: it has learnt from so much text that its guesses can add up to a whole, sensible answer.
It doesn’t read words quite as you do. It splits text into Jargon busterToken: A chunk of text, often part of a word, that a language model reads and writes one at a time., chunks that are often part of a word. Then it predicts a likely next token, adds it, and predicts again, over and over, until the reply is finished.
- Your prompt“Why do onions make you cry?”
- Split into tokensWords become small chunks of text
- Predict the next tokenIts best guess at what comes next
- RepeatAdd it, guess again, until the answer is done
How it learns to be helpful
First comes Jargon busterTraining: The stage where AI learns, by finding patterns in a huge number of examples. on a huge amount of text: books, articles, websites and more. By predicting the next token again and again, it picks up grammar, facts, styles and how arguments fit together.
A model trained like that just carries on whatever you type. So next comes Jargon busterFine-tuning: Extra training that shapes a model for a particular job or style, after its main training.: extra training in which people show it good answers and rate its replies, teaching it to follow instructions, be helpful and polite, and turn down harmful requests. The result is the Jargon busterAI assistant: A chat app such as ChatGPT, Claude or Gemini that answers questions and does tasks you describe in plain language. you chat to.
Why it behaves the way it does
- It can lose track of long chats. There’s a limit to how much it can take in at once, called its Jargon busterContext window: How much text an AI can take into account at once, including your conversation so far.. In a very long chat, the start can slip out of view.
- Ask twice, get two answers. It picks among likely tokens with a little randomness, so replies vary. Handy for ideas; for facts, it’s a reason to check.
- It may not know recent events. Its knowledge stops when its training text was gathered. For anything current, ask it to search the web.
You: This chat is getting long. Sum up what we’ve decided so far, so I can paste it into a fresh chat.
AI: Here’s your summary: garden party for 30 people in late June, £400 budget, vegetarian buffet, gazebo in case it rains. Still to decide: music and invitations.
Search the web for the latest on [topic]. Give me a short summary, and list the date and link for each source you used.
Prediction, not lookup
The big thing to remember: at heart, an LLM predicts what a good answer sounds like rather than looking it up. That’s why it can sound sure and still be wrong. A confident mistake like that is called a Jargon busterHallucination: When AI states something false as if it were true, because it predicts plausible words rather than checking facts.: here’s how to catch one.
The more it knows about what you want, the better its predictions. So the most useful next step is learning how to write a good prompt.
Check yourself
3 quick questions nothing is savedTools in this guide
Spotted a mistake? Tell us and an editor will check it.