Subscribe free
AI 101: start here3 min readBeginner

What is a large language model?

The engine inside ChatGPT, Claude and Gemini is, at heart, a very well-read autocomplete. Here’s how it works, and why that explains its quirks.

3 min read 1:16 video with captions
Video transcript

ChatGPT, Claude and Gemini all run on something called a large language model. It sounds technical, but the idea behind it is surprisingly simple.

You know how your phone suggests the next word as you type? A large language model does the same thing, after learning from an enormous amount of text.

It splits your message into tokens, small chunks of text. Then it predicts a likely next token, adds it, and predicts again, over and over, until the answer is written.

First it’s trained on huge amounts of text, from books to web pages. Then people rate its answers, which fine-tunes it to be helpful, polite and safe.

Knowing how it works explains its quirks. Very long chats can lose track of the start. The same question can get two different answers. And it may miss recent news unless it searches the web.

Here’s a handy trick for long chats: ask for a summary of what you’ve agreed, then paste it into a fresh chat and carry on.

Most of all, remember that it predicts what a good answer sounds like. Predicting isn’t checking, so it can sound sure and still be wrong.

That’s a large language model, in plain English. Read the full guide below to get more from the assistant you use.

In 30 seconds

  • A large language model writes one token at a time: autocomplete on a vast scale.
  • It learns from huge amounts of text, then fine-tuning with human feedback makes it helpful.
  • It can lose track of very long chats, and may not know recent news unless it searches.

Start a text with “Running a bit” and your phone suggests “late”. That’s autocomplete: a guess at the next word, based on what usually comes next. The AI inside ChatGPT, Claude and Gemini works on the same idea, at a scale that’s hard to picture.

Autocomplete on a vast scale

That AI is a , or LLM. Like your phone’s keyboard, its job is to predict what comes next. The difference is scale: it has learnt from so much text that its guesses can add up to a whole, sensible answer.

It doesn’t read words quite as you do. It splits text into , chunks that are often part of a word. Then it predicts a likely next token, adds it, and predicts again, over and over, until the reply is finished.

  1. Your prompt“Why do onions make you cry?”
  2. Split into tokensWords become small chunks of text
  3. Predict the next tokenIts best guess at what comes next
  4. RepeatAdd it, guess again, until the answer is done
Every reply is written this way: one small prediction at a time, repeated very fast.

How it learns to be helpful

First comes on a huge amount of text: books, articles, websites and more. By predicting the next token again and again, it picks up grammar, facts, styles and how arguments fit together.

A model trained like that just carries on whatever you type. So next comes : extra training in which people show it good answers and rate its replies, teaching it to follow instructions, be helpful and polite, and turn down harmful requests. The result is the you chat to.

Why it behaves the way it does

  • It can lose track of long chats. There’s a limit to how much it can take in at once, called its . In a very long chat, the start can slip out of view.
  • Ask twice, get two answers. It picks among likely tokens with a little randomness, so replies vary. Handy for ideas; for facts, it’s a reason to check.
  • It may not know recent events. Its knowledge stops when its training text was gathered. For anything current, ask it to search the web.

You: This chat is getting long. Sum up what we’ve decided so far, so I can paste it into a fresh chat.

AI: Here’s your summary: garden party for 30 people in late June, £400 budget, vegetarian buffet, gazebo in case it rains. Still to decide: music and invitations.

When a chat gets long, carry the important bits into a fresh one, where nothing has slipped out of view.
Get an up-to-date answer

Search the web for the latest on [topic]. Give me a short summary, and list the date and link for each source you used.

Use it for anything that changes: news, rules, prices or opening hours. Then open a link or two to check.

Prediction, not lookup

The big thing to remember: at heart, an LLM predicts what a good answer sounds like rather than looking it up. That’s why it can sound sure and still be wrong. A confident mistake like that is called a : here’s how to catch one.

The more it knows about what you want, the better its predictions. So the most useful next step is learning how to write a good prompt.

Check yourself

3 quick questions nothing is saved
1What is a large language model doing when it writes a reply?

2Why might a very long chat forget something you said at the start?

3Why might it not know about something that happened last week?

Up next in AI 101: start here

Beginner3 min read

Why AI makes things up, and how to catch it

AI can invent facts, quotes and links, and sound completely sure about them. Here’s why it happens, and five simple habits that help you catch it.

More guides

Get AI explained at your level, every weekday

The five stories that matter, in plain English, plus a new guide each week. Free.