Subscribe free
AI 101: start here4 min readAdvanced

How AI models are trained: from raw text to a helpful assistant

Every chatbot goes through the same broad stages, from trillions of words of text to a careful, helpful assistant. Here’s what each stage does, what it costs and why models differ.

4 min read

In 30 seconds

  • Pretraining on trillions of builds broad knowledge; fine-tuning and feedback turn it into an assistant.
  • Each company’s data, feedback and written rules shape its model, so assistants behave differently.
  • Training fixes a knowledge cutoff. Open-weight models give you more control, and more responsibility.

Every assistant began as a model that read an enormous amount of text, was shaped by examples and feedback, then tested before release. Here’s that journey, with published figures. For the basics, see how AI works and what a large language model is.

  1. PretrainingPredict the next token, trillions of times
  2. Fine-tuningImitate examples of good answers
  3. FeedbackLearn which answers people, or AI, prefer
  4. TestingRed-teaming and evaluation before release
The broad stages. Each company runs its own version.

Pretraining: reading at enormous scale

feeds the model huge amounts of text, such as web pages, books and code. For each stretch, it predicts the next token, compares its guess with the real one and adjusts its internal numbers slightly. Nobody labels anything: the text is its own answer key. Repeated trillions of times, this builds grammar, facts, styles and patterns of reasoning.

15 trillion

tokens, at least, of publicly available text went into pretraining Meta’s Llama 3 models, using 7.7 million hours of computing on specialised chips.

Source: Meta, Meta Llama 3 model card, April 2024

That’s the expensive part: one chip would need almost 900 years for 7.7 million hours of work, so companies run thousands side by side for weeks or months. The result knows a lot but isn’t an assistant yet: ask it a question and it may simply carry on writing.

Fine-tuning and feedback: learning to help

trains the model on examples of good answers, many written by people: for Llama 3, Meta used public instruction datasets plus over 10 million human-annotated examples. Then comes : people pick the better of two answers, a separate model learns to predict their picks, and the assistant is trained to score well with it.

100×

smaller, yet preferred: people rated answers from a small version of OpenAI’s InstructGPT, trained with human feedback, above those from GPT-3, a model over 100 times its size.

Source: OpenAI, InstructGPT: Training Language Models to Follow Instructions with Human Feedback

Feedback can come from AI too. In Anthropic’s Constitutional AI method, a model judges which of two answers is better against a written list of principles, so human oversight comes through those rules. Anthropic says Claude now uses its constitution to create many kinds of its own training data, including rankings of possible answers.

Safety, red-teaming and testing

Safety training teaches the model to refuse clearly harmful requests and handle sensitive ones with care. In , specialists try hard to make it misbehave, so weak spots can be fixed. Meta reports extensive red-teaming of Llama 3, and in 2024 the UK and US government AI safety institutes jointly tested a new Claude model before release.

Evaluation runs throughout: benchmark tests of maths, coding and knowledge, safety tests and comparisons by people. Treat scores with care: test questions can leak into training data, and a good score isn’t the same as doing your job well, so test models on your own tasks.

Why models differ, and what open-weight means

Every stage involves choices: which data to include, who writes the examples, what raters reward, and the written rules behind it all. OpenAI publishes a Model Spec for how its models should behave; Anthropic publishes a constitution for Claude. Even the raters matter: OpenAI noted that InstructGPT’s 40 or so contractors weren’t representative of everyone who would use it.

Training also fixes a . Llama 3’s smaller model learnt from data up to March 2023, its larger one up to December 2023. Anthropic lists two dates per Claude model: when its training data ends, and a “reliable knowledge” cutoff, up to which its knowledge is most extensive and reliable.

Open-weight or closed?
Open-weightClosed
Run it on your own computers or cloudUse it through the company’s app or API
Data can stay inside your organisationData goes to the provider, under its terms
Fine-tune it for your own jobCustomise only as far as the provider allows
You handle hosting, security and safetyThe provider handles updates and safety

Examples of an include Meta’s Llama, under its own commercial licence, and OpenAI’s gpt-oss, under the permissive Apache 2.0 licence; the smaller gpt-oss runs within 16GB of memory. Licences vary, so read the terms first. The main models behind ChatGPT, Claude and Gemini are closed. To try one, see run AI on your own computer.

Weigh up open-weight against closed

I’m choosing between an open-weight model and a closed model for [the job, for example summarising customer emails]. Our situation: [how sensitive the data is, expected volume, budget, in-house technical skills]. Compare the two options on data control, running costs, safety responsibilities and licence terms, then list the questions I should ask each provider before deciding.

Treat the answer as a checklist to investigate, and confirm licence terms and prices with each provider.

Check yourself

3 quick questions nothing is saved
1What does a model learn to do during pretraining?

2In reinforcement learning from human feedback, what do the raters do?

3What’s the main trade-off of an open-weight model?

Up next in AI 101: start here

Beginner4 min read

Spot AI fakes: pictures, videos, voices and scams

AI can fake a photo, a video or a loved one’s voice, and the tell-tale glitches are disappearing. Here’s how to check what’s real, starting with where it came from.

More guides

Get AI explained at your level, every weekday

The five stories that matter, in plain English, plus a new guide each week. Free.