How AI models are trained: from raw text to a helpful assistant
Every chatbot goes through the same broad stages, from trillions of words of text to a careful, helpful assistant. Here’s what each stage does, what it costs and why models differ.
In 30 seconds
- Pretraining on trillions of Jargon busterToken: A chunk of text, often part of a word, that a language model reads and writes one at a time. builds broad knowledge; fine-tuning and feedback turn it into an assistant.
- Each company’s data, feedback and written rules shape its model, so assistants behave differently.
- Training fixes a knowledge cutoff. Open-weight models give you more control, and more responsibility.
Every assistant began as a model that read an enormous amount of text, was shaped by examples and feedback, then tested before release. Here’s that journey, with published figures. For the basics, see how AI works and what a large language model is.
- PretrainingPredict the next token, trillions of times
- Fine-tuningImitate examples of good answers
- FeedbackLearn which answers people, or AI, prefer
- TestingRed-teaming and evaluation before release
Pretraining: reading at enormous scale
Jargon busterPretraining: The first and biggest stage of training, where a model learns from vast amounts of text by predicting the next token. feeds the model huge amounts of text, such as web pages, books and code. For each stretch, it predicts the next token, compares its guess with the real one and adjusts its internal numbers slightly. Nobody labels anything: the text is its own answer key. Repeated trillions of times, this builds grammar, facts, styles and patterns of reasoning.
tokens, at least, of publicly available text went into pretraining Meta’s Llama 3 models, using 7.7 million hours of computing on specialised chips.
Source: Meta, Meta Llama 3 model card, April 2024That’s the expensive part: one chip would need almost 900 years for 7.7 million hours of work, so companies run thousands side by side for weeks or months. The result knows a lot but isn’t an assistant yet: ask it a question and it may simply carry on writing.
Fine-tuning and feedback: learning to help
Jargon busterFine-tuning: Extra training that shapes a model for a particular job or style, after its main training. trains the model on examples of good answers, many written by people: for Llama 3, Meta used public instruction datasets plus over 10 million human-annotated examples. Then comes Jargon busterReinforcement learning from human feedback: Training in which people compare a model’s answers, and the model learns to give the kind they prefer. Often shortened to RLHF.: people pick the better of two answers, a separate model learns to predict their picks, and the assistant is trained to score well with it.
smaller, yet preferred: people rated answers from a small version of OpenAI’s InstructGPT, trained with human feedback, above those from GPT-3, a model over 100 times its size.
Source: OpenAI, InstructGPT: Training Language Models to Follow Instructions with Human FeedbackFeedback can come from AI too. In Anthropic’s Constitutional AI method, a model judges which of two answers is better against a written list of principles, so human oversight comes through those rules. Anthropic says Claude now uses its constitution to create many kinds of its own training data, including rankings of possible answers.
Safety, red-teaming and testing
Safety training teaches the model to refuse clearly harmful requests and handle sensitive ones with care. In Jargon busterRed-teaming: Deliberately trying to make an AI misbehave, such as giving harmful or false answers, so the problems can be fixed before release., specialists try hard to make it misbehave, so weak spots can be fixed. Meta reports extensive red-teaming of Llama 3, and in 2024 the UK and US government AI safety institutes jointly tested a new Claude model before release.
Evaluation runs throughout: benchmark tests of maths, coding and knowledge, safety tests and comparisons by people. Treat scores with care: test questions can leak into training data, and a good score isn’t the same as doing your job well, so test models on your own tasks.
Why models differ, and what open-weight means
Every stage involves choices: which data to include, who writes the examples, what raters reward, and the written rules behind it all. OpenAI publishes a Model Spec for how its models should behave; Anthropic publishes a constitution for Claude. Even the raters matter: OpenAI noted that InstructGPT’s 40 or so contractors weren’t representative of everyone who would use it.
Training also fixes a Jargon busterKnowledge cutoff: The date a model’s training data stops. It knows little or nothing about later events unless it can search the web.. Llama 3’s smaller model learnt from data up to March 2023, its larger one up to December 2023. Anthropic lists two dates per Claude model: when its training data ends, and a “reliable knowledge” cutoff, up to which its knowledge is most extensive and reliable.
| Open-weight | Closed |
|---|---|
| Run it on your own computers or cloud | Use it through the company’s app or API |
| Data can stay inside your organisation | Data goes to the provider, under its terms |
| Fine-tune it for your own job | Customise only as far as the provider allows |
| You handle hosting, security and safety | The provider handles updates and safety |
Examples of an Jargon busterOpen-weight model: A model whose numbers are published, so anyone can download and run it themselves. include Meta’s Llama, under its own commercial licence, and OpenAI’s gpt-oss, under the permissive Apache 2.0 licence; the smaller gpt-oss runs within 16GB of memory. Licences vary, so read the terms first. The main models behind ChatGPT, Claude and Gemini are closed. To try one, see run AI on your own computer.
I’m choosing between an open-weight model and a closed model for [the job, for example summarising customer emails]. Our situation: [how sensitive the data is, expected volume, budget, in-house technical skills]. Compare the two options on data control, running costs, safety responsibilities and licence terms, then list the questions I should ask each provider before deciding.
Check yourself
3 quick questions nothing is savedTools in this guide
Sources (9)
- Meta Llama 3 model cardMeta, April 2024
- InstructGPT: Training Language Models to Follow Instructions with Human FeedbackOpenAI
- InstructGPT model cardOpenAI, January 2022
- Constitutional AI: Harmlessness from AI feedbackAnthropic, December 2022
- Claude’s new constitutionAnthropic, January 2026
- Introducing computer use, a new Claude 3.5 Sonnet, and Claude 3.5 HaikuAnthropic, October 2024
- Model SpecOpenAI
- Models overviewAnthropic (Claude documentation)
- gpt-ossOpenAI
Spotted a mistake? Tell us and an editor will check it.