From Raw Model To Assistant: Pretraining, Fine-Tuning, RLHF And Reasoning Models
A freshly pretrained language model only continues text. Several further training stages turn it into an assistant that follows instructions. Here is what each stage does.
Checked against primary sources and independently reviewed on . Sources are listed at the end.
If you type “Write a polite reminder to a supplier about an overdue invoice” into a model that has only been pretrained, you may not get a reminder at all. You might get a list of similar requests, the start of a forum post, or a continuation of the sentence. A raw model has learned to continue text, not to help.
The assistants people use every day go through several more training stages after pretraining. Each stage changes the model’s behaviour in a specific way, and knowing what each one does helps explain both what assistants are good at and where their safeguards come from.
The Training Pipeline At A Glance
- Pretraining
Next-token prediction on very large amounts of text, where the text supplies its own answers. Produces broad knowledge and fluency.
- Supervised Fine-Tuning
Training on written examples of good responses to instructions. Teaches the assistant format.
- Preference Optimisation
Training on judgements of which of two answers is better, using RLHF, DPO or AI feedback.
- Optional: Reasoning Training
For example, rewarding correct final answers to maths or coding problems, which encourages longer step-by-step reasoning.
Stage One: Pretraining
The model reads a huge collection of text and is trained to predict the next token, as described in How An LLM Reads And Writes. Nobody labels anything; the text supplies its own answers, which is why this is called self-supervised learning.
This stage has traditionally taken the bulk of the computing budget. The InstructGPT team, for example, reported that their most expensive follow-on training run used about 60 petaflop/s-days of computing, a unit of total computing work, against roughly 3,640 for pretraining GPT-3.1
The result is a model with wide knowledge of language and facts, and a surprising ability to perform tasks when shown a few examples in the prompt. The GPT-3 paper in 2020 demonstrated this few-shot ability in a model with 175 billion parameters (the adjustable numbers inside the model), without any task-specific retraining.2 What a pretrained model lacks is any built-in sense that it should answer the question in front of it.
Stage Two: Instruction Tuning
The next step is supervised fine-tuning: continuing training on a smaller, carefully chosen set of prompts paired with good responses. When the examples are instructions and helpful replies, this is called instruction tuning.
A 2021 Google study called FLAN fine-tuned a 137 billion parameter model on more than 60 language tasks rewritten as instructions. The tuned model did clearly better on kinds of task it had never been tuned on, and the authors’ tests showed that the number of datasets used for tuning, the size of the model and the instruction wording each played a part.3 Instruction tuning is what teaches a model the basic shape of being an assistant: read the request, then respond to it.
Stage Three: Learning From Preferences
Written examples only go so far. It is often easier for a person to say which of two answers is better than to write the perfect answer from scratch. Preference training uses exactly that kind of judgement.
The best-known method is reinforcement learning from human feedback (RLHF). Reinforcement learning means training by trial and reward: the model produces outputs, each one is scored, and the model is nudged towards outputs that score well. OpenAI’s InstructGPT paper in 2022 described three steps: fine-tune on human-written demonstrations, train a separate reward model to predict which outputs people prefer, then use reinforcement learning to steer the language model towards outputs the reward model scores highly.1
The paper’s headline result came from labellers comparing answers to prompts that customers had sent to OpenAI’s API. They favoured a tuned model with 1.3 billion parameters over the untuned GPT-3, a model more than 100 times larger.
Two later approaches changed parts of this recipe.
| Method | Source Of Judgements | How It Trains | Main Idea |
|---|---|---|---|
| RLHF (InstructGPT, 2022) | People rank model outputs | Reward model, then reinforcement learning | Optimise for what people prefer |
| DPO (2023) | People or models choose the better of two outputs | A single classification-style loss, no separate reward model | Same goal as RLHF with a simpler, more stable procedure |
| Constitutional AI (2022) | A model judges outputs against written principles | Self-critique and revision, then reinforcement learning from AI feedback | Fewer human labels, with the rules written down |
Anthropic’s Constitutional AI, published in December 2022, tackled the harmlessness side without people rating which outputs were harmful. It relied instead on a written list of principles and on an AI model’s judgements against them, an approach it called reinforcement learning from AI feedback.4 Direct Preference Optimisation (DPO), published in 2023 by researchers at Stanford, then showed that the preference goal can be reached by training the language model directly on preferred and rejected answers, without training a separate reward model or running reinforcement learning.5
Stage Four: Reasoning Models
A separate line of work concerns step-by-step reasoning. In 2022, Google researchers showed that prompting large models with a few worked examples that spell out intermediate steps, called chain-of-thought prompting, improved performance on arithmetic, commonsense and symbolic reasoning tasks.7
Newer reasoning models build this behaviour in through training. DeepSeek-AI’s R1 paper, first posted in January 2025 and later published in Nature, reported on a version called R1-Zero. It was trained with reinforcement learning that rewarded correct final answers on checkable problems, such as maths and coding, and it developed longer reasoning with behaviours like checking its own work, without any human-written reasoning examples.8 The released R1 model added some supervised training to that recipe, and the same paper showed that smaller models could pick up reasoning by being fine-tuned on R1’s outputs, a technique called distillation.8 These models typically write out a long stretch of working before their final answer, which tends to help on multi-step problems and adds cost and delay.
That working is generated by the same next-token process as everything else, and it is not a dependable record of how the model reached its answer. In a 2025 study by Anthropic researchers, reasoning models given a hint in the prompt often used it without saying so, and the rate at which the reasoning admitted to using the hint was often below 20 percent.9 The text can still help a reviewer, but read it as a draft explanation rather than a log.
What This Means For Users
Each stage leaves its mark. Pretraining supplies most of the model’s knowledge and largely sets its knowledge cut-off date; one 2024 study found that models picked up new facts through fine-tuning only slowly, and that learning them made the models more prone to hallucinate.10 Instruction tuning sets how it follows requests. Preference training sets its tone, what it declines and how cautious it is. Reasoning training affects how it handles multi-step problems. When an assistant behaves unexpectedly, it is often useful to ask which stage the behaviour most likely comes from. For how these models are then used with tools and given goals, see AI Agents And Multi-Agent Systems.
Footnotes
-
L. Ouyang et al., “Training language models to follow instructions with human feedback”, arXiv:2203.02155, March 2022. arxiv.org ↩ ↩2
-
T. B. Brown et al., “Language Models are Few-Shot Learners”, arXiv:2005.14165, May 2020. arxiv.org ↩
-
J. Wei et al., “Finetuned Language Models Are Zero-Shot Learners”, arXiv:2109.01652, September 2021. arxiv.org ↩
-
Y. Bai et al. (Anthropic), “Constitutional AI: Harmlessness from AI Feedback”, arXiv:2212.08073, December 2022. arxiv.org ↩
-
R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning and C. Finn, “Direct Preference Optimization: Your Language Model is Secretly a Reward Model”, arXiv:2305.18290, May 2023. arxiv.org ↩
-
A. Wei, N. Haghtalab and J. Steinhardt, “Jailbroken: How Does LLM Safety Training Fail?”, arXiv:2307.02483, July 2023. arxiv.org ↩
-
J. Wei et al., “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models”, arXiv:2201.11903, January 2022. arxiv.org ↩
-
DeepSeek-AI, “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning”, arXiv:2501.12948, January 2025; published in Nature, volume 645, 2025. arxiv.org ↩ ↩2
-
Y. Chen et al. (Anthropic), “Reasoning Models Don’t Always Say What They Think”, arXiv:2505.05410, May 2025. arxiv.org ↩
-
Z. Gekhman et al., “Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?”, EMNLP 2024, arXiv:2405.05904. arxiv.org ↩
Knowledge Hub content is general information. It is not legal advice, a compliance certification, a guarantee of security or a substitute for an assessment of your own systems. Standards and rules change; check the sources for the latest position.