AI FUNDAMENTALS / INTERMEDIATE

Three Ways Machines Learn: Supervised, Unsupervised And Reinforcement Learning

Machine learning methods differ in the kind of feedback a model gets. Here is how supervised, unsupervised, self-supervised and reinforcement learning work, with examples of each.

Checked against primary sources and independently reviewed on . Sources are listed at the end.

Every machine learning model learns from some kind of feedback. The clearest way to tell the main approaches apart is to ask one question: what signal tells the model whether it is doing well?

That signal can be a correct answer supplied by a person, the structure hidden in the data itself, or a reward that arrives after the model acts. This article explains each approach with a business example, then shows how a fourth approach, self-supervised learning, made large language models possible.

Supervised Learning: Learning From Labelled Examples

In supervised learning, every training example comes with the right answer, called a label.1 The model sees a loan application and the label “repaid” or “defaulted”, or a photo of a product and the label “damaged” or “intact”. Training adjusts the model until its predictions match the labels as closely as possible.

Many everyday business models are trained this way, including fraud scores, churn prediction, document classification and demand forecasts. Its main cost is the labels. Someone has to produce them, and if those labels are inconsistent or reflect past unfair decisions, the model learns the same pattern.

Unsupervised Learning: Finding Structure Without Answers

In unsupervised learning, there are no labels. The model is asked to find structure in the data on its own.1 A common example is clustering: given a year of customer purchase histories, group customers who behave similarly, without being told in advance what the groups should be.

Another use is anomaly detection. A model learns what normal network traffic or normal expense claims look like and flags items that sit far from that norm. This is useful in security, where examples of new attacks are rare and often unlabelled. The catch is that the model cannot explain why a group or outlier matters; a person still has to interpret the result.

Reinforcement Learning: Learning From Rewards

Reinforcement learning is different again. A learner, usually called an agent in this field, takes actions in an environment and receives rewards or penalties based on the results. Over many attempts, it learns which actions tend to produce the most reward over time. Sutton and Barto’s textbook sets out this framing of an agent, an environment and a reward signal in detail.2

A classic example is a game-playing program that is told only whether it won or lost and gradually discovers good strategies. The reward often arrives late, after a long chain of moves, so the learner has to work out which earlier choices earned it. Reinforcement learning is also used to shape the behaviour of language models. In OpenAI’s InstructGPT work, published in 2022, people ranked several answers to the same prompt, those rankings trained a separate reward model, and the language model was then tuned to score well on it.3 The Large Language Models group explains this in more depth.

ApproachFeedback SignalBusiness ExampleMain Limitation
SupervisedA correct label for every examplePredicting which invoices are fraudulent from past confirmed casesNeeds many accurate labels
UnsupervisedNone; the model finds patterns in the dataGrouping customers by buying behaviourResults need human interpretation
Self-supervisedLabels created from the data itself, such as the next wordPretraining a language model on large amounts of textNeeds very large data and computing power
ReinforcementA reward after acting in an environmentTuning a model so people prefer its answersRewards can be gamed if badly designed
The feedback signal is what separates the four approaches.

Self-Supervised Learning: The Bridge To Language Models

Self-supervised learning sits between the first two approaches. The data has no human labels, but the training process creates labels automatically by hiding part of the data and asking the model to predict it. For text, a common task, and the one GPT-3 was trained on, is: given the words so far, predict the next one.4

  1. Take A Sentence

    The invoice was paid on time.

  2. Hide What Comes Next

    Input: The invoice was paid on. Target: time.

  3. Model Predicts

    The model guesses a word, perhaps late or time.

  4. Compare And Adjust

    The error between the guess and the real word adjusts the parameters.

  5. Repeat At Scale

    Every position in every document becomes a new example.

Self-supervised learning turns ordinary text into millions of training examples with no human labelling.

Because every sentence on the web, in a book or in a code repository contains its own answers, self-supervision scales to amounts of data that no team could ever label by hand. GPT-3, described in 2020, was trained this way to predict text, and its authors then showed it handling many tasks when given only a few examples in the prompt.4 The EU AI Act reflects how central this method has become: its definition of a general-purpose AI model expressly covers models trained on large amounts of data using self-supervision at scale.5

Choosing An Approach

In practice the approaches are combined. InstructGPT is a well-documented example: a model pretrained with self-supervision was fine-tuned on examples of good answers written by people, then shaped with reinforcement learning from their rankings.3 Developers vary this recipe, and not every assistant is trained the same way. In a January 2025 report, later peer-reviewed in Nature, DeepSeek described using reinforcement learning on tasks with checkable answers, such as maths and coding problems, and reported that its DeepSeek-R1 models improved at step-by-step reasoning as a result.6

For a business question, the choice usually comes down to what feedback you can afford to produce. If you have reliable labels, supervised learning is the default. If you only have raw data and want to explore it, unsupervised methods help. If success can only be judged after a sequence of actions, reinforcement learning fits, at the cost of careful reward design.

Footnotes

  1. Google for Developers, “What is Machine Learning?”, Introduction to Machine Learning, last updated 27 January 2026. developers.google.com ↩ ↩2

  2. R. S. Sutton and A. G. Barto, “Reinforcement Learning: An Introduction”, 2nd edition, MIT Press, 2018. incompleteideas.net ↩

  3. L. Ouyang et al., “Training language models to follow instructions with human feedback”, arXiv:2203.02155, March 2022. arxiv.org ↩ ↩2

  4. T. B. Brown et al., “Language Models are Few-Shot Learners”, arXiv:2005.14165, May 2020. arxiv.org ↩ ↩2

  5. Regulation (EU) 2024/1689 (Artificial Intelligence Act), Article 3(63), consolidated text of 27 July 2026 (as amended by Regulation (EU) 2026/1744), EUR-Lex. eur-lex.europa.eu ↩

  6. DeepSeek-AI, “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning”, arXiv:2501.12948, January 2025; peer-reviewed version in Nature 645, 633 to 638, 2025. arxiv.org ↩

Knowledge Hub content is general information. It is not legal advice, a compliance certification, a guarantee of security or a substitute for an assessment of your own systems. Standards and rules change; check the sources for the latest position.