What Is Machine Learning and How Do Models Learn from Data?
Machine learning is a way of building software that learns patterns from data instead of following rules a person wrote by hand. You use it when the relationship between input and output is too complex or too variable to specify directly — recognizing objects in photos, ranking search results, or predicting whether a transaction is fraudulent. The core idea: show the system many examples, let it adjust internal parameters to reduce its errors, then check whether it works on examples it has never seen.
Learning from data vs. hand-coded rules
In traditional programming, a developer writes explicit logic: if the email contains these words, mark it spam. In machine learning, you supply labeled examples and the model derives its own decision boundary. The trade-off is that the model's behavior depends on the data it saw — change the data, and the behavior changes.
The three main learning paradigms
| Paradigm | What the model gets | What it learns | Concrete example |
|---|---|---|---|
| Supervised learning | Inputs paired with correct answers | A mapping from input to output | Predicting house prices from size, location, and age |
| Unsupervised learning | Inputs only, no labels | Structure or groupings in the data | Grouping customers by purchasing behavior |
| Reinforcement learning | A reward signal from acting in an environment | A policy that maximizes cumulative reward | Training a model to solve multi-step reasoning tasks |
Supervised learning covers most everyday applications. Unsupervised learning is used for clustering, compression, and anomaly detection. Reinforcement learning is harder to stabilize but is the approach behind recent work on scaling language-model reasoning — for instance, Qwen's GSPO research explicitly targets "stable and robust training dynamics" for RL at scale, noting that existing algorithms such as GRPO "exhibit severe instability issues during" training.
The core training loop
Every supervised model follows roughly the same cycle:
- Collect and split data. Divide examples into a training set and a held-out test set.
- Define a model. Choose an architecture with adjustable parameters (weights).
- Measure error with a loss function. The loss quantifies how far predictions are from the correct answers.
- Adjust parameters. An optimization algorithm nudges the weights to reduce the loss.
- Repeat. Iterate over the data many times until the loss stops improving.
- Evaluate on unseen data. Measure performance on the test set, not the training set.
The input is the data and the model definition; the action is repeated parameter updates; the expected result is a model whose error on new data is acceptably low.
Overfitting, underfitting, and why splits matter
- Underfitting: the model is too simple to capture the pattern — it performs poorly on both training and test data.
- Overfitting: the model memorizes the training examples, including their noise — it performs well on training data but poorly on test data.
This is why you never judge a model by its training accuracy. A held-out test set (or cross-validation) simulates the real world: data the model has not seen. If training error keeps falling while test error rises, you are overfitting.
Where deep learning and large language models fit
Deep learning is machine learning using neural networks with many layers. It is not a separate field — it is a subcategory that excels when data is abundant and patterns are hierarchical (images, audio, text).
Large language models are deep learning models trained on massive text corpora, usually with a self-supervised objective: predict the next token. That objective needs no human labels, which is why it scales. The Qwen family illustrates the breadth of the umbrella — its releases include a 20B image foundation model (Qwen-Image) for text rendering and editing, a safety classifier (Qwen3Guard) fine-tuned for prompt and response moderation, and RL research (GSPO) for training dynamics. All of these are machine learning systems; they differ in data, objective, and architecture, not in kind.
How to tell the paradigms apart in practice
Ask two questions:
- Does the training data include the correct answer? If yes, it is supervised (or self-supervised, where the answer is derived from the data itself).
- Does the model learn by taking actions and receiving feedback? If yes, it is reinforcement learning.
If neither applies and you are only looking for structure, it is unsupervised. Most real systems combine these — a language model may be pretrained with self-supervision, fine-tuned with supervised examples, and refined with reinforcement learning.