Overfitting, Evaluation And The Basics Of Trustworthy AI
A model that aces its training data can still fail in the real world. Learn how overfitting happens, how held-out testing catches it and how NIST describes trustworthy AI.
Checked against primary sources and independently reviewed on . Sources are listed at the end.
A vendor tells you their model is 99 percent accurate. That number means very little until you know what it was measured on. A model can score perfectly on the data it learned from and still make poor decisions on the cases it meets in use.
This article explains why that happens, how careful evaluation exposes it, and how the US National Institute of Standards and Technology (NIST) places accuracy within a wider picture of what makes an AI system trustworthy.
What Overfitting Is
Overfitting is what happens when a model learns its training examples too well. It latches on to details that only exist in those examples, so it looks impressive during development and then gets new cases wrong, which is how Google’s machine learning course frames the problem.1 The opposite quality, performing well on data the model has never seen, is called generalisation, and it is the whole point of machine learning.
An analogy helps. A student who memorises last year’s exam answers will score perfectly on last year’s paper and badly on this year’s. The student has learned the paper, not the subject. A model with enough parameters can do the same thing: it can memorise quirks and noise in its training examples rather than the underlying pattern.
Suppose a model learns to spot damaged parcels from warehouse photos. If every damaged parcel in the training set happened to be photographed at one loading bay, the model may learn to recognise that bay’s lighting rather than the damage. It will look excellent in training and fail at every other site.
How Held-Out Data Exposes It
The defence is to judge a model on data it did not learn from. That is why data is split into training, validation and test sets before training begins, as explained in How A Machine Learning Model Is Built.
During training, developers track the error, or loss, on both the training set and the validation set. Early on, both usually fall together. In an overfitting model, the training loss keeps falling while the validation loss levels off and may start to rise. Google’s course treats that divergence as a sign of overfitting.1 It is a warning, not proof of exactly when memorising began. Common responses include early stopping, which Google’s course describes as ending training once validation loss starts to climb, along with simplifying the model or collecting more varied data.2
The test set has a special role. It is meant to be kept apart from every decision about the model and used at the end, so that it gives an independent estimate of performance. If a team keeps checking the test score and adjusting the model in response, the model gradually becomes tuned to that test set, and the score says less about new data.3
Accuracy Is Only One Number
A single accuracy figure can hide a lot. If one in a hundred transactions is fraudulent, a model that always says “not fraud” is 99 percent accurate and useless. For problems like this, teams track false positives (good cases wrongly flagged) and false negatives (bad cases missed) separately, because each has a different cost.
NIST’s AI Risk Management Framework makes a similar point. It names false positive and false negative rates among the measures to consider, and says any accuracy figure should come with a description of the test data, which should look like the situations the system will really face, and of the method used.4 It adds that accuracy results may be broken down by data segment. That breakdown is how a model that works well overall but poorly for one group of customers gets caught.
What Trustworthy AI Means In NIST’s Framework
NIST’s framework, published in January 2023 as voluntary guidance, lists seven characteristics of trustworthy AI systems.4 It describes “valid and reliable” as a necessary condition for the others, which is why evaluation sits at the base of the stack below. It treats accountability and transparency as relating to all of the other characteristics.
- Accountable And TransparentRelates to all the others: who is responsible, and what can people see about how the system works.
- SafeUsed within its intended conditions, the system should not put people, property or the environment at risk.
- Secure And ResilientWithstands attacks and unexpected events, and recovers from them.
- Explainable And InterpretablePeople can understand how the system works and what its outputs mean.
- Privacy-EnhancedProtects personal data and human autonomy.
- Fair, With Harmful Bias ManagedAddresses harmful bias and discrimination, recognising that fairness is judged differently across contexts.
- Valid And ReliableDoes what it is meant to do, consistently, including on data it did not train on.
NIST stresses that these qualities have to be balanced against each other for the specific context of use; for example, an organisation may have to trade some predictive accuracy for a model that people can interpret. It also says AI systems should be tested before deployment and regularly while they are in operation, because a model that was valid at launch may not stay valid as conditions change.4 How organisations turn these principles into policy and controls is covered in AI Governance And Regulation, and the security characteristics in AI Security.
Questions To Ask Of Any Model
When you are buying or approving a model, a few questions go a long way. What data was it tested on, and does that data look like your own? Were results broken down by customer group, region or product line? How does it perform on the errors that matter most to you? Who monitors it after launch, and what triggers retraining?
Footnotes
-
Google for Developers, “Overfitting”, Machine Learning Crash Course, last updated 3 December 2025. developers.google.com ↩ ↩2
-
Google for Developers, “Overfitting: L2 regularization”, Machine Learning Crash Course, last updated 9 April 2026. developers.google.com ↩
-
Google for Developers, “Datasets: Dividing the original dataset”, Machine Learning Crash Course, last updated 3 December 2025. developers.google.com ↩
-
NIST, AI 100-1, “Artificial Intelligence Risk Management Framework (AI RMF 1.0)”, January 2023. nvlpubs.nist.gov ↩ ↩2 ↩3
Knowledge Hub content is general information. It is not legal advice, a compliance certification, a guarantee of security or a substitute for an assessment of your own systems. Standards and rules change; check the sources for the latest position.