Grading a student on the exact problems they studied measures memory, not understanding. The same is true of models: performance on the data used for training says almost nothing about performance on data the model has never seen, and the whole point of supervised-learning is the never-seen part. The fix is procedural and non-negotiable: partition your labeled data by role, and never let information from evaluation data leak into the choices that build the model. Professor Scott’s version, repeated like a safety brief, is simply: don’t test on the training set.
The idea
Split the labeled data three ways. The training set fits the model’s parameters. The validation set gives an unbiased look at the fit while you tune hyperparameters and decide when to stop. The test set is independent data, drawn from the same distribution, used exactly once at the end to estimate generalization. Data that influenced any modeling decision can no longer provide an unbiased estimate.
The Three Roles
The training set is what gradient-descent actually sees. The loss is computed on it and the weights are fit to it.
The validation set exists because training involves choices the training loss cannot referee: learning rate, architecture, regularization strength, how long to train. You evaluate candidate settings on the validation set and keep what performs well. CS231n describes this precisely: the validation set is used essentially as a fake test set to tune the hyperparameters. It also drives early stopping, since a validation error that starts climbing while training error keeps falling is the classic sign of overfitting. Note what this implies: the validation set participates in model building. It is spent.
The test set is the one dataset that influenced nothing. It is independent of every fitting and tuning decision but follows the same distribution as the training data, and that independence is what makes its error an unbiased estimate of real-world performance. The discipline is strict and CS231n states it flatly: evaluate on the test set only a single time, at the very end. Every peek before then converts test data into validation data, and your final number stops being trustworthy.
Why Leakage Ruins the Estimate
The failure mode is always the same: information flows from evaluation data into the model, and the evaluation becomes flattery. Testing on training data is the blatant version, but the subtle versions bite harder. Tuning hyperparameters against the test set is training on it, one bit at a time. If you augment data by duplicating and transforming instances, the copies of one original must not straddle the train/test boundary, or the model is being tested on near-duplicates of what it studied (a warning straight from the course slides). The estimator only means something if the test set stayed clean.
When Data Is Scarce: Cross-Validation
A three-way split costs data, and with a small dataset a single small validation set gives noisy, untrustworthy signals. -fold cross-validation is the standard remedy: partition the data into equal folds, then train times, each time holding out a different fold for evaluation and training on the rest, and average the results. Every example gets used for both training and evaluation, just never in the same round. The course leans on this same machinery for comparing two learning algorithms, pairing the per-fold results with a statistical test (see hypothesis-testing) rather than trusting a single split’s difference.
Warning
Validation performance is an honest guide for choosing between models, but it is not an honest final report. The model you picked won the validation set partly by fitting its quirks. That is exactly what the untouched test set is for, and why it only works once.
Related Notes
- bias-variance-tradeoff explains the overfitting this split is designed to catch
- generalization-vs-memorization on why unseen-data performance is the only score that counts
- evaluation-metrics covers what to compute on the test set once you finally use it
- hypothesis-testing for judging whether a measured difference between models is real
- gradient-descent interacts with the validation set through early stopping
- supervised-learning supplies the labeled data being partitioned
Sources
- https://en.wikipedia.org/wiki/Training,_validation,_and_test_data_sets (roles of the three sets; validation as unbiased evaluation during hyperparameter tuning; rising validation error as an overfitting signal; test set independent but same distribution)
- https://cs231n.github.io/classification/ (validation set as a fake test set for hyperparameter tuning; evaluate on the test set only a single time, at the very end; cross-validation when validation data is small)
- https://en.wikipedia.org/wiki/Cross-validation_%28statistics%29 (-fold procedure: partition into folds, rotate the held-out fold, average)
- Course framing: CSCE 479/879 (Stephen Scott, UNL), “Regularization and Performance Evaluation” lecture slides (don’t test on the training set; augmentation duplicates must not span splits; -fold CV with paired tests)