+1 (484) 312-5566

Before the Architecture: Auditing Label Noise and Duplicates in Your Training Set

When a Keras model plateaus two points below the target, the reflex is to reach for architecture: a bigger backbone, a new optimizer, another hyperparameter sweep. In review engagements we usually find the ceiling is somewhere else. A few percent of labels are wrong, the same examples appear on both sides of the split, and one class is defined differently by two annotators. None of that is fixed by EfficientNetV2-L.

This is the audit we run before proposing any modeling change. It takes a day or two, it needs no new labels to start, and it produces a number you can put in a plan: this many examples are suspect, and here is what fixing them is worth.

Why label noise caps your metrics twice

Wrong labels hurt in two distinct places, and teams usually only think about the first.

In training, noisy labels pull weights toward contradictory targets. Modern networks have enough capacity to memorize them, so the damage shows up as worse generalization and longer training, not as visible training loss. Mild noise is survivable; correlated noise — one annotator, one shift, one site systematically mislabeling a class — is not, because the model learns the annotator's mistake as a rule.

In evaluation, wrong labels make your metrics unreadable. If 3% of test labels are wrong, the gap between an 88%-accurate model and a 91%-accurate one is inside the noise floor, and every comparison you run — backbone A vs B, augmentation on vs off — is partly measuring which model better predicts your annotators' errors. This is the more expensive failure, because it silently invalidates the decisions you spent GPU-hours making.

So the rule we apply: fix the evaluation set first, even if you never clean the training set. Cleaning 500 test labels is a day of work that makes every future experiment interpretable.

Step 1: out-of-fold predictions, not training predictions

To find suspect labels you need a prediction for every example from a model that did not train on that example. That means k-fold cross-validation over the full dataset:

from sklearn.model_selection import StratifiedGroupKFold
import numpy as np

oof = np.zeros((len(labels), num_classes), dtype="float32")

for train_idx, val_idx in StratifiedGroupKFold(n_splits=5).split(X, labels, groups):
    model = build_model()           # same architecture every fold
    model.fit(make_ds(train_idx), epochs=EPOCHS, verbose=0)
    oof[val_idx] = model.predict(make_ds(val_idx, shuffle=False))

Two details matter more than the model choice. Use groups — session, patient, machine, capture day — so near-duplicates cannot leak across folds and hide the very problem you are hunting. And use a cheap model: a frozen backbone plus a linear head is enough. You are scoring labels, not chasing state of the art, and five folds of a small model is an hour on one GPU rather than an overnight run.

Calibrate the out-of-fold probabilities before you threshold on them. Temperature scaling on a held-out slice costs minutes and keeps the confidence scores from being systematically overstated.

Step 2: rank the suspects

With oof in hand, three cheap signals find most real problems:

  1. Self-confidence. The probability the model assigned to the given label. Sort ascending; the bottom of that list is where wrong labels live.
  2. Margin. p(top predicted class) - p(given label). A large positive margin means the model is confidently voting for a different class — the classic confident-learning signal.
  3. Class-confusion structure. Build the confusion matrix of predicted vs given labels over out-of-fold predictions. A heavy off-diagonal block between two classes rarely means the model is confused; it usually means your class definitions overlap and annotators split the boundary differently.
given_conf = oof[np.arange(len(labels)), labels]
margin = oof.max(axis=1) - given_conf
suspects = np.argsort(-margin)[:500]

Review the top few hundred by hand. Expect roughly three outcomes: genuinely wrong labels (fix them), genuinely ambiguous examples (this is a taxonomy problem — write the rule down, or add an uncertain class), and correct-but-hard examples (leave them; they are the useful ones). Log the disposition of each, because the ratio tells you what kind of problem you have. Mostly wrong labels means an annotation-quality fix. Mostly ambiguity means a guidelines fix, and no amount of relabeling will help until the guidelines change.

Step 3: duplicates and near-duplicates

Duplicates are the second half of the audit, and they cause the inflated validation numbers that collapse in production. Embed the dataset once with a pretrained backbone and look for close neighbors:

backbone = keras.applications.EfficientNetV2B0(include_top=False, weights="imagenet", pooling="avg")
emb = backbone.predict(dataset)
emb /= np.linalg.norm(emb, axis=1, keepdims=True)
sim = emb @ emb.T                       # use a FAISS/ScaNN index beyond ~50k rows

Pairs above ~0.98 cosine similarity are usually the same underlying thing: burst frames, re-uploads, the same part photographed twice. Two actions follow. If a near-duplicate pair straddles your train/test boundary, your test metric is optimistic — regroup the split. If a near-duplicate pair carries different labels, you have found a labeling inconsistency for free, no model required.

For time-series and tabular data the same idea applies with different machinery: exact-row hashing, then window-overlap checks across split boundaries. Overlapping windows across a split is the single most common way a forecasting backtest lies to you.

Step 4: quantify the fix before you buy it

Relabeling costs money, so bring evidence. Clean the evaluation set first, then re-score your existing candidate models on the cleaned set. Two results are common and both are useful:

  • Metrics rise across the board and the ranking of models changes. That means your previous model-selection decisions were partly noise, and it justifies the cleanup immediately.
  • Metrics rise but the ranking is stable. Your comparisons were sound; the headline number was just pessimistic. Cheaper news, still worth knowing before you promise an accuracy figure to a customer.

Then, for the training set, estimate the payoff on a slice: correct the top 5% most-suspect labels, retrain with fixed seeds and a fixed split, and compare against the same run on uncorrected data across two or three seeds. If cleaning 5% moves the metric more than any architecture change you have tried, you have your roadmap — and a defensible number for the budget conversation.

Make it a standing check, not a one-off

The audit is worth automating because datasets keep growing. We wire three things into the training repo: the out-of-fold scoring job runs on every major dataset version and writes a ranked suspect list; the duplicate scan runs at split time and fails the build if a near-duplicate pair crosses the train/test boundary; and the cleaned evaluation set is versioned and frozen, with every change to it recorded in the run record alongside the code and data hashes.

None of this is modeling work, and that is rather the point. In an hourly engagement, the audit is usually the cheapest day on the invoice and the one that changes the plan most. Baselines before novelty, data quality before architecture — in that order, because the architecture experiments only mean something once the labels underneath them are trustworthy.