What the study found
The study found that the information content of hidden neural network representations changes with label noise and network size. It also found a double descent pattern in this information content as the number of network parameters changes.
Why the authors say this matters
The authors conclude that the relationship between information imbalance, a proxy for conditional mutual information, and test error offers a new perspective on generalization. They also suggest that the results show how training objectives shape internal representations.
What the researchers tested
The researchers compared hidden representations learned by neural networks of different sizes using the Information Imbalance, which they describe as a computationally efficient proxy for conditional mutual information. They trained the networks on datasets with controlled levels of label noise and examined representations across layers.
What worked and what didn't
In the underparameterized regime, representations learned with noisy labels were more informative than those learned with clean labels. In the overparameterized regime, the two were equally informative, and label noise reduced the information content between the penultimate layer and the pre-softmax layer, matching the increase in test error. Representations learned from random labels performed worse than random features when the number of parameters and training samples were scaled proportionally with a fixed ratio.
What to keep in mind
The abstract does not describe limitations beyond the studied settings, so the findings should be read as applying to the networks, dataset conditions, and scaling regimes tested here. The summary also does not report details about specific architectures, datasets, or the magnitude of the observed effects.
Key points
- Hidden representations showed a double descent pattern as network size changed.
- Noisy-label representations were more informative than clean-label ones in the underparameterized regime.
- Overparameterized networks produced representations that were equally informative under noisy and clean labels.
- Label noise lowered information content between the penultimate and pre-softmax layers.
- Random-label training performed worse than random features under proportional scaling of parameters and samples.
Disclosure
- Research title:
- Label noise changes hidden representations in neural networks
- Authors:
- Ali Hussaini Umar, Franky Kevin Nando Tezoh, Jean Barbier, S. Acevedo, Alessandro Laio
- Institutions:
- International Centre for Theoretical Sciences, Scuola Internazionale Superiore di Studi Avanzati, Scuola Internazionale Superiore di Studi Avanzati, Scuola Internazionale Superiore di Studi Avanzati, Scuola Internazionale Superiore di Studi Avanzati
- Publication date:
- 2026-04-22
- OpenAlex record:
- View
Get the weekly research newsletter
Stay current with scholarly research without reading academic papers — one filtered digest, every Friday.