arxiv preprint - Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck

In this episode, we discuss Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck by Nathan Godey, Éric de la Clergerie, Benoît Sagot. This paper investigates the phenomenon of performance saturation in small language models, attributing the issue to a mismatch between the model’s hidden layer size and the complexity of the targeted probability distribution. The softmax bottleneck, a known limitation in neural networks, is identified as a contributing factor to this mismatch, leading to suboptimal performance due to the emergence of degenerate latent representations during late pretraining stages. The study demonstrates that models with fewer than 1000 hidden dimensions are particularly susceptible to this effect, resulting in decreased effectiveness upon evaluation.

arxiv preprint – Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck