Renormalization Group Theory Reveals When Transformer Attention Matters
A new preprint applies Wilsonian renormalization group theory to analyze the role of attention in Transformer models, treating attention as a perturbation to MLP residual stacks. The study finds that attention is crucial ('relevant') for data with long-range correlations, driving a phase transition in representation space, but is largely unnecessary ('irrelevant') for short-range correlations. The first-layer attention head is shown to dominate the representational shift, and attention selectively preserves slow Markov modes in long-correlation regimes.
Why it matters: This work offers a predictive theoretical framework linking the usefulness of attention mechanisms to the spectral properties of input data, potentially guiding model design and application.
Full story at: arXiv Machine Learning ↗