Sharp Stability Threshold and Certification for Designing Stable Residual Architectures
A new preprint introduces the sublinear-growth principle for deep residual architectures, identifying a sharp stability threshold based on the input-magnitude exponent (q ≤ 1) of residual blocks. The authors provide theoretical arguments showing this criterion is both necessary and sufficient for stable training, clarifying the stabilizing role of layer normalization and enabling efficient certification of architectural stability. The work also demonstrates a parameter-free modification that stabilizes the Mamba block without normalization, with experiments confirming stable training for q ≤ 1 variants.
Why it matters: This research offers a principled, theoretically grounded method for designing and certifying stable deep residual networks, potentially reducing reliance on empirical trial-and-error in architecture development.
Full story at: arXiv Machine Learning ↗