CORTEXA
← Browse
arxivcs.LG2026-07-03

Implicit Bias of SGD in Multivariate ReLU Networks: Effective Width Collapse

Shuang Liang, Tom Jacobs, Guido Montúfar

We study the implicit bias of noisy stochastic gradient descent in training wide two-layer ReLU networks for multivariate regression. In a mean-field regime, the training dynamics are approximated by a Wasserstein gradient flow that converges to a unique stationary measure. We characterize the structure of this stationary measure and the predictor it represents. We show that, despite the network being infinitely overparameterized, the learned predictor admits an effectively finite representation: the input weights and biases align along finitely many directions, leading to an effective width collapse. In particular, the solution function is continuous piecewise affine, with affine regions determined by the cells of a finite hyperplane arrangement. The number of learned directions, and hence hyperplanes, is bounded above by $2\mathcal{P}-1$, where $\mathcal{P}$ denotes the number of linear dichotomies realizable on the training inputs. We further establish a non-redundancy property of the learned representation by proving that each learned direction induces a unique ternary activation pattern on the training data. Consequently, the complexity of the learned predictor is governed by the combinatorial geometry of the training data.

View free PDFSource page

Related papers

arxivcs.LG2026-07-08

The Anatomy of Implicit Bias: Information Allocation in Neural Network Training

Zhang Gongyue, Wang Zhiyong, Liu Donghan, Ren Weihong, Sheng Yixuan, Liu Honghai

Implicit bias is usually explained as the preference of an optimization process for certain final solutions and their geometry. This view helps explain where a model finally stops. It gives less direct explanation of how this bias is formed during training. This paper proposes a…

View free PDFSource page
arxivcs.LG2026-07-18

Effects of width-dependent model hyperparameters and $\ell_2$-regularization on the loss landscape of two-layer ReLU networks

Haruka Eshima, Makoto Yamada

Understanding deep neural networks remains a central challenge in machine learning. In particular, the theoretical properties of even two-layer ReLU networks, especially in the presence of weight decay, remain poorly understood. To this end, we derive a sufficient condition on th…

View free PDFSource page
arxivcs.LG2026-07-14

Gradient Flow Dynamics and Implicit Bias of Diagonal Linear Networks under Infinitesimal Initialization

Jiajie Zhao, Jianxing Wang, Junjie Yang, Zhiwei Bai, Yaoyu Zhang

We study the gradient flow dynamics of diagonal linear networks for regression tasks under infinitesimal initialization. Extending Theorem 1 from Pesme & Flammarion (2023), we generalize the analysis to both deep diagonal linear networks and a broader class of two-layer diagonal…

View free PDFSource page
arxivstat.MLcs.ITcs.LGmath.NA2026-07-12

Approximation of Analytic Functions by ReLU Neural Networks with Adjustable Depth and Width

Yanming Lai, Defeng Sun, Yang Wang

In contrast to most studies on neural network approximation theory that characterize results through a single parameter, such as the total number of network parameters, \cite{shen2020deep} pioneered the characterization of approximation rates as a joint function of the width para…

View free PDFSource page
arxivcs.LGcs.AI2026-07-08

On the Principles of Deep Feedforward ReLU Networks

Changcun Huang

The architecture of deep feedforward neural networks is ubiquitous in deep learning, either as a whole system or as a subnetwork of other architectures, and thus its mechanism is a key ingredient of the black box of neural networks. On the basis of the simplest two-layer ReLU net…

View free PDFSource page
arxivcs.LG2026-07-21

Relative Positions Generalize, Absolute Positions Memorize: An Implicit-Bias Account of Length Generalization in Attention

Subham Singh, Ashutosh Mishra, Subha Raut

Transformers with relative positional encodings often extrapolate to sequences longer than those seen during training, whereas transformers with learned absolute encodings typically do not. This is a robust empirical regularity, and the explanations offered for it so far are chie…

View free PDFSource page