CORTEXA
← Browse
arxivcs.LG2026-07-13

How to Tame Grokking: Representation Geometry as a Control Signal

Maksim A Kazanskii

Grokking is a phenomenon in which neural networks initially memorize training data and only later exhibit strong generalization after prolonged optimization. Despite extensive recent study, the factors influencing the emergence and timing of grokking remain incompletely understood. We investigate the relationship between representation geometry and delayed generalization. We find that dimensionality collapse consistently precedes the onset of grokking in all evaluated settings. Motivated by these observations, we introduce Geometric Dimensionality Regularization (GeomDR), a simple spectral regularizer that modifies the effective dimensionality of hidden representations during training. Across modular addition, modular division, and permutation composition tasks, GeomDR consistently alters grokking dynamics and can substantially accelerate the onset of generalization depending on the intervention schedule and target dimensionality. In several settings, grokking is accelerated by up to 52 times relative to standard AdamW training. Similar qualitative effects are observed in both multilayer perceptrons and transformers. Together, these results suggest that representation geometry can serve as an effective control signal for grokking and provide evidence that geometric interventions offer a practical approach for studying and influencing delayed generalization in neural networks.

View free PDFSource page

Related papers

arxivcs.LG2026-07-23

The Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and What Actually Works

Yu Wang

Dense per-step supervision is an appealing remedy for sparse-reward, long-horizon LLM agents: reward the agent for predicting its next observation, and memory should follow. We show that under group-normalized RL (GRPO), this recipe does not merely fail -- it destroys the policy.…

View free PDFSource page
arxivcs.IRcs.LG2026-07-22

Zero-Observation User Reactivation with Gap-Driven Dimensional Gating

Jiandong Ding, Tianying Liu, Fuyuan Liu, Huijie Qin, Tiandeng Wu

Sequential recommendation (SR) models capture continuously observed behavior, but a returning user may have no interactions for months or years. We define this setting as Zero-Observation Reactivation: the user has a pre-gap history, while the platform observes no behavioral sign…

View free PDFSource page
arxivcs.LG2026-07-23

Gradient Concentration, Not Weight Saliency, Explains Representation-Level Class Unlearning

Billel Habbati, Alessio Merlo, Luca Verderame, Meriem Guerar

Machine unlearning aims to remove the influence of specific training data while preserving model utility. Many state-of-the-art approaches pursue this goal by restricting the forgetting update to a subset of parameters selected through gradient-based saliency. Although such metho…

View free PDFSource page
arxivcs.PLcs.LG2026-07-23

Relaxed activation analysis of dataflow networks - A clock calculus for machine learning and real-time scheduling

William Gaudelier, Albert Cohen, Dumitru Potop Butucaru

Previous work has shown that the simple dataflow primitives of the Lustre language allow the natural, semantically unambiguous, and compact representation of machine learning (ML) applications, including models featuring complex conditional execution and recurrent state. The Lust…

View free PDFSource page