Export citation

Export citation

Choose format for download:

Download Citation

    Two-phase perspective on deep learning dynamics

    Robert de Mello Koch1,2,* and Animik Ghosh1,†

    • *Contact author: robert@zjhu.edu.cn
    • †Contact author: animikghosh@gmail.com

    Phys. Rev. E 112, 025307 – Published 12 August, 2025

    DOI: https://doi.org/10.1103/3l3p-vkf2

    Abstract

    We propose that learning in deep neural networks proceeds in two phases: a rapid curve-fitting phase followed by a slower compression or coarse-graining phase. This view is supported by the shared temporal structure of three phenomena—grokking, double descent, and the information bottleneck—all of which exhibit a delayed onset of generalization well after training error reaches zero. We empirically show that the associated timescales align in two rather different settings. Mutual information between hidden layers and input data emerges as a natural progress measure, complementing circuit-based metrics such as local complexity and the linear mapping number. We argue that the second phase is not actively optimized by standard training algorithms and may be unnecessarily prolonged. Drawing on an analogy with the renormalization group, we suggest that this compression phase reflects a principled form of forgetting, critical for generalization.

    Physics Subject Headings (PhySH)

    Authorization Required

    We need you to provide your credentials before accessing this content.

    References (Subscription Required)

    Outline

    Information

    Sign In to Your Journals Account

    Filter

    Filter

    Article Lookup

    Enter a citation