Some networks abruptly generalize long after they seem done learning. Researchers observed ‘grokking,’ where a model masters a task suddenly after prolonged training. The phenomenon probes how learning and memorization differ.
A delayed leap
Timing surprises. Generalization arrives well after memorization. The jump is sharp.
Memorize then generalize
Two phases appear. The model first fits data, then finds the rule. Understanding follows.
A study window
Simplicity helps. Small algorithmic tasks make it observable. Analysis is tractable.
Regularization role
Pressure matters. Weight decay influences when grokking occurs. Dynamics are studied.
Theory interest
The puzzle is deep. It touches how networks learn structure. Theorists engage.
Open questions
Much is unknown. Whether it generalizes to big models is unclear. Work continues.
The bottom line
Grokking is the sudden generalization of a network long after it appears to have learned, revealing a memorize-then-understand transition. It probes learning dynamics. Its scope in large models is unresolved.