Distillation: Compressing One Mind Into a Smaller One
A large model's competence need not stay large.
A large model's competence need not stay large.
Inside a large trained network may hide a tiny one that could have learned the task alone.
Conventional wisdom says a model too large will overfit.
There is no algorithm that is best at everything.
High-dimensional data is rarely as vast as it looks.
To regularize a model is to lean on it, gently, in a chosen direction.
Every model that fits data faces a pull in two directions.
No learner learns from data alone.
To learn is to go beyond what one has seen.