How Feature Learning can Improve Neural Scaling

Last modified: July 21, 2026

#scaling-laws #representations

They mentioned some lazy-training regime?

What is the lazy training regime? The lazy training regime (aka the NTK/kernel regime) is a description of how very wide neural networks behave during gradient descent β€” and it’s a bit counterintuitive: the network trains successfully, but its internal features barely move.