How Feature Learning can Improve Neural Scaling
Last modified: July 21, 2026
#scaling-laws #representations
They mentioned some lazy-training regime?
What is the lazy training regime? The lazy training regime (aka the NTK/kernel regime) is a description of how very wide neural networks behave during gradient descent β and it’s a bit counterintuitive: the network trains successfully, but its internal features barely move.