Low-Rank Emergence (1)
This post begins a series of extensions to our previous post, The Neural Matthew Effect: Low Effective Degrees of Freedom in Training. We want to understand, at least partially, how a full-rank matrix can turn into an effectively low-rank one during training. Here, we use low rank in a broad sense: only a small fraction of the neural connections play a decisive role in the function represented by the network. Several quantities can describe this phenomenon, including effective rank and stable rank; see the previous post for their definitions. ...