Augmented Lagrangian Predictive Coding: training 1000-layer networks without backpropagation
- We introduce PC-ALM, a local alternative to backpropagation.
- PC-ALM trains residual MLPs up to 1000 layers, nearly matching backprop's performance despite using only layer-local dynamics.
- PC-ALM equips each layer with a feedback control dynamical system that distributes and propagates supervision credit throughout a network.
Unverified
- We introduce PC-ALM, a local alternative to backpropagation.
- PC-ALM trains residual MLPs up to 1000 layers, nearly matching backprop's performance despite using only layer-local dynamics.
- PC-ALM equips each layer with a feedback control dynamical system that distributes and propagates supervision credit throughout a network.
Sources: Sakana