Dust: Pretraining Transformers Without Backpropagation
- October 2026Correspondence to s@qlabs.sh·Code· TL;DR We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models.
- Dust perturbs activations (node perturbation) independently at every token, so each token is a virtual population member and one forward pass evaluates them all in parallel.
- Dust approximates backprop closely at large population (i.e.
Unverified
- October 2026Correspondence to s@qlabs.sh·Code· TL;DR We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models.
- Dust perturbs activations (node perturbation) independently at every token, so each token is a virtual population member and one forward pass evaluates them all in parallel.
- Dust approximates backprop closely at large population (i.e.
Sources: Qlabs