Dust: Pretraining Transformers Without Backpropagation
First reported by Qlabs.sh ·
The ability to pretrain transformer models without backpropagation offers a new avenue for training models that were previously difficult to optimize.
Researchers have developed a new pretraining method for transformer language models called Dust, which does not rely on backpropagation. Dust is a zeroth-order optimization algorithm that perturbs activations independently at each token, treating each token as a virtual member of a population. This approach allows for parallel evaluation of all tokens in a single forward pass. The method approximates backpropagation closely, especially with larger computational resources, and in some cases, it has surpassed backprop's performance. Dust is significantly more efficient than traditional weight-space evolutionary strategies, being orders of magnitude faster. Contrary to expectations, larger models trained with Dust demonstrate greater population efficiency. The gradient estimates produced by Dust align well with those from backpropagation, even for models up to 1 billion parameters, indicating its potential for scaling.
Dust's novel approach of perturbing activations instead of weights bypasses the computational overhead associated with traditional evolutionary strategies. By treating each token as an independent member within a single forward pass, Dust achieves remarkable efficiency gains, estimated to be thousands of times faster than existing methods like EGGROLL. This efficiency is particularly pronounced in larger models, which surprisingly become more population-efficient, suggesting a reframing of overparameterization as an advantage for this new training paradigm.
This research challenges the long-standing reliance on backpropagation in deep learning by proposing a viable alternative for transformer pretraining. While Dust is not yet compute-efficient enough to replace backpropagation for current applications, its success in matching and exceeding backprop's performance at scale opens doors for training novel architectures. Future work could explore integrating Dust with recurrent or agentic systems that are currently hindered by backpropagation's limitations, potentially unlocking new frontiers in AI capabilities.
AI-written summary. May contain errors.