Dust: Pretraining Transformers Without Backpropagation
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get school and study supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Q Labs Research describes Dust, a zeroth-order method that perturbs transformer activations to estimate updates without backpropagation. The researchers report competitive results in small-scale pretraining experiments, while saying that matching backprop more closely requires larger populations and substantially more compute. The findings are from the authors’ report; large-scale language-model performance and practical costs remain unproven.

Q Labs Research has reported a method called Dust for training transformer language models without backpropagation, the standard technique for calculating how model parameters should change. The October 2026 report says Dust produced competitive results in the researchers’ pretraining tests by perturbing activations and estimating updates from the resulting changes in loss. The claim is significant as a research result, but the evidence presented is not proof that Dust can replace backprop in large-scale production training.

Dust is a zeroth-order optimization method: rather than calculating analytic gradients through the network, it perturbs intermediate activations and uses the loss response to estimate which changes may improve performance. The authors describe this as a way to search without the backward pass required by backpropagation.

The report’s central efficiency idea is a virtual population. Dust perturbs activations independently at each token, treating tokens as separate members of a population while processing them together in a forward pass. This differs from evolution-strategy methods that perturb model weights and must evaluate separate candidate models. Q Labs says Dust is much more efficient than a transformer implementation of EGGROLL in its estimates at training scales from one million tokens upward; that comparison is an author-reported extrapolation, not a measured result across large production runs.

The researchers say Dust’s gradient estimates align more closely with backprop as the population grows and remain well aligned in tests up to one billion tokens. They also report that a 243-million-parameter model outperformed a model 120 times smaller at most tested population sizes. These findings concern the tested setups and metrics; the supplied report does not establish that the method improves general language-model quality or beats backprop under matched real-world compute budgets.

At a glance
reportWhen: Report dated October 2026; the claims d…
The developmentQ Labs Research has published a report describing Dust, an activation-perturbation method for pretraining transformer language models without a backward pass.

A Different Route to Transformer Updates

If the approach scales beyond the reported experiments, Dust could broaden the methods researchers can use to train neural networks. It asks whether more parallel search can compensate for giving up backprop’s direct gradient calculation, potentially making training less dependent on differentiability and the hardware and architecture choices built around it.

That possibility is conditional. The report says Dust approximates backprop more closely at larger populations, which also means substantially more computation. The efficiency comparisons against weight-space evolution strategies do not show that Dust is cheaper or more effective than backprop at the same training scale. For readers and practitioners, the main near-term importance is a research result that challenges assumptions about zeroth-order methods—not evidence that existing training systems should be replaced.

Amazon

transformer model training hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why the Backward Pass Matters

Backpropagation calculates how a change in each parameter affects a model’s loss by propagating information backward through the network. It has been central to modern deep learning because it provides useful training signals without testing a large number of independent alternatives. Evolution strategies, by comparison, use perturbations and observed outcomes to guide updates, but weight-space searches can require many costly model evaluations.

Q Labs’ report positions Dust between those approaches: it uses perturbations, but applies them to activations at token level so candidate variations can be evaluated in parallel during a forward pass. The paper argues that zeroth-order methods may become more attractive when compute is abundant. That is the authors’ motivation, not a finding that brute-force search will eventually outperform gradient-based training. The supplied material describes a research report and experiments, not independent replication or a peer-reviewed assessment.

“We present the first zeroth-order method that is competitive with backprop at pretraining transformer language models.”

— Q Labs Research, report summary

Amazon

GPU for neural network training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Scale, Cost and Independent Checks

The supplied report summary does not provide enough detail to establish how Dust compares with backprop under matched compute, data and training budgets, or whether the reported performance holds across different model architectures and evaluation tasks. The phrase “competitive” is the authors’ description; the material provided does not specify a broad benchmark suite or independent replication.

It is also unclear how the method behaves in larger training runs, whether its update estimates remain useful as models and datasets grow, and what engineering costs arise from implementing it. The claimed efficiency advantage over EGGROLL relies on extrapolations for the stated token range. The report’s suggestion that Dust might surpass backprop in compute-rich settings is a possibility raised by the researchers, not a demonstrated outcome.

Amazon

machine learning optimization tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Replication and Larger Training Runs

The next useful evidence would be independent replication and larger, directly measured comparisons against backprop and weight-space evolution strategies. Those tests would need to report model quality alongside total compute, training time and data use, so readers can distinguish better optimization from simply using a larger population and more resources.

Q Labs’ report identifies scaling as an encouraging direction but the supplied material does not name a planned release date, follow-up experiment or external evaluation. Until those details emerge, Dust is best understood as a proposed training method with promising results in the authors’ tests and substantial open questions about cost and generalization.

Amazon

AI research development kit

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Dust?

Dust is a zeroth-order method that perturbs a transformer’s activations and uses changes in loss to estimate training updates, rather than calculating gradients with backpropagation.

Does Dust eliminate backpropagation?

The method is designed to train without a backward pass, and Q Labs reports experiments using that approach. The report does not establish that Dust can replace backprop in large-scale or production training.

How does Dust use tokens as a population?

Dust perturbs activations independently at each token. The authors treat each token as a virtual population member and evaluate these perturbations together in a forward pass.

Is Dust more efficient than backpropagation?

The supplied report claims efficiency gains compared with an implementation of EGGROLL in extrapolations, but does not show that Dust is more efficient than backprop under a matched compute budget. Larger populations also require substantially more compute.

Has Dust been independently validated?

The source material is Q Labs Research’s own report. It does not provide evidence of independent replication or a peer-reviewed assessment, so the findings should be treated as author-reported results.

Source: hn

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Stop Anthropomorphizing Intermediate Tokens As Reasoning/Thinking Traces (2025)

Researchers caution against attributing human-like reasoning to intermediate tokens in AI models, emphasizing the need for clearer interpretation methods.

Corvus ISR Publishes Synthetic Benchmark Showing Tracker Improvements

AIThis post was created with the assistance of artificial intelligence (AI).The published…

A War Room for Your Next Idea: Inside IdeaClyst

Discover how IdeaClyst transforms idea validation with a local-first, AI-driven war room. Learn how to make smarter decisions faster today.

A Design Space Exploration Of Async/Await

A recent trend analysis examines the design options and implications of async/await in programming, amid rising interest and ongoing debates.