I could buy that this method scales better than previous zeroth-order methods, and that's interesting, but it doesn't seem like enough of a moat to keep improved first-order methods from drinking its milkshake, except in cases where a zeroth-order method is already a primary option: the network needs to call a simulator that doesn't expose gradient-like information. (In cases where gradients don't exist, I'd still argue for other options, e.g., Clarke-generalized gradients where applicable, so long as those can be computed with the available information. I know this technology has been published for automatic differentiation, so I would imagine it could be incorporated into backprop and used with a suitable optimization algorithm.)
Though, I suppose RL has non-gradient based methods too.
As a mixture, could activation-space search produce useful teaching targets for backprop?
Zeroth-order search would discover candidates, first-order learning would consolidate them. The potentially valuable step is converting a sparse judgment into a reusable training target.
This also changes the relevance of convexity.
Both algorithms are bound by the same Pareto frontier based on the Empirical Risk Minimisation Principle, so they’re already on the same trajectory. Interestingly backprop is limited by conditioning of the Hessian matrix in order to converge (differentiate correctly). So removing this limitation is actually a great step. I’m excited to see a comeback of evolutionary methods because they’re much more general, albeit costly and naive. We’re now very close to what can be described best as brute forcing the Pareto frontier out of our datasets. Not sure that’s what we want but I have no better ideas either.
What does this mean?
Something like Dust skips the backward pass on backprop. But other techniques like Neural Predictive Coding can be completely asynchronous, each "weight" can fire independent of those far away from it. Innocenti, et. al have shown that NPC gradients converge to backprop within a certain "regime".
The win with asynchronous techniques like NPC is that you do not need the extreme co-ordination that backprop requires and hence should be computationally much easier given the right device.
Although at this point the industry has so much money in the forward-backward pass system that I doubt a backprop successor would win unless someone makes NPC hardware feasible and can prove scaling up to billions of params
That idea was taken further by N'dri et al in PCL, in which "activation energy" was minimized as well, and inhibitory neurons added https://www.nature.com/articles/s41467-025-64234-z.pdf
While trying to find the link for that I stumbled upon
https://arxiv.org/pdf/2605.12732
Which also looks pretty interesting
If we can figure out how the brain's learning dynamics function well enough? We could figure out how to interface with them and extend them.
"Much easier given the right device" - the "right" there just isn't shaped like the devices we actually build. And the price of "not having backprop" is usually expending more FLOPs, getting worse sample efficiency, etc.
The biggest "device" that doesn't do backprop is the brain, and that's because the brain doesn't have the connectivity or the coordination to pull it off. Both of those are "expensive" for something like it to implement. Cheap for us though. We aren't stuck with neurons that only get locally available information and have to implement learning rules based on that. So, skill issue?
The trick about comparing the two is that different things are expensive to different substrates.
Coordination is cheap for GPGPU and expensive for brain. When you have a fixed number of reusable general purpose computational units, coordinating execution is more natural than not coordinating execution, and the power cost is nil. When your computational units are independent, purpose specific, and fully embedded into the data path, coordinating them can get less natural and, frankly, optional. When wiring is expensive, coordination can become expensive in turn.
Another thing in the same "cheap for GPGPU but expensive for brain" regime is bandwidth. Look no further than optic nerve to see just how hard it is for nerves to push any appreciable amount of data. Another thing is connectivity. For GPGPU, global connectivity is natural - but the brain has to pay in physical wires for all the connectivity it has, and, see "bandwidth": it doesn't have any good wires. Yet another thing is weight reuse: a big part of why humans get "handedness" is that the brain can't just reuse the motion control circuitry for one hand for another nearly identical hand.
And the final thing I can name off the top of my head is memory - but specifically, memory capable of fast R/W. The capacity of human "working memory" is a disgrace, and not because there was no use for more. Humans rapidly lose visual fidelity of representations for objects they aren't directly looking at, and not at all because "being able to check how things looked 2 seconds ago" is useless. Those capabilities were just too expensive for the substrate to afford them easily.
It's why brains, broadly, favor dataflow-like and SSM-like dynamics, with largely fixed asynchronous dataflows and recurrence over updated local information - instead of something that would require a lot of global connectivity and transformer-like many-to-many attention ops. SSM is not necessarily the "best" tool for the job in ML land, for most jobs - but when you struggle to fit "attention" into your connectivity/bandwidth budget, and your memory is extremely expensive but hard-coupled to processing, SSM starts looking very appealing.
Now, something that might be expensive for GPGPU but cheap for brain, for once? Online learning. Maybe it's substrate dependent, or maybe it's going to get cheap in GPU land too once we figure out the trick. But so far? No one figured out how to make it cheap, stable and usable. You'd be lucky to get "pick one".
If we could have success with spiking neural networks in silico they would take even less energy, because they don't require global co-ordination. Co-ordination is information and "information = energy by the second law of thermodynamics" is my crank proof
Also the brain has way more parameters than LLMs and also has different neurotransmitters, loops, branching etc so they probably have WAY more capacity than LLMs.
But coding output/W LLMs have us beat
I frankly don't believe in spiking neural networks giving any advantages over what we have. It's a different way to implement ANNs, but "different" isn't "better". It's how the brain does things, sure, but the answer to "why the brain does what it does" is "workarounds for being made of flesh issues" at least half the time.
I can believe in brain having more capacity than frontier LLMs quite easily. We know a single BNN neuron can have the expressiveness of many ANN neurons. And well leveraged overparametrization + compute overhang could explain a decent chunk of the apparent sample efficiency edge.
But that apparent "extra capacity" could also be tied up in things like neurons having to contend with metabolism, in brain's learning algorithms being noisy, in brain having to use neuron circuits to implement "hot memory", etc - instead of contributing only to performance.
Then A waits for B to compute its backwards pass, which is waiting for C to do the same thing. Again you are sending around potentially gigabytes of data.
This is in contrast to mining bitcoins for example which doesn’t require any coordination from miners because their work is completely independent, and the answer is very small compared to the work needed to get it.
But the cool thing is that if your NN is split into mostly self contained chunks then you can go widthwise parallel.
An architecture like MOE exploits this fact so that the active weights during pre-training you're backproping only through active experts
If something more bio inspired ie. predictive coding and in-memory compute fundamentally makes continual learning and much lower energy consumption possible there will be specialised hardware for it at some point
FWIW I think the brain has multiple “learning rules” and operates at multiple timescales
DUST does have an advantage specifically along those lines because it doesn't have to save a ton of intermediate state other than each layer's input activations during a single forward pass.
There are many other issues that this algorithm does not address thoigh like catastrophic forgetting. it's still operating on a transformer which contains no inherent mechanism for selecting the relative value of a training step based on current knowledge, nor does it have segmentation of functionalities with specialized areas used for specific things that can be sequestered off and ignore new updates (we do not risk forgetting how to walk as we increase our French vocabulary)
Genuine question due to unfamiliarity with the subject.
But I guess the industry is littered with techniques for computing the same thing but vastly slower that some people find interesting. Homomorphic encryption. Zero knowledge proofs. Blockchain computing. Except in those cases there might be a legitimate reason to use it occasionally.
- isnt the whole point of back propoagation so that you dont guess weights by brute forcing them since that is computationally infeasible once you go beyond a dozen weights?
- if you dont use backpropogation, how exactly are the initial weights assigned them if they are not random values?
Derivative-free optimization can be useful for genuinely discontinuous objectives [1], but common neural network objectives are smooth and/or Lipschitz.
The gradient is useful. Instead of trying random directions and hoping that one of them is an improvement, it tells you where to go. The more parameters you have, the more useful it becomes.
A strict complexity gap between gradient-based and derivative-free Lipschitz convex optimization has been suspected for decades and recently proved (using AI, [2]). Neural net optimization is nonconvex, but not radically different.
IMO, a more promising direction is gradient-based optimizers specialized to the neural network structure, like Muon [3].
[1] https://arxiv.org/abs/2202.00817
[2] https://arxiv.org/abs/2607.13335
[3] https://jeremybernste.in/writing/deriving-muon
Hell, animals are an even better example. Many animals pop out of the womb and start walking and eating and acting just like an adult!
A fun but silly exercise.
It's much harder to argue with the math and empirical results.
For discontinuous objectives, I know there's been work on using envelope approximations, but the little I'm aware of in that work was in low-dimensional settings where the structure of the discontinuity was known explicitly. On the other extreme, lack of continuity comes up all the time in infinite-dimensional, PDE-constrained optimization, and some methods rely on tangent cones or various generalized notions of subdifferentiability (e.g., Mordukhovich, Bouligand) to demonstrate convergence. Admittedly, that work was somewhat outside my area of expertise, so I may be getting the details there slightly wrong, but the broad point stands that even in those settings, some directional information can be obtained and used profitably without resorting to zeroth-order methods.
Gradients can be calculated numerically, meaning that any method that samples the cost function and makes optimization decisions based on that can actually compute gradients if it needs to.
https://arxiv.org/pdf/2603.12228
I wonder though if someone tested mutating training objective though, like keeping original loss/goal and somehow defining loss differently and then comparing against original. This intuitively feels like how mind tries to handle difficult tasks.
There is even an analog to the continuous derivative for discrete binary functions, called "Boolean variation": https://proceedings.neurips.cc/paper_files/paper/2024/hash/7...
Like for derivatives, there is a chain rule for Boolean variations, so you can use something like backpropagation, but without needing any expensive floating point math. Though I don't think this has been used much so far. There must be some other downside.
The disparity was so large that I was certain I must have made a mistake and I spent a few hours debugging, and then a few hours more trying different neural architectures.
Turns out this is just a super common experience for anyone in NNs who would also try the more established learning algorithms.