Deep learning has been built around backprop, the only credit assignment algorithm capable of training modern neural nets, including transformer-based language models. Backprop requires differentiability and produces first-order gradients, and deep learning’s architectures, optimizers, and hardware have co-evolved around this constraint.

However, as the amount of compute available in the world increases, we might prefer more generic and brute-force learning algorithms based on search over inductive biases like differentiability, backprop, and approximations of higher-order gradients. The bitter lesson (Sutton, 2019Richard S. Sutton. The bitter lesson. http://www.incompleteideas.net/IncIdeas/BitterLesson.html, 2019. Blog post.) is that general methods that scale with compute eventually win, and AlphaGo Zero (Silver et al., 2017David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis. Mastering the game of Go without human knowledge. Nature, 550 (7676): 354–359, 2017. doi: 10.1038/nature24270.) is the obvious example. Bootstrapping AlphaGo on human data helped the network learn faster initially, but with a lot of computation the purely self-play network overtook it. Similarly, differentiability and backprop might be good inductive biases in the low-compute regime, where they make learning efficient, but in the high-compute regime they limit the space of architectures that work. Even within an architecture, gradient-based methods fail to explore the loss landscape optimally (Liu et al., 2020Shengchao Liu, Dimitris Papailiopoulos, and Dimitris Achlioptas. Bad global minima exist and SGD can reach them. In Advances in Neural Information Processing Systems, volume 33, 2020.). This might also explain why current neural nets require massive amounts of data to generalize. A more flexible credit assignment algorithm based on search is likely an important step towards much better generalization.