chapter seven
7 The pivot to reasoning
This chapter covers
- Variational lossy autoencoders
- Relation networks on CLEVR and bAbI
- Message passing neural networks on QM9
- Relational memory core for sequential reasoning across benchmarks
- Paper versus living doubts
- Variational Lossy Autoencoder (2016) Chen et al.
- A Simple Neural Network Module for Relational Reasoning (2017) Santoro et al.
- Neural Message Passing for Quantum Chemistry (2017) Gilmer et al.
- Relational Recurrent Neural Networks (2018) Santoro et al.
Rather than scaling to GPT-5 with trillions of parameters, OpenAI invested in techniques such as “test-time compute,” which involves models “thinking” more during inference [1]. This pivot reflected a broader realization across the field that scaling pretraining alone was yielding diminishing returns, particularly on tasks requiring reasoning. As a matter of fact, even the loudest champions of scale began discussing its limits. In 2024, Ilya said that “the 2010s were the age of scaling; now we’re back in the age of wonder and discovery once again.” Sutskever added, “Scaling the right thing matters more now than ever” [2].