chapter nine

9 The Calculus of Decisions: Policy Gradient Methods

 

This chapter covers

  • Why optimize policies directly instead of relying on value functions
  • The theoretical foundations of policy gradient methods
  • The REINFORCE algorithm and its extension with attention mechanisms
  • Optimizing data center cooling using an attention-based REINFORCE approach
It’s better to be approximately right than exactly wrong.  

John Maynard Keynes, British Economist

Imagine you are responsible for cooling a massive data center—thousands of servers humming with computation, generating heat that must be precisely managed across dozens of interconnected cooling zones. At every moment, you must decide how fast to spin each fan and how far to open each chilled-water valve. Push the cooling too hard, and you waste enormous amounts of energy. Pull back too much, and temperatures spike, threatening equipment failure. The zones are not independent: cool air pushed into one zone leaks into its neighbors, and a sudden computational surge in one rack ripples across the entire floor.

9.1 Why optimizing the policy directly?

9.2 The mathematics of policy gradients

9.3 The REINFORCE algorithm

9.4 Variance reduction through baselines

9.5 Attention mechanism: learning what to look at

9.6 The Transformer architecture

9.7 Case study: data center cooling optimization

9.8 Summary