1 Introduction
1.1 Why entropy?
Modern machine learning has long been nurtured by scientific disciplines. A canonical example is the Boltzmann machine, rooted in statistical physics and motivated by the Maximum Entropy Principle. This book takes entropy — and its close relative, energy — as the organizing thread for a family of learning methods that are both principled and practically effective.
That lineage is unusually direct. In 1982 John Hopfield borrowed the energy function of a magnetic spin system to build an associative memory (Hopfield 1982); within a year, simulated annealing had turned the competition between energy and entropy into an optimization schedule (Kirkpatrick et al. 1983); and by 1985 the Boltzmann machine placed Hopfield’s network at a finite temperature and learned its couplings from data (Ackley et al. 1985). The same trade-off resurfaced as the evidence lower bound (Dayan et al. 1995; Neal and Hinton 1998), as the contrastive objectives of energy-based learning (LeCun et al. 2006), as the temperature parameter of every language-model decoder, and as the KL penalty with which language models are fine-tuned from human feedback (Ziegler et al. 2019; Ouyang et al. 2022). In 2024 the Nobel Prize in Physics was awarded to Hopfield and Hinton for this line of work (The Royal Swedish Academy of Sciences 2024). Chapter 3 develops the single equation, \(F = U - TS\), that runs beneath all of it.
We study three directions that are connected:
- Maximizing entropy to obtain the least-biased model consistent with given constraints. This yields exponential-family / Gibbs distributions and, as we will see, a principled framework for valuation problems in machine learning, in which the energy is given and nothing is learned (Bian et al. 2022). When the energy is itself learned rather than given, the same principle becomes a learning method on discrete domains — neural set functions trained from their optimal subsets alone (Ou et al. 2022; Xie et al. 2024) — and, on trajectories, the inverse reinforcement learning that infers a reward from demonstrations (Ziebart et al. 2008). With a pretrained model as the base measure, the resulting Gibbs distribution is the common target of energy-based guidance at inference time and of post-training, which is how RLHF, GRPO, and DPO enter this book.
- Minimizing entropy to sharpen a model’s own predictions. When applied to large language models in a latent semantic space, this idea enables fully unsupervised elicitation of reasoning capabilities, as in EMPO (Zhang et al. 2025). The idea has two limits, and both are structural rather than defects of an implementation: it elicits capability but cannot add it, since its training loop never touches ground truth (Section 6.7), and, left unchecked, it collapses the model onto one confident answer through a first-order transition in a mean-field free energy (Section 6.6).
- Minimaxing entropy to decide which features a model should include at all, and what it should commit to. The optimal features are those whose maximum-entropy model has the minimum entropy, i.e., gives the shortest description of the data (Zhu et al. 1997; Carcamo et al. 2025). The same nesting covers prediction: a softmax classifier with a linear head is a conditional maximum-entropy model, and the cross-entropy it minimizes is the largest conditional entropy its features allow (Section 7.2.2). Without labels, a temperature on the label moves the problem from maximum-entropy learning with missing data, at \(T = 1\), to the classification likelihood, as \(T \to 0\) (Section 7.3).
What connects the three is the free energy. Read at the level of algorithms, they share one inner problem, namely the minimization of a free energy over distributions, which Chapter 5 calls the representation and inference levels, and they differ only in the outer criterion: i) fit the data, ii) sharpen the model’s own output, or iii) choose what to measure and what to commit to. In the terms of Chapter 3, they are three settings of one temperature dial.
1.2 Roadmap
- Part I — Foundations surveys the thermodynamic, information-theoretic, quantum, and psychological notions of entropy, states the axioms that single out these formulas and loosens them one at a time, carries the same method over to the diversity of a set, and develops the maximum-entropy principle together with the minimum-description-length view that leads to minimax entropy (Chapter 2). It then derives the free-energy relationship \(F = U - TS\) that connects entropy to energy (Chapter 3), and introduces energy-based models together with three classical ways of training them without computing the partition function (Chapter 4).
- Part II — Methods develops maximum-entropy learning, read through the three levels of representation, inference, and learning (Chapter 5), minimum-entropy learning (Chapter 6), and minimax-entropy learning (Chapter 7). It then applies the maximum-entropy machinery to pretrained models in two ways: energy-based guidance and refinement, which sample from a Gibbs tilt of a pretrained model at inference time, with the model itself frozen (Chapter 8), and maximum-entropy post-training, which reads RLHF, GRPO, DPO, and their relatives as the same tilt and pays for it once in training rather than at every query (Chapter 9).
- Part III — Applications connects these ideas to scientific foundation models and AI-for-science (Chapter 10).