11 Summary
This book introduced entropy-based learning through three complementary principles and their shared energy-based substrate:
- Maximum entropy yields the least-biased model consistent with given constraints — grounding exponential families, Boltzmann machines, maximum-entropy reinforcement learning, and a principled family of game-theoretic valuations (Bian et al. 2022). When the energy is learned rather than given, the same principle trains neural set functions from optimal subsets (Ou et al. 2022; Xie et al. 2024) and infers rewards from demonstrations (Ziebart et al. 2008). Applied to a pretrained model as base measure, it becomes energy-based guidance at inference time and maximum-entropy post-training — RLHF, GRPO, DPO — at training time, one Gibbs tilt paid for in two currencies. Read through the three levels of representation, inference, and learning, all of these methods belong to one family.
- Minimum entropy sharpens a capable model’s own predictions, enabling unsupervised elicitation of reasoning (EMPO) in a latent semantic space (Zhang et al. 2025). It elicits but does not add capability, since the training loop never touches ground truth, and its collapse onto one confident answer is a first-order transition in a mean-field free energy, which is why the durable remedies hold the entropy above a floor rather than tune a coefficient.
- Minimax entropy couples the two: maximize entropy to model the world, minimize entropy to commit to what matters. Its first question is what to measure, and the answer is the features whose maximum-entropy model carries the least entropy, simultaneously the shortest, most accurate, and most informative description (Zhu et al. 1997; Carcamo et al. 2025). Its second question is what to commit to: a softmax classifier with a linear head is a conditional maximum-entropy model whose cross-entropy is the largest conditional entropy its features allow, and without labels a temperature on the label moves the problem from maximum-entropy learning with missing data to the classification likelihood.
Binding the three together is the free energy \(F = U - TS\): temperature sets the exchange rate between energy and entropy, so the principles are three settings of one dial rather than three separate doctrines. Read at the level of algorithms, they share one inner problem, the minimization of a free energy over distributions, and differ only in the outer criterion. It is also the equation that carried statistical physics into machine learning, from the Hopfield network and the Boltzmann machine through to the evidence lower bound — a lineage recognized by the 2024 Nobel Prize in Physics (The Royal Swedish Academy of Sciences 2024).
This is an early scaffold; future revisions will add worked examples, code, and experiments. For updates, see https://yataobian.com/ and the Blue Whale Lab.