2 Entropy and the Maximum Entropy Principle
2.1 The many faces of entropy
Entropy was invented three separate times — for heat engines in the 1860s, for quantum ensembles in the 1920s, and for telephone lines in the 1940s — and has been rediscovered many times since. The names, the units, and the motivating problems all differ, but each version measures the same thing: given what we have specified about a system, how much remains unspecified. This section walks through those notions before the rest of the book settles on one of them.
2.1.1 Information: Shannon entropy
For a discrete random variable \(X\) with probability mass function \(p(x)\), the Shannon entropy is
\[ H(p) = -\sum_{x} p(x) \log p(x), \]
which quantifies the average uncertainty (in nats or bits) of \(X\) (Shannon 1948). Entropy is maximized by the uniform distribution, where it equals \(\log |\mathcal{X}|\), and minimized (equal to zero) by a deterministic one.
This functional form is not an arbitrary choice. Shannon showed that three requirements — continuity in the probabilities, monotonicity in the number of equally likely outcomes, and consistency when outcomes are grouped — determine \(H\) uniquely up to the base of the logarithm; the requirements, their reformulation as the Shannon–Khinchin axioms, and the uniqueness theorem are stated in Section 2.2. It also carries a concrete operational meaning: \(H(p)\) is the expected code length of an optimal lossless encoding of samples from \(p\) — exactly, in the limit of coding long blocks, and to within one bit if each symbol must be coded on its own. Uncertainty and description length are one quantity seen from two sides, a duality that returns in Chapter 7.
2.1.2 Heat: thermodynamic entropy
The original entropy is older and, on its face, unrelated to information. Studying the irreversibility of heat flow, Clausius defined a state function whose differential along a reversible path is
\[ dS = \frac{\delta Q_{\text{rev}}}{T}, \]
and named it from the Greek τροπή, transformation, deliberately shaping the word to rhyme with energy because the two quantities seemed to him so closely related (Clausius 1865). That 1865 paper closes with the two sentences that became the standard statement of the laws of thermodynamics: the energy of the universe is constant, and its entropy tends towards a maximum.
Boltzmann supplied the microscopic reading. A macrostate — a temperature, a pressure — is compatible with an enormous number \(W\) of microstates, and entropy simply counts them (Boltzmann 1877):
\[ S = k_B \log W . \]
Gibbs then generalized the count to microstates that are not equally likely (Gibbs 1902):
\[ S = -k_B \sum_i p_i \log p_i . \]
The Gibbs expression and the Shannon expression are the same formula, separated only by the constant \(k_B\) that carries the physical units. The coincidence is not accidental, and Jaynes turned it into the argument developed in Section 2.3: thermodynamic entropy is the Shannon entropy of our ignorance about the microstate, which makes equilibrium statistical mechanics a problem of inference rather than of mechanics (Jaynes 1957).
2.1.3 Quantum states: von Neumann entropy
Quantum mechanics replaces a probability distribution with a density matrix \(\rho\), a positive semidefinite operator of unit trace. Von Neumann extended entropy to this setting in 1927, two decades before Shannon (Neumann 1927):
\[ S(\rho) = -\operatorname{Tr}(\rho \log \rho). \]
Written in the eigenbasis of \(\rho\), this reduces to the Shannon entropy of its eigenvalues, so the same functional form reappears with a different argument. It vanishes exactly for pure states, which are completely specified, and is maximal for the maximally mixed state.
The quantum case does break the analogy in one instructive place. Classically, conditional entropy is never negative: learning \(Y\) cannot leave us knowing more than everything about \(X\). Quantum mechanically it can be — the joint state may be more sharply specified than either of its parts. A negative value therefore certifies entanglement, though the implication runs one way only: entangled states with non-negative conditional entropy exist.
2.1.4 One knob, a family: Rényi and Tsallis
Shannon’s requirements can be relaxed, and each relaxation produces a one-parameter family. Rényi kept additivity over independent systems but dropped the requirement that entropy be a linear average of \(-\log p\) (Rényi 1961):
\[ H_\alpha(p) = \frac{1}{1-\alpha} \log \sum_x p(x)^\alpha , \qquad \alpha > 0, \; \alpha \neq 1 . \]
The order \(\alpha\) is a knob controlling how much weight rare outcomes receive. The limit \(\alpha \to 1\) recovers Shannon entropy; \(\alpha = 0\) gives the Hartley entropy \(\log|\operatorname{supp} p|\), which counts possibilities and ignores their probabilities; \(\alpha = 2\) gives the collision entropy
\[ H_2(p) = -\log \sum_x p(x)^2 , \]
whose argument \(\sum_x p(x)^2\) is the probability that two independent draws coincide — a quantity that returns as a training objective in Section 6.5; and \(\alpha \to \infty\) gives the min-entropy
\[ H_\infty(p) = -\log \max_x p(x), \]
the worst-case measure used in cryptography, where what matters is not average uncertainty but the chance that an adversary guesses right on the first try.
A warning about that last name, since both readings appear later in this book. The min-entropy \(H_\infty\) is a specific functional, the most pessimistic member of the Rényi family. It should not be confused with the minimum-entropy principle of Chapter 6, which is not a functional at all but a class of learning objectives that reduce the entropy — usually the Shannon or the collision entropy — of a model’s predictive distribution. The two are unrelated, and the shared prefix is an accident of terminology in two literatures that developed apart.
Tsallis relaxed the other requirement, replacing the chain rule by a \(q\)-deformed version (Section 2.2), so that additivity itself is lost (Tsallis 1988):
\[ S_q(p) = \frac{1}{q-1}\Big(1 - \sum_x p(x)^q\Big). \]
Here \(q \to 1\) again recovers the Boltzmann–Gibbs form, but for \(q \neq 1\) two independent systems combine as \(S_q(A,B) = S_q(A) + S_q(B) + (1-q)\,S_q(A)\,S_q(B)\). This non-extensivity was introduced to describe systems with long-range interactions and multifractal structure, where the usual assumption that entropy scales with system size fails.
Both families are useful in their own domains, and both make the same point by contrast: Shannon entropy is the unique member satisfying the full set of Shannon–Khinchin axioms, including the chain rule that makes conditional entropy and mutual information well behaved, as Section 2.2 spells out. That is why it is the form used throughout the remainder of this book.
2.1.5 Minds: psychological entropy
The most surprising reappearance is in psychology. Hirsh, Mar, and Peterson propose the entropy model of uncertainty, which applies Shannon’s formula directly to the human information system (Hirsh et al. 2012). At any moment an organism faces a set of competing affordances: possible interpretations of the sensory input, and possible actions to take. The model treats these as a subjectively weighted probability distribution and defines psychological entropy as its Shannon entropy.
The reading is unusually literal. A familiar situation places nearly all the weight on one dominant affordance, giving low entropy and a fast habitual response. An unfamiliar or ambiguous one spreads the weight across many competing frames, giving high entropy and sustained neural competition — which, the authors argue, is what anxiety is subjectively, with anterior cingulate activity and noradrenaline release as its physiological correlates.
Two features of the account matter for what follows. First, goals act as constraints: adopting a goal reweights the distribution towards the actions that serve it, and thereby lowers entropy. This is structurally the same move as imposing moment constraints in Section 2.3, since a constraint is whatever narrows the space of live possibilities. Second, the model makes entropy reduction a drive rather than a calculation. That idea runs from Schrödinger’s argument that organisms persist by exporting entropy to their environment (Schrödinger 1944) through to the free-energy principle (Friston 2010), and it is the biological precedent for the minimum-entropy learning objectives of Chapter 6.
2.1.6 One form, many readings
| Notion | Probabilities over | Definition | Origin |
|---|---|---|---|
| Thermodynamic | no explicit distribution | \(dS = \delta Q_{\text{rev}} / T\) | Clausius, 1865 |
| Statistical | microstates, equally likely | \(k_B \log W\) | Boltzmann, 1877 |
| Gibbs | microstates, unequally likely | \(-k_B \sum_i p_i \log p_i\) | Gibbs, 1902 |
| Information | a random variable or message source | \(-\sum_x p(x) \log p(x)\) | Shannon, 1948 |
| Quantum | eigenvalues of a density matrix | \(-\operatorname{Tr}(\rho \log \rho)\) | von Neumann, 1927 |
| Rényi | any distribution, at order \(\alpha\) | \(\frac{1}{1-\alpha} \log \sum_x p(x)^\alpha\) | Rényi, 1961 |
| Tsallis | non-extensive systems, at order \(q\) | \(\frac{1}{q-1}(1 - \sum_x p(x)^q)\) | Tsallis, 1988 |
| Psychological | competing perceptions and actions | \(-\sum_x p(x) \log p(x)\) | Hirsh et al., 2012 |
Two observations follow. The first is mathematical: almost every row is the same functional \(-\sum p \log p\), differing only in what the probabilities range over and what constant multiplies the result. Entropy is less a family of analogies than a single quantity applied to different objects.
The second is interpretive, and it is the one that does work in this book. The same number reads as dispersal of energy to a physicist, as expected code length to an engineer, and as felt uncertainty to a psychologist. Those readings are what make the two principles that follow feel necessary rather than arbitrary: maximizing entropy is being honest about what we have not specified, and minimizing it is committing to what we have.
2.2 Why these formulas and no others: the axioms
In two places so far, this chapter has referenced the uniqueness of Shannon entropy without stating the basis for that claim: first, in attributing the form of \(H\) to Shannon’s trio of requirements (which specify \(H\) up to the logarithm’s base), and again at the end of Section 2.1.4, which asserts that Shannon entropy alone satisfies the full set of Shannon–Khinchin axioms. In this section we state the axioms, give the uniqueness theorem in its sharpest known form, and then loosen one axiom at a time; \(H_\alpha\) and \(S_q\) are as defined in Section 2.1.4.
The same method is then applied to a second quantity, which this book needs from Section 5.4 onward: the diversity of a set of objects, with or without a notion of similarity between them. Diversity has its own axiomatic literature, in ecology, economics and recently machine learning, and it is related to entropy by an exponential; we give the correspondence axiom by axiom and record what the axiomatic method cannot decide.
2.2.1 The load-bearing axiom
Let us start with a numerical example. Consider a distribution on five outcomes,
\[ p = \Big(\tfrac14, \tfrac14, \tfrac16, \tfrac16, \tfrac16\Big), \]
and think of drawing from it in two stages: first a coin toss \(w = (\tfrac12, \tfrac12)\) decides between a left block and a right block, then within the left block one draws from \(p_1 = (\tfrac12, \tfrac12)\) and within the right block from \(p_2 = (\tfrac13, \tfrac13, \tfrac13)\). Written as a composite, \(p = w \circ (p_1, p_2)\). Computed directly, the Shannon entropy of \(p\) is \(\tfrac12 \log 4 + \tfrac12 \log 6 = \tfrac32 \log 2 + \tfrac12 \log 3\), about \(2.29\) bits. The two-stage computation gives \(H(w) + \tfrac12 H(p_1) + \tfrac12 H(p_2) = \log 2 + \tfrac12 \log 2 + \tfrac12 \log 3\), which is the same number. The equality is an instance of the chain rule,
\[ H\big(w \circ (p_1, \dots, p_n)\big) = H(w) + \sum_{i=1}^{n} w_i\, H(p_i), \]
valid for every \(w \in \Delta_n\) (the simplex of distributions on \(n\) outcomes) and every choice of \(p_i \in \Delta_{k_i}\). With \(X\) the block and \(Y\) the outcome within the block, it reads \(H(X,Y) = H(X) + H(Y \mid X)\); for independent \(X\) and \(Y\) (all \(p_i\) equal to one \(p\), a composite written \(w \otimes p\)) it reduces to additivity, \(H(X,Y) = H(X) + H(Y)\), which is a strictly weaker statement.
Shannon’s requirements. Shannon’s own theorem (Shannon 1948, Theorem 2) assumes i) continuity in the \(p_i\); ii) that \(A(n) = H(1/n, \dots, 1/n)\) increases in \(n\); iii) consistency under grouping, which is the two-stage computation above, and concludes that \(H = -c \sum_i p_i \log p_i\) for a constant \(c > 0\).
Khinchin’s axioms. Most later work uses Khinchin’s reformulation (Khinchin 1957) (the 1957 translation of his 1953 Uspekhi paper). Khinchin himself isolates three properties, namely maximality at the uniform distribution, the chain rule \(H(AB) = H(A) + H_A(B)\) and expansibility, and proves that any \(H\) continuous in its arguments and satisfying them is \(-\lambda \sum_i p_i \log p_i\) with \(\lambda > 0\); later authors package continuity together with these three as the four Shannon–Khinchin axioms, with labels SK1–SK4. For a sequence of functions \(H: \Delta_n \to \mathbb{R}\), \(n \ge 1\), they are:
- SK1 (continuity): \(H\) is continuous in \(p\);
- SK2 (maximality): \(H(p) \le H(u_n)\) for all \(p \in \Delta_n\), where \(u_n\) is the uniform distribution;
- SK3 (expansibility): \(H(p_1, \dots, p_n, 0) = H(p_1, \dots, p_n)\), i.e., adding an outcome of probability zero changes nothing;
- SK4 (strong additivity): \(H(X,Y) = H(X) + \sum_x p(x)\, H(Y \mid X = x)\), i.e., the chain rule.
Any sequence satisfying SK1–SK4 is \(-c \sum_i p_i \log p_i\) with \(c \ge 0\), where \(c = 0\) is the trivial \(H \equiv 0\) that Khinchin’s \(\lambda > 0\) tacitly excludes (Khinchin 1957). Note that SK4 is strictly stronger than additivity for independent systems, which is the special case in which every conditional distribution \(p(y \mid x)\) equals the marginal \(p(y)\). SK4 is the load-bearing axiom in the following sense: every generalized entropy discussed below keeps SK1–SK3 and modifies SK4. This is what the phrase “unique member satisfying the full set of Shannon–Khinchin axioms” in Section 2.1.4 means.
Faddeev’s theorem. Once SK4 is assumed, only continuity is needed from SK1–SK3, as Leinster’s variant of Faddeev’s theorem shows (Faddeev 1956; Leinster 2021, Theorem 2.5.1). Let \((I: \Delta_n \to \mathbb{R})_{n \ge 1}\) be a sequence of functions. Then the following are equivalent: i) the functions \(I\) are continuous and satisfy the chain rule \(I(w \circ (p_1, \dots, p_n)) = I(w) + \sum_i w_i I(p_i)\) for all \(n\), \(k_1, \dots, k_n \ge 1\), \(w \in \Delta_n\) and \(p_i \in \Delta_{k_i}\); ii) \(I = cH\) for some \(c \in \mathbb{R}\).
Proof sketch. Firstly, the chain rule applied to \(u_{mn} = u_m \circ (u_n, \dots, u_n)\) gives \(I(u_{mn}) = I(u_m) + I(u_n)\); together with \(I(u_{n+1}) - I(u_n) \to 0\), which follows from continuity, this gives \(I(u_n) = c \log n\). Secondly, a distribution with non-zero rational entries \(p = (k_1/N, \dots, k_n/N)\) satisfies \(p \circ (u_{k_1}, \dots, u_{k_n}) = u_N\), so the chain rule gives \(c \log N = I(p) + \sum_i p_i\, c \log k_i\), i.e., \(I(p) = c H(p)\). Lastly, continuity extends the identity to all distributions (Leinster 2021, sec. 2.5).
We add some remarks on the hypotheses (Leinster 2021, Remarks 2.5.2; Faddeev 1956). Firstly, continuity is the only regularity condition; SK3 is not assumed, it follows from the conclusion \(I = cH\) for every \(c\), and SK2 follows as soon as the sign of \(c\) is fixed, e.g., by a normalization such as \(I(u_2) = \log 2\). That Khinchin’s maximality and expansibility axioms can be dispensed with is Faddeev’s own observation: his system replaces the former by positivity of \(I\) at a single point and the latter, together with strong additivity, by a single recursion. Secondly, Faddeev’s original 1956 theorem assumed in addition that \(I\) is symmetric, and assumed only the two-block special case \(I(t w_1, (1-t) w_1, w_2, \dots, w_n) = I(w) + w_1 I(t, 1-t)\) (Faddeev split the last block, which is equivalent under symmetry), while requiring continuity only of the binary function \(t \mapsto I(t, 1-t)\); under symmetry the two forms of the chain rule are equivalent by induction, so the general form allows symmetry to be removed (if symmetry is kept, continuity can instead be weakened to measurability, which is a 1964 theorem of Lee). Lastly, some regularity condition is necessary: for an additive but nonlinear \(f: \mathbb{R} \to \mathbb{R}\) (which exists under the axiom of choice), \(p \mapsto -\sum_{i \in \operatorname{supp}(p)} p_i f(\log p_i)\) satisfies the chain rule without being a multiple of \(H\).
2.2.2 Loosening one axiom at a time
The five-outcome example makes the failure of SK4 concrete. For the collision entropy \(H_2\) of Section 2.1.4, one computes directly \(H_2(p) = -\log\big(2 \cdot \tfrac{1}{16} + 3 \cdot \tfrac{1}{36}\big) = \log \tfrac{24}{5} \approx 1.569\) nats, whereas the chain-rule prediction is \(H_2(w) + \tfrac12 H_2(p_1) + \tfrac12 H_2(p_2) = \tfrac32 \log 2 + \tfrac12 \log 3 \approx 1.589\) nats. The two disagree, so \(H_2\) violates SK4, although one can verify that it is still additive for independent \(X, Y\). Now take the Tsallis entropy \(S_2(p) = 1 - \sum_i p_i^2 = \tfrac{19}{24}\). The chain-rule prediction fails again, but the modified rule
\[ S_q\big(w \circ (p_1, \dots, p_n)\big) = S_q(w) + \sum_{i=1}^{n} w_i^{\,q}\, S_q(p_i) \]
gives \(S_2(w) + \tfrac14 S_2(p_1) + \tfrac14 S_2(p_2) = \tfrac12 + \tfrac18 + \tfrac16 = \tfrac{19}{24}\), exactly. We now describe the two modifications one at a time.
Rényi: a different mean. Rényi kept additivity on independent systems (his Postulate 4) and replaced the arithmetic mean of \(-\log p_i\) under \(p\) by a Kolmogorov–Nagumo quasi-arithmetic mean \(g^{-1}\big(\sum_i p_i\, g(-\log p_i)\big)\) with \(g\) continuous and strictly monotone (Rényi 1961). Two choices of \(g\) are compatible with additivity: the linear functions, which return Shannon entropy, and the exponential functions \(g(x) \propto 2^{(1-\alpha)x}\), which return \(H_\alpha\) (Rényi’s Theorem 2; his printed formula has the exponent \((\alpha-1)x\), but with his convention \(H(\{p\}) = \log_2(1/p)\) the generator for \(H_\alpha\) is \(2^{(1-\alpha)x}\)). Rényi himself called it an open question whether any other \(g\) is admissible for entropy, and proved the linear-or-exponential dichotomy only for the information gain \(I(Q \,\|\, P)\) (his Theorem 3); for the entropy of incomplete distributions the dichotomy was proved by Aczél and Daróczy (Aczél and Daróczy 1963). Therefore the Rényi entropies satisfy SK1–SK3 and additivity on independent systems, and fail SK4 for \(\alpha \ne 1\). Jizba and Arimitsu derive the family from SK-type axioms with a quasi-linear conditional entropy in place of SK4 (Jizba and Arimitsu 2004). Lazarev revisits Rényi’s question without assuming additivity (Lazarev 2026) (a preprint at the time of writing): under the admissibility condition that a measure never has larger entropy than a reference measure dominating it, a continuous strictly monotone \(g\) is admissible exactly when \(t \mapsto g(1/t)\) is strictly convex for increasing \(g\), or strictly concave for decreasing \(g\), and additivity on products then singles out the Rényi family.
Tsallis: a different chain rule. The functional \(S_q\) was introduced by Havrda and Charvát and by Daróczy in information theory, and independently by Tsallis in statistical physics (Havrda and Charvát 1967; Daróczy 1970; Tsallis 1988). Its defining property is the \(q\)-chain rule displayed above, in which the weights \(w_i\) of SK4 are replaced by \(w_i^q\); for independent systems, where all \(p_i\) are equal, it reduces to the pseudo-additivity \(S_q(A,B) = S_q(A) + S_q(B) + (1-q)\, S_q(A)\, S_q(B)\) stated in Section 2.1.4. The wording used there can accordingly be made more precise: rather than “giving up additivity itself”, Tsallis replaced the chain rule by its \(q\)-deformation, and non-additivity on independent systems is the consequence. Uniqueness theorems with generalized SK axioms were given by Santos and by Abe (Santos 1997; Abe 2000). The shortest statement is Leinster’s (Leinster 2021, Theorem 4.1.5): for real \(q \ne 1\) and a sequence of functions \((I: \Delta_n \to \mathbb{R})_{n \ge 1}\), the following are equivalent: i) \(I\) is symmetric and satisfies
\[ I(w \otimes p) = I(w) + \Big(\sum_{i \in \operatorname{supp}(w)} w_i^{\,q}\Big)\, I(p) \]
for all \(n, k \ge 1\), \(w \in \Delta_n\) and \(p \in \Delta_k\); ii) \(I = c S_q\) for some \(c \in \mathbb{R}\). Notice that no regularity condition appears, in contrast to the Shannon case.
Hanel–Thurner: no SK4. One can also keep SK1–SK3 only and ask what freedom remains in a trace-form entropy \(S = \sum_i g(p_i)\). Hanel and Thurner showed that the admissible \(g\) are classified by two asymptotic scaling exponents \((c, d)\), with a representative
\[ S_{c,d}(p) \propto \sum_i \Gamma\big(1 + d,\, 1 - c \log p_i\big) \]
for each class, up to constants, where \(\Gamma(\cdot,\cdot)\) is the incomplete gamma function (Hanel and Thurner 2011). Shannon entropy is the point \((1,1)\), the Tsallis family is the line \((c, 0)\), and the stretched-exponential entropies are the line \((1, d)\); the associated maximum-entropy distributions are Lambert-\(W\) exponentials. Tempesta’s composability axiom requires, for statistically independent subsystems \(A\) and \(B\), that \(S(A \cup B) = \Phi(S(A), S(B))\) with \(\Phi\) symmetric, associative and satisfying \(\Phi(x, 0) = x\), i.e., a commutative formal group law over the reals; the axiom replaces SK4 and generalizes the group entropies that Tempesta had introduced in 2011 (Tempesta 2016). Korbel’s review surveys these frameworks (Shannon–Khinchin, Tempesta’s group composability, Hanel–Thurner scaling, Shore–Johnson, Lieb–Yngvason) (Korbel 2026) (also a preprint). Once SK4 is dropped, choosing a trace-form entropy amounts, up to its asymptotic class, to choosing a point \((c, d)\), i.e., a modelling decision, to which Chapter 7 returns.
The following table records which axiom each family keeps (Khinchin 1957; Faddeev 1956; Rényi 1961; Aczél and Daróczy 1963; Daróczy 1970; Abe 2000; Leinster 2021; Hanel and Thurner 2011).
| axiom | Shannon \(H\) | Rényi \(H_\alpha\) | Tsallis \(S_q\) | Hanel–Thurner \(S_{c,d}\) |
|---|---|---|---|---|
| SK1 continuity | yes | yes | yes | yes |
| SK2 maximal at uniform | yes | yes | yes | yes |
| SK3 expansibility | yes | yes | yes | yes |
| SK4 chain rule | yes | no | no (\(q\)-chain rule) | no |
| additive on independent systems | yes | yes | no (pseudo-additive) | no in general |
| arithmetic mean of \(-\log p\) | yes | no (quasi-arithmetic) | — | — |
| uniqueness theorem | Khinchin; Faddeev | Rényi, 1961 (given a Kolmogorov–Nagumo mean); Aczél–Daróczy, 1963 | Daróczy, 1970; Abe, 2000; Leinster, Theorem 4.1.5 | Hanel–Thurner, 2011 (classification into asymptotic \((c,d)\)-classes; \(S_{c,d}\) is a representative) |
The rows SK1–SK3 hold for Rényi and Tsallis when \(\alpha, q > 0\); for negative orders the uniform distribution is not the maximizer (Leinster 2021, Remarks 4.4.4(ii)). Hanel and Thurner assume the trace form \(S = \sum_i g(p_i)\) with \(g\) continuous and concave, \(g(0) = 0\), and \(c \in (0, 1]\).
2.2.3 The other faces, axiomatized
For three readings in the table closing the preceding section, the axiomatic method gives an independent derivation.
Entropy as information loss. Return to the five-outcome distribution and merge the two outcomes of probability \(\tfrac14\) into one. The result is \(q = (\tfrac12, \tfrac16, \tfrac16, \tfrac16)\) with \(H(q) = \log 2 + \tfrac12 \log 3\), so the merge loses \(H(p) - H(q) = \tfrac12 \log 2\), exactly the term \(w_1 H(p_1)\) that the chain rule attributes to the left block. Baez, Fritz and Leinster take this loss as the object of the axioms (Baez et al. 2011). On the category FinProb, whose objects are finite probability spaces \((X, p)\) and whose morphisms \(f: (X,p) \to (Y,q)\) are measure-preserving maps, an information loss assigns a number \(F(f) \in [0,\infty)\) to each morphism and is i) functorial, \(F(g \circ f) = F(g) + F(f)\); ii) convex-linear, \(F(\lambda f \oplus (1-\lambda) g) = \lambda F(f) + (1-\lambda) F(g)\); iii) continuous. Their theorem states that every such \(F\) equals \(c\,[H(p) - H(q)]\) for a constant \(c \ge 0\); replacing the weights \(\lambda, 1-\lambda\) in convex-linearity by \(\lambda^{\beta}, (1-\lambda)^{\beta}\), \(\beta > 0\), gives \(c\,[S_\beta(p) - S_\beta(q)]\) with the Tsallis entropy \(S_\beta\) instead (Baez et al. 2011, Theorem 7). Baez and Fritz characterize relative entropy in the same style on a category FinStat (Baez and Fritz 2014).
Quantum states. Ochs gave a list of postulates for a quantity measuring the intrinsic dispersion of a quantum state, including invariance under unitaries and partial isometries, additivity on product states, subadditivity, continuity, and some technical conditions, and proved that these single out \(S(\rho) = -k \operatorname{Tr} \rho \log \rho\) up to the scale \(k\) (Ochs 1975). Section 2.2.5 applies this functional to a normalized similarity kernel.
Heat. Earlier in this chapter, the argument connecting the thermodynamic row to the rest was Jaynes’s interpretive one (Jaynes 1957); Lieb and Yngvason give an axiomatic one, which assumes nothing about heat, temperature or microscopic constituents (Lieb and Yngvason 1999). The primitive is the relation of adiabatic accessibility \(X \prec Y\) between equilibrium states, meaning that \(Y\) can be reached from \(X\) by a process whose only net effect on the surroundings is that a weight has been raised or lowered. The axioms are properties of this relation: reflexivity, transitivity, consistency under composition of systems, scaling invariance, splitting and recombination, and stability, together with a comparison hypothesis (any two states of the same system are comparable under \(\prec\)). From these they prove that there is a function \(S\) on states, unique up to affine transformations, with \(X \prec Y\) if and only if \(S(X) \le S(Y)\), additive under composition and extensive under scaling. Weilenmann, Krämer, Faist and Renner apply the framework, including its 2013 extension to non-equilibrium states, to finite-dimensional quantum systems ordered by majorization; the resulting entropy coincides with the von Neumann entropy on the equilibrium states, in this setting those with flat spectrum, and with the min- and max-entropies of one-shot information theory, \(H_{\min}(\rho) = -\log \|\rho\|_\infty\) and \(H_{\max}(\rho) = \log \operatorname{rank} \rho\), away from them (Weilenmann et al. 2016). The result gives an axiomatic identification of the Clausius and von Neumann rows of the table, where the chapter so far had only Jaynes’s interpretive argument; the thermodynamic side returns in Chapter 3.
The inference rule. Shore and Johnson axiomatize the rule that takes a prior and a set of constraints on expected values to a posterior, rather than the entropy functional. They require of such a rule, realized by minimizing a functional \(H(q, p)\), uniqueness, invariance under coordinate changes, system independence and subset independence, and prove that every rule satisfying these is equivalent to minimization of relative entropy (their “cross-entropy”, \(\int q \log (q/p)\)), which reduces to the maximum-entropy principle of Section 2.3 for a uniform prior (Shore and Johnson 1980). Chapter 5 uses this result, which axiomatizes a procedure rather than \(H\). Csiszár’s survey (Csiszár 2008) treats the characterizations of functionals, of entropic vectors and of inference rules in one place.
2.2.4 From entropy to diversity
The following example is due to Jost (Jost 2009; Jost et al. 2010), in the form given by Leinster (Leinster 2021, Example 2.4.11); Jost’s original has 20 islands with Barro Colorado abundance data, and the oil company, the lawyers and the figures below are Leinster’s re-casting. An oil company plans work on 16 equally sized islands which will destroy all wildlife on 8 of them. The islands share no species and each has diversity 4, in the sense of the effective number defined below, so the group has diversity \(16 \times 4 = 64\) before the work and \(32\) after, a loss of 50%. With Shannon entropy as the index, as ecologists have long used it, the company’s lawyers compute \(\log 64\) before and \(\log 32\) after, hence \(\log 32 / \log 64 = 5/6 \approx 83\%\) of the diversity preserved; the environmentalists’ lawyers compute the entropy of the islands to be destroyed, \(\log 32\) out of \(\log 64\), and claim that 83% is lost. The two conclusions contradict each other, and both are wrong, since half of the diversity is lost; the reason is that \(H\) fails the replication principle stated below.
Effective numbers. The remedy, argued by Jost under the slogan that entropies are not diversities (Jost 2006), is to exponentiate. For an abundance distribution \(p\) on \(n\) species and the Rényi entropy \(H_q\) of Section 2.1.4, the Hill number of order \(q\) is
\[ D_q(p) = \Big(\sum_{i=1}^{n} p_i^{\,q}\Big)^{1/(1-q)} = \exp H_q(p), \]
introduced into ecology by Hill and, for \(q = 1\), by MacArthur (Hill 1973; MacArthur 1965). \(D_q(p)\) is the effective number of species: a community with \(D_q = m\) counts as diverse as one of \(m\) equally abundant, completely distinct species. Specifically, \(D_0\) is species richness (the Hartley count of possibilities of Section 2.1.4), \(D_1 = e^{H}\), \(D_2 = 1/\sum_i p_i^2\) is the inverse of Simpson’s concentration (Simpson 1949), and \(D_\infty = 1/\max_i p_i\). Patil and Taillie read a diversity index as an average rarity \(\sum_i p_i R(p_i)\), with \(R\) a rarity function (Patil and Taillie 1982); the choice \(R(p) = (1 - p^{q-1})/(q-1)\) gives the \(q\)-logarithmic entropies \(S_q\), whose numbers equivalents are the Hill numbers. \(D_q\) itself is an average rarity, namely the power mean \(M_{1-q}(p, 1/p)\) of the rarities \(1/p_i\) weighted by \(p\) (Leinster 2021, Definition 4.3.4), which is Hill’s own reading of \(D_q\) as the reciprocal of a mean proportional abundance (Hill 1973). The property on which the oil-company example depends is the replication principle: if \(n\) communities of equal size, sharing no species and each of diversity \(D\), are pooled, the pooled community has diversity \(nD\). Formally, \(D(u_n \otimes p) = n\, D(p)\), whereas \(H(u_n \otimes p) = \log n + H(p)\).
Leinster’s characterization. The Hill numbers are characterized by this property plus a few elementary ones (Leinster 2021, Theorem 7.4.3). One condition needs a definition (Leinster 2021, Definition 7.4.1): a sequence \(D: \Delta_n \to (0, \infty)\) is modular-monotone if \(D(p_i) \le D(\tilde p_i)\) for all \(i\) implies \(D(w \circ (p_1, \dots, p_n)) \le D(w \circ (\tilde p_1, \dots, \tilde p_n))\), i.e., if no subcommunity’s diversity decreases then neither does that of the whole. It implies modularity, i.e., dependence on the \(p_i\) only through the \(D(p_i)\). The theorem states that for a sequence of functions \((D: \Delta_n \to (0, \infty))_{n \ge 1}\) the following are equivalent: i) \(D\) is symmetric, absence-invariant (deleting a species of abundance zero does not change \(D\)), continuous in the positive probabilities, normalized (\(D(u_1) = 1\)), modular-monotone, and satisfies the replication principle; ii) \(D = D_q\) for some \(q \in [-\infty, \infty]\). A refinement in the same section excludes negative \(q\) under an additional condition.
Compared with the SK axioms, the correspondence holds exactly for SK3 and for additivity, and up to a deliberate weakening for SK4. Expansibility (SK3) corresponds to absence-invariance. Additivity on independent systems corresponds to multiplicativity \(D(w \otimes p) = D(w)\, D(p)\), of which the replication principle is the special case \(w = u_n\); the theorem assumes only this special case and derives multiplicativity from it (Leinster 2021, Lemma 7.4.7). The chain rule (SK4), whose exact diversity form is \(D(w \circ (p_1, \dots, p_n)) = D(w) \prod_i D(p_i)^{w_i}\) (Leinster 2021, Corollary 2.5.8), corresponds to the strictly weaker modular-monotonicity, i.e., the diversity of the whole depends monotonically only on \(w\) and the diversities of the parts; this weakening is the reason the theorem produces the whole family \(D_q\) instead of \(D_1\) alone. Finally, the effective-number property \(D(u_n) = n\), which is the exponential-scale counterpart of the normalization \(H(u_n) = \log n\) rather than of the maximality inequality SK2, is not assumed at all, because it follows from normalization and replication: \(D(u_n) = D(u_n \otimes u_1) = n\, D(u_1) = n\) (Leinster 2021, Lemma 7.4.4). Maximality at the uniform distribution itself is not a consequence of the six hypotheses: the theorem admits Hill numbers of negative order, which are not maximized at \(u_n\) (Leinster 2021, Remarks 4.4.4(ii)); it holds for the refinement \(q \ge 0\). Since \(H_q = \log D_q\), the theorem is also a new characterization of the Rényi entropies, without Rényi’s quasi-arithmetic mean among the hypotheses (Leinster 2021, Remark 7.4.15).
2.2.5 Similarity
Hill numbers treat every pair of distinct species as equally distinct. A similarity matrix removes this restriction, and the two constructions below use it differently. Let two objects have similarity \(s \in [0,1]\) and equal weight, so that \(Z = K = \begin{pmatrix} 1 & s \\ s & 1 \end{pmatrix}\). The eigenvalues of \(K/2\) are \((1 \pm s)/2\), and the exponential of their Shannon entropy is \(\exp H\big(\tfrac{1+s}{2}, \tfrac{1-s}{2}\big)\), about \(1.75\) for \(s = 0.5\). Alternatively, the expected similarity of a random individual to either object is \((Zp)_i = (1+s)/2\), and \(1/(Zp)_i = 2/(1+s)\), which is \(1.33\) for \(s = 0.5\). The two functionals coincide when \(K\) is block-diagonal with all-ones blocks, in which case the non-zero eigenvalues of \(K/n\) are the block masses (as noted in Section 6.5), and they differ in general.
Similarity-sensitive diversity. Rao’s quadratic entropy \(\sum_{i,j} p_i p_j d_{ij}\), the expected dissimilarity between two randomly drawn individuals, is the best-known early index that admits inter-species distances (the same form appears as Nei and Li’s 1979 nucleotide diversity) (Rao 1982); Stirling decomposes diversity into variety, balance and disparity and proposes the heuristic \(\sum_{i \ne j} d_{ij}^{\alpha} (p_i p_j)^{\beta}\), whose \(\alpha = \beta = 1\) case is Rao’s index (Stirling 2007). The family that contains the Hill numbers is Leinster and Cobbold’s similarity-sensitive diversity of order \(q\) (Leinster and Cobbold 2012): for a similarity matrix \(Z\) with \(Z_{ii} = 1\) and \(0 \le Z_{ij} \le 1\), and \(p \in \Delta_n\),
\[ D_q^{Z}(p) = \Big(\sum_{i:\, p_i > 0} p_i\, \big((Zp)_i\big)^{q-1}\Big)^{1/(1-q)}, \qquad q \ne 1 , \]
where \((Zp)_i = \sum_j Z_{ij} p_j\) is the expected similarity between species \(i\) and an individual chosen at random, and \(q = 1, \infty\) are taken as limits. \(Z = I\) recovers \(D_q\). The properties they establish are of three kinds: i) elementary ones (symmetry, absent species, identical species, effective number); ii) the effect of similarity (monotonicity in \(Z\), the naive model \(Z = I\) as an upper bound, and the range \(1 \le D_q^Z(p) \le n\)); iii) partitioning (modularity and replication for mutually dissimilar subcommunities).
Maximum diversity. For a symmetric \(Z\), i) there exists a distribution that maximizes \(D_q^Z\) for all \(q \in [0,\infty]\) simultaneously, and ii) the maximum value \(\sup_p D_q^Z(p)\) is independent of \(q\) (Leinster and Meckes 2016; Leinster 2021, Theorem 6.3.2). The maximum equals the magnitude \(|Z_B|\) of some principal submatrix \(Z_B\) of \(Z\) (for invertible \(Z_B\), the sum of all entries of \(Z_B^{-1}\)) (Leinster 2021, Theorem 6.3.13), and in important classes of cases the magnitude of \(Z\) itself. The same quantity had appeared in 1994, when Solow and Polasky proposed \(V(S) = \mathbf{1}^{\top} F^{-1} \mathbf{1}\) with \(F_{ij} = f(d_{ij})\) for a positive-definite, decreasing \(f\) with \(f(0) = 1\) and \(f(\infty) = 0\), their leading example being \(f(d) = \exp(-\theta d)\), and called it the effective number of species (Solow and Polasky 1994); for the exponential choice it is exactly the magnitude of the scaled metric space, a coincidence that Leinster records as a historical surprise (Leinster 2021). Leinster’s earlier note on the result is titled “A maximum entropy theorem with applications to the measurement of biodiversity” (Leinster 2009). The title is meant in a broader sense than the principle of Section 2.3: there are no moment constraints, the maximized quantity is the similarity-sensitive entropy \(\log D_q^Z\), and the maximizer does not depend on \(q\).
Diversity of a set: the economics tradition. Economists measure the diversity of a plain set, without abundances. Weitzman defines a diversity function recursively from pairwise distances, \(V(S) = \max_{i \in S} \big[V(S \setminus \{i\}) + d(i, S \setminus \{i\})\big]\) with \(d(i, Q) = \min_{j \in Q} d_{ij}\), and shows that it has monotonicity in species, the twin property, and continuity and monotonicity in distances (Weitzman 1992). Solow and Polasky set out three natural requirements: monotonicity in species (adding a new distinct species does not decrease diversity), twinning (adding an exact copy of a present species does not increase it; the term is Weitzman’s), and monotonicity in distance; their \(V(S)\) satisfies the first two, but they note that it is not monotone in distance in general, and conjecture that it is for exponential \(f\) when the distances satisfy the triangle inequality (Solow and Polasky 1994; as read by Emmerich et al. 2026b). Nehring and Puppe model diversity as the total weight of the attributes realized by some member of a set, \(v(S) = \sum_{A:\, A \cap S \ne \emptyset} \lambda_A\) with \(\lambda_A \ge 0\), and show that these are exactly the set functions with non-negative conjugate Möbius inverse, equivalently the monotone, totally submodular set functions, which are the duals of the totally monotone belief functions (Nehring and Puppe 2002). In particular such functions are submodular, which places set diversity among the set functions learned in Section 5.4.
Kernels in machine learning. The Vendi Score of Friedman and Dieng is the eigenvalue-based construction of the opening example (Friedman and Dieng 2023). Given a positive semidefinite kernel \(k\) with \(k(x,x) = 1\) and samples \(x_1, \dots, x_n\), let \(K_{ij} = k(x_i, x_j)\) and \(\rho = K/n\), a positive semidefinite matrix of unit trace, i.e., a density matrix. The Vendi Score is
\[ \mathrm{VS}_k(x_1, \dots, x_n) = \exp\Big(-\sum_i \lambda_i \log \lambda_i\Big) = \exp\big(-\operatorname{Tr} \rho \log \rho\big), \]
with \(\lambda_i\) the eigenvalues of \(\rho\), i.e., the exponential of the von Neumann entropy of \(\rho\), hence the name. Their Theorem 3.1 lists four properties: 1) effective number: if \(k(x_i, x_j) = 0\) for all \(i \ne j\) then \(\mathrm{VS}\) is maximal and equals \(n\), and if \(k(x_i,x_j) = 1\) for all \(i,j\) it is minimal and equals 1; 2) identical elements: if \(k(x_i, x_j) = 1\) for some \(i \ne j\), then merging the weights, \(p'_i = p_i + p_j\) and \(p'_j = 0\) in the weighted form of the score, leaves \(\mathrm{VS}\) unchanged; 3) partitioning: if \(S_1, \dots, S_m\) are mutually dissimilar collections (\(k = 0\) across groups) of relative sizes \(p_i\), then \(\mathrm{VS}(S_1, \dots, S_m) = \exp\big(H(p_1, \dots, p_m)\big) \prod_i \mathrm{VS}(S_i)^{p_i}\); 4) symmetry under permutation. No reference dataset or labels are required. Property 3) is the chain rule in exponential form. Pasarkar and Dieng replace the Shannon entropy of the spectrum by \(H_q\), giving \(\mathrm{VS}_q = \exp H_q(\lambda)\) (Pasarkar and Dieng 2024); this is the matrix-based Rényi entropy of Giraldo et al. (Sánchez Giraldo et al. 2015), as Section 6.5 notes. At \(q = 2\) no eigendecomposition is needed, since \(\operatorname{Tr}\rho^2 = n^{-2}\sum_{i,j} K_{ij}^2\); the entropy \(-\log\big(n^{-2} \sum_{i,j} K_{ij}^2\big)\) is the RKE of Jalali, Li and Farnia, computable in \(O(n^2)\) (Jalali et al. 2023). FKEA accelerates the general case with random features (Ospanov et al. 2024), and quality-weighted variants exist (Nguyen and Dieng 2024). A ninth row could be added to the table of readings in the preceding section: diversity, over the eigenvalues of a normalized similarity kernel, \(\exp(-\operatorname{Tr}\rho\log\rho)\), Friedman and Dieng, 2023.
DCScore. A cheaper summarization of the same kernel reads diversity as a classification problem (Zhu et al. 2025). With embeddings \(h_i\), kernel matrix \(K\) and temperature \(\tau > 0\), form \(P = \mathrm{softmax}(K/\tau)\) row-wise, \(P_{ij} = \exp(K_{ij}/\tau) / \sum_{j'} \exp(K_{ij'}/\tau)\), and set \(\mathrm{DCScore} = \operatorname{tr}(P) = \sum_i P_{ii}\), the total probability that each sample is classified back into its own class. Zhu et al. prove four properties, named after those of Leinster and Cobbold; stated precisely, they are the following. i) Effective number: for a kernel whose rows are maximal on the diagonal (cosine similarity of normalized embeddings, or an RBF kernel), \(P_{ii} \ge 1/n\), so \(1 \le \mathrm{DCScore} < n\); the value is 1 exactly when all samples coincide, and the value \(n\) for distinct samples is reached only in the limit, which the proof states as “tending to \(n\)”, i.e., as \(\tau \to 0\). ii) Identical samples: merging a dataset with an identical copy of itself leaves the score unchanged. This is a dataset-level property, not Leinster and Cobbold’s merging of two identical species, and it holds exactly for every kernel and every \(\tau\), since duplication halves each \(P_{ii}\) and doubles their number. iii) Symmetry under permutation. iv) Monotonicity, a comparative statement: two datasets of \(n\) mutually distinct samples, both with DCScore \(n\), receive the same new sample, which is more similar to the samples of the second; then the first augmented dataset scores higher. Since DCScore \(= n\) holds only as \(\tau \to 0\), the version at a fixed \(\tau\) is the elementwise one: for symmetric \(K\) and \(i \ne j\),
\[ \frac{\partial\, \mathrm{DCScore}}{\partial K_{ij}} = -\frac{1}{\tau}\big(P_{ii} P_{ij} + P_{jj} P_{ji}\big) < 0 , \]
where \(K_{ij} = K_{ji}\) varies jointly, so the score strictly decreases whenever an off-diagonal similarity increases, and iv) holds whenever the two augmented kernel matrices differ only in the row and column of the new sample, with the entries of the second at least as large and one of them larger. Given \(K\), the summarization costs \(O(n^2)\) for DCScore against \(O(n^3)\) for the eigendecomposition in the Vendi Score; with an inner-product kernel and \(n \gg d\), the Vendi Score can use the \(d \times d\) matrix \(X^\top X\) instead, at total cost \(O(d^2 n)\), against \(O(n^2 d)\) for DCScore (Zhu et al. 2025, Table 2). The following table collects the properties of these measures and of the constructions of the next subsection, where the rows A1–A3 are defined.
| property | Hill \(D_q\) | Leinster–Cobbold \(D_q^Z\) (equal weights, \(q \in [0, \infty)\), \(q \ne 1\)) | Vendi / \(q\)-Vendi | DCScore | Mironov–Prokhorenkova constructions |
|---|---|---|---|---|---|
| symmetry | yes | yes | yes | yes | yes |
| absence-invariance / identical elements | yes | yes | yes (weighted form) | dataset duplication only | — |
| effective number | yes | yes | yes | 1 for identical samples; \(n\) only as \(\tau \to 0\) | no (not normalized; inferred from the definitions) |
| modularity / partitioning | yes | yes | yes | no | — |
| replication | yes | yes | yes | no | — |
| strict monotonicity in distances (A1) | n/a | yes | no for \(q = 1\); yes for RKE (\(q = 2\), non-negative similarities) | yes, if \(K_{ij} = f(d_{ij})\) with \(f\) strictly decreasing | yes |
| uniqueness (A2) | n/a | no | no | no | yes |
| continuity (A3) | yes | yes | yes | yes | yes |
| cost | \(O(n)\) | \(O(n^2)\) | \(O(n^3)\) given \(K\), \(O(d^2 n)\) with a linear kernel on \(d\)-dimensional embeddings; \(O(n^2)\) for RKE | \(O(n^2)\) given \(K\), \(O(n^2 d)\) with a linear kernel | NP-hard |
Several cells carry qualifications. The A2 failure of \(D_q^Z\) is established in Mironov and Prokhorenkova (2025) numerically, on one three-point example for \(q \in [0, 100]\), and not by a proof for all \(q\); Leinster and Cobbold’s own monotonicity in \(Z\) (Leinster 2021, Lemma 6.2.2) is non-strict, and the strict A1 is Mironov and Prokhorenkova’s statement for the equal-weight case. “Not normalized” is not stated in Mironov and Prokhorenkova (2025) but follows from the definitions: MultiDimVolume equals \(n - 1\) for \(n\) points at unit distance and \(0\) for coincident points, so it is not an effective number in \([1, n]\). The DCScore entries for A1–A3, partitioning and replication are not stated in Zhu et al. (2025) and follow from the definition. A1 holds whenever the similarity is a strictly decreasing function of the distance, e.g., cosine similarity of normalized embeddings, \(1 - \|h_i - h_j\|^2/2\), by the derivative above; the argument is unaffected by duplicates and, unlike that for RKE, needs no sign condition on the similarities. A2 fails: on the three points of Mironov and Prokhorenkova (2025) used there against RKE (on a circle with cosine similarity, arcs \(1.1\) and \(0.4\) radians from \(x_1\) to \(x_2\) and from \(x_2\) to \(x_3\)), making \(x_2\) a copy of \(x_3\) raises DCScore from \(1.337\) to \(1.394\) at \(\tau = 1\). A smaller \(\tau\) does not restore A2: by Proposition 6.1 of Mironov and Prokhorenkova (2025), a continuous measure satisfying A2 cannot depend on which element is duplicated, whereas for points at angles \(0, 0.3, 1.0, 2.2\) and \(\tau = 0.1\), adding a copy of \(x_2\) or of \(x_3\) gives \(2.998\) or \(3.095\). Partitioning and replication fail at every \(\tau > 0\), since a similarity of 0 still contributes \(e^0 = 1\) to the softmax denominator: two samples with \(K_{11} = K_{22} = 1\) and \(K_{12} = 0\) score \(2e^{1/\tau}/(e^{1/\tau} + 1) < 2\), e.g., \(1.462\) at \(\tau = 1\). The two limits are degenerate: as \(\tau \to \infty\) every dataset scores 1, and as \(\tau \to 0\) DCScore tends to the number of distinct samples, provided \(K_{ij} < K_{ii}\) whenever \(x_j \ne x_i\), i.e., to the measure Unique of Mironov and Prokhorenkova (2025), which satisfies A2 only.
2.2.6 What axioms cannot settle
Consider three points on a circle with cosine similarity, with arc \(0.6\) radians from \(x_1\) to \(x_2\) and \(1.4\) radians from \(x_2\) to \(x_3\). Move \(x_3\) a further \(0.1\) radians away from the other two. No pairwise distance decreases, and the Vendi Score falls from \(1.941\) to \(1.916\). Now let the arcs be \(0.2\) and \(0.3\) radians and replace \(x_2\) by a copy of \(x_1\). The set now contains a duplicate, and the Vendi Score rises from \(1.187\) to \(1.233\). Both computations are from (Mironov and Prokhorenkova 2025, sec. 3 and Appendix B).
Three axioms for the diversity of a set. Mironov and Prokhorenkova formalize the intuitions violated here (Mironov and Prokhorenkova 2025). The setting is \(n\) objects, possibly repeated, with pairwise distances \(d_{ij}\) satisfying \(d_{ij} \ge 0\), \(d_{ii} = 0\), symmetry, and the consistency condition that \(d_{ij} = 0\) implies \(d_{ik} = d_{jk}\) for all \(k\); the triangle inequality is not required. A diversity function is a permutation-invariant map from this space \(\mathcal{D}_n\) to \(\mathbb{R}\), with \(n\) fixed. The axioms are: A1 (monotonicity), the function is strictly increasing in every pairwise distance, and this must hold also when the set contains duplicates; A2 (uniqueness), replacing an element that coincides with some other element by any element that coincides with none strictly increases diversity; A3 (continuity) in the distances. Their Table 2 checks the existing measures. Average pairwise distance and its sum, RKE, and Species(\(q\)), the latter being Leinster–Cobbold diversity with equal weights, satisfy A1 and A3 and fail A2, at cost \(O(n^2)\) (for RKE, which depends on the similarities only through \(s_{ij}^2\), A1 presupposes \(s_{ij} \ge 0\)); Diameter, Bottleneck and the Energy measures fail A1 and A2; the count of unique elements satisfies A2 only; the Vendi Score and the DPP determinant fail A1 and A2 and satisfy A3, at cost \(O(n^3)\); #Circles(\(t\)) fails all three and HamDiv fails A1 and A2, and both are NP-hard. Two new constructions, MultiDimVolume and IntegralMaxClique, satisfy all three axioms, and both are NP-hard to compute. The paper leaves open whether a polynomial-time measure satisfying A1–A3 exists. Velikonivtsev et al. state slightly weaker variants of A1 and A2 (they assume all objects pairwise distinct), and the two-object “dissimilarity” axiom of Xie et al., i.e., non-strict monotonicity of the diversity of a pair in its distance, is implied by A1 (Velikonivtsev et al. 2024; Xie et al. 2023, Axiom 4.3; cf. Mironov and Prokhorenkova 2025, sec. 4.4). Emmerich et al. show, in two preprints, that selecting a fixed-cardinality subset of maximum Solow–Polasky diversity is NP-hard in general metric spaces, and remains so for points in the Euclidean plane (Emmerich et al. 2026b, 2026a).
It is important to read the Vendi row correctly. That the Vendi Score fails A2 does not contradict its own identical-elements property. The latter is an invariance of the weighted representation (two identical samples may be merged into one of double weight without changing the score), whereas A2 asks for a strict increase when a duplicate is replaced by a new element; the two axiom sets are about different properties. Propositions 6.1 and 6.2 of Mironov and Prokhorenkova (2025) explain the difficulty. Proposition 6.1 states that under A2 and A3 the diversity of a set does not depend on which of its elements are duplicated. Proposition 6.2 states that no measure satisfying A1–A3 decomposes as \(F(d_{12}, \dots, d_{1n}) + G(\text{distances among } x_2, \dots, x_n)\), so no additive split into one element’s contribution plus the diversity of the rest is available.
Which similarity. The axiomatic method fixes a functional only after the similarity, or the embedding that induces it, has been chosen. Friedman and Dieng note the dependence themselves: a too sensitive similarity makes every set look diverse, and one that is not sensitive enough gives every set a low score (Friedman and Dieng 2023). Zhao et al., reviewing 135 image and text datasets presented as diverse, borrow the vocabulary of measurement theory (Zhao et al. 2024): conceptualization, operationalization, and evaluation of the indicators for reliability (inter-annotator agreement, test–retest) and validity (convergent and discriminant); scale, they note, is not diversity. In these terms the choice of \(k\) or \(Z\) is a construct-validity question rather than a theorem, and, since diversity metrics depend on the embedding space, they recommend benchmarking a dataset across several embedding spaces suited to the chosen definition of diversity.
Several words in this section have different meanings in different literatures. i) “Uniqueness” is a theorem in the axiomatics of entropy (the functional is unique given the axioms) and an axiom in Mironov and Prokhorenkova (a duplicate is worth less than a new element). ii) “Monotonicity” has at least four senses: Shannon’s increasing \(A(n)\); Solow and Polasky’s monotonicity in species and in distance, both non-strict; Mironov and Prokhorenkova’s strict monotonicity in every distance, duplicates included; and the comparative property proved for DCScore. iii) “Effective number”, \(D(u_n) = n\) for all \(n\), is stronger than “normalized”, \(D(u_1) = 1\); Leinster’s theorem assumes the latter and derives the former from replication. Compare the warning on min-entropy in Section 2.1.4.
In summary, this section supplies the following for learning. The chain rule is the reason the remainder of this book works with \(H\) by default; when a later chapter writes a Rényi or Tsallis quantity instead, one named axiom has been traded for a computational or modelling advantage, and Chapter 7 treats that trade as part of model selection. The kernel relaxation of collision entropy in Section 6.5 is the \(q = 2\) member of the Vendi family, i.e., RKE up to a monotone transformation, and the table above records that it keeps A1 for non-negative similarities and gives up A2. Finally, the drop of output diversity after RLHF (Kirk et al. 2024) and the diversity collapse under RLVR (Yuan et al. 2026) (a preprint) are statements about the effective number of a model’s outputs under some similarity, and Chapter 9 discusses them in these terms.
2.3 The maximum entropy principle
Given constraints — typically moment constraints of the form \(\mathbb{E}_p[f_k(X)] = \mu_k\) — the maximum entropy distribution solves
\[ \max_{p}\; H(p) \quad \text{s.t.} \quad \sum_x p(x) = 1,\;\; \mathbb{E}_p[f_k(X)] = \mu_k . \]
Its solution is the Gibbs / Boltzmann distribution
\[ p(x) = \frac{1}{Z(\boldsymbol{\lambda})}\, \exp\!\Big(\textstyle\sum_k \lambda_k f_k(x)\Big), \]
where \(\boldsymbol{\lambda}\) are Lagrange multipliers and \(Z\) is the partition function. This is the least-biased distribution consistent with what we know (Jaynes 1957).
The single most important instance of this framework is the one in which the constrained feature is the energy of a physical system. That case, worked out in Chapter 3, is what turns the multipliers above into a temperature and connects the whole construction to thermodynamics.
2.4 Minimum description length and minimax entropy
The maximum-entropy principle can be derived rather than postulated. Under the minimum description length (MDL) view (Grünwald 2007), a model \(P\) encodes each state with a code word of length \(-\log P(x)\), and its description length is the expected code length across all datasets consistent with the measured features. Among models matching those features, description length is minimized by the maximum-entropy model (Feder 1986).
MDL also poses a question that maximum entropy alone cannot answer: which features should be measured in the first place? For maximum-entropy models the description length coincides with the model’s own entropy,
\[ L(P_F) = S(P_F), \]
so the best feature set \(F\) is the one whose maximum-entropy model has the smallest entropy. This is the minimax entropy principle, developed in Chapter 7.
2.5 Why it matters for learning
The maximum-entropy principle explains why exponential families are ubiquitous, grounds the Boltzmann machine in statistical physics, and — as developed in Bian et al. (2022) — provides a principled basis for defining valuations in machine learning. The next chapter supplies the missing half of the picture: what happens when entropy is traded off against energy, and what the exchange rate between them turns out to be.