7  Minimax-Entropy Learning

Released

August 4, 2026

Last updated

October 6, 2026

Maximum entropy (Chapter 5) answers a precise question: given a set of measured features, which model is least biased? It is silent on a prior question — which features should we measure at all?

Every model is a lossy compression that trades complexity against uncertainty. Include too few features and the model stays vague; include too many and it ceases to be a compression. For a fixed budget of features, we want the description that leaves the least uncertainty.

Many learning problems raise a second question as well. A classifier, a test-time adapter, or a reasoning model has to commit: it must put its probability on one value of a designated variable \(Y\), e.g., a label or an answer, and describing its inputs well is not enough. The same nesting answers this second question. The inner maximum-entropy model describes the data as honestly as the features allow, and the outer minimization acts on the entropy of the prediction.

The chapter takes the two questions in turn. We state the principle for the first in Section 7.1 and answer the second on the joint space of inputs and labels, with labels in Section 7.2 and without them in Section 7.3; Section 7.4 then marks the boundaries of the principle and weighs the evidence.

7.1 The minimax entropy principle: what to measure

Starting from the minimum description length (MDL) principle (Grünwald 2007), two results follow in sequence. First, among all models consistent with a feature set \(F\), the shortest description is the maximum-entropy model (Feder 1986). Second, for such models the description length equals the model’s own entropy, \(L(P_F) = S(P_F)\). Choosing features to minimize description length therefore means choosing them to minimize that entropy.

The optimal features are the ones producing the maximum-entropy model with minimum entropy (Carcamo et al. 2025):

\[ F^{*} = \arg\min_{F}\; S(P_F), \qquad P_F = \arg\max_{P}\;\big\{\, S(P) \;:\; \mathbb{E}_P[f_\mu] = \langle f_\mu \rangle_{\exp},\; f_\mu \in F \,\big\}. \]

The nesting gives the principle its name: maximize entropy on the inside, minimize it on the outside. The best features are those that maximally constrain our expectations while leaving the model maximally random about everything unobserved. The principle was first proposed for texture modeling in computer vision (Zhu et al. 1997). Throughout this chapter, minimax entropy is meant in this sense of Zhu, Wu, and Mumford; a method of the same name in domain adaptation is unrelated, as we note in Section 7.4.

7.1.1 Three equivalent views

Minimizing \(S(P_F)\) optimizes three quantities at once. In what follows, \(P_{\text{true}}\) denotes the distribution that actually generated the data, and \(P_{\text{ind}}\) the featureless baseline model in which the variables are treated as independent.

  • Shortest description. Since \(L(P_F) = S(P_F)\), the optimum is the most compressed account of the data.
  • Closest to the truth. The entropy gap to the true distribution is exactly a divergence, \(S(P_F) - S(P_{\text{true}}) = D_{\mathrm{KL}}(P_{\text{true}} \Vert P_F)\), so the optimum is also the most accurate model.
  • Most informative features. The information carried by the selected features, \(I_F = S(P_{\text{ind}}) - S(P_F)\), is maximized at the same point.

Compression, accuracy, and informativeness coincide — a reassuring sign that the criterion is the right one.

7.1.2 Solving it

The difficulty is combinatorial. Each candidate feature set requires solving a maximum-entropy problem for its parameters, and the number of candidate sets grows super-exponentially, so brute-force search is infeasible beyond small systems (Carcamo et al. 2025). The practical workhorse is a greedy algorithm that repeatedly adds whichever remaining feature reduces the entropy most.

Structure sometimes restores exactness. For trees of pairwise correlations the entropy reduction decomposes into a sum of mutual informations, the problem becomes a maximum-spanning-tree problem, and greedy is optimal. More generally, if entropy reduction is submodular — diminishing returns as features accumulate — greedy comes within \(1 - 1/e\) of the optimum (Nemhauser et al. 1978). Whether the minimax entropy objective is submodular in general remains open; where it is not, guarantees can still be recovered in terms of how far the objective departs from submodularity (Bian et al. 2017).

7.1.3 Implicit minimax entropy in energy-based models

The principle is closer to everyday practice than its combinatorial statement suggests. Let an energy-based model (Chapter 4) be linear in its last layer, \(E_\theta(x) = -\sum_\mu w_\mu\, \phi_\mu(x; \theta')\), so that for fixed \(\theta'\) the functions \(\phi_\mu(\cdot\,; \theta')\) play the role of the feature set \(F\). For fixed \(\theta'\) the maximum-likelihood \(w\) is the maximum-entropy model \(P_F\) of these features, and at that point the cross-entropy of the data equals \(S(P_F)\); this is the identity Equation 5.18 of Section 5.6.2, whose hypotheses are a linear last layer, an exact fit of \(w\), and a uniform base measure. Maximum likelihood over \(\theta'\) is therefore \(\min_F S(P_F)\) over a continuous family of feature sets: minimax entropy, with gradient descent in place of greedy selection. In this sense every neural energy-based model trained by maximum likelihood is a minimax-entropy learner, whether or not anyone says so. What this chapter adds, and what the implicit version lacks, is the explicit selection of a feature set under a description-length budget — a combinatorial outer problem whose solution can be inspected, bounded, and compared across budgets. It also adds a second kind of outer problem, which acts on the entropy of a prediction rather than on the entropy of the whole model (Section 7.2).

7.2 Minimax entropy on the joint space: modelling and predicting

So far the outer problem has chosen features, i.e., it has decided what to measure. We now turn to the second question raised at the beginning of the chapter, namely what a model should commit to. Let us consider the most common case, a classifier over \(K\) labels, with labelled data in this section and without labels in Section 7.3. As in Chapter 5 and Chapter 6, we write \(H\) for the entropy that Section 7.1 denotes by \(S\), and \(\Delta_K\) for the probability simplex over the labels.

7.2.1 Two questions, one principle

Let \(f_\theta : \mathcal{X} \to \mathbb{R}^K\) be the logits of a classifier and \(h\) a base measure on \(\mathcal{X}\). The joint energy-based model (JEM) of Grathwohl et al. reads the logits as a negative joint energy (Grathwohl et al. 2020), as in Section 4.9:

\[ p_\theta(x, y) = \frac{h(x)\, e^{f_\theta(x)[y]}}{Z(\theta)}, \qquad p_\theta(x) = \frac{h(x)\, e^{-E_\theta(x)}}{Z(\theta)}, \tag{7.1}\]

where \(E_\theta(x) = -\log \sum_{y} e^{f_\theta(x)[y]}\) is the marginal energy and \(Z(\theta) = \int h(x) \sum_y e^{f_\theta(x)[y]}\, dx\). The normalizer cancels in the conditional, \(p_\theta(y \mid x) = \operatorname{softmax}\big(f_\theta(x)\big)_y\), which is the usual softmax classifier. Note that adding the same input-dependent scalar to all the logits, a redundant degree of freedom for an ordinary classifier, now changes \(p_\theta(x)\).

Let the last layer be linear, \(f_\theta(x)[y] = w_y^{\top} \phi_{\theta'}(x) + b_y\), with \(\theta = (\theta', W, b)\). For a fixed representation \(\theta'\), the model Equation 7.1 is the exponential family with the sufficient statistics

\[ \mathsf{T}_{\theta'}(x, y) = \big(e_y \otimes \phi_{\theta'}(x),\; e_y\big), \tag{7.2}\]

where \(e_y\) is the \(y\)-th unit vector of \(\mathbb{R}^K\); it is the maximum-entropy distribution relative to \(h\) that matches the moments of these statistics (Equation 5.1). For a classifier read in this way, the phrase “the model takes maximum entropy” is therefore exact, at the representation level of Section 5.1. Over such a model the outer problem can minimize two different entropies, and we term them two types.

NoteDefinition: minimax-entropy learning

A method belongs to minimax-entropy learning if i) inner: its model is the maximum-entropy (Gibbs) distribution over the modelled variables that matches the moments of a set of features, possibly relative to a base measure; ii) outer: the features, and the assignment of a designated variable \(Y\) where \(Y\) is unobserved, are chosen to minimize an entropy of that model, namely

  • Type I (what to measure): the entropy of the model itself, with the features ranging over a budgeted combinatorial family, as in the principle of Section 7.1, or over a continuous family, as in the implicit version of Section 7.1.3; or
  • Type II (what to commit to): the predictive entropy of \(Y\), measured by the cross-entropy where \(Y\) is observed (Section 7.2.2), and through a posterior sharpened at a label temperature \(T < 1\) where it is not (Section 7.3.1);

and iii) anchor: the inner model is a maximum-entropy learner in the sense of Section 5.1.4, fitted to the same observed data on which the outer entropy is minimized; a frozen reference model does not suffice. When the inner model is fitted by score matching or noise-contrastive estimation instead of maximum likelihood, we speak, as in Section 5.1.4, of minimax-entropy learning in the generalized sense.

Type I is the principle of Section 7.1. Type II is what the slogan of this chapter, minimize entropy to commit to what matters, refers to, and the next two subsections show that, for classification, the slogan can be taken literally. Condition iii) is what separates the present chapter from Chapter 6, and we return to it in Section 7.4.

7.2.2 Classification is conditional minimax entropy

Let us start with labelled data \(\{(x_i, y_i)\}_{i=1}^{n}\) with empirical distribution \(\hat p\), and write \(\hat H_q(Y \mid X) = \frac{1}{n} \sum_i H\big(q(\cdot \mid x_i)\big)\) for the conditional entropy of a conditional distribution \(q\) on the observed inputs.

The conditional bridge. Fix the representation \(\theta'\), let \((W^\star, b^\star)\) be the maximum-likelihood softmax head, and write \(\theta = (\theta', W^\star, b^\star)\). Let \(\mathcal{Q}_{\theta'}\) be the set of conditional distributions \(q\) that match the data on the statistics of Equation 7.2 under the empirical input marginal, i.e., \(\mathbb{E}_{\hat p(x)\, q(y \mid x)}\big[\mathsf{T}_{\theta'}\big] = \mathbb{E}_{\hat p(x, y)}\big[\mathsf{T}_{\theta'}\big]\). Then

\[ \underbrace{\mathbb{E}_{\hat p}\big[-\log p_\theta(y \mid x)\big]}_{\text{cross-entropy}} = \hat H_{p_\theta}(Y \mid X) = \max_{q \in \mathcal{Q}_{\theta'}} \hat H_q(Y \mid X) . \tag{7.3}\]

Minimizing the cross-entropy over the representation is therefore \(\min_{\theta'} \max_{q \in \mathcal{Q}_{\theta'}} \hat H_q(Y \mid X)\): the outer problem minimizes, over the features, the largest conditional entropy that the features allow.

Write \(\lambda = (W, b)\), so that \(\log p_\theta(y \mid x) = \lambda^{\top} \mathsf{T}_{\theta'}(x, y) - \log Z_\lambda(x)\) with \(Z_\lambda(x) = \sum_{y} \exp\big(\lambda^{\top} \mathsf{T}_{\theta'}(x, y)\big)\). The likelihood equations in \(\lambda\) are the moment conditions

\[ \frac{1}{n} \sum_{i} \mathsf{T}_{\theta'}(x_i, y_i) \;=\; \frac{1}{n} \sum_{i} \mathbb{E}_{p_\theta(y \mid x_i)} \big[\mathsf{T}_{\theta'}(x_i, y)\big] . \]

Substituting them into the cross-entropy \(\frac{1}{n} \sum_i \big[\log Z_\lambda(x_i) - \lambda^{\top} \mathsf{T}_{\theta'}(x_i, y_i)\big]\) replaces the data moments by the model moments, and gives \(\frac{1}{n} \sum_i \mathbb{E}_{p_\theta(y \mid x_i)}\big[-\log p_\theta(y \mid x_i)\big] = \hat H_{p_\theta}(Y \mid X)\), which is the first equality. The second is the duality of Berger, Della Pietra, and Della Pietra (Berger et al. 1996): among the conditional distributions that satisfy these moment constraints under the empirical input marginal, the one with the largest conditional entropy is the conditional exponential family with the maximum-likelihood parameters.

It is worth noting that Equation 7.3 removes an apparent conflict of signs. Berger et al. (Berger et al. 1996) and Farnia and Tse (Farnia and Tse 2016) describe classification as maximizing a conditional entropy, and Farnia and Tse identify the maximum-conditional-entropy rule with the robust Bayes decision rule under the logarithmic loss, i.e., with the game against nature of Section 5.6. Training a classifier, on the other hand, is commonly described as minimizing a cross-entropy, which at the fitted head equals the conditional entropy of its predictions. Both descriptions are correct, at different levels of the nesting. The inner problem commits as little as the features allow, and the outer problem chooses the features such that this least committal prediction is still as sharp as possible. Zhou et al. use the name minimax conditional entropy for an analogous nesting in crowdsourcing (Zhou et al. 2015).

The hypotheses are the analogues of those of Equation 5.18: a linear head, an exact fit of the head, and no regularization; no base measure is needed, since the conditional does not involve one. A deep network is trained jointly rather than by nested optimization, so the identity holds at stationary points of the head. With weight decay it becomes an identity of regularized maximum entropy, in which the moment constraints are relaxed rather than exact (Section 5.6). If the features separate the labelled data, the maximum-likelihood head does not exist. The constraint set then contains only the conditional that puts all its mass on the observed labels, and both sides of Equation 7.3 equal zero, the cross-entropy as an infimum.

7.2.3 A classifier is a joint maximum-entropy model

Let us now fit the joint model Equation 7.1 instead of the conditional one.

The joint bridge. Fix \(\theta'\), let \((W^\star, b^\star)\) be the maximum-likelihood head of the joint model, and let \(h\) be uniform. Then

\[ \begin{aligned} & \underbrace{\mathbb{E}_{\hat p}\big[-\log p_\theta(x)\big]}_{\text{generative}} + \underbrace{\mathbb{E}_{\hat p}\big[-\log p_\theta(y \mid x)\big]}_{\text{discriminative}} \\[4pt] & \qquad =\; H\big(p_\theta(X, Y)\big) \;=\; \underbrace{H_\theta(X)}_{\text{modelling}} + \underbrace{H_\theta(Y \mid X)}_{\text{predicting}} , \end{aligned} \tag{7.4}\]

where \(H_\theta(Y \mid X) = \mathbb{E}_{p_\theta(x)} H\big(p_\theta(\cdot \mid x)\big)\) and the entropy of \(X\) is differential; for a non-uniform \(h\), the entropies are taken relative to \(h\). Training JEM with equal weights on its two terms is therefore the principle of Section 7.1 on the joint space \((X, Y)\). The outer problem minimizes the joint entropy of the maximum-entropy model, and the chain rule splits this entropy into a part that models the inputs and a part that predicts the label. In the terms of the definition above, JEM is of both types at once.

With \(\lambda = (W, b)\), the joint negative log-likelihood is \(-\log p_\theta(x, y) = \log Z(\theta) - \lambda^{\top} \mathsf{T}_{\theta'}(x, y) - \log h(x)\), which, for uniform \(h\), is affine in the statistics of Equation 7.2. The likelihood equations in \(\lambda\) equate the data and model moments of \(\mathsf{T}_{\theta'}\), so the expected negative log-likelihood under \(\hat p\) equals that under \(p_\theta\), i.e., \(H\big(p_\theta(X, Y)\big)\); this is Equation 5.18 for the joint exponential family. Splitting \(p_\theta(x, y) = p_\theta(x)\, p_\theta(y \mid x)\) on the left gives the two data terms, and the chain rule of entropy gives the right-hand side.

One should be careful about what is split. The identity Equation 7.4 contains two different splits of one total: a data-side split into a generative negative log-likelihood and a cross-entropy, and a model-side split into \(H_\theta(X)\) and \(H_\theta(Y \mid X)\). Only the totals are equal. The reason is that \(-\log p_\theta(x) = E_\theta(x) + \log Z(\theta) + \text{const}\), and the marginal energy \(E_\theta\) is a log-sum-exp of the statistics rather than a linear function of them, so the moment conditions do not equate its data and model expectations. One can verify on small, exactly solvable examples that the two splits indeed differ term by term. The precise statement behind “JEM minimizes the entropy of its predictions” is therefore that it minimizes the total \(H_\theta(X) + H_\theta(Y \mid X)\), and the two parts can trade against each other.

Grathwohl et al. train JEM with the factorization \(\log p_\theta(x, y) = \log p_\theta(x) + \log p_\theta(y \mid x)\): the second term is the exact cross-entropy, and the gradient of the first is estimated by stochastic gradient Langevin dynamics (SGLD) with persistent chains (Grathwohl et al. 2020). On CIFAR-10 the model reaches an accuracy of 92.9%, against 95.8% for the purely discriminative Wide ResNet, with an Inception Score of 8.76 and an FID of 38.4. On CIFAR-100 it reaches 72.2% against a baseline of 74.2%, and it is, in the authors’ words, nearly perfectly calibrated, whereas the baseline is very poorly calibrated. The authors also note that their gradient estimators are quite unstable. Later work narrows the gaps in accuracy and in sample quality with sharpness-aware minimization and by excluding data augmentation from the maximum-likelihood estimate (Yang et al. 2023), improves accuracy, stability, and speed with a proximal variant of SGLD and an informative initialization (Yang and Ji 2021), and addresses the insufficient quality of the SGLD negative samples (Sustek et al. 2023).

One ablation of the original paper is instructive for our reading. The factorization above was chosen so that \(\log p_\theta(y \mid x)\), which is computed exactly, receives no bias from the sampler. Prior work factorizes the objective instead as \(\log p_\theta(x \mid y) + \log p_\theta(y)\), with a separate normalizer for every class, so that neither \(p(y \mid x)\) nor \(p(x)\) can be computed. Trained to maximize this objective, the model reaches only 30.1% accuracy (Grathwohl et al. 2020, Table 1). In the authors’ terms, the two objectives are alternative factorings of the same likelihood. The principle fixes the criterion, the joint likelihood, and leaves the estimator open, and the estimator can matter as much as the criterion.

The observation that a discriminative network contains a generative model predates JEM, and it comes from the same line of work as the minimax entropy principle. Dai, Lu, and Wu construct a generative model of a ConvNet as an exponential tilting of a reference distribution (Dai et al. 2015). Xie et al. derive the generative ConvNet from the discriminative ConvNet, by assuming that one of the categories is a base category generated by a reference distribution, and present it as a hierarchical FRAME model (Xie et al. 2016); FRAME is the texture model for which the minimax entropy principle was proposed (Zhu et al. 1997, 1998). A second line learns generative models through discriminative classifiers (Tu 2007; Jin et al. 2017). What JEM adds is a single normalizer for \(p_\theta(x)\), \(p_\theta(x, y)\), and an exact \(p_\theta(y \mid x)\), which makes the joint form trainable with modern architectures. Placing JEM in the present chapter is thus a return to the original line rather than a graft.

7.3 When the label is missing

Without labels, the data moments in Equation 7.3 and Equation 7.4 cannot be computed, and one has to decide how the missing label enters the objective. A natural device is a distribution \(q(\cdot \mid x)\) over the label, with a temperature on it. We develop this device first, and then turn to the term that distinguishes a joint model from a purely discriminative one.

7.3.1 A temperature on the label

Free energy at a label temperature. For an unlabelled input \(x\) and a label temperature \(T > 0\), let

\[ \ell_T(\theta; x) \;=\; \min_{q \in \Delta_K} \Big\{ \mathbb{E}_{q}\big[-\log p_\theta(x, y)\big] - T\, H(q) \Big\} . \tag{7.5}\]

The minimizer is \(q^\star \propto p_\theta(y \mid x)^{1/T}\), and

\[ \begin{aligned} \ell_T(\theta; x) \;&=\; -T \log \sum_{y} p_\theta(x, y)^{1/T} \\[4pt] \;&=\; -\log p_\theta(x) + (1 - T)\, H_{1/T}\big(p_\theta(\cdot \mid x)\big), \end{aligned} \tag{7.6}\]

where \(H_\alpha\) is the Rényi entropy of order \(\alpha\) (Section 2.1.4). In particular, \(\ell_1 = -\log p_\theta(x)\), and as \(T \to 0\),

\[ \ell_T(\theta; x) \;\to\; -\log \max_y p_\theta(x, y) \;=\; -\log p_\theta(x) + H_\infty\big(p_\theta(\cdot \mid x)\big), \]

with \(H_\infty\) the min-entropy.

The objective Equation 7.5 is the free energy Equation 5.2 on the label, with the energy \(-\log p_\theta(x, y)\) and the counting measure as the base measure. Its minimizer is the Gibbs distribution \(q^\star \propto p_\theta(x, y)^{1/T} \propto p_\theta(y \mid x)^{1/T}\), and its minimum is \(-T\) times the log-partition function, \(-T \log \sum_y p_\theta(x, y)^{1/T}\). Writing \(p_\theta(x, y) = p_\theta(x)\, p_\theta(y \mid x)\) and \(\alpha = 1/T\), one has \(\log \sum_y p_\theta(y \mid x)^{\alpha} = (1 - \alpha) H_\alpha\), hence \(-T \log \sum_y p_\theta(y \mid x)^{1/T} = -T (1 - 1/T) H_{1/T} = (1 - T) H_{1/T}\). The value at \(T = 1\) is immediate, and the limit \(T \to 0\) follows from \(-T \log \sum_y a_y^{1/T} \to -\log \max_y a_y\) for positive \(a_y\).

We note that Equation 7.6 places maximum-entropy and minimum-entropy learning of the missing label on one dial, and that the entropy it produces is a Rényi entropy whose order is the reciprocal of the label temperature. At \(T = 1\) the missing label is assigned by the posterior with its full entropy, and the objective reduces to the marginal likelihood. This is maximum-entropy learning with missing data, i.e., EM, or the semi-supervised JEM of Zhao et al. (Zhao et al. 2020), and it belongs to Chapter 5. For \(T < 1\) the entropy enters with the positive weight \(1 - T\), and the objective penalizes the uncertainty of the prediction. This is where the problem becomes minimax: the input side is still a maximum-entropy learner, while the label side is pushed towards low entropy. As \(T \to 0\) the objective becomes the classification likelihood. Lowering \(T\) along a path is deterministic annealing (Rose 1998), here applied to the label, and Table 7.1 places known methods on the dial.

Table 7.1: The label-temperature dial on an unlabelled input, where \(T\) is the label temperature. The objective is \(\ell_T(\theta; x) = -\log p_\theta(x) + (1 - T)\, H_{1/T}\) of Equation 7.6, and the third column gives its label-side term, with \(H_{1/T}\) the Rényi entropy of order \(1/T\) of \(p_\theta(\cdot \mid x)\). The instances include joint and purely discriminative methods; dropping the term \(-\log p_\theta(x)\) gives the purely discriminative version of each row.
\(T\) entropy of the missing label label-side term known instances
\(T > 1\) rewarded \(-(T - 1)\, H_{1/T}\) discriminative analogue: the confidence penalty (Pereyra et al. 2017)
\(T = 1\) full, from the posterior \(0\) EM; semi-supervised JEM (Zhao et al. 2020); TEA on test inputs (Yuan et al. 2024)
\(0 < T < 1\) penalized, from a sharpened posterior \((1 - T)\, H_{1/T}\) deterministic-annealing EM for entropy regularization (Grandvalet and Bengio 2006)
\(T \to 0\) none, a hard label \(H_\infty\) classification EM (Celeux and Govaert 1992); self-training with pseudo-labels (Lee 2013)

Relation to entropy regularization. Grandvalet and Bengio solve their entropy-regularized criterion with a deterministic-annealing EM (Grandvalet and Bengio 2006). Its E-step distributes an unlabelled input over the classes with the Gibbs weights \(g_m \propto P(m \mid x)^{1/(1 - \lambda)}\); \(\lambda = 0\) recovers EM, \(\lambda = 1\) gives the hard assignments of classification EM (Celeux and Govaert 1992), and the authors describe \(T = 1 - \lambda\) as the analogue of a temperature. These weights are exactly the minimizer \(q^\star\) of Equation 7.5 at \(T = 1 - \lambda\). Two remarks are in order. Firstly, their criterion involves only \(P(y \mid x)\) and favours low density separation, in their words, “without any modeling of the density of input features”. The extra term \(-\log p_\theta(x)\) in Equation 7.6 is precisely the maximum-entropy learner on the inputs that a joint model supplies, so Type II can be described as entropy regularization with a maximum-entropy model of the inputs. Secondly, their criterion penalizes the Shannon entropy, \(\lambda H\), and they present it as the subproblem that deterministic annealing solves. The iterations with the weights \(g_m\), however, are block-coordinate descent on Equation 7.5, with \(-\log P(y \mid x)\) in place of \(-\log p_\theta(x, y)\), and the value at the optimal \(q\) is the Rényi penalty \(\lambda H_{1/(1 - \lambda)}\). The two penalties agree to first order in \(\lambda\) and differ beyond it. At \(\lambda = 1\), which they identify with self-training, the iterations reach the classification likelihood, whose penalty is the min-entropy \(H_\infty\) rather than \(H\). The Rényi form of a label-only objective of this kind also appears in the deep clustering of Chen et al. (Chen et al. 2021).

The plug-in version. If \(q\) is not optimized but set to the posterior, one obtains

\[ \mathbb{E}_{p_\theta(y \mid x)}\big[-\log p_\theta(x, y)\big] \;=\; -\log p_\theta(x) + H\big(p_\theta(\cdot \mid x)\big), \tag{7.7}\]

i.e., the marginal negative log-likelihood plus a Shannon entropy of weight one. This is the decomposition behind Hathaway’s interpretation of EM for mixtures as coordinate descent on a single objective (Hathaway 1986): the mixture log-likelihood equals the fuzzy classification log-likelihood plus the entropy of the posterior partition. Classification EM instead maximizes the classification likelihood over hard partitions (Celeux and Govaert 1992). A general weight \(\lambda\) gives \(-\log p_\theta(x) + \lambda H\), whose discriminative version is the entropy regularization of Grandvalet and Bengio (Grandvalet and Bengio 2004).

Minimax over the missing labels. Let the soft labels \(q(\cdot \mid x_i)\) of the unlabelled inputs be outer variables as well, and let \(p^\star_{\theta', q}\) be the maximum-entropy joint model obtained by fitting the head to the pseudo-labelled data \(\hat p(x)\, q(y \mid x)\). Under the hypotheses of the joint bridge,

\[ \min_{q} \min_{\theta'} H\big(p^\star_{\theta', q}\big) = \min_{\theta} \tfrac{1}{n} \sum_{i=1}^{n} \min_{y} \big[-\log p_\theta(x_i, y)\big], \tag{7.8}\]

which is the classification likelihood, the \(T \to 0\) end of the dial. Indeed, by the joint bridge, \(H\big(p^\star_{\theta', q}\big) = \min_{W, b} \mathbb{E}_{\hat p(x)\, q(y \mid x)}\big[-\log p_\theta(x, y)\big]\) for fixed \((\theta', q)\). After exchanging the minimizations, the objective is linear in \(q\), so the optimal \(q\) sits at a vertex of the simplex, i.e., at the label that minimizes \(-\log p_\theta(x_i, y)\) for each \(x_i\). The minimax entropy principle with the missing labels among the outer variables is thus the hard end of the dial, and the label temperature interpolates between it and maximum-entropy learning at \(T = 1\).

The structure has a direct precedent in crowdsourcing. Zhou et al. estimate the true labels from the noisy labels of many workers by a minimax entropy principle (Zhou et al. 2012). The inner problem maximizes the entropy of a distribution over workers, items, and labels, under per-item constraints suggested by majority voting and per-worker constraints suggested by the model of Dawid and Skene, and the outer problem minimizes the entropy of this maximum-entropy model over the unknown true labels. Their Theorem 2.1 shows that, when the true measurements are known, the Kullback–Leibler loss to the true distributions equals the entropy of the maximum-entropy model minus that of the truth, which is the view “closest to the truth” in the three views of Section 7.1.1. Since each item is usually labelled only a few times, they relax the constraints to prevent overfitting, i.e., they use the regularized maximum entropy of Section 5.6; a conditional version of the principle followed (Zhou et al. 2015).

7.3.2 What the density term buys, and what it does not

A joint model differs from a purely discriminative one by the term \(-\log p_\theta(x)\). Let us look at what this term can and cannot do. In this subsection we write \(p_x = p_\theta(\cdot \mid x)\) for short.

The energy–entropy identity. For every input \(x\),

\[ \begin{aligned} -E_\theta(x) \;&=\; \log \sum_{y} e^{f_\theta(x)[y]} \\[4pt] \;&=\; \max_{q \in \Delta_K} \Big\{ \mathbb{E}_{q}\big[f_\theta(x)[y]\big] + H(q) \Big\} \\[4pt] \;&=\; \big\langle f_\theta(x) \big\rangle_{p_x} + H(p_x), \end{aligned} \tag{7.9}\]

where \(\langle f_\theta(x) \rangle_{p}\) denotes the expected logit under \(p\). Hence, for uniform \(h\), \(-\log p_\theta(x) = -\langle f_\theta(x) \rangle_{p_x} - H(p_x) + \log Z(\theta) + \text{const}\). This is the Gibbs variational principle of the chapter on free energy, Chapter 3, applied to the label: the marginal energy of an input is the free energy of its label ensemble, as noted in Section 4.9. Park et al. state it as a Fenchel duality between the energy and the negative entropy (Park et al. 2025, Eq. 5). Three consequences are worth spelling out.

Firstly, the energy and the entropy of a prediction can move independently. Park et al. read Equation 7.9 in one direction: as the entropy goes to zero, the energy approaches minus the largest logit, so that, in their words, minimizing entropy “does not provide a clear momentum to reduce the overall energy”. Read in the other direction, the identity says that lowering the energy need not sharpen the prediction. For the plug-in objective with a weight \(\lambda\) on the entropy, it gives

\[ \begin{aligned} -\log p_\theta(x) + \lambda\, H(p_x) \;&=\; -\big\langle f_\theta(x) \big\rangle_{p_x} + (\lambda - 1)\, H(p_x) \\[4pt] &\qquad + \log Z(\theta) + \text{const}, \end{aligned} \]

so at \(\lambda = 1\) the explicit entropy cancels, and what remains is the soft self-labelled joint negative log-likelihood of Equation 7.7. This is an identity and not a statement that \(\lambda = 1\) is a neutral point: the expected logit and the entropy are not independent coordinates, and, as the third consequence shows, the density term has no preference for any split of the mass among the labels.

Secondly, the density term makes the scale of the logits identifiable. For purely discriminative entropy minimization on unlabelled data, multiplying the logits by \(c > 1\) changes no decision and lowers the entropy, which tends to zero as \(c \to \infty\); only other terms, e.g., weight decay, can stop the growth. With the density term this is no longer the case. Let \(g_c(x) = \frac{1}{c} \log \sum_y e^{c f(x)[y]}\), which converges uniformly to \(g(x) = \max_y f(x)[y]\) as \(c \to \infty\), and let \(h\) have a positive continuous density on a compact \(\mathcal{X}\), with \(g\) continuous. Write \(A_c = \frac{1}{c} \log \int h(x')\, e^{c\, g_c(x')}\, dx'\), which tends to \(\max_{x'} g(x')\) by a Laplace argument. Then

\[ \begin{aligned} -\log p_{cf}(x) \;&=\; c\, \big( A_c - g_c(x) \big) - \log h(x) \\[4pt] \;&=\; c\, \Big( \max_{x'} g(x') - g(x) \Big) + o(c), \end{aligned} \tag{7.10}\]

so the density term grows linearly in \(c\) at every input that is not a maximizer of \(g\). The scale can no longer be sent to infinity for free; it is set by the fit of \(p_\theta(x)\). This does not contradict the remark in Section 5.1.4 that the temperature of a learned energy is not identifiable when the family is closed under rescaling. There the temperature is absorbed into the scale of the energy, and here that common scale, which governs both \(p_\theta(x)\) and \(p_\theta(y \mid x)\), is what the fit of \(p_\theta(x)\) determines. The improved calibration of JEM (Grathwohl et al. 2020) is consistent with this mechanism, although it has not been shown to be caused by it.

Thirdly, the density term does not prevent collapse. By Equation 7.1, the term \(-\log p_\theta(x)\) depends on the logits only through their log-sum-exp, at \(x\) and, through \(Z(\theta)\), at every other input. Any change that keeps the log-sum-exp fixed at every input leaves the density term unchanged, including, in the limit, a change that moves all the posterior mass to one class. The density term shapes the landscape of the marginal energy and does not decide the split of the labels; it can influence the split only indirectly, through the shared network. Whether the labels collapse is therefore decided by the label-side terms, and a fairness term is needed, exactly as in Section 6.3.

In the language of joint maximum entropy, the fairness term has a clean place. The label statistic \(e_y\) of Equation 7.2 has the class frequencies as its moment, and its multiplier is the bias \(b_y\). Without labels, a fairness term prescribes this moment, \(\mathbb{E}_{\hat p(x)\, q(y \mid x)}[e_y] = \pi\), instead of estimating it. Adding this constraint to the sum of Equation 7.5 over the unlabelled inputs gives

\[ q^\star(y \mid x) \;\propto\; \exp\Big( \frac{f_\theta(x)[y] - \gamma_y}{T} \Big), \tag{7.11}\]

where \(\gamma\) is the Lagrange multiplier of the constraint, chosen such that the constraint holds. This is the per-class offset that the history in Section 6.3 finds behind every fairness mechanism, in its hard-constraint form, i.e., Sinkhorn equipartition run to convergence. In practice the moment is computed with the batch-averaged posterior, as in distribution alignment (Berthelot et al. 2020) and in the diversity term of SHOT (Liang et al. 2020).

These observations place the three terms of Section 6.3 in one frame. Firmness, the conditional entropy down, is the outer minimization of \(H(Y \mid X)\), i.e., a label temperature \(T < 1\). Fairness is a prescribed moment of the inner model, Equation 7.11. Capacity, the term that makes the classes mean something, is where the maximum-entropy model of the inputs may contribute, since the shared features must then also explain \(p_\theta(x)\); we state this as a hypothesis rather than a result. What the density term cannot do is replace fairness, since it is blind to collapse. Notably, Sustek et al. add the mutual information \(I_\theta(X; Y) = H_\theta(Y) - H_\theta(Y \mid X)\) under the model distribution to the objective of JEM, which, in their words, encourages the classifier to make more certain decisions about output labels (Sustek et al. 2023). In the terms of this section, this puts a further weight on the Type II part \(H_\theta(Y \mid X)\) together with a fairness term \(H_\theta(Y)\), i.e., it places the mutual-information objective of Section 6.3 on top of a joint maximum-entropy learner.

Finally, it is worth writing the losses with and without labels in one form, because a sign is easily lost. With \(\mathrm{CE}(a, b) = -\sum_y a_y \log b_y\), an entropy is a cross-entropy against the best target,

\[ \begin{aligned} H(p) \;&=\; \min_{r \in \Delta_K} \mathrm{CE}(p, r), \\[4pt] H_\infty(p) \;&=\; -\log \max_y p_y \;=\; \min_{r \in \Delta_K} \mathrm{CE}(r, p), \end{aligned} \tag{7.12}\]

attained at \(r = p\) and at the one-hot mode of \(p\), respectively. With labels the target is given by the data, and without labels the model chooses it. Note that hard pseudo-labelling minimizes the min-entropy \(H_\infty\); this is one place where the min-entropy and minimum-entropy learning, which Section 2.1.4 keeps apart, coincide in form. The common forms are collected in Table 7.2, which shows in particular that the Kullback–Leibler divergence to the uniform distribution raises the entropy when it is minimized.

Table 7.2: Losses with and without labels. Here \(p = p_\theta(\cdot \mid x)\), \(\bar p = \mathbb{E}_{x}\, p_\theta(\cdot \mid x)\) is the batch-averaged prediction, \(u\) is the uniform distribution on the labels, and \(\operatorname{sg}\) denotes a stop-gradient, i.e., the target is held fixed. In DINO the sharpened target comes from a teacher network.
case loss as a cross-entropy or KL effect of minimizing
labelled \(-\log p_\theta(y \mid x)\) \(\mathrm{CE}(\delta_y, p) = \mathrm{KL}(\delta_y \Vert p)\) fits the label
unlabelled, Shannon \(H(p)\) \(\min_r \mathrm{CE}(p, r)\) sharpens (Grandvalet and Bengio 2004)
unlabelled, hard pseudo-label \(H_\infty(p)\) \(\min_r \mathrm{CE}(r, p)\) hard self-training (Lee 2013)
unlabelled, sharpened target \(\mathrm{CE}\big(\operatorname{sg}[r_\tau], p\big)\) \(r_\tau \propto p^{1/\tau}\), a target at temperature \(\tau\) soft self-training (Berthelot et al. 2020; Caron et al. 2021)
towards uniform \(\mathrm{KL}(p \Vert u)\) \(\log K - H(p)\) raises the entropy: the confidence penalty (Pereyra et al. 2017)
fairness \(\mathrm{KL}(\bar p \Vert \pi)\) \(\log K - H(\bar p)\) for \(\pi = u\) resists collapse (Bridle et al. 1992; Liang et al. 2020)

7.4 Boundaries, evidence, and a name clash

Why condition iii) is needed. Without it, every objective of the form “entropy minimization plus a Kullback–Leibler trust region to a reference” would be of Type II, since the optimum of such an objective is a Gibbs tilt of the reference (Equation 6.2). This would include EMPO, which minimizes the semantic entropy of the model’s own answers (Zhang et al. 2025), and reinforcement learning with the negative entropy as the only reward (Agarwal et al. 2025). In these methods, however, the reference, if any, is frozen. It is a maximum-entropy model in the sense of i)–ii) of the definition in Section 5.1.4, not a learner on the current data. Moreover, the analysis in Section 6.6 shows that the released EMPO code switches the KL anchor off altogether, and that the resulting dynamics are a rich-get-richer process with collapse as its asymptotic behaviour. Condition iii) draws the line at whether the maximum-entropy side learns from the data on which the entropy is minimized, and these methods stay in Chapter 6.

The same condition sorts the semi-supervised and test-time methods. The entropy regularization of Grandvalet and Bengio is of Type II on the labelled inputs, through the conditional bridge, and belongs to the minimum-entropy learning of Chapter 6 on the unlabelled ones, where no model of the inputs is fitted; a joint model makes it Type II on both. Entropy minimization at test time (Wang et al. 2021) has no learner on the test inputs at all. TEA (Yuan et al. 2024) fits the marginal energy of the test inputs by contrastive divergence, which is maximum-entropy learning at \(T = 1\). ReTTA (Park et al. 2025) pairs entropy minimization with a sliced-score-matching fit of the energy on the test inputs, and with a cross-entropy that takes the most probable class as its target. The latter equals \(-\log \max_y p_\theta(y \mid x) = H_\infty\), so the objective of ReTTA contains both terms of the \(T \to 0\) end of Equation 7.6, together with a Shannon entropy, and its density is fitted by score matching: ReTTA is of Type II in the generalized sense. Its entropy term taken alone is the self-referential move of Section 9.4, and this is how Table 9.2 lists it.

Why \(T = 1\) is not of Type II. At \(T = 1\) the missing label keeps its full entropy, and nothing on the label side asks the model to commit; the objective is the marginal likelihood, and EM is coordinate descent on it (Hathaway 1986). This is the boundary between Chapter 5 and the present chapter. In the language of Chapter 3, the label temperature is the dial of that chapter applied to the missing label only, while the inputs stay at \(T = 1\). The placement is summarized in Table 7.3.

Table 7.3: Methods placed by the definition of Section 7.2.1. The second column gives the data to which a maximum-entropy learner is fitted, and the third the entropy that the outer problem minimizes. The decisive column is the second, i.e., whether a maximum-entropy learner is fitted to the data on which the entropy is minimized.
method maximum-entropy fit on outer entropy place
FRAME (Zhu et al. 1997) images the entropy of the model, over a filter bank Type I
energy-based model by maximum likelihood (Section 5.6.2) inputs the entropy of the model, over a continuous family Type I, implicit
softmax classifier (Equation 7.3) labelled pairs, conditionally \({\hat H(Y \mid X)}\), through the cross-entropy Type II
JEM (Grathwohl et al. 2020) (Equation 7.4) labelled pairs, jointly \(H_\theta(X)\) and \({H_\theta(Y \mid X)}\) Types I and II
semi-supervised JEM (Zhao et al. 2020); TEA (Yuan et al. 2024) inputs none on the missing label, at \(T = 1\) maximum-entropy learning
joint model at \({T < 1}\); classification EM (Celeux and Govaert 1992) inputs \((1 - T)\, H_{1/T}\), or \(H_\infty\) as \({T \to 0}\) Type II
ReTTA (Park et al. 2025) test inputs, by score matching \(H\) and \(H_\infty\) Type II, generalized sense
entropy regularization (Grandvalet and Bengio 2004) labelled pairs only \(H\) on the unlabelled inputs Type II on the labelled inputs only
Tent (Wang et al. 2021); EMPO (Zhang et al. 2025) none on the current data \(H\), or a semantic entropy minimum-entropy learning

Published evidence. The evidence on the two sides of the boundary is indirect, and it should be read with care. At \(T = 1\), Zhao et al. report that semi-supervised JEM reaches 95.4% on MNIST with 100 labels, against 86.0% for the baseline classifier and 98.4% for virtual adversarial training, and 66.0% on SVHN with 1000 labels, against 62.7% and 62.8%, respectively; in their MNIST experiments, entropy regularization was not found to be helpful for either method (Zhao et al. 2020). Grathwohl et al. train JEM on CIFAR-10 with 4,000 labels and the rest of the training set as unlabelled data, and find a noticeable improvement in calibration but, surprisingly in their words, no improvement in generalization (Grathwohl et al. 2020, Appendix E.2). TEA reports an average gain of 4.7% over the best-performing test-time adaptation methods on corruption and domain-generalization benchmarks (Yuan et al. 2024). Of Type II, ReTTA reaches an average accuracy of 49.2% on ImageNet-C at severity 5 in its mild scenario, at least 0.6% above all compared methods. In its ablation under online label shift, adding the score-matching term to entropy minimization raises the accuracy from 43.9% to 45.6% for a ResNet-50 with group normalization, and from 62.3% to 64.7% for a ViT-Base (Park et al. 2025, Table 3). In summary, the published results are consistent with a density term that helps when it is fitted on the inputs where the prediction is made, but none of them tests one end of the dial against the other with everything else held fixed. We are not aware of a study that varies the label temperature of a joint energy-based model systematically; this is an open question rather than a result.

WarningTerminology

Three neighbours share words with this chapter. i) The minimax entropy (MME) of Saito et al. is a method of semi-supervised domain adaptation, in which the conditional entropy of the unlabelled target data is maximized with respect to the classifier and minimized with respect to the feature extractor, through a gradient-reversal layer (Saito et al. 2019). It is an adversarial game on a single entropy rather than the nesting of a maximum-entropy model inside an entropy minimization, and it is unrelated to the principle of this chapter. ii) The minimax approach of Farnia and Tse (Farnia and Tse 2016), like the game against nature of Section 5.6, is a minimum over decision rules of a maximum over distributions, i.e., the inner problem of this chapter alone; in the same spirit, Section 5.6 calls the primal–dual form of maximum likelihood a saddle-point form and not minimax. iii) Maximum entropy discrimination (Jaakkola et al. 1999) and the principled hybrids of Lasserre et al. (Lasserre et al. 2006) also combine maximum entropy with discrimination, and generative with discriminative training, respectively, but along different axes: the former seeks a minimum-relative-entropy distribution over parameters under margin constraints, and the latter interpolates between a generative and a discriminative model through a prior that couples two sets of parameters.

What is known, and what may be new. The pieces of Section 7.2 and Section 7.3 are known. The conditional and joint bridges are instances of the duality between maximum likelihood and maximum entropy (Berger et al. 1996); JEM and its predecessors read a classifier as an energy-based model (Grathwohl et al. 2020; Xie et al. 2016); the label temperature is the deterministic annealing of Rose (Rose 1998), and its weights are those of Grandvalet and Bengio (Grandvalet and Bengio 2006); the plug-in identity is Hathaway’s (Hathaway 1986); the energy–entropy identity is stated by Park et al. (Park et al. 2025); the minimax entropy principle over unknown labels appears in crowdsourcing (Zhou et al. 2012); and the Rényi form of a label-only objective appears in deep clustering (Chen et al. 2021). What we add is a reading that puts them together: a definition with the anchor condition iii), under which the softmax classifier, JEM, the label-temperature dial, and recent test-time adaptation sit in one nesting, and the observation that, written as Equation 7.6, the dial separates a maximum-entropy learner on the inputs from a Rényi penalty on the label. We make no claim of priority for any of the identities.

7.5 Takeaway

Minimax entropy couples the two preceding principles into a single criterion: maximize entropy to model the world, minimize entropy to commit to what matters. Maximum entropy supplies possibility, minimum entropy supplies certainty, and the minimax formulation decides which details deserve to be modelled at all.

On the joint space of inputs and labels, the slogan can be taken literally. A classifier with a linear head is a conditional maximum-entropy model, and the cross-entropy it minimizes is the largest conditional entropy that its features allow. Read as a joint energy-based model, its two training terms add up to the joint entropy of a maximum-entropy model, one part modelling the inputs and one part predicting the label. Without labels, a temperature on the label is the dial: at \(T = 1\) the problem is maximum-entropy learning with missing data, below it the label side pays a Rényi entropy of order \(1/T\), and as \(T \to 0\) it becomes the classification likelihood. The density term makes the scale of the logits identifiable but is blind to collapse, so fairness remains a separate, prescribed moment.

The principle has recently yielded exactly solvable models of large neuronal populations (Lynn et al. 2025). Its implicit form runs inside every energy-based model trained by maximum likelihood (Section 5.6.2), and its second type is implicit in joint energy-based models and in test-time adaptation that pairs entropy minimization with a learned energy, as shown in Section 7.2 and Section 7.4. Its explicit form — the selection of a feature set under a description-length budget, with guarantees on sample complexity — remains largely untapped in machine learning, and so does an explicit schedule for the label temperature. Open directions include scalability and sample complexity in high dimensions, realizations for LLMs, agents, and multimodal models, and the relationship to the free-energy principle (Friston 2010).