References
Abbeel, Pieter, and Andrew Y. Ng. 2004. “Apprenticeship Learning
via Inverse Reinforcement Learning.” International Conference
on Machine Learning (ICML).
Abe, Sumiyoshi. 2000. “Axioms and Uniqueness Theorem for Tsallis
Entropy.” Physics Letters A 271 (1–2): 74–79. https://doi.org/10.1016/S0375-9601(00)00337-6.
Ackley, David H., Geoffrey E. Hinton, and Terrence J. Sejnowski. 1985.
“A Learning Algorithm for Boltzmann Machines.”
Cognitive Science 9 (1): 147–69. https://doi.org/10.1207/s15516709cog0901_7.
Aczél, János, and Zoltán Daróczy. 1963. “Über Verallgemeinerte
Quasilineare Mittelwerte, Die Mit Gewichtsfunktionen Gebildet
Sind.” Publicationes Mathematicae Debrecen 10: 171–90.
https://doi.org/10.5486/PMD.1963.10.1-4.24.
Agarwal, Shivam, Zimin Zhang, Lifan Yuan, Jiawei Han, and Hao Peng.
2025. “The Unreasonable Effectiveness of Entropy Minimization in
LLM Reasoning.” Advances in Neural Information
Processing Systems (NeurIPS).
Ahmadian, Arash, Chris Cremer, Matthias Gallé, et al. 2024. “Back
to Basics: Revisiting REINFORCE-Style Optimization for
Learning from Human Feedback in LLMs.”
Proceedings of the 62nd Annual Meeting of the Association for
Computational Linguistics (Volume 1: Long Papers), 12248–67.
Ahuja, Ravindra K., and James B. Orlin. 2001. “Inverse
Optimization.” Operations Research 49 (5): 771–83.
Altun, Yasemin, and Alexander J. Smola. 2006. “Unifying Divergence
Minimization and Statistical Inference via Convex Duality.”
Conference on Learning Theory (COLT).
Asano, Yuki M., Christian Rupprecht, and Andrea Vedaldi. 2020.
“Self-Labelling via Simultaneous Clustering and Representation
Learning.” International Conference on Learning
Representations (ICLR). https://arxiv.org/abs/1911.05371.
Baez, John C., and Tobias Fritz. 2014. “A Bayesian
Characterization of Relative Entropy.” Theory and
Applications of Categories 29 (16): 422–56. https://doi.org/10.70930/tac/myyolhqw.
Baez, John C., Tobias Fritz, and Tom Leinster. 2011. “A
Characterization of Entropy in Terms of Information Loss.”
Entropy 13 (11): 1945–57. https://doi.org/10.3390/e13111945.
Balcan, Maria-Florina, and Nicholas J. A. Harvey. 2018.
“Submodular Functions: Learnability, Structure, and
Optimization.” SIAM Journal on Computing 47 (3): 703–54.
https://doi.org/10.1137/120888909.
Balcerak, Michal, Tamaz Amiranashvili, Antonio Terpin, et al. 2025.
“Energy Matching: Unifying Flow Matching and Energy-Based Models
for Generative Modeling.” Advances in Neural Information
Processing Systems 38 (NeurIPS).
Barbera, Elvira. 1999. “On the Principle of Minimal Entropy
Production for Navier–Stokes–Fourier Fluids.” Continuum
Mechanics and Thermodynamics 11 (5): 327–30. https://doi.org/10.1007/s001610050127.
Beirami, Ahmad, Alekh Agarwal, Jonathan Berant, et al. 2024.
“Theoretical Guarantees on the Best-of-n Alignment Policy.”
arXiv Preprint arXiv:2401.01879.
Belanger, David, and Andrew McCallum. 2016. “Structured Prediction
Energy Networks.” International Conference on Machine
Learning (ICML).
Belanger, David, Bishan Yang, and Andrew McCallum. 2017.
“End-to-End Learning for Structured Prediction Energy
Networks.” International Conference on Machine Learning
(ICML), 429–39.
Bengio, Yoshua, and Olivier Delalleau. 2009. “Justifying and
Generalizing Contrastive Divergence.” Neural Computation
21 (6): 1601–21.
Bennett, Charles H. 1976. “Efficient Estimation of Free Energy
Differences from Monte Carlo Data.”
Journal of Computational Physics 22 (2): 245–68. https://doi.org/10.1016/0021-9991(76)90078-4.
Berger, Adam L., Stephen A. Della Pietra, and Vincent J. Della Pietra.
1996. “A Maximum Entropy Approach to Natural Language
Processing.” Computational Linguistics 22 (1): 39–71.
Berthelot, David, Nicholas Carlini, Ekin D. Cubuk, et al. 2020.
“ReMixMatch: Semi-Supervised Learning with Distribution Alignment
and Augmentation Anchoring.” International Conference on
Learning Representations (ICLR). https://arxiv.org/abs/1911.09785.
Besag, Julian. 1974. “Spatial Interaction and the Statistical
Analysis of Lattice Systems.” Journal of the Royal
Statistical Society: Series B (Methodological) 36 (2): 192–236.
Bian, Andrew An, Joachim M. Buhmann, Andreas Krause, and Sebastian
Tschiatschek. 2017. “Guarantees for Greedy Maximization of
Non-Submodular Functions with Applications.” International
Conference on Machine Learning (ICML), 498–507.
Bian, Yatao. 2020. Awesome Energy-Based Models/Learning: A
Comprehensive List of Energy-Based Learning Papers and Materials.
Https://github.com/yataobian/awesome-ebm.
Bian, Yatao, Joachim M. Buhmann, and Andreas Krause. 2019.
“Optimal Continuous DR-Submodular Maximization and
Applications to Provable Mean Field Inference.” International
Conference on Machine Learning (ICML), 644–53.
Bian, Yatao, Yu Rong, Tingyang Xu, Jiaxiang Wu, Andreas Krause, and
Junzhou Huang. 2022. “Energy-Based Learning for Cooperative Games,
with Applications to Valuation Problems in Machine Learning.”
International Conference on Learning Representations (ICLR). https://openreview.net/forum?id=xLfAgCroImw.
Bilmes, Jeffrey A., and Wenruo Bai. 2017. “Deep Submodular
Functions.” arXiv Preprint arXiv:1701.08939.
Black, Kevin, Michael Janner, Yilun Du, Ilya Kostrikov, and Sergey
Levine. 2024. “Training Diffusion Models with Reinforcement
Learning.” International Conference on Learning
Representations (ICLR).
Bloem-Reddy, Benjamin, and Yee Whye Teh. 2020. “Probabilistic
Symmetries and Invariant Neural Networks.” Journal of Machine
Learning Research 21 (90): 1–61.
Boltzmann, Ludwig. 1877. “Über Die Beziehung Zwischen Dem Zweiten
Hauptsatze Der Mechanischen Wärmetheorie Und Der
Wahrscheinlichkeitsrechnung Respektive Den Sätzen Über Das
Wärmegleichgewicht.” Sitzungsberichte Der Kaiserlichen
Akademie Der Wissenschaften in Wien 76: 373–435.
Bradley, Ralph Allan, and Milton E. Terry. 1952. “Rank Analysis of
Incomplete Block Designs: I. The Method of Paired
Comparisons.” Biometrika 39 (3/4): 324–45.
Bridle, John S., Anthony J. R. Heading, and David J. C. MacKay. 1992.
“Unsupervised Classifiers, Mutual Information and ’Phantom
Targets’.” Advances in Neural Information Processing Systems
(NIPS) 4, 1096–101.
Burton, Didier, and Philippe L. Toint. 1992. “On an Instance of
the Inverse Shortest Paths Problem.” Mathematical
Programming 53: 45–61.
Calinescu, Gruia, Chandra Chekuri, Martin Pál, and Jan Vondrák. 2007.
“Maximizing a Submodular Set Function Subject to a Matroid
Constraint.” Integer Programming and Combinatorial
Optimization (IPCO), 182–96.
Callen, Herbert B. 1985. Thermodynamics and an Introduction to
Thermostatistics. 2nd ed. Wiley.
Carcamo, David P., Nicholas J. Weaver, Purushottam D. Dixit, and
Christopher W. Lynn. 2025. “Minimax Entropy: The Statistical
Physics of Optimal Models.” Physical Review E 112 (6):
061001. https://doi.org/10.1103/kr9x-q59y.
Caron, Mathilde, Ishan Misra, Julien Mairal, Priya Goyal, Piotr
Bojanowski, and Armand Joulin. 2020. “Unsupervised Learning of
Visual Features by Contrasting Cluster Assignments.” Advances
in Neural Information Processing Systems (NeurIPS). https://arxiv.org/abs/2006.09882.
Caron, Mathilde, Hugo Touvron, Ishan Misra, et al. 2021. “Emerging
Properties in Self-Supervised Vision Transformers.” IEEE/CVF
International Conference on Computer Vision (ICCV). https://arxiv.org/abs/2104.14294.
Carreira-Perpiñán, Miguel Á., and Geoffrey E. Hinton. 2005. “On
Contrastive Divergence Learning.” International Workshop on
Artificial Intelligence and Statistics (AISTATS), Proceedings of
machine learning research, vol. R5: 33–40.
Celeux, Gilles, and Gérard Govaert. 1992. “A Classification
EM Algorithm for Clustering and Two Stochastic
Versions.” Computational Statistics & Data Analysis
14 (3): 315–32.
Chen, Minhua, Badrinath Jayakumar, Padmasundari Gopalakrishnan, Qiming
Huang, Michael Johnston, and Patrick Haffner. 2021. “Deep
Clustering with Measure Propagation.” arXiv Preprint
arXiv:2104.08967.
Christiano, Paul F., Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg,
and Dario Amodei. 2017. “Deep Reinforcement Learning from Human
Preferences.” Advances in Neural Information Processing
Systems 30 (NeurIPS).
Clark, Kevin, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning.
2020a. “ELECTRA: Pre-Training Text Encoders as
Discriminators Rather Than Generators.” International
Conference on Learning Representations (ICLR).
Clark, Kevin, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning.
2020b. “Pre-Training Transformers as Energy-Based Cloze
Models.” Proceedings of the 2020 Conference on Empirical
Methods in Natural Language Processing (EMNLP), 285–94.
Clausius, Rudolf. 1865. “Über Verschiedene Für Die Anwendung
Bequeme Formen Der Hauptgleichungen Der Mechanischen
Wärmetheorie.” Annalen Der Physik Und Chemie 125:
353–400.
Cobbe, Karl, Vineet Kosaraju, Mohammad Bavarian, et al. 2021.
“Training Verifiers to Solve Math Word Problems.” arXiv
Preprint arXiv:2110.14168.
Cohen, Jacob. 1960. “A Coefficient of Agreement for Nominal
Scales.” Educational and Psychological Measurement 20
(1): 37–46. https://doi.org/10.1177/001316446002000104.
Collins, Michael. 2002. “Discriminative Training Methods for
Hidden Markov Models: Theory and Experiments with
Perceptron Algorithms.” Conference on Empirical Methods in
Natural Language Processing (EMNLP).
Csiszár, Imre. 1975. “I-Divergence Geometry of
Probability Distributions and Minimization Problems.” Annals
of Probability 3 (1): 146–58.
Csiszár, Imre. 2008. “Axiomatic Characterizations of Information
Measures.” Entropy 10 (3): 261–73. https://doi.org/10.3390/e10030261.
Cui, Ganqu, Yuchen Zhang, Jiacheng Chen, et al. 2025. “The Entropy
Mechanism of Reinforcement Learning for Reasoning Language
Models.” arXiv Preprint arXiv:2505.22617. https://arxiv.org/abs/2505.22617.
Dai, Bo, Zhen Liu, Hanjun Dai, et al. 2019. “Exponential Family
Estimation via Adversarial Dynamics Embedding.” Advances in
Neural Information Processing Systems (NeurIPS) 32.
Dai, Jifeng, Yang Lu, and Ying Nian Wu. 2015. “Generative Modeling
of Convolutional Neural Networks.” International Conference
on Learning Representations (ICLR).
Daróczy, Zoltán. 1970. “Generalized Information Functions.”
Information and Control 16 (1): 36–51. https://doi.org/10.1016/S0019-9958(70)80040-7.
Dawid, A. Philip. 2007. “The Geometry of Proper Scoring
Rules.” Annals of the Institute of Statistical
Mathematics 59 (1): 77–93. https://doi.org/10.1007/s10463-006-0099-8.
Dawid, Anna, and Yann LeCun. 2024. “Introduction to Latent
Variable Energy-Based Models: A Path Towards Autonomous Machine
Intelligence.” Journal of Statistical Mechanics: Theory and
Experiment 2024 (10): 104011.
Dayan, Peter, Geoffrey E. Hinton, Radford M. Neal, and Richard S. Zemel.
1995. “The Helmholtz Machine.” Neural Computation
7 (5): 889–904. https://doi.org/10.1162/neco.1995.7.5.889.
DeepSeek-AI. 2025. “DeepSeek-R1 Incentivizes Reasoning in LLMs
Through Reinforcement Learning.” Nature 645. https://doi.org/10.1038/s41586-025-09422-z.
Della Pietra, Stephen, Vincent Della Pietra, and John Lafferty. 1997.
“Inducing Features of Random Fields.” IEEE Transactions
on Pattern Analysis and Machine Intelligence 19 (4): 380–93. https://doi.org/10.1109/34.588021.
Deng, Yuntian, Anton Bakhtin, Myle Ott, Arthur Szlam, and Marc’Aurelio
Ranzato. 2020. “Residual Energy-Based Models for Text
Generation.” International Conference on Learning
Representations (ICLR).
Dhariwal, Prafulla, and Alexander Nichol. 2021. “Diffusion Models
Beat GANs on Image Synthesis.” Advances in
Neural Information Processing Systems (NeurIPS).
Domingo-Enrich, Carles, Michal Drozdzal, Brian Karrer, and Ricky T. Q.
Chen. 2025. “Adjoint Matching: Fine-Tuning Flow and Diffusion
Generative Models with Memoryless Stochastic Optimal Control.”
International Conference on Learning Representations (ICLR).
Domke, Justin. 2012. “Generic Methods for Optimization-Based
Modeling.” International Conference on Artificial
Intelligence and Statistics (AISTATS), 318–26.
Domke, Justin. 2013. “Learning Graphical Model Parameters with
Approximate Marginal Inference.” IEEE Transactions on Pattern
Analysis and Machine Intelligence 35 (10): 2454–67. https://doi.org/10.1109/TPAMI.2013.31.
Dong, Hanze, Wei Xiong, Deepanshu Goyal, et al. 2023.
“RAFT: Reward rAnked
FineTuning for Generative Foundation Model
Alignment.” Transactions on Machine Learning Research.
Du, Yilun, Conor Durkan, Robin Strudel, et al. 2023. “Reduce,
Reuse, Recycle: Compositional Generation with Energy-Based Diffusion
Models and MCMC.” Proceedings of the 40th
International Conference on Machine Learning (ICML), Proceedings of
machine learning research, vol. 202.
Du, Yilun, Shuang Li, and Igor Mordatch. 2020. “Compositional
Visual Generation with Energy Based Models.” Advances in
Neural Information Processing Systems (NeurIPS).
Du, Yilun, and Igor Mordatch. 2019. “Implicit Generation and
Modeling with Energy-Based Models.” Advances in Neural
Information Processing Systems (NeurIPS).
Dudı́k, Miroslav, Steven J. Phillips, and Robert E. Schapire. 2007.
“Maximum Entropy Density Estimation with Generalized
Regularization and an Application to Species Distribution
Modeling.” Journal of Machine Learning Research 8:
1217–60.
Efron, Bradley. 1975. “Defining the Curvature of a Statistical
Problem (with Applications to Second Order Efficiency).”
Annals of Statistics 3 (6): 1189–242. https://doi.org/10.1214/aos/1176343282.
Emmerich, Michael T. M., Ksenia Pereverdieva, and André H. Deutz. 2026a.
“Maximum Solow–Polasky Diversity Subset Selection Is NP-Hard Even
in the Euclidean Plane.” arXiv Preprint
arXiv:2604.19484. https://arxiv.org/abs/2604.19484.
Emmerich, Michael T. M., Ksenia Pereverdieva, and André H. Deutz. 2026b.
“Selecting a Maximum Solow–Polasky Diversity Subset in General
Metric Spaces Is NP-Hard.” arXiv Preprint
arXiv:2604.05495. https://arxiv.org/abs/2604.05495.
Faddeev, D. K. 1956. “On the Concept of Entropy of a Finite
Probabilistic Scheme.” Uspekhi Matematicheskikh Nauk 11
(1(67)): 227–31. https://www.mathnet.ru/eng/rm7756.
Fan, Ying, Olivia Watkins, Yuqing Du, et al. 2023.
“DPOK: Reinforcement Learning for Fine-Tuning
Text-to-Image Diffusion Models.” Advances in Neural
Information Processing Systems 36 (NeurIPS).
Farnia, Farzan, and David Tse. 2016. “A Minimax Approach to
Supervised Learning.” Advances in Neural Information
Processing Systems (NeurIPS) 29.
Farquhar, Sebastian, Jannik Kossen, Lorenz Kuhn, and Yarin Gal. 2024.
“Detecting Hallucinations in Large Language Models Using Semantic
Entropy.” Nature 630: 625–30. https://doi.org/10.1038/s41586-024-07421-0.
Feder, M. 1986. “Maximum Entropy as a Special Case of the Minimum
Description Length Criterion.” IEEE Transactions on
Information Theory 32: 847–49.
Feng, Ruiqi, Chenglei Yu, Wenhao Deng, Peiyan Hu, and Tailin Wu. 2025.
“On the Guidance of Flow Matching.” Proceedings of the
42nd International Conference on Machine Learning (ICML),
Proceedings of machine learning research, vol. 267: 16993–7029.
Fox, Roy, Ari Pakman, and Naftali Tishby. 2016. “Taming the Noise
in Reinforcement Learning via Soft Updates.” Conference on
Uncertainty in Artificial Intelligence (UAI).
Frieden, B. Roy. 1990. “Fisher Information, Disorder, and the
Equilibrium Distributions of Physics.” Physical Review A
41 (8): 4265–76. https://doi.org/10.1103/PhysRevA.41.4265.
Friedman, Dan, and Adji Bousso Dieng. 2023. “The Vendi Score: A
Diversity Evaluation Metric for Machine Learning.”
Transactions on Machine Learning Research. https://arxiv.org/abs/2210.02410.
Friston, Karl. 2010. “The Free-Energy Principle: A Unified Brain
Theory?” Nature Reviews Neuroscience 11 (2): 127–38.
Geiger, Dan, David Heckerman, Henry King, and Christopher Meek. 2001.
“Stratified Exponential Families: Graphical Models and Model
Selection.” The Annals of Statistics 29 (2): 505–29. https://doi.org/10.1214/aos/1009210550.
Geist, Matthieu, Bruno Scherrer, and Olivier Pietquin. 2019. “A
Theory of Regularized Markov Decision Processes.”
International Conference on Machine Learning (ICML).
Gibbs, Josiah Willard. 1902. Elementary Principles in Statistical
Mechanics. Charles Scribner’s Sons.
Gladstone, Alexi, Ganesh Nanduru, Md Mofijul Islam, et al. 2026.
“Energy-Based Transformers Are Scalable Learners and
Thinkers.” International Conference on Learning
Representations (ICLR).
Gnaiger, Erich. 2009. “Open and Closed Systems: Styles of Thinking
Explain Controversies on the ’Negative Entropy’ Concept of Ludwig
Boltzmann and Erwin Schrödinger.” Mitochondrial Physiology
Network. https://www.mitophysiology.org/images/2/23/Gnaiger_2009_OCESHS.pdf.
Go, Dongyoung, Tomasz Korbak, Germán Kruszewski, Jos Rozen, Nahyeon Ryu,
and Marc Dymetman. 2023. “Aligning Language Models with
Preferences Through f-Divergence Minimization.”
Proceedings of the 40th International Conference on Machine Learning
(ICML), Proceedings of machine learning research, vol. 202.
Gomes, Ryan, Andreas Krause, and Pietro Perona. 2010.
“Discriminative Clustering by Regularized Information
Maximization.” Advances in Neural Information Processing
Systems (NeurIPS) 23.
Grandvalet, Yves, and Yoshua Bengio. 2004. “Semi-Supervised
Learning by Entropy Minimization.” Advances in Neural
Information Processing Systems (NeurIPS) 17.
Grandvalet, Yves, and Yoshua Bengio. 2006. “Entropy
Regularization.” In Semi-Supervised Learning, edited by
Olivier Chapelle, Bernhard Schölkopf, and Alexander Zien. MIT Press.
Grathwohl, Will, Kevin Swersky, Milad Hashemi, David Duvenaud, and Chris
Maddison. 2021. “Oops I Took a Gradient: Scalable
Sampling for Discrete Distributions.” International
Conference on Machine Learning (ICML), 3831–41.
Grathwohl, Will, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud,
Mohammad Norouzi, and Kevin Swersky. 2020. “Your Classifier Is
Secretly an Energy Based Model and You Should Treat It Like One.”
International Conference on Learning Representations (ICLR).
Gregor, Karol, and Yann LeCun. 2010. “Learning Fast Approximations
of Sparse Coding.” International Conference on Machine
Learning (ICML), 399–406.
Grünwald, Peter D. 2007. The Minimum Description Length
Principle. MIT Press.
Grünwald, Peter D., and A. Philip Dawid. 2004. “Game Theory,
Maximum Entropy, Minimum Discrepancy and Robust Bayesian
Decision Theory.” Annals of Statistics 32 (4): 1367–433.
Gui, Lin, Cristina Gârbacea, and Victor Veitch. 2024.
“BoNBoN Alignment for Large Language Models and the
Sweetness of Best-of-n
Sampling.” Advances in Neural Information Processing Systems
37 (NeurIPS).
Gulcehre, Caglar, Tom Le Paine, Srivatsan Srinivasan, et al. 2023.
“Reinforced Self-Training (ReST) for Language
Modeling.” arXiv Preprint arXiv:2308.08998.
Gutmann, Michael U., and Aapo Hyvärinen. 2012. “Noise-Contrastive
Estimation of Unnormalized Statistical Models, with Applications to
Natural Image Statistics.” Journal of Machine Learning
Research 13: 307–61.
Gutmann, Michael, and Jun-ichiro Hirayama. 2011. “Bregman
Divergence as General Framework to Estimate Unnormalized Statistical
Models.” Conference on Uncertainty in Artificial Intelligence
(UAI), 283–90.
Gutmann, Michael, and Aapo Hyvärinen. 2010. “Noise-Contrastive
Estimation: A New Estimation Principle for Unnormalized Statistical
Models.” International Conference on Artificial Intelligence
and Statistics (AISTATS), 297–304.
GX-Chen, Anthony, Jatin Prakash, Jeff Guo, Rob Fergus, and Rajesh
Ranganath. 2025. “KL-Regularized Reinforcement Learning Is
Designed to Mode Collapse.” arXiv Preprint
arXiv:2510.20817. https://arxiv.org/abs/2510.20817.
Haarnoja, Tuomas, Haoran Tang, Pieter Abbeel, and Sergey Levine. 2017.
“Reinforcement Learning with Deep Energy-Based Policies.”
International Conference on Machine Learning (ICML), 1352–61.
Haarnoja, Tuomas, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018.
“Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement
Learning with a Stochastic Actor.” International Conference
on Machine Learning (ICML), 1861–70.
Hanel, Rudolf, and Stefan Thurner. 2011. “A Comprehensive
Classification of Complex Statistical Systems and an Axiomatic
Derivation of Their Entropy and Distribution Functions.” EPL
(Europhysics Letters) 93 (2): 20006. https://doi.org/10.1209/0295-5075/93/20006.
Hassani, Hamed, Mahdi Soltanolkotabi, and Amin Karbasi. 2017.
“Gradient Methods for Submodular Maximization.”
Advances in Neural Information Processing Systems (NeurIPS) 30.
Hathaway, Richard J. 1986. “Another Interpretation of the
EM Algorithm for Mixture Distributions.”
Statistics & Probability Letters 4 (2): 53–56. https://doi.org/10.1016/0167-7152(86)90016-7.
Havrda, Jan, and František Charvát. 1967. “Quantification Method
of Classification Processes. Concept of Structural a-Entropy.”
Kybernetika 3 (1): 30–35. https://www.kybernetika.cz/content/1967/1/30.
Hershey, John R., Jonathan Le Roux, and Felix Weninger. 2014.
“Deep Unfolding: Model-Based Inspiration of Novel Deep
Architectures.” arXiv Preprint arXiv:1409.2574.
Heuberger, Clemens. 2004. “Inverse Combinatorial Optimization: A
Survey on Problems, Methods, and Results.” Journal of
Combinatorial Optimization 8 (3): 329–61.
Hill, M. O. 1973. “Diversity and Evenness: A Unifying Notation and
Its Consequences.” Ecology 54 (2): 427–32. https://doi.org/10.2307/1934352.
Hinton, Geoffrey E. 2002. “Training Products of Experts by
Minimizing Contrastive Divergence.” Neural Computation
14 (8): 1771–800. https://doi.org/10.1162/089976602760128018.
Hirsh, Jacob B., Raymond A. Mar, and Jordan B. Peterson. 2012.
“Psychological Entropy: A Framework for Understanding
Uncertainty-Related Anxiety.” Psychological Review 119
(2): 304–20. https://doi.org/10.1037/a0026767.
Ho, Jonathan, Ajay Jain, and Pieter Abbeel. 2020. “Denoising
Diffusion Probabilistic Models.” Advances in Neural
Information Processing Systems (NeurIPS).
Ho, Jonathan, and Tim Salimans. 2022. “Classifier-Free Diffusion
Guidance.” arXiv Preprint arXiv:2207.12598.
Hopfield, John J. 1982. “Neural Networks and Physical Systems with
Emergent Collective Computational Abilities.” Proceedings of
the National Academy of Sciences 79 (8): 2554–58. https://doi.org/10.1073/pnas.79.8.2554.
Hu, Edward J., Moksh Jain, Eric Elmoznino, et al. 2024.
“Amortizing Intractable Inference in Large Language
Models.” International Conference on Learning Representations
(ICLR).
Hu, Weihua, Takeru Miyato, Seiya Tokui, Eiichi Matsumoto, and Masashi
Sugiyama. 2017. “Learning Discrete Representations via Information
Maximizing Self-Augmented Training.” International Conference
on Machine Learning (ICML). https://arxiv.org/abs/1702.08720.
Hyvärinen, Aapo. 2005. “Estimation of Non-Normalized Statistical
Models by Score Matching.” Journal of Machine Learning
Research 6: 695–709.
Hyvärinen, Aapo. 2007. “Connections Between Score Matching,
Contrastive Divergence, and Pseudolikelihood for Continuous-Valued
Variables.” IEEE Transactions on Neural Networks 18 (5):
1529–31.
Jaakkola, Tommi, Marina Meila, and Tony Jebara. 1999. “Maximum
Entropy Discrimination.” Advances in Neural Information
Processing Systems (NeurIPS) 12: 470–76.
Jalali, Mohammad, Cheuk Ting Li, and Farzan Farnia. 2023. “An
Information-Theoretic Evaluation of Generative Models in Learning
Multi-Modal Distributions.” Advances in Neural Information
Processing Systems (NeurIPS) 36: 9931–43. https://doi.org/10.52202/075280-0434.
Jaques, Natasha, Shixiang Gu, Dzmitry Bahdanau, José Miguel
Hernández-Lobato, Richard E. Turner, and Douglas Eck. 2017.
“Sequence Tutor: Conservative Fine-Tuning of Sequence Generation
Models with KL-Control.” Proceedings of the 34th
International Conference on Machine Learning (ICML), Proceedings of
machine learning research, vol. 70: 1645–54.
Jaynes, Edwin T. 1957. “Information Theory and Statistical
Mechanics.” Physical Review 106: 620.
Jaynes, Edwin T. 1968. “Prior Probabilities.” IEEE
Transactions on Systems Science and Cybernetics 4 (3): 227–41.
Jaynes, Edwin T. 1980. “The Minimum Entropy Production
Principle.” Annual Review of Physical Chemistry 31:
579–601. https://doi.org/10.1146/annurev.pc.31.100180.003051.
Jiang, Eric H., Haozheng Luo, Shengyuan Pang, et al. 2025.
“Learning to Rank Chain-of-Thought: Using a Small Model.”
arXiv Preprint arXiv:2505.14999.
Jin, Long, Justin Lazarow, and Zhuowen Tu. 2017. “Introspective
Classification with Convolutional Nets.” Advances in Neural
Information Processing Systems (NeurIPS) 30.
Jizba, Petr, and Toshihico Arimitsu. 2004. “The World According to
Rényi: Thermodynamics of Multifractal Systems.” Annals of
Physics 312 (1): 17–59. https://doi.org/10.1016/j.aop.2004.01.002.
Jordan, Michael I., Zoubin Ghahramani, Tommi S. Jaakkola, and Lawrence
K. Saul. 1999. “An Introduction to Variational Methods for
Graphical Models.” Machine Learning 37 (2): 183–233.
Jost, Lou. 2006. “Entropy and Diversity.” Oikos
113 (2): 363–75. https://doi.org/10.1111/j.2006.0030-1299.14714.x.
Jost, Lou. 2009. “Mismeasuring Biological Diversity: Response to
Hoffmann and Hoffmann (2008).”
Ecological Economics 68 (4): 925–28. https://doi.org/10.1016/j.ecolecon.2008.10.015.
Jost, Lou, Philip DeVries, Thomas Walla, Harold Greeney, Anne Chao, and
Carlo Ricotta. 2010. “Partitioning Diversity for Conservation
Analyses.” Diversity and Distributions 16 (1): 65–76. https://doi.org/10.1111/j.1472-4642.2009.00626.x.
Kappen, Hilbert J. 2005. “Path Integrals and Symmetry Breaking for
Optimal Control Theory.” Journal of Statistical Mechanics:
Theory and Experiment 2005 (11): P11011.
Kappen, Hilbert J., Vicenç Gómez, and Manfred Opper. 2012.
“Optimal Control as a Graphical Model Inference Problem.”
Machine Learning 87 (2): 159–82.
Karalias, Nikolaos, Joshua Robinson, Andreas Loukas, and Stefanie
Jegelka. 2022. “Neural Set Function Extensions: Learning with
Discrete Functions in High Dimensions.” Advances in Neural
Information Processing Systems (NeurIPS) 35: 15338–52.
Khalifa, Muhammad, Hady Elsahar, and Marc Dymetman. 2021. “A
Distributional Approach to Controlled Text Generation.”
International Conference on Learning Representations (ICLR).
Khinchin, A. I. 1957. Mathematical Foundations of Information
Theory. Dover Publications.
Kianercy, Ardeshir, and Aram Galstyan. 2012. “Dynamics of
Boltzmann q Learning in Two-Player Two-Action Games.”
Physical Review E 85 (4): 041145. https://doi.org/10.1103/PhysRevE.85.041145.
Kim, Taesup, and Yoshua Bengio. 2016. “Deep Directed Generative
Models with Energy-Based Probability Estimation.” arXiv
Preprint arXiv:1606.03439.
Kirk, Robert, Ishita Mediratta, Christoforos Nalmpantis, et al. 2024.
“Understanding the Effects of RLHF on LLM Generalisation and
Diversity.” International Conference on Learning
Representations (ICLR), 20620–53.
Kirkpatrick, Scott, C. Daniel Gelatt, and Mario P. Vecchi. 1983.
“Optimization by Simulated Annealing.” Science 220
(4598): 671–80. https://doi.org/10.1126/science.220.4598.671.
Koller, Daphne, and Nir Friedman. 2009. Probabilistic Graphical
Models: Principles and Techniques. MIT Press.
Korbak, Tomasz, Ethan Perez, and Christopher L. Buckley. 2022.
“RL with KL Penalties Is Better Viewed
as Bayesian Inference.” Findings of the
Association for Computational Linguistics: EMNLP 2022.
Korbel, Jan. 2026. “Foundations of Entropy in Complex
Systems.” arXiv Preprint arXiv:2606.16312. https://arxiv.org/abs/2606.16312.
Krähenbühl, Philipp, and Vladlen Koltun. 2013. “Parameter Learning
and Convergent Inference for Dense Random Fields.”
International Conference on Machine Learning (ICML), 513–21.
Kuhn, Lorenz, Yarin Gal, and Sebastian Farquhar. 2023. “Semantic
Uncertainty: Linguistic Invariances for Uncertainty Estimation in
Natural Language Generation.” International Conference on
Learning Representations (ICLR). https://arxiv.org/abs/2302.09664.
Kullback, Solomon. 1959. Information Theory and Statistics.
Wiley.
Kumar, Rithesh, Sherjil Ozair, Anirudh Goyal, Aaron Courville, and
Yoshua Bengio. 2019. “Maximum Entropy Generators for Energy-Based
Models.” arXiv Preprint arXiv:1901.08508.
Lafferty, John, Andrew McCallum, and Fernando Pereira. 2001.
“Conditional Random Fields: Probabilistic Models for Segmenting
and Labeling Sequence Data.” International Conference on
Machine Learning (ICML), 282–89.
Landauer, Rolf. 1975. “Inadequacy of Entropy and Entropy
Derivatives in Characterizing the Steady State.” Physical
Review A 12 (2): 636–38. https://doi.org/10.1103/PhysRevA.12.636.
Lasserre, Julia A., Christopher M. Bishop, and Thomas P. Minka. 2006.
“Principled Hybrids of Generative and Discriminative
Models.” IEEE Conference on Computer Vision and Pattern
Recognition (CVPR), 87–94.
Lazarev, Daniel. 2026. “A Structural Characterization of Entropy
Functionals.” arXiv Preprint arXiv:2608.13917. https://arxiv.org/abs/2608.13917.
LeCun, Yann. 2022. “A Path Towards Autonomous Machine
Intelligence.” OpenReview Preprint.
LeCun, Yann, Sumit Chopra, Raia Hadsell, Marc’Aurelio Ranzato, and Fu
Jie Huang. 2006. “A Tutorial on Energy-Based Learning.” In
Predicting Structured Data, edited by Gökhan Bakir, Thomas
Hofmann, Bernhard Schölkopf, Alexander J. Smola, and Ben Taskar. MIT
Press.
Lee, Dong-Hyun. 2013. “Pseudo-Label: The Simple and Efficient
Semi-Supervised Learning Method for Deep Neural Networks.”
ICML Workshop on Challenges in Representation Learning.
Lee, Jonghyun, Dahuin Jung, Saehyung Lee, et al. 2024. “Entropy Is
Not Enough for Test-Time Adaptation: From the Perspective of
Disentangled Factors.” International Conference on Learning
Representations (ICLR). https://arxiv.org/abs/2403.07366.
Leinster, Tom. 2009. “A Maximum Entropy Theorem with Applications
to the Measurement of Biodiversity.” arXiv Preprint
arXiv:0910.0906. https://arxiv.org/abs/0910.0906.
Leinster, Tom. 2021. Entropy and Diversity: The Axiomatic
Approach. Cambridge University Press. https://doi.org/10.1017/9781108963558.
Leinster, Tom, and Christina A. Cobbold. 2012. “Measuring
Diversity: The Importance of Species Similarity.”
Ecology 93 (3): 477–89. https://doi.org/10.1890/10-2402.1.
Leinster, Tom, and Mark W. Meckes. 2016. “Maximizing Diversity in
Biology and Beyond.” Entropy 18 (3): 88. https://doi.org/10.3390/e18030088.
Levine, Sergey. 2018. “Reinforcement Learning and Control as
Probabilistic Inference: Tutorial and Review.” arXiv Preprint
arXiv:1805.00909.
Liang, Jian, Dapeng Hu, and Jiashi Feng. 2020. “Do We Really Need
to Access the Source Data? Source Hypothesis Transfer for Unsupervised
Domain Adaptation.” International Conference on Machine
Learning (ICML). https://arxiv.org/abs/2002.08546.
Lieb, Elliott H., and Jakob Yngvason. 1999. “The Physics and
Mathematics of the Second Law of Thermodynamics.” Physics
Reports 310 (1): 1–96. https://doi.org/10.1016/S0370-1573(98)00082-9.
Lightman, Hunter, Vineet Kosaraju, Yuri Burda, et al. 2024. “Let’s
Verify Step by Step.” International Conference on Learning
Representations (ICLR).
Lin, Hui, and Jeff A. Bilmes. 2012. “Learning Mixtures of
Submodular Shells with Application to Document Summarization.”
Conference on Uncertainty in Artificial Intelligence (UAI),
479–90.
Liu, Weitang, Xiaoyun Wang, John D. Owens, and Yixuan Li. 2020.
“Energy-Based Out-of-Distribution Detection.” Advances
in Neural Information Processing Systems (NeurIPS).
Lu, Cheng, Huayu Chen, Jianfei Chen, Hang Su, Chongxuan Li, and Jun Zhu.
2023. “Contrastive Energy Prediction for Exact Energy-Guided
Diffusion Sampling in Offline Reinforcement Learning.”
Proceedings of the 40th International Conference on Machine Learning
(ICML), Proceedings of machine learning research, vol. 202.
Lynn, Christopher W., Qiwei Yu, Rich Pang, Stephanie E. Palmer, and
William Bialek. 2025. “Exact Minimax Entropy Models of Large-Scale
Neuronal Activity.” Physical Review E 111 (5): 054411.
https://doi.org/10.1103/PhysRevE.111.054411.
Lyu, Siwei. 2009. “Interpretation and Generalization of Score
Matching.” Conference on Uncertainty in Artificial
Intelligence (UAI), 359–66.
Ma, Xueguang, Qian Liu, Dongfu Jiang, Ge Zhang, Zejun Ma, and Wenhu
Chen. 2025. “General-Reasoner: Advancing LLM Reasoning Across All
Domains.” arXiv Preprint arXiv:2505.14652. https://arxiv.org/abs/2505.14652.
MacArthur, Robert H. 1965. “Patterns of Species Diversity.”
Biological Reviews 40 (4): 510–33. https://doi.org/10.1111/j.1469-185X.1965.tb00815.x.
Maes, Christian, and Karel Netočný. 2007. “Minimum Entropy
Production Principle from a Dynamical Fluctuation Law.”
Journal of Mathematical Physics 48 (5): 053306. https://doi.org/10.1063/1.2738753.
Maes, Christian, and Karel Netočný. 2013. “Minimum Entropy
Production Principle.” Scholarpedia 8 (7): 9664. https://doi.org/10.4249/scholarpedia.9664.
Martyushev, Leonid M., A. S. Nazarova, and Vladimir D. Seleznev. 2007.
“On the Problem of the Minimum Entropy Production in the
Nonequilibrium Stationary State.” Journal of Physics A:
Mathematical and Theoretical 40 (3): 371–80. https://doi.org/10.1088/1751-8113/40/3/002.
Martyushev, Leonid M., and Vladimir D. Seleznev. 2006. “Maximum
Entropy Production Principle in Physics, Chemistry and Biology.”
Physics Reports 426 (1): 1–45. https://doi.org/10.1016/j.physrep.2005.12.001.
Martyushev, Leonid M., and Vladimir D. Seleznev. 2013. “Entropy
and Entropy Production: Old Misconceptions and New
Breakthroughs.” Entropy 15 (4): 1152–70. https://doi.org/10.3390/e15041152.
Menon, Aditya Krishna, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu
Jain, Andreas Veit, and Sanjiv Kumar. 2021. “Long-Tail Learning
via Logit Adjustment.” International Conference on Learning
Representations (ICLR). https://arxiv.org/abs/2007.07314.
Mironov, Mikhail, and Liudmila Prokhorenkova. 2025. “Measuring
Diversity: Axioms and Challenges.” Proceedings of the 42nd
International Conference on Machine Learning 267: 44396–411. https://proceedings.mlr.press/v267/mironov25a.html.
Mnih, Andriy, and Yee Whye Teh. 2012. “A Fast and Simple Algorithm
for Training Neural Probabilistic Language Models.”
International Conference on Machine Learning (ICML), 419–26.
Moussouris, John. 1974. “Gibbs and Markov Random Systems with
Constraints.” Journal of Statistical Physics 10 (1):
11–33. https://doi.org/10.1007/BF01011714.
Mudgal, Sidharth, Jong Lee, Harish Ganapathy, et al. 2024.
“Controlled Decoding from Language Models.” Proceedings
of the 41st International Conference on Machine Learning (ICML),
Proceedings of machine learning research, vol. 235.
Neal, Radford M., and Geoffrey E. Hinton. 1998. “A View of the EM
Algorithm That Justifies Incremental, Sparse, and Other
Variants.” In Learning in Graphical Models, edited by
Michael I. Jordan. Kluwer Academic Publishers.
Nehring, Klaus, and Clemens Puppe. 2002. “A Theory of
Diversity.” Econometrica 70 (3): 1155–98. https://doi.org/10.1111/1468-0262.00321.
Nemhauser, George L., Laurence A. Wolsey, and Marshall L. Fisher. 1978.
“An Analysis of Approximations for Maximizing Submodular Set
Functions—i.” Mathematical Programming 14: 265–94.
Neu, Gergely, Anders Jonsson, and Vicenç Gómez. 2017. “A Unified
View of Entropy-Regularized Markov Decision
Processes.” arXiv Preprint arXiv:1705.07798.
Neumann, John von. 1927. “Thermodynamik Quantenmechanischer
Gesamtheiten.” Nachrichten von Der Gesellschaft Der
Wissenschaften Zu Göttingen, Mathematisch-Physikalische Klasse
1927: 273–91.
Ng, Andrew Y., and Stuart J. Russell. 2000. “Algorithms for
Inverse Reinforcement Learning.” International Conference on
Machine Learning (ICML).
Ngiam, Jiquan, Zhenghao Chen, Pang Wei Koh, and Andrew Y. Ng. 2011.
“Learning Deep Energy Models.” International Conference
on Machine Learning (ICML), 1105–12.
Nguyen, Quan, and Adji Bousso Dieng. 2024. “Quality-Weighted Vendi
Scores and Their Application to Diverse Experimental Design.”
Proceedings of the 41st International Conference on Machine
Learning 235: 37667–82. https://proceedings.mlr.press/v235/nguyen24d.html.
Nijkamp, Erik, Mitch Hill, Tian Han, Song-Chun Zhu, and Ying Nian Wu.
2020. “On the Anatomy of MCMC-Based Maximum
Likelihood Learning of Energy-Based Models.” AAAI Conference
on Artificial Intelligence 34: 5272–80. https://doi.org/10.1609/aaai.v34i04.5973.
Nijkamp, Erik, Mitch Hill, Song-Chun Zhu, and Ying Nian Wu. 2019.
“Learning Non-Convergent Non-Persistent Short-Run
MCMC Toward Energy-Based Model.” Advances in
Neural Information Processing Systems (NeurIPS).
Nikitin, Alexander, Jannik Kossen, Yarin Gal, and Pekka Marttinen. 2024.
“Kernel Language Entropy: Fine-Grained Uncertainty Quantification
for LLMs from Semantic Similarities.” Advances
in Neural Information Processing Systems (NeurIPS). https://arxiv.org/abs/2405.20003.
Niu, Shuaicheng, Jiaxiang Wu, Yifan Zhang, et al. 2022. “Efficient
Test-Time Model Adaptation Without Forgetting.” International
Conference on Machine Learning (ICML). https://arxiv.org/abs/2204.02610.
Niu, Shuaicheng, Jiaxiang Wu, Yifan Zhang, et al. 2023. “Towards
Stable Test-Time Adaptation in Dynamic Wild World.”
International Conference on Learning Representations (ICLR). https://arxiv.org/abs/2302.12400.
Ochs, W. 1975. “A New Axiomatic Characterization of the von
Neumann Entropy.” Reports on Mathematical Physics 8 (1):
109–20. https://doi.org/10.1016/0034-4877(75)90022-1.
Ospanov, Azim, Jingwei Zhang, Mohammad Jalali, Xuenan Cao, Andrej
Bogdanov, and Farzan Farnia. 2024. “Towards a Scalable
Reference-Free Evaluation of Generative Models.” Advances in
Neural Information Processing Systems (NeurIPS) 37: 120892–927. https://doi.org/10.52202/079017-3841.
Ou, Zijing, Tingyang Xu, Qinliang Su, Yingzhen Li, Peilin Zhao, and
Yatao Bian. 2022. “Learning Neural Set Functions Under the Optimal
Subset Oracle.” Advances in Neural Information Processing
Systems (NeurIPS) 35: 35021–34.
Ouyang, Long, Jeff Wu, Xu Jiang, et al. 2022. “Training Language
Models to Follow Instructions with Human Feedback.” Advances
in Neural Information Processing Systems 35 (NeurIPS).
Owen, Guillermo. 1972. “Multilinear Extensions of Games.”
Management Science 18 (5): P64–79. https://doi.org/10.1287/mnsc.18.5.P64.
Owen, Guillermo. 1975. “Multilinear Extensions and the Banzhaf
Value.” Naval Research Logistics Quarterly 22 (4):
741–50. https://doi.org/10.1002/nav.3800220409.
Özcan, Gözde, Chengzhi Shi, and Stratis Ioannidis. 2025. “Learning
Set Functions with Implicit Differentiation.” AAAI Conference
on Artificial Intelligence, 19777–85.
Paltridge, Garth W. 1975. “Global Dynamics and Climate — a System
of Minimum Entropy Exchange.” Quarterly Journal of the Royal
Meteorological Society 101 (429): 475–84. https://doi.org/10.1002/qj.49710142906.
Paninski, Liam. 2003. “Estimation of Entropy and Mutual
Information.” Neural Computation 15 (6): 1191–253. https://doi.org/10.1162/089976603321780272.
Park, Mincheol, Heeji Won, Won Woo Ro, and Suhyun Kim. 2025.
“Rethinking Entropy in Test-Time Adaptation: The Missing Piece
from Energy Duality.” Advances in Neural Information
Processing Systems 38 (NeurIPS).
Parry, Matthew, A. Philip Dawid, and Steffen Lauritzen. 2012.
“Proper Local Scoring Rules.” Annals of Statistics
40 (1): 561–92. https://doi.org/10.1214/12-AOS971.
Parshakova, Tetiana, Jean-Marc Andreoli, and Marc Dymetman. 2019.
“Distributional Reinforcement Learning for Energy-Based Sequential
Models.” Optimization Foundations for Reinforcement Learning
Workshop at NeurIPS 2019.
Pasarkar, Amey P., and Adji Bousso Dieng. 2024. “Cousins of the
Vendi Score: A Family of Similarity-Based Diversity Metrics for Science
and Machine Learning.” International Conference on Artificial
Intelligence and Statistics (AISTATS), PMLR, vol. 238. https://arxiv.org/abs/2310.12952.
Patil, G. P., and C. Taillie. 1982. “Diversity as a Concept and
Its Measurement.” Journal of the American Statistical
Association 77 (379): 548–61. https://doi.org/10.1080/01621459.1982.10477845.
Pearl, Judea. 1988. Probabilistic Reasoning in Intelligent Systems:
Networks of Plausible Inference. Morgan Kaufmann.
Pereyra, Gabriel, George Tucker, Jan Chorowski, Łukasz Kaiser, and
Geoffrey Hinton. 2017. “Regularizing Neural Networks by Penalizing
Confident Output Distributions.” ICLR 2017 Workshop
Track.
Prigogine, Ilya. 1945. “Modération Et Transformations
Irréversibles Des Systèmes Ouverts.” Bulletin de La Classe
Des Sciences, Académie Royale de Belgique 31: 600–606.
Prigogine, Ilya. 1947. Étude Thermodynamique Des Phénomènes
Irréversibles. Dunod, Paris; Desoer, Liège.
Prigogine, Ilya. 1978. “Time, Structure, and Fluctuations.”
Science 201 (4358): 777–85. https://doi.org/10.1126/science.201.4358.777.
Prigogine, Ilya, and Jean-Marie Wiame. 1946. “Biologie Et
Thermodynamique Des Phénomènes Irréversibles.”
Experientia 2 (11): 451–53. https://doi.org/10.1007/BF02153597.
Qin, Lianhui, Sean Welleck, Daniel Khashabi, and Yejin Choi. 2022.
“COLD Decoding: Energy-Based Constrained Text
Generation with Langevin Dynamics.” Advances in
Neural Information Processing Systems (NeurIPS).
Rafailov, Rafael, Joey Hejna, Ryan Park, and Chelsea Finn. 2024.
“From r to Q*: Your Language
Model Is Secretly a Q-Function.” arXiv Preprint
arXiv:2404.12358.
Rafailov, Rafael, Archit Sharma, Eric Mitchell, Stefano Ermon,
Christopher D. Manning, and Chelsea Finn. 2023. “Direct Preference
Optimization: Your Language Model Is Secretly a Reward Model.”
Advances in Neural Information Processing Systems 36 (NeurIPS).
Rao, C. Radhakrishna. 1982. “Diversity and Dissimilarity
Coefficients: A Unified Approach.” Theoretical Population
Biology 21 (1): 24–43. https://doi.org/10.1016/0040-5809(82)90004-1.
Ratliff, Nathan D., J. Andrew Bagnell, and Martin A. Zinkevich. 2006.
“Maximum Margin Planning.” International Conference on
Machine Learning (ICML), 729–36.
Rényi, Alfréd. 1961. “On Measures of Entropy and
Information.” Proceedings of the Fourth Berkeley Symposium on
Mathematical Statistics and Probability, Volume 1: Contributions to the
Theory of Statistics, 547–61.
Rose, Kenneth. 1998. “Deterministic Annealing for Clustering,
Compression, Classification, Regression, and Related Optimization
Problems.” Proceedings of the IEEE 86 (11): 2210–39.
Rose, Kenneth, Eitan Gurewitz, and Geoffrey C. Fox. 1990.
“Statistical Mechanics and Phase Transitions in
Clustering.” Physical Review Letters 65 (8): 945–48. https://doi.org/10.1103/PhysRevLett.65.945.
Sahin, Aytunc, Yatao Bian, Joachim M. Buhmann, and Andreas Krause. 2020.
“From Sets to Multisets: Provable Variational Inference for
Probabilistic Integer Submodular Models.” International
Conference on Machine Learning (ICML), 8388–97.
Saito, Kuniaki, Donghyun Kim, Stan Sclaroff, Trevor Darrell, and Kate
Saenko. 2019. “Semi-Supervised Domain Adaptation via Minimax
Entropy.” IEEE/CVF International Conference on Computer
Vision (ICCV), 8050–58.
Sánchez Giraldo, Luis Gonzalo, Murali Rao, and José C. Príncipe. 2015.
“Measures of Entropy from Data Using Infinitely Divisible
Kernels.” IEEE Transactions on Information Theory 61
(1): 535–48. https://doi.org/10.1109/TIT.2014.2370058.
Santos, Roberto J. V. dos. 1997. “Generalization of Shannon’s
Theorem for Tsallis Entropy.” Journal of Mathematical
Physics 38 (8): 4104–7. https://doi.org/10.1063/1.532107.
Schrödinger, Erwin. 1944. What Is Life? The Physical Aspect of the
Living Cell. Cambridge University Press.
Schulman, John, Sergey Levine, Pieter Abbeel, Michael I. Jordan, and
Philipp Moritz. 2015. “Trust Region Policy Optimization.”
International Conference on Machine Learning (ICML).
Schulman, John, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg
Klimov. 2017. “Proximal Policy Optimization Algorithms.”
arXiv Preprint arXiv:1707.06347.
Shannon, Claude E. 1948. “A Mathematical Theory of
Communication.” Bell System Technical Journal 27 (3):
379–423.
Shao, Zhihong, Peiyi Wang, Qihao Zhu, et al. 2024. “DeepSeekMath:
Pushing the Limits of Mathematical Reasoning in Open Language
Models.” arXiv Preprint arXiv:2402.03300. https://arxiv.org/abs/2402.03300.
Shi, Chence, Shitong Luo, Minkai Xu, and Jian Tang. 2021.
“Learning Gradient Fields for Molecular Conformation
Generation.” International Conference on Machine Learning
(ICML).
Shi, Yongquan, Zijing Ou, Shiping Wang, and Yatao Bian. 2026.
“Advancing Optimal Subset Oracle via Learning Relaxation of Neural
Set Functions.” arXiv Preprint arXiv:2607.11555.
Shirts, Michael R., Eric Bair, Giles Hooker, and Vijay S. Pande. 2003.
“Equilibrium Free Energies from Nonequilibrium Measurements Using
Maximum-Likelihood Methods.” Physical Review Letters 91:
140601. https://doi.org/10.1103/PhysRevLett.91.140601.
Shore, John E., and Rodney W. Johnson. 1980. “Axiomatic Derivation
of the Principle of Maximum Entropy and the Principle of Minimum
Cross-Entropy.” IEEE Transactions on Information Theory
26 (1): 26–37.
Simpson, E. H. 1949. “Measurement of Diversity.”
Nature 163 (4148): 688. https://doi.org/10.1038/163688a0.
Sipos, Ruben, Pannaga Shivaswamy, and Thorsten Joachims. 2012.
“Large-Margin Learning of Submodular Summarization Models.”
Conference of the European Chapter of the Association for
Computational Linguistics (EACL), 224–33.
Sohl-Dickstein, Jascha, Eric Weiss, Niru Maheswaranathan, and Surya
Ganguli. 2015. “Deep Unsupervised Learning Using Nonequilibrium
Thermodynamics.” International Conference on Machine Learning
(ICML), 2256–65.
Solow, Andrew R., and Stephen Polasky. 1994. “Measuring Biological
Diversity.” Environmental and Ecological Statistics 1
(2): 95–103. https://doi.org/10.1007/BF02426650.
Song, Yang, Conor Durkan, Iain Murray, and Stefano Ermon. 2021.
“Maximum Likelihood Training of Score-Based Diffusion
Models.” Advances in Neural Information Processing Systems
(NeurIPS) 34.
Song, Yang, and Stefano Ermon. 2019. “Generative Modeling by
Estimating Gradients of the Data Distribution.” Advances in
Neural Information Processing Systems (NeurIPS).
Song, Yang, Sahaj Garg, Jiaxin Shi, and Stefano Ermon. 2019.
“Sliced Score Matching: A Scalable Approach to Density and Score
Estimation.” Uncertainty in Artificial Intelligence
(UAI).
Song, Yang, and Diederik P. Kingma. 2021. “How to Train Your
Energy-Based Models.” arXiv Preprint arXiv:2101.03288.
Song, Yang, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar,
Stefano Ermon, and Ben Poole. 2021. “Score-Based Generative
Modeling Through Stochastic Differential Equations.”
International Conference on Learning Representations (ICLR).
Stirling, Andy. 2007. “A General Framework for Analysing Diversity
in Science, Technology and Society.” Journal of the Royal
Society Interface 4 (15): 707–19. https://doi.org/10.1098/rsif.2007.0213.
Stoyanov, Veselin, Alexander Ropson, and Jason Eisner. 2011.
“Empirical Risk Minimization of Graphical Model Parameters Given
Approximate Inference, Decoding, and Model Structure.”
International Conference on Artificial Intelligence and Statistics
(AISTATS).
Sustek, Martin, Samik Sadhu, Lukáš Burget, et al. 2023.
“Stabilized Training of Joint Energy-Based Models and Their
Practical Applications.” arXiv Preprint
arXiv:2303.04187.
Sutskever, Ilya, and Tijmen Tieleman. 2010. “On the Convergence
Properties of Contrastive Divergence.” International
Conference on Artificial Intelligence and Statistics (AISTATS),
Proceedings of machine learning research, vol. 9: 789–95.
Tanaka, Daiki, Daiki Ikami, Toshihiko Yamasaki, and Kiyoharu Aizawa.
2018. “Joint Optimization Framework for Learning with Noisy
Labels.” IEEE Conference on Computer Vision and Pattern
Recognition (CVPR). https://arxiv.org/abs/1803.11364.
Taskar, Ben, Carlos Guestrin, and Daphne Koller. 2003. “Max-Margin
Markov Networks.” Advances in Neural Information
Processing Systems (NeurIPS) 16.
Tempesta, Piergiulio. 2016. “Beyond the Shannon–Khinchin
Formulation: The Composability Axiom and the Universal-Group
Entropy.” Annals of Physics 365: 180–97. https://doi.org/10.1016/j.aop.2015.08.013.
The Royal Swedish Academy of Sciences. 2024. Scientific Background
to the Nobel Prize in Physics 2024: For Foundational Discoveries and
Inventions That Enable Machine Learning with Artificial Neural
Networks. https://www.nobelprize.org/prizes/physics/2024/advanced-information/.
Tieleman, Tijmen. 2008. “Training Restricted Boltzmann Machines
Using Approximations to the Likelihood Gradient.”
International Conference on Machine Learning (ICML), 1064–71.
Todorov, Emanuel. 2006. “Linearly-Solvable Markov
Decision Problems.” Advances in Neural Information Processing
Systems (NeurIPS) 19: 1369–76.
Todorov, Emanuel. 2009. “Efficient Computation of Optimal
Actions.” Proceedings of the National Academy of
Sciences 106 (28): 11478–83.
Topsøe, Flemming. 1979. “Information-Theoretical Optimization
Techniques.” Kybernetika 15 (1): 8–27.
Toussaint, Marc. 2009. “Robot Trajectory Optimization Using
Approximate Inference.” International Conference on Machine
Learning (ICML).
Tsallis, Constantino. 1988. “Possible Generalization of
Boltzmann–Gibbs Statistics.” Journal of Statistical
Physics 52 (1–2): 479–87.
Tschiatschek, Sebastian, Aytunc Sahin, and Andreas Krause. 2018.
“Differentiable Submodular Maximization.” International
Joint Conference on Artificial Intelligence (IJCAI), 2731–38.
Tsochantaridis, Ioannis, Thorsten Joachims, Thomas Hofmann, and Yasemin
Altun. 2005. “Large Margin Methods for Structured and
Interdependent Output Variables.” Journal of Machine Learning
Research 6: 1453–84.
Tu, Zhuowen. 2007. “Learning Generative Models via Discriminative
Approaches.” IEEE Conference on Computer Vision and Pattern
Recognition (CVPR).
Uehara, Masatoshi, Yulai Zhao, Kevin Black, et al. 2024.
“Fine-Tuning of Continuous-Time Diffusion Models as
Entropy-Regularized Control.” arXiv Preprint
arXiv:2402.15194.
Velikonivtsev, Fedor, Mikhail Mironov, and Liudmila Prokhorenkova. 2024.
“Challenges of Generating Structurally Diverse Graphs.”
Advances in Neural Information Processing Systems (NeurIPS) 37:
57993–8022. https://doi.org/10.52202/079017-1849.
Verschaffelt, Jules-Émile. 1954. “Sur Les Minima de Production
d’entropie Et de Dissipation d’énergie.” Bulletin de La
Classe Des Sciences, Académie Royale de Belgique 40: 779–83.
Vincent, Pascal. 2011. “A Connection Between Score Matching and
Denoising Autoencoders.” Neural Computation 23 (7):
1661–74.
Wainwright, Martin J., and Michael I. Jordan. 2008. “Graphical
Models, Exponential Families, and Variational Inference.”
Foundations and Trends in Machine Learning 1 (1–2): 1–305. https://doi.org/10.1561/2200000001.
Wallace, Bram, Meihua Dang, Rafael Rafailov, et al. 2024.
“Diffusion Model Alignment Using Direct Preference
Optimization.” Proceedings of the IEEE/CVF Conference on
Computer Vision and Pattern Recognition (CVPR), 8228–38.
Wang, Dequan, Evan Shelhamer, Shaoteng Liu, Bruno Olshausen, and Trevor
Darrell. 2021. “Tent: Fully Test-Time Adaptation by Entropy
Minimization.” International Conference on Learning
Representations (ICLR). https://arxiv.org/abs/2006.10726.
Wang, Xudong, Zhirong Wu, Long Lian, and Stella X. Yu. 2022.
“Debiased Learning from Naturally Imbalanced
Pseudo-Labels.” IEEE/CVF Conference on Computer Vision and
Pattern Recognition (CVPR), 14647–57. https://doi.org/10.1109/CVPR52688.2022.01424.
Weilenmann, Mirjam, Lea Kraemer, Philippe Faist, and Renato Renner.
2016. “Axiomatic Relation Between Thermodynamic and
Information-Theoretic Entropies.” Physical Review
Letters 117 (26): 260601. https://doi.org/10.1103/PhysRevLett.117.260601.
Weitzman, Martin L. 1992. “On Diversity.” The Quarterly
Journal of Economics 107 (2): 363–405. https://doi.org/10.2307/2118476.
Williams, Ronald J. 1992. “Simple Statistical Gradient-Following
Algorithms for Connectionist Reinforcement Learning.” Machine
Learning 8 (3–4): 229–56.
Wu, Fa-Yueh. 1982. “The Potts Model.” Reviews of Modern
Physics 54 (1): 235–68. https://doi.org/10.1103/RevModPhys.54.235.
Wu, Jiaxiang, Tao Shen, Haidong Lan, Yatao Bian, and Junzhou Huang.
2021. “SE(3)-Equivariant Energy-Based Models for
End-to-End Protein Folding.” bioRxiv Preprint.
Wu, Tailin, and Ian Fischer. 2020. “Phase Transitions for the
Information Bottleneck in Representation Learning.”
International Conference on Learning Representations (ICLR). https://arxiv.org/abs/2001.01878.
Wu, Tailin, Ian Fischer, Isaac L. Chuang, and Max Tegmark. 2019.
“Learnability for the Information Bottleneck.”
Uncertainty in Artificial Intelligence (UAI). https://arxiv.org/abs/1907.07331.
Xie, Binghui, Yatao Bian, Kaiwen Zhou, et al. 2024. “Enhancing
Neural Subset Selection: Integrating Background Information into Set
Representations.” International Conference on Learning
Representations (ICLR).
Xie, Binghui, Yixuan Wang, Yongqiang Chen, et al. 2024.
“HORSE: Hierarchical Representation for Large-Scale
Neural Subset Selection.” Advances in Neural Information
Processing Systems (NeurIPS), 4852–77.
Xie, Jianwen, Yang Lu, Ruiqi Gao, and Ying Nian Wu. 2018.
“Cooperative Learning of Energy-Based Model and Latent Variable
Model via MCMC Teaching.” AAAI Conference on
Artificial Intelligence.
Xie, Jianwen, Yang Lu, Song-Chun Zhu, and Ying Nian Wu. 2016. “A
Theory of Generative ConvNet.” International
Conference on Machine Learning (ICML), Proceedings of machine
learning research, vol. 48: 2635–44.
Xie, Yutong, Ziqiao Xu, Jiaqi Ma, and Qiaozhu Mei. 2023. “How Much
Space Has Been Explored? Measuring the Chemical Space Covered by
Databases and Machine-Generated Molecules.” International
Conference on Learning Representations (ICLR). https://openreview.net/forum?id=Yo06F8kfMa1.
Xu, Minkai, Tomas Geffner, Karsten Kreis, et al. 2025.
“Energy-Based Diffusion Language Models for Text
Generation.” International Conference on Learning
Representations (ICLR).
Yan, Yuchen, Minkai Xu, Zaiquan Yang, and Yatao Bian. 2026.
“Unified Energy for Invariant and Independent Decoding in
Diffusion Language Models.” arXiv Preprint
arXiv:2606.09159.
Yang, Xiulong, and Shihao Ji. 2021. “JEM++: Improved
Techniques for Training JEM.” IEEE/CVF
International Conference on Computer Vision (ICCV), 6494–503.
Yang, Xiulong, Qing Su, and Shihao Ji. 2023. “Towards Bridging the
Performance Gaps of Joint Energy-Based Models.” IEEE/CVF
Conference on Computer Vision and Pattern Recognition (CVPR),
15732–41.
Yuan, Suqin, Jinkun Chen, Jiyang Zheng, et al. 2026.
“Understanding Diversity Collapse in RLVR via the Lens of
Overtraining.” arXiv Preprint arXiv:2606.15455. https://arxiv.org/abs/2606.15455.
Yuan, Yige, Bingbing Xu, Liang Hou, Fei Sun, Huawei Shen, and Xueqi
Cheng. 2024. “TEA: Test-Time Energy
Adaptation.” IEEE/CVF Conference on Computer Vision and
Pattern Recognition (CVPR), 23901–11.
Zaheer, Manzil, Satwik Kottur, Siamak Ravanbakhsh, Barnabás Póczos,
Ruslan Salakhutdinov, and Alexander J. Smola. 2017. “Deep
Sets.” Advances in Neural Information Processing Systems
(NeurIPS) 30: 3391–401.
Zhang, Qingyang, Haitao Wu, Changqing Zhang, Peilin Zhao, and Yatao
Bian. 2025. “Right Question Is Already Half the Answer: Fully
Unsupervised LLM Reasoning Incentivization.” Advances in
Neural Information Processing Systems (NeurIPS). https://arxiv.org/abs/2504.05812.
Zhao, Dora, Jerone T. A. Andrews, Orestis Papakyriakopoulos, and Alice
Xiang. 2024. “Position: Measure Dataset Diversity, Don’t Just
Claim It.” Proceedings of the 41st International Conference
on Machine Learning 235: 60644–73. https://proceedings.mlr.press/v235/zhao24a.html.
Zhao, Stephen, Rob Brekelmans, Alireza Makhzani, and Roger Grosse. 2024.
“Probabilistic Inference in Language Models via Twisted Sequential
Monte Carlo.” Proceedings of the
41st International Conference on Machine Learning (ICML),
Proceedings of machine learning research, vol. 235.
Zhao, Stephen, Jörn-Henrik Jacobsen, and Will Grathwohl. 2020.
“Joint Energy-Based Models for Semi-Supervised
Classification.” ICML 2020 Workshop on Uncertainty and
Robustness in Deep Learning.
Zhao, Xuandong, Zhewei Kang, Aosong Feng, Sergey Levine, and Dawn Song.
2025. “Learning to Reason Without External Rewards.”
arXiv Preprint arXiv:2505.19590. https://arxiv.org/abs/2505.19590.
Zheng, Shuai, Sadeep Jayasumana, Bernardino Romera-Paredes, et al. 2015.
“Conditional Random Fields as Recurrent Neural Networks.”
IEEE International Conference on Computer Vision (ICCV),
1529–37.
Zhou, Dengyong, Qiang Liu, John C. Platt, Christopher Meek, and Nihar B.
Shah. 2015. “Regularized Minimax Conditional Entropy for
Crowdsourcing.” arXiv Preprint arXiv:1503.07240.
Zhou, Dengyong, John C. Platt, Sumit Basu, and Yi Mao. 2012.
“Learning from the Wisdom of Crowds by Minimax Entropy.”
Advances in Neural Information Processing Systems (NeurIPS) 25:
2195–203.
Zhu, Song Chun, Ying Nian Wu, and David Mumford. 1997. “Minimax
Entropy Principle and Its Application to Texture Modeling.”
Neural Computation 9 (8): 1627–60. https://doi.org/10.1162/neco.1997.9.8.1627.
Zhu, Song Chun, Ying Nian Wu, and David Mumford. 1998. “Filters,
Random Fields and Maximum Entropy (FRAME): Towards a
Unified Theory for Texture Modeling.” International Journal
of Computer Vision 27 (2): 107–26.
Zhu, Yuchang, Huizhe Zhang, Bingzhe Wu, et al. 2025. “Measuring
Diversity in Synthetic Datasets.” Proceedings of the 42nd
International Conference on Machine Learning 267: 80373–97. https://proceedings.mlr.press/v267/zhu25ac.html.
Ziebart, Brian D., J. Andrew Bagnell, and Anind K. Dey. 2010.
“Modeling Interaction via the Principle of Maximum Causal
Entropy.” International Conference on Machine Learning
(ICML).
Ziebart, Brian D., Andrew Maas, J. Andrew Bagnell, and Anind K. Dey.
2008. “Maximum Entropy Inverse Reinforcement Learning.”
AAAI Conference on Artificial Intelligence, 1433–38.
Ziegler, Daniel M., Nisan Stiennon, Jeffrey Wu, et al. 2019.
“Fine-Tuning Language Models from Human Preferences.”
arXiv Preprint arXiv:1909.08593.
Zuo, Yuxin, Kaiyan Zhang, Li Sheng, et al. 2025.
“TTRL: Test-Time Reinforcement Learning.”
Advances in Neural Information Processing Systems 38 (NeurIPS).