References
Ackley, David H., Geoffrey E. Hinton, and Terrence J. Sejnowski. 1985.
“A Learning Algorithm for Boltzmann Machines.”
Cognitive Science 9 (1): 147–69. https://doi.org/10.1207/s15516709cog0901_7.
Asano, Yuki M., Christian Rupprecht, and Andrea Vedaldi. 2020.
“Self-Labelling via Simultaneous Clustering and Representation
Learning.” International Conference on Learning
Representations (ICLR). https://arxiv.org/abs/1911.05371.
Barbera, Elvira. 1999. “On the Principle of Minimal Entropy
Production for Navier–Stokes–Fourier Fluids.” Continuum
Mechanics and Thermodynamics 11 (5): 327–30. https://doi.org/10.1007/s001610050127.
Berger, Adam L., Stephen A. Della Pietra, and Vincent J. Della Pietra.
1996. “A Maximum Entropy Approach to Natural Language
Processing.” Computational Linguistics 22 (1): 39–71.
Berthelot, David, Nicholas Carlini, Ekin D. Cubuk, et al. 2020.
“ReMixMatch: Semi-Supervised Learning with Distribution Alignment
and Augmentation Anchoring.” International Conference on
Learning Representations (ICLR). https://arxiv.org/abs/1911.09785.
Besag, Julian. 1974. “Spatial Interaction and the Statistical
Analysis of Lattice Systems.” Journal of the Royal
Statistical Society: Series B (Methodological) 36 (2): 192–236.
Bian, Andrew An, Joachim M. Buhmann, Andreas Krause, and Sebastian
Tschiatschek. 2017. “Guarantees for Greedy Maximization of
Non-Submodular Functions with Applications.” International
Conference on Machine Learning (ICML), 498–507.
Bian, Yatao. 2020. Awesome Energy-Based Models/Learning: A
Comprehensive List of Energy-Based Learning Papers and Materials.
Https://github.com/yataobian/awesome-ebm.
Bian, Yatao, Yu Rong, Tingyang Xu, Jiaxiang Wu, Andreas Krause, and
Junzhou Huang. 2022. “Energy-Based Learning for Cooperative Games,
with Applications to Valuation Problems in Machine Learning.”
International Conference on Learning Representations (ICLR). https://openreview.net/forum?id=xLfAgCroImw.
Boltzmann, Ludwig. 1877. “Über Die Beziehung Zwischen Dem Zweiten
Hauptsatze Der Mechanischen Wärmetheorie Und Der
Wahrscheinlichkeitsrechnung Respektive Den Sätzen Über Das
Wärmegleichgewicht.” Sitzungsberichte Der Kaiserlichen
Akademie Der Wissenschaften in Wien 76: 373–435.
Bridle, John S., Anthony J. R. Heading, and David J. C. MacKay. 1992.
“Unsupervised Classifiers, Mutual Information and ’Phantom
Targets’.” Advances in Neural Information Processing Systems
(NIPS) 4, 1096–101.
Callen, Herbert B. 1985. Thermodynamics and an Introduction to
Thermostatistics. 2nd ed. Wiley.
Carcamo, David P., Nicholas J. Weaver, Purushottam D. Dixit, and
Christopher W. Lynn. 2025. “Minimax Entropy: The Statistical
Physics of Optimal Models.” Physical Review E 112 (6):
061001. https://doi.org/10.1103/kr9x-q59y.
Caron, Mathilde, Ishan Misra, Julien Mairal, Priya Goyal, Piotr
Bojanowski, and Armand Joulin. 2020. “Unsupervised Learning of
Visual Features by Contrasting Cluster Assignments.” Advances
in Neural Information Processing Systems (NeurIPS). https://arxiv.org/abs/2006.09882.
Caron, Mathilde, Hugo Touvron, Ishan Misra, et al. 2021. “Emerging
Properties in Self-Supervised Vision Transformers.” IEEE/CVF
International Conference on Computer Vision (ICCV). https://arxiv.org/abs/2104.14294.
Clausius, Rudolf. 1865. “Über Verschiedene Für Die Anwendung
Bequeme Formen Der Hauptgleichungen Der Mechanischen
Wärmetheorie.” Annalen Der Physik Und Chemie 125:
353–400.
Cohen, Jacob. 1960. “A Coefficient of Agreement for Nominal
Scales.” Educational and Psychological Measurement 20
(1): 37–46. https://doi.org/10.1177/001316446002000104.
Cui, Ganqu, Yuchen Zhang, Jiacheng Chen, et al. 2025. “The Entropy
Mechanism of Reinforcement Learning for Reasoning Language
Models.” arXiv Preprint arXiv:2505.22617. https://arxiv.org/abs/2505.22617.
Dawid, Anna, and Yann LeCun. 2024. “Introduction to Latent
Variable Energy-Based Models: A Path Towards Autonomous Machine
Intelligence.” Journal of Statistical Mechanics: Theory and
Experiment 2024 (10): 104011.
Dayan, Peter, Geoffrey E. Hinton, Radford M. Neal, and Richard S. Zemel.
1995. “The Helmholtz Machine.” Neural Computation
7 (5): 889–904. https://doi.org/10.1162/neco.1995.7.5.889.
DeepSeek-AI. 2025. “DeepSeek-R1 Incentivizes Reasoning in LLMs
Through Reinforcement Learning.” Nature 645. https://doi.org/10.1038/s41586-025-09422-z.
Della Pietra, Stephen, Vincent Della Pietra, and John Lafferty. 1997.
“Inducing Features of Random Fields.” IEEE Transactions
on Pattern Analysis and Machine Intelligence 19 (4): 380–93. https://doi.org/10.1109/34.588021.
Deng, Yuntian, Anton Bakhtin, Myle Ott, Arthur Szlam, and Marc’Aurelio
Ranzato. 2020. “Residual Energy-Based Models for Text
Generation.” International Conference on Learning
Representations (ICLR).
Dhariwal, Prafulla, and Alexander Nichol. 2021. “Diffusion Models
Beat GANs on Image Synthesis.” Advances in
Neural Information Processing Systems (NeurIPS).
Du, Yilun, Shuang Li, and Igor Mordatch. 2020. “Compositional
Visual Generation with Energy Based Models.” Advances in
Neural Information Processing Systems (NeurIPS).
Du, Yilun, and Igor Mordatch. 2019. “Implicit Generation and
Modeling with Energy-Based Models.” Advances in Neural
Information Processing Systems (NeurIPS).
Farquhar, Sebastian, Jannik Kossen, Lorenz Kuhn, and Yarin Gal. 2024.
“Detecting Hallucinations in Large Language Models Using Semantic
Entropy.” Nature 630: 625–30. https://doi.org/10.1038/s41586-024-07421-0.
Feder, M. 1986. “Maximum Entropy as a Special Case of the Minimum
Description Length Criterion.” IEEE Transactions on
Information Theory 32: 847–49.
Friedman, Dan, and Adji Bousso Dieng. 2023. “The Vendi Score: A
Diversity Evaluation Metric for Machine Learning.”
Transactions on Machine Learning Research. https://arxiv.org/abs/2210.02410.
Friston, Karl. 2010. “The Free-Energy Principle: A Unified Brain
Theory?” Nature Reviews Neuroscience 11 (2): 127–38.
Geiger, Dan, David Heckerman, Henry King, and Christopher Meek. 2001.
“Stratified Exponential Families: Graphical Models and Model
Selection.” The Annals of Statistics 29 (2): 505–29. https://doi.org/10.1214/aos/1009210550.
Gibbs, Josiah Willard. 1902. Elementary Principles in Statistical
Mechanics. Charles Scribner’s Sons.
Gnaiger, Erich. 2009. “Open and Closed Systems: Styles of Thinking
Explain Controversies on the ’Negative Entropy’ Concept of Ludwig
Boltzmann and Erwin Schrödinger.” Mitochondrial Physiology
Network. https://www.mitophysiology.org/images/2/23/Gnaiger_2009_OCESHS.pdf.
Gomes, Ryan, Andreas Krause, and Pietro Perona. 2010.
“Discriminative Clustering by Regularized Information
Maximization.” Advances in Neural Information Processing
Systems (NeurIPS) 23.
Grandvalet, Yves, and Yoshua Bengio. 2004. “Semi-Supervised
Learning by Entropy Minimization.” Advances in Neural
Information Processing Systems (NeurIPS) 17.
Grathwohl, Will, Kevin Swersky, Milad Hashemi, David Duvenaud, and Chris
Maddison. 2021. “Oops I Took a Gradient: Scalable
Sampling for Discrete Distributions.” International
Conference on Machine Learning (ICML), 3831–41.
Grathwohl, Will, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud,
Mohammad Norouzi, and Kevin Swersky. 2020. “Your Classifier Is
Secretly an Energy Based Model and You Should Treat It Like One.”
International Conference on Learning Representations (ICLR).
Grünwald, Peter D. 2007. The Minimum Description Length
Principle. MIT Press.
Gutmann, Michael, and Aapo Hyvärinen. 2010. “Noise-Contrastive
Estimation: A New Estimation Principle for Unnormalized Statistical
Models.” International Conference on Artificial Intelligence
and Statistics (AISTATS), 297–304.
GX-Chen, Anthony, Jatin Prakash, Jeff Guo, Rob Fergus, and Rajesh
Ranganath. 2025. “KL-Regularized Reinforcement Learning Is
Designed to Mode Collapse.” arXiv Preprint
arXiv:2510.20817. https://arxiv.org/abs/2510.20817.
Haarnoja, Tuomas, Haoran Tang, Pieter Abbeel, and Sergey Levine. 2017.
“Reinforcement Learning with Deep Energy-Based Policies.”
International Conference on Machine Learning (ICML), 1352–61.
Haarnoja, Tuomas, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018.
“Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement
Learning with a Stochastic Actor.” International Conference
on Machine Learning (ICML), 1861–70.
Hill, M. O. 1973. “Diversity and Evenness: A Unifying Notation and
Its Consequences.” Ecology 54 (2): 427–32. https://doi.org/10.2307/1934352.
Hinton, Geoffrey E. 2002. “Training Products of Experts by
Minimizing Contrastive Divergence.” Neural Computation
14 (8): 1771–800. https://doi.org/10.1162/089976602760128018.
Hirsh, Jacob B., Raymond A. Mar, and Jordan B. Peterson. 2012.
“Psychological Entropy: A Framework for Understanding
Uncertainty-Related Anxiety.” Psychological Review 119
(2): 304–20. https://doi.org/10.1037/a0026767.
Ho, Jonathan, Ajay Jain, and Pieter Abbeel. 2020. “Denoising
Diffusion Probabilistic Models.” Advances in Neural
Information Processing Systems (NeurIPS).
Hopfield, John J. 1982. “Neural Networks and Physical Systems with
Emergent Collective Computational Abilities.” Proceedings of
the National Academy of Sciences 79 (8): 2554–58. https://doi.org/10.1073/pnas.79.8.2554.
Hu, Weihua, Takeru Miyato, Seiya Tokui, Eiichi Matsumoto, and Masashi
Sugiyama. 2017. “Learning Discrete Representations via Information
Maximizing Self-Augmented Training.” International Conference
on Machine Learning (ICML). https://arxiv.org/abs/1702.08720.
Hyvärinen, Aapo. 2005. “Estimation of Non-Normalized Statistical
Models by Score Matching.” Journal of Machine Learning
Research 6: 695–709.
Hyvärinen, Aapo. 2007. “Connections Between Score Matching,
Contrastive Divergence, and Pseudolikelihood for Continuous-Valued
Variables.” IEEE Transactions on Neural Networks 18 (5):
1529–31.
Jaynes, Edwin T. 1957. “Information Theory and Statistical
Mechanics.” Physical Review 106: 620.
Jaynes, Edwin T. 1980. “The Minimum Entropy Production
Principle.” Annual Review of Physical Chemistry 31:
579–601. https://doi.org/10.1146/annurev.pc.31.100180.003051.
Kianercy, Ardeshir, and Aram Galstyan. 2012. “Dynamics of
Boltzmann q Learning in Two-Player Two-Action Games.”
Physical Review E 85 (4): 041145. https://doi.org/10.1103/PhysRevE.85.041145.
Kirkpatrick, Scott, C. Daniel Gelatt, and Mario P. Vecchi. 1983.
“Optimization by Simulated Annealing.” Science 220
(4598): 671–80. https://doi.org/10.1126/science.220.4598.671.
Koller, Daphne, and Nir Friedman. 2009. Probabilistic Graphical
Models: Principles and Techniques. MIT Press.
Kuhn, Lorenz, Yarin Gal, and Sebastian Farquhar. 2023. “Semantic
Uncertainty: Linguistic Invariances for Uncertainty Estimation in
Natural Language Generation.” International Conference on
Learning Representations (ICLR). https://arxiv.org/abs/2302.09664.
Lafferty, John, Andrew McCallum, and Fernando Pereira. 2001.
“Conditional Random Fields: Probabilistic Models for Segmenting
and Labeling Sequence Data.” International Conference on
Machine Learning (ICML), 282–89.
Landauer, Rolf. 1975. “Inadequacy of Entropy and Entropy
Derivatives in Characterizing the Steady State.” Physical
Review A 12 (2): 636–38. https://doi.org/10.1103/PhysRevA.12.636.
LeCun, Yann. 2022. “A Path Towards Autonomous Machine
Intelligence.” OpenReview Preprint.
LeCun, Yann, Sumit Chopra, Raia Hadsell, Marc’Aurelio Ranzato, and Fu
Jie Huang. 2006. “A Tutorial on Energy-Based Learning.” In
Predicting Structured Data, edited by Gökhan Bakir, Thomas
Hofmann, Bernhard Schölkopf, Alexander J. Smola, and Ben Taskar. MIT
Press.
Lee, Dong-Hyun. 2013. “Pseudo-Label: The Simple and Efficient
Semi-Supervised Learning Method for Deep Neural Networks.”
ICML Workshop on Challenges in Representation Learning.
Lee, Jonghyun, Dahuin Jung, Saehyung Lee, et al. 2024. “Entropy Is
Not Enough for Test-Time Adaptation: From the Perspective of
Disentangled Factors.” International Conference on Learning
Representations (ICLR). https://arxiv.org/abs/2403.07366.
Leinster, Tom, and Christina A. Cobbold. 2012. “Measuring
Diversity: The Importance of Species Similarity.”
Ecology 93 (3): 477–89. https://doi.org/10.1890/10-2402.1.
Liang, Jian, Dapeng Hu, and Jiashi Feng. 2020. “Do We Really Need
to Access the Source Data? Source Hypothesis Transfer for Unsupervised
Domain Adaptation.” International Conference on Machine
Learning (ICML). https://arxiv.org/abs/2002.08546.
Liu, Weitang, Xiaoyun Wang, John D. Owens, and Yixuan Li. 2020.
“Energy-Based Out-of-Distribution Detection.” Advances
in Neural Information Processing Systems (NeurIPS).
Lynn, Christopher W., Qiwei Yu, Rich Pang, Stephanie E. Palmer, and
William Bialek. 2025. “Exact Minimax Entropy Models of Large-Scale
Neuronal Activity.” Physical Review E 111 (5): 054411.
https://doi.org/10.1103/PhysRevE.111.054411.
Ma, Xueguang, Qian Liu, Dongfu Jiang, Ge Zhang, Zejun Ma, and Wenhu
Chen. 2025. “General-Reasoner: Advancing LLM Reasoning Across All
Domains.” arXiv Preprint arXiv:2505.14652. https://arxiv.org/abs/2505.14652.
Maes, Christian, and Karel Netočný. 2007. “Minimum Entropy
Production Principle from a Dynamical Fluctuation Law.”
Journal of Mathematical Physics 48 (5): 053306. https://doi.org/10.1063/1.2738753.
Maes, Christian, and Karel Netočný. 2013. “Minimum Entropy
Production Principle.” Scholarpedia 8 (7): 9664. https://doi.org/10.4249/scholarpedia.9664.
Martyushev, Leonid M., A. S. Nazarova, and Vladimir D. Seleznev. 2007.
“On the Problem of the Minimum Entropy Production in the
Nonequilibrium Stationary State.” Journal of Physics A:
Mathematical and Theoretical 40 (3): 371–80. https://doi.org/10.1088/1751-8113/40/3/002.
Martyushev, Leonid M., and Vladimir D. Seleznev. 2006. “Maximum
Entropy Production Principle in Physics, Chemistry and Biology.”
Physics Reports 426 (1): 1–45. https://doi.org/10.1016/j.physrep.2005.12.001.
Martyushev, Leonid M., and Vladimir D. Seleznev. 2013. “Entropy
and Entropy Production: Old Misconceptions and New
Breakthroughs.” Entropy 15 (4): 1152–70. https://doi.org/10.3390/e15041152.
Menon, Aditya Krishna, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu
Jain, Andreas Veit, and Sanjiv Kumar. 2021. “Long-Tail Learning
via Logit Adjustment.” International Conference on Learning
Representations (ICLR). https://arxiv.org/abs/2007.07314.
Mnih, Andriy, and Yee Whye Teh. 2012. “A Fast and Simple Algorithm
for Training Neural Probabilistic Language Models.”
International Conference on Machine Learning (ICML), 419–26.
Moussouris, John. 1974. “Gibbs and Markov Random Systems with
Constraints.” Journal of Statistical Physics 10 (1):
11–33. https://doi.org/10.1007/BF01011714.
Neal, Radford M., and Geoffrey E. Hinton. 1998. “A View of the EM
Algorithm That Justifies Incremental, Sparse, and Other
Variants.” In Learning in Graphical Models, edited by
Michael I. Jordan. Kluwer Academic Publishers.
Nemhauser, George L., Laurence A. Wolsey, and Marshall L. Fisher. 1978.
“An Analysis of Approximations for Maximizing Submodular Set
Functions—i.” Mathematical Programming 14: 265–94.
Neumann, John von. 1927. “Thermodynamik Quantenmechanischer
Gesamtheiten.” Nachrichten von Der Gesellschaft Der
Wissenschaften Zu Göttingen, Mathematisch-Physikalische Klasse
1927: 273–91.
Nijkamp, Erik, Mitch Hill, Song-Chun Zhu, and Ying Nian Wu. 2019.
“Learning Non-Convergent Non-Persistent Short-Run
MCMC Toward Energy-Based Model.” Advances in
Neural Information Processing Systems (NeurIPS).
Nikitin, Alexander, Jannik Kossen, Yarin Gal, and Pekka Marttinen. 2024.
“Kernel Language Entropy: Fine-Grained Uncertainty Quantification
for LLMs from Semantic Similarities.” Advances
in Neural Information Processing Systems (NeurIPS). https://arxiv.org/abs/2405.20003.
Niu, Shuaicheng, Jiaxiang Wu, Yifan Zhang, et al. 2022. “Efficient
Test-Time Model Adaptation Without Forgetting.” International
Conference on Machine Learning (ICML). https://arxiv.org/abs/2204.02610.
Niu, Shuaicheng, Jiaxiang Wu, Yifan Zhang, et al. 2023. “Towards
Stable Test-Time Adaptation in Dynamic Wild World.”
International Conference on Learning Representations (ICLR). https://arxiv.org/abs/2302.12400.
Owen, Guillermo. 1972. “Multilinear Extensions of Games.”
Management Science 18 (5): P64–79. https://doi.org/10.1287/mnsc.18.5.P64.
Owen, Guillermo. 1975. “Multilinear Extensions and the Banzhaf
Value.” Naval Research Logistics Quarterly 22 (4):
741–50. https://doi.org/10.1002/nav.3800220409.
Paltridge, Garth W. 1975. “Global Dynamics and Climate — a System
of Minimum Entropy Exchange.” Quarterly Journal of the Royal
Meteorological Society 101 (429): 475–84. https://doi.org/10.1002/qj.49710142906.
Paninski, Liam. 2003. “Estimation of Entropy and Mutual
Information.” Neural Computation 15 (6): 1191–253. https://doi.org/10.1162/089976603321780272.
Pasarkar, Amey P., and Adji Bousso Dieng. 2024. “Cousins of the
Vendi Score: A Family of Similarity-Based Diversity Metrics for Science
and Machine Learning.” International Conference on Artificial
Intelligence and Statistics (AISTATS), PMLR, vol. 238. https://arxiv.org/abs/2310.12952.
Pearl, Judea. 1988. Probabilistic Reasoning in Intelligent Systems:
Networks of Plausible Inference. Morgan Kaufmann.
Prigogine, Ilya. 1945. “Modération Et Transformations
Irréversibles Des Systèmes Ouverts.” Bulletin de La Classe
Des Sciences, Académie Royale de Belgique 31: 600–606.
Prigogine, Ilya. 1947. Étude Thermodynamique Des Phénomènes
Irréversibles. Dunod, Paris; Desoer, Liège.
Prigogine, Ilya. 1978. “Time, Structure, and Fluctuations.”
Science 201 (4358): 777–85. https://doi.org/10.1126/science.201.4358.777.
Prigogine, Ilya, and Jean-Marie Wiame. 1946. “Biologie Et
Thermodynamique Des Phénomènes Irréversibles.”
Experientia 2 (11): 451–53. https://doi.org/10.1007/BF02153597.
Qin, Lianhui, Sean Welleck, Daniel Khashabi, and Yejin Choi. 2022.
“COLD Decoding: Energy-Based Constrained Text
Generation with Langevin Dynamics.” Advances in
Neural Information Processing Systems (NeurIPS).
Rényi, Alfréd. 1961. “On Measures of Entropy and
Information.” Proceedings of the Fourth Berkeley Symposium on
Mathematical Statistics and Probability, Volume 1: Contributions to the
Theory of Statistics, 547–61.
Rose, Kenneth, Eitan Gurewitz, and Geoffrey C. Fox. 1990.
“Statistical Mechanics and Phase Transitions in
Clustering.” Physical Review Letters 65 (8): 945–48. https://doi.org/10.1103/PhysRevLett.65.945.
Sánchez Giraldo, Luis Gonzalo, Murali Rao, and José C. Príncipe. 2015.
“Measures of Entropy from Data Using Infinitely Divisible
Kernels.” IEEE Transactions on Information Theory 61
(1): 535–48. https://doi.org/10.1109/TIT.2014.2370058.
Schrödinger, Erwin. 1944. What Is Life? The Physical Aspect of the
Living Cell. Cambridge University Press.
Shannon, Claude E. 1948. “A Mathematical Theory of
Communication.” Bell System Technical Journal 27 (3):
379–423.
Shao, Zhihong, Peiyi Wang, Qihao Zhu, et al. 2024. “DeepSeekMath:
Pushing the Limits of Mathematical Reasoning in Open Language
Models.” arXiv Preprint arXiv:2402.03300. https://arxiv.org/abs/2402.03300.
Shi, Chence, Shitong Luo, Minkai Xu, and Jian Tang. 2021.
“Learning Gradient Fields for Molecular Conformation
Generation.” International Conference on Machine Learning
(ICML).
Sohl-Dickstein, Jascha, Eric Weiss, Niru Maheswaranathan, and Surya
Ganguli. 2015. “Deep Unsupervised Learning Using Nonequilibrium
Thermodynamics.” International Conference on Machine Learning
(ICML), 2256–65.
Song, Yang, and Stefano Ermon. 2019. “Generative Modeling by
Estimating Gradients of the Data Distribution.” Advances in
Neural Information Processing Systems (NeurIPS).
Song, Yang, Sahaj Garg, Jiaxin Shi, and Stefano Ermon. 2019.
“Sliced Score Matching: A Scalable Approach to Density and Score
Estimation.” Uncertainty in Artificial Intelligence
(UAI).
Song, Yang, and Diederik P. Kingma. 2021. “How to Train Your
Energy-Based Models.” arXiv Preprint arXiv:2101.03288.
Song, Yang, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar,
Stefano Ermon, and Ben Poole. 2021. “Score-Based Generative
Modeling Through Stochastic Differential Equations.”
International Conference on Learning Representations (ICLR).
Tanaka, Daiki, Daiki Ikami, Toshihiko Yamasaki, and Kiyoharu Aizawa.
2018. “Joint Optimization Framework for Learning with Noisy
Labels.” IEEE Conference on Computer Vision and Pattern
Recognition (CVPR). https://arxiv.org/abs/1803.11364.
The Royal Swedish Academy of Sciences. 2024. Scientific Background
to the Nobel Prize in Physics 2024: For Foundational Discoveries and
Inventions That Enable Machine Learning with Artificial Neural
Networks. https://www.nobelprize.org/prizes/physics/2024/advanced-information/.
Tieleman, Tijmen. 2008. “Training Restricted Boltzmann Machines
Using Approximations to the Likelihood Gradient.”
International Conference on Machine Learning (ICML), 1064–71.
Tsallis, Constantino. 1988. “Possible Generalization of
Boltzmann–Gibbs Statistics.” Journal of Statistical
Physics 52 (1–2): 479–87.
Verschaffelt, Jules-Émile. 1954. “Sur Les Minima de Production
d’entropie Et de Dissipation d’énergie.” Bulletin de La
Classe Des Sciences, Académie Royale de Belgique 40: 779–83.
Vincent, Pascal. 2011. “A Connection Between Score Matching and
Denoising Autoencoders.” Neural Computation 23 (7):
1661–74.
Wainwright, Martin J., and Michael I. Jordan. 2008. “Graphical
Models, Exponential Families, and Variational Inference.”
Foundations and Trends in Machine Learning 1 (1–2): 1–305. https://doi.org/10.1561/2200000001.
Wang, Dequan, Evan Shelhamer, Shaoteng Liu, Bruno Olshausen, and Trevor
Darrell. 2021. “Tent: Fully Test-Time Adaptation by Entropy
Minimization.” International Conference on Learning
Representations (ICLR). https://arxiv.org/abs/2006.10726.
Wang, Xudong, Zhirong Wu, Long Lian, and Stella X. Yu. 2022.
“Debiased Learning from Naturally Imbalanced
Pseudo-Labels.” IEEE/CVF Conference on Computer Vision and
Pattern Recognition (CVPR), 14647–57. https://doi.org/10.1109/CVPR52688.2022.01424.
Wu, Fa-Yueh. 1982. “The Potts Model.” Reviews of Modern
Physics 54 (1): 235–68. https://doi.org/10.1103/RevModPhys.54.235.
Wu, Jiaxiang, Tao Shen, Haidong Lan, Yatao Bian, and Junzhou Huang.
2021. “SE(3)-Equivariant Energy-Based Models for
End-to-End Protein Folding.” bioRxiv Preprint.
Wu, Tailin, and Ian Fischer. 2020. “Phase Transitions for the
Information Bottleneck in Representation Learning.”
International Conference on Learning Representations (ICLR). https://arxiv.org/abs/2001.01878.
Wu, Tailin, Ian Fischer, Isaac L. Chuang, and Max Tegmark. 2019.
“Learnability for the Information Bottleneck.”
Uncertainty in Artificial Intelligence (UAI). https://arxiv.org/abs/1907.07331.
Zhang, Qingyang, Haitao Wu, Changqing Zhang, Peilin Zhao, and Yatao
Bian. 2025. “Right Question Is Already Half the Answer: Fully
Unsupervised LLM Reasoning Incentivization.” Advances in
Neural Information Processing Systems (NeurIPS). https://arxiv.org/abs/2504.05812.
Zhao, Xuandong, Zhewei Kang, Aosong Feng, Sergey Levine, and Dawn Song.
2025. “Learning to Reason Without External Rewards.”
arXiv Preprint arXiv:2505.19590. https://arxiv.org/abs/2505.19590.
Zhu, Song Chun, Ying Nian Wu, and David Mumford. 1997. “Minimax
Entropy Principle and Its Application to Texture Modeling.”
Neural Computation 9 (8): 1627–60. https://doi.org/10.1162/neco.1997.9.8.1627.
Ziebart, Brian D., Andrew Maas, J. Andrew Bagnell, and Anind K. Dey.
2008. “Maximum Entropy Inverse Reinforcement Learning.”
AAAI Conference on Artificial Intelligence, 1433–38.