References

Ackley, David H., Geoffrey E. Hinton, and Terrence J. Sejnowski. 1985. “A Learning Algorithm for Boltzmann Machines.” Cognitive Science 9 (1): 147–69. https://doi.org/10.1207/s15516709cog0901_7.
Asano, Yuki M., Christian Rupprecht, and Andrea Vedaldi. 2020. “Self-Labelling via Simultaneous Clustering and Representation Learning.” International Conference on Learning Representations (ICLR). https://arxiv.org/abs/1911.05371.
Barbera, Elvira. 1999. “On the Principle of Minimal Entropy Production for Navier–Stokes–Fourier Fluids.” Continuum Mechanics and Thermodynamics 11 (5): 327–30. https://doi.org/10.1007/s001610050127.
Berger, Adam L., Stephen A. Della Pietra, and Vincent J. Della Pietra. 1996. “A Maximum Entropy Approach to Natural Language Processing.” Computational Linguistics 22 (1): 39–71.
Berthelot, David, Nicholas Carlini, Ekin D. Cubuk, et al. 2020. “ReMixMatch: Semi-Supervised Learning with Distribution Alignment and Augmentation Anchoring.” International Conference on Learning Representations (ICLR). https://arxiv.org/abs/1911.09785.
Besag, Julian. 1974. “Spatial Interaction and the Statistical Analysis of Lattice Systems.” Journal of the Royal Statistical Society: Series B (Methodological) 36 (2): 192–236.
Bian, Andrew An, Joachim M. Buhmann, Andreas Krause, and Sebastian Tschiatschek. 2017. “Guarantees for Greedy Maximization of Non-Submodular Functions with Applications.” International Conference on Machine Learning (ICML), 498–507.
Bian, Yatao. 2020. Awesome Energy-Based Models/Learning: A Comprehensive List of Energy-Based Learning Papers and Materials. Https://github.com/yataobian/awesome-ebm.
Bian, Yatao, Yu Rong, Tingyang Xu, Jiaxiang Wu, Andreas Krause, and Junzhou Huang. 2022. “Energy-Based Learning for Cooperative Games, with Applications to Valuation Problems in Machine Learning.” International Conference on Learning Representations (ICLR). https://openreview.net/forum?id=xLfAgCroImw.
Boltzmann, Ludwig. 1877. “Über Die Beziehung Zwischen Dem Zweiten Hauptsatze Der Mechanischen Wärmetheorie Und Der Wahrscheinlichkeitsrechnung Respektive Den Sätzen Über Das Wärmegleichgewicht.” Sitzungsberichte Der Kaiserlichen Akademie Der Wissenschaften in Wien 76: 373–435.
Bridle, John S., Anthony J. R. Heading, and David J. C. MacKay. 1992. “Unsupervised Classifiers, Mutual Information and ’Phantom Targets’.” Advances in Neural Information Processing Systems (NIPS) 4, 1096–101.
Callen, Herbert B. 1985. Thermodynamics and an Introduction to Thermostatistics. 2nd ed. Wiley.
Carcamo, David P., Nicholas J. Weaver, Purushottam D. Dixit, and Christopher W. Lynn. 2025. “Minimax Entropy: The Statistical Physics of Optimal Models.” Physical Review E 112 (6): 061001. https://doi.org/10.1103/kr9x-q59y.
Caron, Mathilde, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. 2020. “Unsupervised Learning of Visual Features by Contrasting Cluster Assignments.” Advances in Neural Information Processing Systems (NeurIPS). https://arxiv.org/abs/2006.09882.
Caron, Mathilde, Hugo Touvron, Ishan Misra, et al. 2021. “Emerging Properties in Self-Supervised Vision Transformers.” IEEE/CVF International Conference on Computer Vision (ICCV). https://arxiv.org/abs/2104.14294.
Clausius, Rudolf. 1865. “Über Verschiedene Für Die Anwendung Bequeme Formen Der Hauptgleichungen Der Mechanischen Wärmetheorie.” Annalen Der Physik Und Chemie 125: 353–400.
Cohen, Jacob. 1960. “A Coefficient of Agreement for Nominal Scales.” Educational and Psychological Measurement 20 (1): 37–46. https://doi.org/10.1177/001316446002000104.
Cui, Ganqu, Yuchen Zhang, Jiacheng Chen, et al. 2025. “The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.” arXiv Preprint arXiv:2505.22617. https://arxiv.org/abs/2505.22617.
Dawid, Anna, and Yann LeCun. 2024. “Introduction to Latent Variable Energy-Based Models: A Path Towards Autonomous Machine Intelligence.” Journal of Statistical Mechanics: Theory and Experiment 2024 (10): 104011.
Dayan, Peter, Geoffrey E. Hinton, Radford M. Neal, and Richard S. Zemel. 1995. “The Helmholtz Machine.” Neural Computation 7 (5): 889–904. https://doi.org/10.1162/neco.1995.7.5.889.
DeepSeek-AI. 2025. “DeepSeek-R1 Incentivizes Reasoning in LLMs Through Reinforcement Learning.” Nature 645. https://doi.org/10.1038/s41586-025-09422-z.
Della Pietra, Stephen, Vincent Della Pietra, and John Lafferty. 1997. “Inducing Features of Random Fields.” IEEE Transactions on Pattern Analysis and Machine Intelligence 19 (4): 380–93. https://doi.org/10.1109/34.588021.
Deng, Yuntian, Anton Bakhtin, Myle Ott, Arthur Szlam, and Marc’Aurelio Ranzato. 2020. “Residual Energy-Based Models for Text Generation.” International Conference on Learning Representations (ICLR).
Dhariwal, Prafulla, and Alexander Nichol. 2021. “Diffusion Models Beat GANs on Image Synthesis.” Advances in Neural Information Processing Systems (NeurIPS).
Du, Yilun, Shuang Li, and Igor Mordatch. 2020. “Compositional Visual Generation with Energy Based Models.” Advances in Neural Information Processing Systems (NeurIPS).
Du, Yilun, and Igor Mordatch. 2019. “Implicit Generation and Modeling with Energy-Based Models.” Advances in Neural Information Processing Systems (NeurIPS).
Farquhar, Sebastian, Jannik Kossen, Lorenz Kuhn, and Yarin Gal. 2024. “Detecting Hallucinations in Large Language Models Using Semantic Entropy.” Nature 630: 625–30. https://doi.org/10.1038/s41586-024-07421-0.
Feder, M. 1986. “Maximum Entropy as a Special Case of the Minimum Description Length Criterion.” IEEE Transactions on Information Theory 32: 847–49.
Friedman, Dan, and Adji Bousso Dieng. 2023. “The Vendi Score: A Diversity Evaluation Metric for Machine Learning.” Transactions on Machine Learning Research. https://arxiv.org/abs/2210.02410.
Friston, Karl. 2010. “The Free-Energy Principle: A Unified Brain Theory?” Nature Reviews Neuroscience 11 (2): 127–38.
Geiger, Dan, David Heckerman, Henry King, and Christopher Meek. 2001. “Stratified Exponential Families: Graphical Models and Model Selection.” The Annals of Statistics 29 (2): 505–29. https://doi.org/10.1214/aos/1009210550.
Gibbs, Josiah Willard. 1902. Elementary Principles in Statistical Mechanics. Charles Scribner’s Sons.
Gnaiger, Erich. 2009. “Open and Closed Systems: Styles of Thinking Explain Controversies on the ’Negative Entropy’ Concept of Ludwig Boltzmann and Erwin Schrödinger.” Mitochondrial Physiology Network. https://www.mitophysiology.org/images/2/23/Gnaiger_2009_OCESHS.pdf.
Gomes, Ryan, Andreas Krause, and Pietro Perona. 2010. “Discriminative Clustering by Regularized Information Maximization.” Advances in Neural Information Processing Systems (NeurIPS) 23.
Grandvalet, Yves, and Yoshua Bengio. 2004. “Semi-Supervised Learning by Entropy Minimization.” Advances in Neural Information Processing Systems (NeurIPS) 17.
Grathwohl, Will, Kevin Swersky, Milad Hashemi, David Duvenaud, and Chris Maddison. 2021. “Oops I Took a Gradient: Scalable Sampling for Discrete Distributions.” International Conference on Machine Learning (ICML), 3831–41.
Grathwohl, Will, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud, Mohammad Norouzi, and Kevin Swersky. 2020. “Your Classifier Is Secretly an Energy Based Model and You Should Treat It Like One.” International Conference on Learning Representations (ICLR).
Grünwald, Peter D. 2007. The Minimum Description Length Principle. MIT Press.
Gutmann, Michael, and Aapo Hyvärinen. 2010. “Noise-Contrastive Estimation: A New Estimation Principle for Unnormalized Statistical Models.” International Conference on Artificial Intelligence and Statistics (AISTATS), 297–304.
GX-Chen, Anthony, Jatin Prakash, Jeff Guo, Rob Fergus, and Rajesh Ranganath. 2025. “KL-Regularized Reinforcement Learning Is Designed to Mode Collapse.” arXiv Preprint arXiv:2510.20817. https://arxiv.org/abs/2510.20817.
Haarnoja, Tuomas, Haoran Tang, Pieter Abbeel, and Sergey Levine. 2017. “Reinforcement Learning with Deep Energy-Based Policies.” International Conference on Machine Learning (ICML), 1352–61.
Haarnoja, Tuomas, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018. “Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor.” International Conference on Machine Learning (ICML), 1861–70.
Hill, M. O. 1973. “Diversity and Evenness: A Unifying Notation and Its Consequences.” Ecology 54 (2): 427–32. https://doi.org/10.2307/1934352.
Hinton, Geoffrey E. 2002. “Training Products of Experts by Minimizing Contrastive Divergence.” Neural Computation 14 (8): 1771–800. https://doi.org/10.1162/089976602760128018.
Hirsh, Jacob B., Raymond A. Mar, and Jordan B. Peterson. 2012. “Psychological Entropy: A Framework for Understanding Uncertainty-Related Anxiety.” Psychological Review 119 (2): 304–20. https://doi.org/10.1037/a0026767.
Ho, Jonathan, Ajay Jain, and Pieter Abbeel. 2020. “Denoising Diffusion Probabilistic Models.” Advances in Neural Information Processing Systems (NeurIPS).
Hopfield, John J. 1982. “Neural Networks and Physical Systems with Emergent Collective Computational Abilities.” Proceedings of the National Academy of Sciences 79 (8): 2554–58. https://doi.org/10.1073/pnas.79.8.2554.
Hu, Weihua, Takeru Miyato, Seiya Tokui, Eiichi Matsumoto, and Masashi Sugiyama. 2017. “Learning Discrete Representations via Information Maximizing Self-Augmented Training.” International Conference on Machine Learning (ICML). https://arxiv.org/abs/1702.08720.
Hyvärinen, Aapo. 2005. “Estimation of Non-Normalized Statistical Models by Score Matching.” Journal of Machine Learning Research 6: 695–709.
Hyvärinen, Aapo. 2007. “Connections Between Score Matching, Contrastive Divergence, and Pseudolikelihood for Continuous-Valued Variables.” IEEE Transactions on Neural Networks 18 (5): 1529–31.
Jaynes, Edwin T. 1957. “Information Theory and Statistical Mechanics.” Physical Review 106: 620.
Jaynes, Edwin T. 1980. “The Minimum Entropy Production Principle.” Annual Review of Physical Chemistry 31: 579–601. https://doi.org/10.1146/annurev.pc.31.100180.003051.
Kianercy, Ardeshir, and Aram Galstyan. 2012. “Dynamics of Boltzmann q Learning in Two-Player Two-Action Games.” Physical Review E 85 (4): 041145. https://doi.org/10.1103/PhysRevE.85.041145.
Kirkpatrick, Scott, C. Daniel Gelatt, and Mario P. Vecchi. 1983. “Optimization by Simulated Annealing.” Science 220 (4598): 671–80. https://doi.org/10.1126/science.220.4598.671.
Koller, Daphne, and Nir Friedman. 2009. Probabilistic Graphical Models: Principles and Techniques. MIT Press.
Kuhn, Lorenz, Yarin Gal, and Sebastian Farquhar. 2023. “Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation.” International Conference on Learning Representations (ICLR). https://arxiv.org/abs/2302.09664.
Lafferty, John, Andrew McCallum, and Fernando Pereira. 2001. “Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data.” International Conference on Machine Learning (ICML), 282–89.
Landauer, Rolf. 1975. “Inadequacy of Entropy and Entropy Derivatives in Characterizing the Steady State.” Physical Review A 12 (2): 636–38. https://doi.org/10.1103/PhysRevA.12.636.
LeCun, Yann. 2022. “A Path Towards Autonomous Machine Intelligence.” OpenReview Preprint.
LeCun, Yann, Sumit Chopra, Raia Hadsell, Marc’Aurelio Ranzato, and Fu Jie Huang. 2006. “A Tutorial on Energy-Based Learning.” In Predicting Structured Data, edited by Gökhan Bakir, Thomas Hofmann, Bernhard Schölkopf, Alexander J. Smola, and Ben Taskar. MIT Press.
Lee, Dong-Hyun. 2013. “Pseudo-Label: The Simple and Efficient Semi-Supervised Learning Method for Deep Neural Networks.” ICML Workshop on Challenges in Representation Learning.
Lee, Jonghyun, Dahuin Jung, Saehyung Lee, et al. 2024. “Entropy Is Not Enough for Test-Time Adaptation: From the Perspective of Disentangled Factors.” International Conference on Learning Representations (ICLR). https://arxiv.org/abs/2403.07366.
Leinster, Tom, and Christina A. Cobbold. 2012. “Measuring Diversity: The Importance of Species Similarity.” Ecology 93 (3): 477–89. https://doi.org/10.1890/10-2402.1.
Liang, Jian, Dapeng Hu, and Jiashi Feng. 2020. “Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain Adaptation.” International Conference on Machine Learning (ICML). https://arxiv.org/abs/2002.08546.
Liu, Weitang, Xiaoyun Wang, John D. Owens, and Yixuan Li. 2020. “Energy-Based Out-of-Distribution Detection.” Advances in Neural Information Processing Systems (NeurIPS).
Lynn, Christopher W., Qiwei Yu, Rich Pang, Stephanie E. Palmer, and William Bialek. 2025. “Exact Minimax Entropy Models of Large-Scale Neuronal Activity.” Physical Review E 111 (5): 054411. https://doi.org/10.1103/PhysRevE.111.054411.
Ma, Xueguang, Qian Liu, Dongfu Jiang, Ge Zhang, Zejun Ma, and Wenhu Chen. 2025. “General-Reasoner: Advancing LLM Reasoning Across All Domains.” arXiv Preprint arXiv:2505.14652. https://arxiv.org/abs/2505.14652.
Maes, Christian, and Karel Netočný. 2007. “Minimum Entropy Production Principle from a Dynamical Fluctuation Law.” Journal of Mathematical Physics 48 (5): 053306. https://doi.org/10.1063/1.2738753.
Maes, Christian, and Karel Netočný. 2013. “Minimum Entropy Production Principle.” Scholarpedia 8 (7): 9664. https://doi.org/10.4249/scholarpedia.9664.
Martyushev, Leonid M., A. S. Nazarova, and Vladimir D. Seleznev. 2007. “On the Problem of the Minimum Entropy Production in the Nonequilibrium Stationary State.” Journal of Physics A: Mathematical and Theoretical 40 (3): 371–80. https://doi.org/10.1088/1751-8113/40/3/002.
Martyushev, Leonid M., and Vladimir D. Seleznev. 2006. “Maximum Entropy Production Principle in Physics, Chemistry and Biology.” Physics Reports 426 (1): 1–45. https://doi.org/10.1016/j.physrep.2005.12.001.
Martyushev, Leonid M., and Vladimir D. Seleznev. 2013. “Entropy and Entropy Production: Old Misconceptions and New Breakthroughs.” Entropy 15 (4): 1152–70. https://doi.org/10.3390/e15041152.
Menon, Aditya Krishna, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain, Andreas Veit, and Sanjiv Kumar. 2021. “Long-Tail Learning via Logit Adjustment.” International Conference on Learning Representations (ICLR). https://arxiv.org/abs/2007.07314.
Mnih, Andriy, and Yee Whye Teh. 2012. “A Fast and Simple Algorithm for Training Neural Probabilistic Language Models.” International Conference on Machine Learning (ICML), 419–26.
Moussouris, John. 1974. “Gibbs and Markov Random Systems with Constraints.” Journal of Statistical Physics 10 (1): 11–33. https://doi.org/10.1007/BF01011714.
Neal, Radford M., and Geoffrey E. Hinton. 1998. “A View of the EM Algorithm That Justifies Incremental, Sparse, and Other Variants.” In Learning in Graphical Models, edited by Michael I. Jordan. Kluwer Academic Publishers.
Nemhauser, George L., Laurence A. Wolsey, and Marshall L. Fisher. 1978. “An Analysis of Approximations for Maximizing Submodular Set Functions—i.” Mathematical Programming 14: 265–94.
Neumann, John von. 1927. “Thermodynamik Quantenmechanischer Gesamtheiten.” Nachrichten von Der Gesellschaft Der Wissenschaften Zu Göttingen, Mathematisch-Physikalische Klasse 1927: 273–91.
Nijkamp, Erik, Mitch Hill, Song-Chun Zhu, and Ying Nian Wu. 2019. “Learning Non-Convergent Non-Persistent Short-Run MCMC Toward Energy-Based Model.” Advances in Neural Information Processing Systems (NeurIPS).
Nikitin, Alexander, Jannik Kossen, Yarin Gal, and Pekka Marttinen. 2024. “Kernel Language Entropy: Fine-Grained Uncertainty Quantification for LLMs from Semantic Similarities.” Advances in Neural Information Processing Systems (NeurIPS). https://arxiv.org/abs/2405.20003.
Niu, Shuaicheng, Jiaxiang Wu, Yifan Zhang, et al. 2022. “Efficient Test-Time Model Adaptation Without Forgetting.” International Conference on Machine Learning (ICML). https://arxiv.org/abs/2204.02610.
Niu, Shuaicheng, Jiaxiang Wu, Yifan Zhang, et al. 2023. “Towards Stable Test-Time Adaptation in Dynamic Wild World.” International Conference on Learning Representations (ICLR). https://arxiv.org/abs/2302.12400.
Owen, Guillermo. 1972. “Multilinear Extensions of Games.” Management Science 18 (5): P64–79. https://doi.org/10.1287/mnsc.18.5.P64.
Owen, Guillermo. 1975. “Multilinear Extensions and the Banzhaf Value.” Naval Research Logistics Quarterly 22 (4): 741–50. https://doi.org/10.1002/nav.3800220409.
Paltridge, Garth W. 1975. “Global Dynamics and Climate — a System of Minimum Entropy Exchange.” Quarterly Journal of the Royal Meteorological Society 101 (429): 475–84. https://doi.org/10.1002/qj.49710142906.
Paninski, Liam. 2003. “Estimation of Entropy and Mutual Information.” Neural Computation 15 (6): 1191–253. https://doi.org/10.1162/089976603321780272.
Pasarkar, Amey P., and Adji Bousso Dieng. 2024. “Cousins of the Vendi Score: A Family of Similarity-Based Diversity Metrics for Science and Machine Learning.” International Conference on Artificial Intelligence and Statistics (AISTATS), PMLR, vol. 238. https://arxiv.org/abs/2310.12952.
Pearl, Judea. 1988. Probabilistic Reasoning in Intelligent Systems: Networks of Plausible Inference. Morgan Kaufmann.
Prigogine, Ilya. 1945. “Modération Et Transformations Irréversibles Des Systèmes Ouverts.” Bulletin de La Classe Des Sciences, Académie Royale de Belgique 31: 600–606.
Prigogine, Ilya. 1947. Étude Thermodynamique Des Phénomènes Irréversibles. Dunod, Paris; Desoer, Liège.
Prigogine, Ilya. 1978. “Time, Structure, and Fluctuations.” Science 201 (4358): 777–85. https://doi.org/10.1126/science.201.4358.777.
Prigogine, Ilya, and Jean-Marie Wiame. 1946. “Biologie Et Thermodynamique Des Phénomènes Irréversibles.” Experientia 2 (11): 451–53. https://doi.org/10.1007/BF02153597.
Qin, Lianhui, Sean Welleck, Daniel Khashabi, and Yejin Choi. 2022. COLD Decoding: Energy-Based Constrained Text Generation with Langevin Dynamics.” Advances in Neural Information Processing Systems (NeurIPS).
Rényi, Alfréd. 1961. “On Measures of Entropy and Information.” Proceedings of the Fourth Berkeley Symposium on Mathematical Statistics and Probability, Volume 1: Contributions to the Theory of Statistics, 547–61.
Rose, Kenneth, Eitan Gurewitz, and Geoffrey C. Fox. 1990. “Statistical Mechanics and Phase Transitions in Clustering.” Physical Review Letters 65 (8): 945–48. https://doi.org/10.1103/PhysRevLett.65.945.
Sánchez Giraldo, Luis Gonzalo, Murali Rao, and José C. Príncipe. 2015. “Measures of Entropy from Data Using Infinitely Divisible Kernels.” IEEE Transactions on Information Theory 61 (1): 535–48. https://doi.org/10.1109/TIT.2014.2370058.
Schrödinger, Erwin. 1944. What Is Life? The Physical Aspect of the Living Cell. Cambridge University Press.
Shannon, Claude E. 1948. “A Mathematical Theory of Communication.” Bell System Technical Journal 27 (3): 379–423.
Shao, Zhihong, Peiyi Wang, Qihao Zhu, et al. 2024. “DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.” arXiv Preprint arXiv:2402.03300. https://arxiv.org/abs/2402.03300.
Shi, Chence, Shitong Luo, Minkai Xu, and Jian Tang. 2021. “Learning Gradient Fields for Molecular Conformation Generation.” International Conference on Machine Learning (ICML).
Sohl-Dickstein, Jascha, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015. “Deep Unsupervised Learning Using Nonequilibrium Thermodynamics.” International Conference on Machine Learning (ICML), 2256–65.
Song, Yang, and Stefano Ermon. 2019. “Generative Modeling by Estimating Gradients of the Data Distribution.” Advances in Neural Information Processing Systems (NeurIPS).
Song, Yang, Sahaj Garg, Jiaxin Shi, and Stefano Ermon. 2019. “Sliced Score Matching: A Scalable Approach to Density and Score Estimation.” Uncertainty in Artificial Intelligence (UAI).
Song, Yang, and Diederik P. Kingma. 2021. “How to Train Your Energy-Based Models.” arXiv Preprint arXiv:2101.03288.
Song, Yang, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. 2021. “Score-Based Generative Modeling Through Stochastic Differential Equations.” International Conference on Learning Representations (ICLR).
Tanaka, Daiki, Daiki Ikami, Toshihiko Yamasaki, and Kiyoharu Aizawa. 2018. “Joint Optimization Framework for Learning with Noisy Labels.” IEEE Conference on Computer Vision and Pattern Recognition (CVPR). https://arxiv.org/abs/1803.11364.
The Royal Swedish Academy of Sciences. 2024. Scientific Background to the Nobel Prize in Physics 2024: For Foundational Discoveries and Inventions That Enable Machine Learning with Artificial Neural Networks. https://www.nobelprize.org/prizes/physics/2024/advanced-information/.
Tieleman, Tijmen. 2008. “Training Restricted Boltzmann Machines Using Approximations to the Likelihood Gradient.” International Conference on Machine Learning (ICML), 1064–71.
Tsallis, Constantino. 1988. “Possible Generalization of Boltzmann–Gibbs Statistics.” Journal of Statistical Physics 52 (1–2): 479–87.
Verschaffelt, Jules-Émile. 1954. “Sur Les Minima de Production d’entropie Et de Dissipation d’énergie.” Bulletin de La Classe Des Sciences, Académie Royale de Belgique 40: 779–83.
Vincent, Pascal. 2011. “A Connection Between Score Matching and Denoising Autoencoders.” Neural Computation 23 (7): 1661–74.
Wainwright, Martin J., and Michael I. Jordan. 2008. “Graphical Models, Exponential Families, and Variational Inference.” Foundations and Trends in Machine Learning 1 (1–2): 1–305. https://doi.org/10.1561/2200000001.
Wang, Dequan, Evan Shelhamer, Shaoteng Liu, Bruno Olshausen, and Trevor Darrell. 2021. “Tent: Fully Test-Time Adaptation by Entropy Minimization.” International Conference on Learning Representations (ICLR). https://arxiv.org/abs/2006.10726.
Wang, Xudong, Zhirong Wu, Long Lian, and Stella X. Yu. 2022. “Debiased Learning from Naturally Imbalanced Pseudo-Labels.” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 14647–57. https://doi.org/10.1109/CVPR52688.2022.01424.
Wu, Fa-Yueh. 1982. “The Potts Model.” Reviews of Modern Physics 54 (1): 235–68. https://doi.org/10.1103/RevModPhys.54.235.
Wu, Jiaxiang, Tao Shen, Haidong Lan, Yatao Bian, and Junzhou Huang. 2021. SE(3)-Equivariant Energy-Based Models for End-to-End Protein Folding.” bioRxiv Preprint.
Wu, Tailin, and Ian Fischer. 2020. “Phase Transitions for the Information Bottleneck in Representation Learning.” International Conference on Learning Representations (ICLR). https://arxiv.org/abs/2001.01878.
Wu, Tailin, Ian Fischer, Isaac L. Chuang, and Max Tegmark. 2019. “Learnability for the Information Bottleneck.” Uncertainty in Artificial Intelligence (UAI). https://arxiv.org/abs/1907.07331.
Zhang, Qingyang, Haitao Wu, Changqing Zhang, Peilin Zhao, and Yatao Bian. 2025. “Right Question Is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization.” Advances in Neural Information Processing Systems (NeurIPS). https://arxiv.org/abs/2504.05812.
Zhao, Xuandong, Zhewei Kang, Aosong Feng, Sergey Levine, and Dawn Song. 2025. “Learning to Reason Without External Rewards.” arXiv Preprint arXiv:2505.19590. https://arxiv.org/abs/2505.19590.
Zhu, Song Chun, Ying Nian Wu, and David Mumford. 1997. “Minimax Entropy Principle and Its Application to Texture Modeling.” Neural Computation 9 (8): 1627–60. https://doi.org/10.1162/neco.1997.9.8.1627.
Ziebart, Brian D., Andrew Maas, J. Andrew Bagnell, and Anind K. Dey. 2008. “Maximum Entropy Inverse Reinforcement Learning.” AAAI Conference on Artificial Intelligence, 1433–38.