scieee AI-readable full text Open interactive document viewer

Maxwell-Boltzmann Distribution and Minimization of Information Versus Extremization

Ruggeri, Francesco R.

Abstract

Extremization is a mathematical procedure which involves taking the functional derivative of a general function of functions subject to constraints. Formally, if the general function only involves the appearance of the constituents functions without any derivatives of these functions, then the constituent functions appear as variables and one may take the functional derivative as one might take a usual variable derivative. The idea of extremization is that the general function has a certain maximum or minimum value and slight changes in the values of the constituent functions do not change it to first order. This is a formal math procedure, but nevertheless one that is used very often. One example of a quantity which may be minimized is information, i.e. the general function is a measure of information. At this point, information is simply a word, but it is linked a priori with constraints which might involve an increase in information. Thus, minimization of information should be associated with the elimination of all possible constraints, keeping only a constraint(s) which is absolutely necessary. A math process which ensures that all unnecessary constraints are removed should produce the same result as the extremization of a general function representing information (subject to the necessary constraints). As examples, we consider the extremization of Shanon’s entropy - Sum over i p(ei) ln(p(ei)) subject to Sum over i ei p(ei) = Etotal and reaction balance, e1+e2= e3+e4 with p(e1)p(e2) = p(e3)p(e4). We have pointed this out in a previous note, but did not give a formal argument as to why this is the case. We try to do so here and show that no extremization process is needed even in the a priori derivation of Shannon’s entropy from Product over i p(ei) power (Np(ei)).

Full text

Maxwell-Boltzmann Distribution and Minimization of Information Versus Extremization Francesco R. Ruggeri Hanwell, N.B. Dec.4 , 2025 Extremization is a mathematical procedure which involves taking the functional derivative of a general function of functions subject to constraints. Formally, if the general function only involves the appearance of the constituents functions without any derivatives of these functions, then the constituent functions appear as variables and one may take the functional derivative as one might take a usual variable derivative. The idea of extremization is that the general function has a certain maximum or minimum value and slight changes in the values of the constituent functions do not change it to first order. This is a formal math procedure, but nevertheless one that is used very often. One example of a quantity which may be minimized is information, i.e. the general function is a measure of information. At this point, information is simply a word, but it is linked a priori with constraints which might involve an increase in information. Thus, minimization of information should be associated with the elimination of all possible constraints, keeping only a constraint(s) which is absolutely necessary. A math process which ensures that all unnecessary constraints are removed should produce the same result as the extremization of a general function representing information (subject to the necessary constraints). As examples, we consider the extremization of Shanon’s entropy - Sum over i p(ei) ln(p(ei)) subject to Sum over i ei p(ei) = Etotal and reaction balance, e1+e2= e3+e4 with p(e1)p(e2) = p(e3)p(e4). We have pointed this out in a previous note, but did not give a formal argument as to why this is the case. We try to do so here and show that no extremization process is needed even in the a priori derivation of Shannon’s entropy from Product over i p(ei) power (Np(ei)). A Priori Definition of a General Function Linked to Overall Information We discuss above the maximization of a general function linked to a minimal amount of information which we take to mean a maximum number of states that a system may be in. Thus: Probability( particular arrangement of p(ei)s ) = 1/ Number of states ((1)) Here p(ei) is the probability to have a particle with energy ei. It is assumed that the total number of states N is extremely large so: p(ei)N = # particles with ei ((2)) Probability(particular arrangement of p(ei)s) = Product over i p(ei) power ( N p(ei) ((3)) Or -ln(# states) = Sum over i p(ei) ln(p(ei) ((4)) ((4)) is a statement of Shannon’s entropy and represents a priori a function which measures the number of states (ln(states) and so p(ei)’s may be adjusted to ensure that one has a maximum number, i..e minimum 1/ (#states) subject to Sum over i p(ei) ei = Etotal ((5)) This is the usual maximization of Shannon’s entropy subject to the constraint Sum over i ei p(ei). We argue, however, that one does not have to find an a priori function ((4)) and maximize it in order to find the solution exp(-ei/T) of the maximization of entropy subject to a constraint. Elastic Scattering Approach To show that one may find the minimal information approach without an extremization process, one may consider time reversal reaction balance: E1+e2 = e3+e4 ((5a)) and p(e1)p(e2) = p(e3)p(e4) ((5b)) Because any initial set has the same probability as a final set as long as energy is conserved, one has removed any bias other than conservation of energy and so ((5a)) and ((5b)) should represent a minimal information solution, i.e. exp(-ei/T) ((6)) This is the same solution as found in the above section and we noted that the two approaches must thus be equivalent in previous notes, but have not explained why. General Ideas Associated with Minimal Information in an Ideal Gas Given a gas, one has the additive conserved quantity called single particle energy, i.e. Sum over i=1, N n(ei) ei = E total Here n(ei) = #particles with energy ei ((7)) (Particle number is also additive, but we take this as a given.) We note that in classical probability: p1* p2 = probability of an AND scenario ((8)) Given the assumption that p(ei), we note that there exists an additive quantity: ln(p(ei)p(ej)) = ln(p(ei)) + ln(p(ej)) ((9)) (This is called information in information theory.) As a result, ei and ln(p(ei)) seem to be two separate additive quantities based on ei. This assumes extra information in the system because one has to account for both ei and ln(p(ei)).The simplest scenario (least information) would be to assume that: -ei/T = ln(p(ei) ((10)) (Here -1/T is simply a constant) Thus, one arrives at a minimal information solution without any need for the concept of reaction balance or maximization of entropy subject to a constraint. We wish, however, to see why ((10)) should be equivalent to maximization of entropy subject to a constraint. If one considers the maximization of entropy H(p(ei)) linked to the number of states, then the a priori constraint one requires is: Sum over i ei p(ei) = E total ((11)) Thus, dH/dp = constant * ei ((12)) Given that one deals with probabilities, ln(p(ei)) existing as an additive entity must exist. Thus, one would again have to assume that ei and ln(p(ei)) are independent additive quantities and this is not the simplest (least constrained) solution. The least constrained statement is ((10)) which dictates the form of H(p(ei)). Thus, H(p(ei)) is actually created in order to equal ln(p(ei)) when one takes dH/dp if one wishes to have minimal information. An issue, however, is that H(p(ei)) was obtained a priori in the first section as the ln of the AND probability of N single particle probabilities. No assumption of ((10)) was used in obtaining Shannon’s entropy form. It simply appeared by assuming a p(ei) power (Np(ei)) for each ei. We note, however, that this approach assumes independent probabilities and so uses probability to a power, i.e. p(ei) power f(N,ei) → exp ( f(N,ei) ln(p(ei)) ((13)) Here f(N,ei) is the number of occurrences of ei. ln(p(ei)), however, is an additive quantity by probability theory and the existing additive quantity in the problem is ei. Thus, the simplest solution (least information) is one in which ln(p(ei)) = -ei/T. Given the product of all ((13))’s one has: Sum over i f(N,ei) (-ei/T) ((14)) There is already a constraint Sum over i ei p(ei) = E total so in order to have the minimal amount of information: f(N,ei) = p(ei) ((15)) Thus, one does not have to really use extremization even in the a priori definitional approach of Shannon’s entropy presented in section 1, we argue. Conclusion In conclusion, one often sees the Maxwell-Boltzmann distribution derived by maximizing Shannon’s entropy subject to the a priori constraint Sum over i ei p(ei). Shannon’s entropy may be obtained a priori using: 1/ number of states = Product over i p(ei) power (Np(ei)) for a very large N set of particles. As a result, one has the impression that the mathematical extremization process subject to the a priori constraint is required. On the other hand, the same Maxwell-Boltzmann solution exp(-ei/T) is obtained using time reversal reaction balance through e1+e2 = e3+e4 and p(e1)p(e2) = p(e3)p(e4). This begs the question: Why was extremization not required? The reason is that it is not necessary. We note that in the problem of an ideal gas, one has the additive entity energy. Probability p(ei) is also additive, but p(ei) cannot equal ei. From properties of an AND probability situation, namely p(ei)p(ej) → ln(p(ei)p(ej)) = ln(p(ei)) + ln(p(ej)), one sees that a second additive quantity exists in the problem (which is not probability or particle number). To have the simplest scenario, i.e. least information, one should have ln(p(ei)) = -ei/T so there is essentially only one additive quantity. Thus, one obtains the MB distribution with no formal extremization process. We also show that these same arguments may be used in the a priori derivation of Shannon’s entropy from 1/number of states = Product over i p(ei) power (Np(ei)). We note that even before one obtains Shannon’s entropy and extremizes one may observe that ln(p(ei)) is an additive quantity from probability theory. In the simplest case (least information) it should be identified with -ei/T. The full form is then: Sum over i p(ei) ln(p(ei)) which matches the constraint. As a result, no extremization of entropy is needed in this case to find the MB distribution, we argue. References 1. Hosoya, A. et al (2015) Operational Derivation of Boltzmann distribution with Maxwell’s demon model Scientific Reports volume 5, Article number: 17011 (2015) https://www.nature.com/articles/srep17011