scieee AI-readable full text Open interactive document viewer

Derivation of Shannon's Entropy for Constraints Sum over i g(ei) P(ei) with Probability P(ei)

Ruggeri, Francesco R.

Abstract

Dec. 27, 2025 Note: For P(ei)P(ej) = P(ei+ej) and P1(g(ei)) P1(g(ej)) = P1(g(ei) + g(ej)), P and P1 are unnormalized. Dec. 27, 2025 The constraint really being used in this note is Sum over i g(ei) P(ei) and not Sum over i f(ei). In the note, the latter is converted into the former, but this means having an unknown P(ei) within g(ei). Thus, the constraint should be Sum over i g(ei) P(ei), where g(ei) is a known function. In (1), it is suggested that there are two fundamental approaches to obtaining a probability distribution. The first, called the principle of indifference, is associated with a uniform distribution as in the toss of a coin or die. If a constraint on probabilities is present, however, then there is bias and (1) suggests that Shannon’s entropy -Sum over i p(ei) ln(p(ei)) be maximized subject to the constraint. In (2), we argued that for a constraint associated with a conserved quantity, i.e. Sum over i ei p(ei) = Eave, one may take a subset of p(ei)s ( 0 <= ei <= E1) and apply the notion of a uniform distribution to p(ei)p(ej) in this subset without any other constraint. Then one immediately has a solution from p(ei)p(ej) = p(ei+ej), i.e. p(ei) = C exp(-ei/T) for a probability which drops with rising ei. Here we argue that one may extend this idea to a constraint of the form Sum over i f(ei) as this may be rewritten in terms of conserved quantity, i.e. Sum over i f(ei)/P(ei) P(ei) with g(ei) = f(ei)/P(ei) being seen as the conserved quantity. In such a case, one may write P(ei) as P1( g(ei)) and use P1(g(ei)) P1(g(ej)) = P1( g(ei)+g(ej)). Then the solution is P1(g(ei)) = exp(-g(ei)/T) or ln(P1(g(ei)) = g(ei). The above ideas may then be applied to the full P(ei) set because one may argue that one seeks a special function F(P(ei)) which is maximized under P(ei)--> P(ei) + dP(ei) subject to the constraint. This means that dF/dP1 = g(ei) from the subset uniform product probability reasoning. Thus, F= -Sum over i P1(g(ei)) ln(P1(g(ei)). Now P1(g(ei)) = P(ei) so F = - Sum over i P(ei) ln(P(ei)) and so we argue that Shannon’s entropy arises in such a manner. In other words, for a constraint of the form Sum over i f(ei) = number, there is one fundamental idea, namely that of a uniform distribution. The notion of Shannon’s entropy and its maximization subject to a constraint follow from the uniform distribution on a subset of P(ei)s when one focuses on P(ei)P(ej).

Full text

Derivation of Shannon’s Entropy for Constraints Sum over i f(ei) with Probability P(ei) Francesco R. Ruggeri Hanwell, N.B. Dec. 27, 2025 In (1), it is suggested that there are two fundamental approaches to obtaining a probability distribution. The first, called the principle of indifference, is associated with a uniform distribution as in the toss of a coin or die. If a constraint on probabilities is present, however, then there is bias and (1) suggests that Shannon’s entropy -Sum over i p(ei) ln(p(ei)) be maximized subject to the constraint. In (2), we argued that for a constraint associated with a conserved quantity, i.e. Sum over i ei p(ei) = Eave, one may take a subset of p(ei)s ( 0 <= ei <= E1) and apply the notion of a uniform distribution to p(ei)p(ej) in this subset without any other constraint. Then one immediately has a solution from p(ei)p(ej) = p(ei+ej), i.e. p(ei) = C exp(-ei/T) for a probability which drops with rising ei. Here we argue that one may extend this idea to a constraint of the form Sum over i f(ei) as this may be rewritten in terms of conserved quantity, i.e. Sum over i f(ei)/P(ei) P(ei) with g(ei) = f(ei)/P(ei) being seen as the conserved quantity. In such a case, one may write P(ei) as P1( g(ei)) and use P1(g(ei)) P1(g(ej)) = P1( g(ei)+g(ej)). Then the solution is P1(g(ei)) = exp(-g(ei)/T) or ln(P1(g(ei)) = g(ei). The above ideas may then be applied to the full P(ei) set because one may argue that one seeks a special function F(P(ei)) which is maximized under P(ei)--> P(ei) + dP(ei) subject to the constraint. This means that dF/dP1 = g(ei) from the subset uniform product probability reasoning. Thus, F= -Sum over i P1(g(ei)) ln(P1(g(ei)). Now P1(g(ei)) = P(ei) so F = - Sum over i P(ei) ln(P(ei)) and so we argue that Shannon’s entropy arises in such a manner. In other words, for a constraint of the form Sum over i f(ei) = number, there is one fundamental idea, namely that of a uniform distribution. The notion of Shannon’s entropy and its maximization subject to a constraint follow from the uniform distribution on a subset of P(ei)s when one focuses on P(ei)P(ej). Two Principles for Determining a Probability P(ei) In (1), it is suggested that there are two general math approaches to finding probability. The first is called the principle of indifference and is associated with a uniform distribution as in the toss of a coin or die. In particular, there is no bias in such a scenario. If, however, there exists a constraint in P(ei)’s, then there is a bias and a uniform distribution no longer holds. In such a case, (1) suggests that one should create Shannon’s entropy -Sum over i P(ei) ln(P(ei)) and maximize it subject to the constraint. We suggest for constraints of the form Sum over i f(ei) = number (where f is an arbitrary function), that one may create a conserved quantity g(ei) and apply a uniform distribution to a P(ei)P(ej) for a subset of P(ei)’s. In other words, Shannon’s entropy and its maximization subject to a constraint on all P(ei)’s follow from the uniform distribution applied to a subset of P(ei)s through the product probability P(ei)P(ej) being uniform for all ei+ej=E1. (E1 may then be changed to take on different values.) Conservation of ei In (2), we argued that if one has the constraint Sum over i ei P(ei) = Eave, one may use the notion of conservation of ei. In particular, one may consider an E1 and two body elastic interactions such that P(ei)P(ej) = P(ek)P(el) if ei+ej = ek+el = E1. This is equivalent to P(ei)P(ej) = P(ei+ej) so P(ei) = C exp(-ei/T) to have P(ei) drop with rising ei. In such a case, one works with a subset of P(ei)s, i.e. only ones associated with 0 <= ei <= E1. Within this subset, the product P(ei)P(ej) is uniform as there is no global constraint on this product probability. In this special case, Cexp(-ei/T) is already known without maximizing Shannon’s entropy subject to the constraint Sum over i ei P(ei). There is, however, nothing wrong with maximizing Shannon’s entropy because from the subset uniform distribution one knows that: ln(P(ei)) = -ei/T (form P(ei) = Cexp(-ei/T) ((1)) If one argues that one seeks a function F(P(ei)) which is maximized under P(ei)-> P(ei)+dP(ei) subject to the constraint Sum over i ei P(ei), then one knows that: dF/dP = -ei/T = ln(P(ei)) + constant ((2)) Note: One may add a constant to ln(P(ei)) which becomes part of the normalization constant. Thus, F is automatically Shannon’s entropy: S = - Sum over i P(ei) ln(P(ei)) ((3)) We wish to generalize this idea to constraints of the form Sum over i f(ei) = number, where f(ei) is an arbitrary function. Constraints of the Form Sum over i f(ei) = number Given a constraint of the form Sum over i f(ei) = number, we argue that one may write it in terms of a conserved quantity: Sum over i f(ei) = Sum over i { f(ei)/P(ei)} P(ei) because P(ei)>0 ((4)) g(ei) = f(ei)/P(ei) is the conserved quantity and: P1(g(ei)) = P(ei) ((5)) One may use the subset of P(ei)’s approach together with the uniform distribution, i.e. P1(g(ei)) P1(g(ej)) = P1(g(ei)+g(ej)) so P1(g(ei)) =C exp(-g(ei)/T) ((6)) or ln(P1(g(ei)) = -g(ei)/T ((7)) If one moves out of the subset of P(ei)’s to the full set of P(ei)’s and the full constraint Sum over i f(ei) = number, one may imagine a function F(P1(ei)) which may be maximized subject to the constraint to produce the probability distribution, i.e. dF/dP1 = -g(ei)/T + constant ((8)) One may add the constant because it becomes part of the normalization constant. Then: F = - Sum over i P1(g(ei)) ln(P1(g(ei)) ((9)) Now P1(g(ei)) = P(ei) so F = - Sum over i P(ei) ln(P(ei)) ((10)) Thus, one obtains the Shannon’s entropy form and its maximization with respect to the full constraint Sum over i f(ei) = number from considerations of a uniform distribution P1(g(ei)P1(g(ej)) on a subject of P(ei)s in general. As a result, for constraint of the form of Sum over i f(ei), maximization of Shannon’s entropy subject to a constraint is not a process which is independent of the uniform distribution, i.e principle of indifference, we argue Conclusion In conclusion, in (1) it is suggested that there are two general mathematical principles for obtaining probability distributions. The first, called the principle of indifference, is based on the uniform distribution as there is no bias and applies to tosses of coins and dice. If a constraint on P(ei)’s (probabilities) exists, then there is a bias and (1) suggests that one maximize Shannon’s entropy - Sum over i P(ei) ln(P(ei)) subject to the constraint. We argue that a constraint of the form Sum over i f(ei) = number, may always be written in terms of a conserved quantity, i.e. Sum over i {f(ei)/P(ei)} P(ei) with g(ei)=conserved quantity. One may then write P1(g(ei)) = P(ei) and consider a subject of P1(ei)s with g(ei)+g(ej) = E1 such that P1(g(ei))P1(g(ej)) = P1(g(ei)+g(ej)) or P(ei) = Cexp(-g(ei)/T). Thus, one only needs the concept of a uniform distribution in such a case, and not the notion of Shannon’s entropy and its maximization subject to a constraint. One may, however, derive Shannon’s entropy and its maximization from the subset-uniform distribution approach based on the result ln(P1(g(ei)) = -g(ei)/T. If one suggests that there is a function F(P1(ei)) which may be maximized subject to the constraint Sum over i f(ei), then dF/dP1 = -g(ei)/T = ln(P1(g(ei))) + constant, with the ln(P1(ei)) following from the subset solution. Then F= - Sum over i P1(g(eI) ln(P1(g(ei)), but P1(g(ei)= P(ei) so F = Shannon’s entropy = - Sum over i P(ei) ln(P(ei)). Thus, the maximization of Shannon’s entropy subject to constraints is not a concept independent of the uniform distribution result for constraints of the form Sum over i f(ei) = number. References 1.Conrad, K. Probability Distributions and Maximum Entropy https://kconrad.math.uconn.edu/blurbs/analysis/entropypost.pdf 2. Ruggeri, Francesco R. Solving for a Probability Distribution P(ei) from a Simpler Constraint for a Subset of P(ei) Derived from a Constraint Over All P(ei) (preprint, zenodo, 2025)