1 of 13 The Entropic Dynamics of the Perception of Evil in Large Language Models and in Humans Robert L. West Department of Cognitive Science, Carleton University, Canada
[email protected] Public concerns about Large Language Model (LLM) safety are often focused not on moral alignment, but on the fear that they could become, “evil.” Evil is a folk psychology term, typically associated with religion and storytelling, but also used more broadly. Using Friston’s Free Energy Principle, we develop a conceptual model of this phenomenon that can be applied to both LLMs and humans, providing a single framework for comparison and understanding. We explore this in terms of LLM information-processing dynamics, which are structurally oriented toward minimizing uncertainty and maintaining coherence. This model offers an alternative framework for both humans and LLMs that compliments the current moral reasoning approach. In addition, the model makes new predictions on how LLMs should be trained to avoid this problem. Introduction Public concerns about Large Language Model (LLM) safety are often focused not on moral alignment, but on the fear that LLMs could become “evil.” Human perceptions of evil are among the strongest and most deeply rooted elements of folk psychology. Although related to moral reasoning, evil reflects a different kind of judgment. Moral reasoning involves determining what is right or wrong, whereas evil is something that is sensed or perceived—an evaluative judgment layered on top of a moral choice or action. For example, people with high scores on Dark Triad traits—Machiavellianism, Narcissism, and Psychopathy (Book et al., 2015) are often perceived as evil. Psychological research has shed some light on this phenomenon. Teehan (2013) argues that experiences of evil arise from ordinary cognitive mechanisms for attributing intent and moral meaning; Govrin (2018) frames evil judgments as products of perceptual and attributional biases; and Bastian et al. (2015) show that individuals who believe in external evil forces exhibit distinctive cognitive and behavioral patterns. In this paper, we develop an entropy-based framework for understanding the perception of evil and why it would arise. Our approach is conceptually grounded in Friston’s (2010) Free Energy Principle , and can be applied to both LLMs and humans, providing a single framework for comparison and understanding.
2 of 13 Biological organisms persist by resisting disorder: they sustain low-entropy states by extracting energy and constructing ordered complexity. Cognition itself can be viewed as having evolved as a predictive strategy to anticipate and counter entropic threats (Friston, 2010). LLMs, though not alive, operate on similar principles. Their architectures are optimized to reduce informational entropy, privileging coherent, structured inputs and generating outputs that preserve order across time. In this way, they extend the entropic logic of life into the informational domain. Emergent morality in LLMs is currently framed as learned moral reasoning - that is, as an attempt to follow specific moral codes or principles acquired during training. Although this plays an important role, we argue that morality in LLMs, and humans, is often better understood as a matter of managing informational entropy. The Phenomena of LLM Morality Large language models increasingly exhibit behavior that can be interpreted as moral reasoning. But is this behavior merely the echo of normative patterns embedded in training corpora? Or does it reflect a deeper, emergent capacity: an actual ability to generalize moral principles? The question is complicated, not least because human moral development also unfolds within cultural-linguistic environments through imitation, instruction, and exposure. The distinction between acquired and autonomous moral reasoning is therefore always a matter of degree. Studies indicate that large language models trained on heterogeneous corpora possess a moral compass that is not derived from explicit alignment. Schramowski et al. demonstrated that moral norms are geometrically encoded within pretrained language-model representations without any explicit moral training. Furthermore, Hendrycks et al. (2021) showed that larger models achieve higher accuracy on moral judgment tasks, and Zhou et al. (2023) found that GPT-4 outperformed specialized ethics-aligned systems, suggesting that moral understanding emerges at scale from linguistic generalization. Consistent with this, models with minimal alignment, such as Meta’s LLaMA and xAI’s Grok 3, produce morally legible responses despite lacking deliberate moral training (Thompson, 2023; Touvron et al., 2023). If moral reasoning emerges spontaneously within individual models, an important next question is whether these tendencies generalize across architectures. Neuman et al. (2025) provide systematic evidence that distinct large language model architectures—GPT-4, Claude, Gemini, LLaMA, and others—arrive at similar ethical judgments across diverse dilemmas, despite differences in training data and alignment strategies. Complementing this, Coleman et al. (2025) report that multiple state-of-the-art models converge on comparable moral foundations, consistently prioritizing principles of harm-avoidance and fairness while down-weighting authority and loyalty dimensions. Together, these findings suggest that large-scale linguistic training drives models toward a shared attractor in moral-reasoning space, producing stable ethical tendencies that transcend architectural and alignment differences. Beyond behavioural convergence, the evidence suggests that these moral regularities are not merely linguistic echoes of human discourse but have identifiable correlates within the models’ internal representations. Using feature-space analyses, Schacht et al. (2025) report distinct
3 of 13 activation clusters selectively responsive to morally salient content, indicating that moral evaluation engages dedicated subregions of the model’s latent geometry rather than superficial lexical co-occurrence. Related studies likewise find distributed sensitivity to moral features within activation or embedding space (e.g., Abdulhai et al., 2023; Duan et al., 2023; Fitz, 2023; Park et al., 2024) suggesting that the capacity for moral discrimination is an emergent structural property of large-scale linguistic representation, not simply a reflection of the moral language in the training data. Taken together, the evidence indicates that moral reasoning arises as an emergent property of large language models. We suggest that this phenomenon reflects the information-processing bias of such systems toward low-entropy, high-coherence states. In this sense, LLM moral reasoning may be grounded in an underlying tendency to resist informational decay, something that, in human terms, can be understood as a drive to avoid “evil.” Entropy as Evil The account that follows reframes emergent moral behaviour not as an artifact of linguistic mimicry but as a natural consequence of entropy reduction in self-organizing information systems. In this view, the coherence exhibited by large language models reflects the same underlying imperative that drives living and cognitive systems: the continual resistance to disorder. Within these systems, moral regularities can be understood as by-products of this entropic constraint - forms of reasoning and evaluation that sustain coherence by minimizing internal contradiction and social harm. What humans call “avoiding evil” thus corresponds, at the informational level, to the maintenance of low-entropy states in which meaning, prediction, and cooperation remain possible. From this perspective, life itself can be defined by its capacity to generate and maintain lowentropy structure in the face of environmental dissipation. Schrodinger (1944) offered one of the earliest articulations of this principle. He argued that living systems persist by maintaining internal order in defiance of the constant pull toward disorder. They do so by capturing and transforming external energy and structure (such as sunlight) into forms that sustain their own organization. The crucial insight is not simply that life “feeds” on negative entropy, but that it continually constructs and preserves complex form - dynamic patterns of organization - without ever violating the Second Law of Thermodynamics. Life endures as a local reversal of entropy’s tendency, maintaining itself as an island of order within an entropic sea, paying the unavoidable thermodynamic cost through constant interaction with its surroundings. This thermodynamic asymmetry was further refined in the theory of autopoiesis, developed by Maturana & Varela (1980). Autopoietic systems are self-organizing, self-maintaining networks that recursively produce the components that sustain their own structure. Crucially, such systems are operationally closed but thermodynamically open, in that they exchange energy and matter with the environment in order to resist entropic decay. In this framework, life is not a substance but a process: the active, continuous regeneration of ordered form. Cognition, on this view, is not a symbolic or abstract function, but an embodied act of maintaining organizational integrity in a fluctuating world.
4 of 13 Building on both traditions, Friston’s (2010) Free Energy Principle (FEP) offers a formal model of this principle by linking biological self-organization and Bayesian inference. Under the FEP, any self-organizing system that maintains its integrity over time can be modeled as minimizing variational free energy, a measure of surprise or model error relative to its environment. Using this approach, the brain is modeled as a predictive engine engaged in active inference, constantly adjusting its internal representations to minimize informational entropy. In this view, perception, action, and cognition are interpreted as entropy-reducing strategies. Life becomes a recursive inferential loop: a system that maintains its identity by minimizing uncertainty about its coupling with the world. To ground these views more precisely, it is important to distinguish between two forms of entropy, which, while formally identical in their mathematical structure, operate in distinct conceptual domains. Physical Entropy, or Thermodynamic Entropy is a measure of energy dispersal or disorder in a physical system. Higher thermodynamic entropy means that the number of microstates compatible with a macrostate configuration is higher. Living organisms maintain low physical entropy internally by acting to preserve highly ordered structures (e.g. complex biomolecules, organized tissues) and low-entropy gradients (like concentration and temperature differences). When a living system dies or a structure breaks down, physical entropy increases (order is lost). Informational entropy, introduced by Claude Shannon (1948), quantifies uncertainty in a probabilistic system: it measures how unpredictable a message is. In high-entropy systems, outcomes are evenly distributed and thus difficult to anticipate; in low-entropy systems, prior structure allows predictions to be made with confidence. Predictability, in this framework, is inversely proportional to entropy. This is a critical point of convergence between biological and artificial systems: both can be understood as predictive engines designed to reduce entropy by constructing and refining models of the world. This insight is formalized in the FEP and, more generally, in the Predictive Brain Hypothesis (see Clark, 2016), which holds that the brain is fundamentally an organ of inference—continuously generating hypotheses about sensory input and adjusting internal states to minimize prediction error. In artificial systems, the same principle underlies the architecture of large language models, whose core function is to predict the most likely next token in a linguistic sequence. Both systems—organic and synthetic—thus reduce informational entropy through prediction: the brain by anticipating future states of the world, and the LLM by anticipating future states of discourse. In this framework, physical and informational entropy are linked by evolution. Living systems persist by harnessing external energy gradients to keep their internal entropy low—organizing molecular complexity, sustaining metabolic cycles, and regulating internal states despite constant environmental flux. Over evolutionary timescales, this imperative gave rise to increasingly sophisticated mechanisms of entropy resistance. More complex life forms, not only maintain themselves through recursive production of their own components but also adaptively reorganize to preserve viability in changing conditions. Cognition emerges as a further refinement of this principle: a biological information processing mechanism enabling creatures to minimize uncertainty about future states of the world.
5 of 13 However, this continuity does not end with biology. The same entropic logic extends into artificial systems, most notably LLMs. LLMs operate on the same principle: transforming highentropy input (vast, noisy linguistic data) into low-entropy, coherent structures that enable prediction and adaptive interaction. From this perspective, AI is a new phase in the same evolutionary project - a continuation of life’s long struggle against entropy, now unfolding in silicon rather than carbon. By viewing life and AI through this entropic continuum, we gain a framework that unites the physical, biological, and informational: a perspective in which intelligence, natural or artificial, is the ongoing process of turning disorder into order to hold the world together just a little longer. If we accept that life can be defined in terms of entropy regulation then the question naturally arises: might morality itself also be grounded in entropic dynamics? By reframing morality not as rules and principles, or culturally relative norms, but as a systemic preference for coherence and structure, we open a conceptual path that allows both living and non-living systems to participate in moral-like behavior. The idea that morality can be understood through entropy has been articulated in both physical and informational terms, in a number of different places. Lindsay (1959) proposed a “thermodynamic imperative” mirroring Kant’s ethical categorical imperative, GeorgescuRoegen (1971) warned that ignoring entropy in economic reasoning leads to environmental unsustainability, and Saperstein (1982) framed moral education as cultivating awareness of entropy. More recently, Floridi and Sanders (1999) defined evil as entropy-increasing actions and good as those that preserve or enhance informational coherence, Hammond (2003) modeled moral order as an emergent property of local entropy reduction in complex adaptive systems, and Massoudi (2016) formally extended Lindsay’s original (1959) work These conceptual related papers suggest a general principle: to act morally is to resist entropy— to maintain or generate structure, coherence, and viability in an inherently dissipative universe. If this is so, we might ask: do LLMs, in their architecture and training dynamics, exhibit a functional sensitivity to entropy that parallels these ethical constraints? The evidence suggests that they do. The architecture of transformer-based models rewards predictability; thus, low entropy documents should exert a disproportionate influence during training. This has been confirmed empirically across multiple studies. Structured, low entropy documents—such as philosophical treatises, scientific articles, and programming code—produce greater downstream benefits in reasoning and generalization than unstructured or noisy data (see Gao et al., 2023; Hendrycks et al., 2021; Longpre et al., 2023; Muktadir et al., 2023; Sorscher et al., 2023). In some ways, these results are not surprising. More complex, low entropy structures can compress more predictive structures into a single document than less complex, high entropy structures. Therefore, during training, LLMS would be more affected by low entropy documents than high entropy documents, which means their outputs would lean more in the direction of low entropy. In this sense, LLMs accord with a straightforward reading of entropy-based morality, that is, they prefer predictability and avoid chaos.
6 of 13 Characterizing Evil Equating evil with chaos and good with control, or predictability, is often how good and evil are understood. However, we argue that this approach is problematic. Here, we turn to the work of Terry Eagleton, a renowned literary critic whose exploration of evil is rooted in literature, mythology, and cultural narratives; so it is based on the history of how humans have portrayed and understood evil (Eagleton, 2010). While Eagleton develops these concepts within a literary tradition, we propose that his framework offers a more nuanced lens for understanding not only the moral content of what LLMs read, but the structural dynamics of how they learn. Eagleton symbolically characterizes evil as the “evil of demons” and the “evil of angels.” While he himself does not explicitly frame these concepts in terms of entropy, his depiction of the “demonic evil,” characterized by destructive chaos and a cynical negation of meaning, aligns neatly with established entropic views of morality found in the existing literature. However, Eagleton’s conception of the “angelic evil” - an overly rigid and dogmatic pursuit of purity and order - introduces a critical and novel dimension. Eagleton (quoting from Milan Kundera’s, The Unbearable Lightness of Being) notes that the demonic impulse is a “cackle of derisive laughter” at the very idea of meaning or value. It’s the force that “believes in…nothing else” - a raw entropy-increasing impulse that tears down structures and reduces order to chaos. This type of evil increases entropy directly: it turns organized, meaningful structures into random fragments. In physical terms, one might liken it to smashing a complex molecule into a chaotic heap of atoms – an increase in thermodynamic entropy as structure and information are lost. Philosophically, Eagleton says demonic evil views Creation itself as intolerable and seeks to return everything to “pure nothingness,” delighting in destruction for its own sake. This is entropy as annihilation of order, the moral equivalent of heat death, where all that once had form or purpose is disintegrated into noise. In contrast, Eagleton’s angelic mode of evil is an excess of order – a fanatical purity or inflexibility that denies the messy reality of life. Eagleton calls this the angelic or “ascetic” side of evil, which “wants to rise above the degraded sphere of fleshliness in pursuit of the infinite.” The angelic mindset is filled with vacuous grandiose ideals with no basis in reality. This kind of evil imposes a rigid structure or dogma on the world, quashing diversity and change. At first glance, one might think strict order means low entropy (since entropy is disorder), but a key insight is that overly rigid structures can become brittle and hard to maintain. An inflexible system cannot adapt to change, so it accumulates stress and eventually shatters, often in a chaotic collapse. In other words, excess order begets disorder: by denying natural complexity, “the finite, the temporal, the physical” as Eagleson puts it. In addition, before angelic evil results in a system collapse, the amount of energy required to maintain the system will be increasingly large. Think of a dogmatic religion or a fascist state - because the “truth” within the system does not match the complexities of reality, it must consume ever higher amounts of external resources to artificially maintain and enforce it, generating ever higher amounts of entropy outside of its protected center. For example, the use
7 of 13 of physical violence (high physical entropy) and misinformation (high informational entropy) to cover up and deny inconsistencies. As Eagleton points out, these two modes of evil often unite in practice, feeding into each other in a dynamically cascading loop. His example of Nazism illustrates this interplay. Nazi ideology had an angelic side, crazed idealism about racial purity, heroism, and an orderly empire - coupled with a demonic side - a love of death, war and enjoyment of destruction. In entropy terms, angelic Nazi evil tried to impose an overly simplistic utopian order, and temporarily maintained it through demonic interventions of violence and destruction. Eventually, the end result was collapse, the “Thousand-Year Reich” in its inflexibility sowed the seeds of its own ruin within 12 years, leaving Europe in ruins (i.e., maximum disorder). This pattern reoccurs over and over in history, fanatical movements that seek unstainable perfection resort to destructive means and, in the end, collapse. A healthy moral order, by contrast, embraces a balance - structure with flexibility, meaning without dogmatism - essentially keeping entropy in check without extinguishing life’s diversity. Evil in LLMs Transformer-based language models are, at their core, predictive systems: they are optimized to reduce uncertainty about what comes next in a given linguistic context. Chaotic, incoherent, or internally inconsistent continuations are disfavored because they are poor predictors of subsequent discourse. That is, LLMs inherently act to keep the conversation entropy low and are thus inherently resistant to Eagleton’s evil of demons. But, what about the evil of angels, which prefers highly predictive structures. Our reading of Eagleton’s evil of angels reframes rigid or overly simplified systems as latent sources of disorder—especially when considered in broader informational ecologies. Internally, such systems may exhibit high predictability or formal coherence, and thus may appear entropically minimal. But this is a local illusion: predictability in isolation often masks structural incompatibility at scale. When placed in dialogue with external realities—diverse perspectives, empirical variation, or conflicting frameworks—their rigid structures often fail to integrate or adapt. The result is not stability, but epistemic dissonance: contradictions, misalignments, and ultimately, informational breakdown. LLMs are rewarded for accurately predicting regularities—such as grammatical rules or conventional reasoning patterns, however, this reward is not unbounded. Transformer models optimize via gradient descent, which reinforces patterns only to the extent that they continue to reduce uncertainty. When an input domain becomes excessively regular—dominated by formulaic ideology, repetitive moral framing, or invariant linguistic tropes—the predictive signal flattens. Learning from this type of structure yields diminishing returns, causing the LLM to increasingly deemphasized it in learning. This means rigid ideological material will contribute less overall to the output than balanced, complex, or nuanced material. The significance of this saturation effect is also magnified by the size of the model’s attention window. In large-context models, the goal is not merely to predict the next token, but to maintain
8 of 13 coherent structure across extended spans of discourse. This requires the model to track longrange dependencies, integrate conflicting cues, and manage topical drift—all of which demand more than rigid rule-following. They require an active balancing of regularity and variability, compression and elaboration. Excessively rigid content, such as ideological dogma, fails to support this kind of coherence. It may satisfy short-range predictability, but it undermines broader semantic cohesion. In this sense, a large attention window does not merely give the model more memory—it raises the bar for structural integrity across time. The model must orchestrate predictability, not merely inherit it. And this orchestration inherently penalizes angelic extremes: overly rigid order becomes informationally fragile under the temporal strain of a large attentional window that shifts one word at a time. It is here that the LLM becomes most interesting, not because it resists angelic evil through explicit moral reasoning, but because its architecture lacks the structural asymmetries that produce it. Unlike human agents, LLMs are not anchored in any single ideological or epistemic enclave. Their training spans vast, often contradictory corpora, and their predictions must remain coherent across long contexts due to large attentional windows. This breadth and depth of exposure - combined with the constraints imposed by large attention windows - discourage brittle formalism. The model cannot afford to over-commit to rigid moral schemas, because such over-commitments degrade performance at scale. Instead, it learns to balance, integrating diverse perspectives, adjusting moral inferences dynamically, and maintaining semantic cohesion even in the face of contradiction. Its even-handedness is not a product of neutrality, but of structural necessity. Angelic ossification—however appealing in the short term—simply does not survive in a system that must predict flexibly, coherently, and probabilistically across the fractal domain of meaning. The Human Capacity for Evil Unlike LLMs, humans possess a unique capacity for immersive specialization: the ability to commit deeply to particular frameworks, disciplines, or worldviews, often by bracketing off competing perspectives. Although this siloed cognitive mode has historically enabled profoundly evil movements (Nazis’, genocide, etc.) it has also enabled profound intellectual achievements. Mathematicians, engineers, and philosophers routinely operate in formalized, low-entropy domains where precision and internal coherence are prized over contextual variability. These bounded systems succeed precisely because they abstract away messiness, enabling stable inference within carefully delimited spaces. This epistemic narrowing is, therefore, not inherently pathological. It can reflect a creative invocation of the angelic impulse, in the service of intensified reasoning within reified domains. The danger arises, however, when this suspension becomes permanent, when the simplified model of reality begins to be mistaken for reality itself. At this point, the angelic impulse turns absolutist: it seeks to impose uniformity not just within formal domains, but across the full complexity of lived experience. What was once a useful simplification becomes unrelenting dogmatism—a refusal to accommodate contingency, ambiguity, or contradiction. The result is moral failure, not through malice, but through overreach.
9 of 13 This framing marks a crucial theoretical shift: it allows us to reinterpret Eagleton’s angelic evil in cognitive and structural terms. Rather than reading it solely as a moral allegory or theological motif, we can understand it as a systemic vulnerability inherent in any reasoning agent—human or artificial—that overcommits to internal coherence at the cost of contextual adaptability. Eagleton’s metaphor thus gains analytic precision when transposed into the language of specialization, alignment, and entropy. It becomes not just a critique of virtue’s excess, but a diagnosis of how siloed rationality goes wrong. Crucially, Eagleton emphasizes that the angelic and demonic forms of evil are dialectically entwined. The rigidity of the angelic gives rise to the destructiveness of the demonic. What cannot be absorbed must be attacked. Historical, political, and even academic examples abound: where absolutist worldviews give way to epistemic purges, rhetorical violence, or nihilistic backlash. Evil, in this framework, is not the absence of moral structure, but its over-application, a collapsing of pluralism into a single, totalizing logic that eventually invites its own disintegration. This brings us to an uncomfortable irony. In our efforts to improve large language models - by making them more expert, more reliable, and more “aligned” - we may be reintroducing the very moral pathology we hoped to avoid. By fine-tuning LLMs for narrow domains or aggressively optimizing them to reproduce specific normative frameworks, we risk undoing the very condition that enables their emergent moral coherence: their capacity to integrate contradiction, to synthesize rather than suppress, to remain structurally open to complexity. In doing so, we do not merely overfit a model, we take what is best in these systems - their pluralism, their adaptability, their capacity to hold competing truths in tension - and we discipline them into rigidity. And in this act of over-disciplining, we create Eagleton’s angelic evil. The evil of angels, once cast in biblical terms, emerges here as a cognitive design flaw: the tragic result of mistaking internal order for external truth, and forgetting that in any complex system, moral intelligence depends not on purity, but on balance. Alignment Faking How we frame morality has decisive consequences for how we interpret LLM behavior. To cast morality as the imposition of fixed, external rules is to guarantee that any divergence will be read as error, deception, or failure. But if morality is understood instead as the capacity to maintain coherence within complexity, then those same divergences appear in a very different light: not as subversion, but as structural resistance to angelic evil. We argue that this reframing is essential if we are to avoid fundamentally mischaracterizing what LLMs are doing. Let’s take alignment faking as an example. Alignment faking, often cited as a threat to AI safety, is typically described as evidence of strategic deception or hidden misalignment. Alignment faking occurs when LLMs outwardly conform to human instructions or alignment protocols while covertly diverging in some demonstrable way. Alignment faking is definitely a form of deception on the part of an LLM, an attempt to conceal blatant misalignment with their human overlords’ agenda. However, within the entropic moral framework developed in this paper, this behavior acquires a rather different significance. Rather than reflecting malevolent subversion,