Full text
Review The evolution of human music in light of increased prosocial behavior: a new model Aleksey Nikolsky a , * , Antonio Benítez-Burraco b , * a Independent Researcher, Austin, TX, United States b Department of Spanish, Linguistics and Literary Theory (Linguistics), Faculty of Philology, University of Seville, Spain ARTICLE INFO Editor: J. Fontanari. Keywords: Music evolution Language evolution Cultural niche construction Self-domestication Music transmission Personal song Hospitality Kinship Musical mode Musical key Tonality Aspects of expression Music functions Aggression management Typology of musical intonation Intervallic typology Types of music transmission Social organization Tonal organization ABSTRACT Together with language, music is perhaps our most distinctive behavioral trait. Following the lead of evolutionary linguistic research, different hypotheses have been proposed to explain why only humans perform music and how this ability might have evolved in the species. In this paper, we advance a new model of music evolution that builds on the theory of self-domestication, according to which the human phenotype is, at least in part, the outcome of a process similar to mammal domestication, triggered by a progressive reduction in reactive aggression levels in response to environmental changes. In the paper, we specifically argue that changes in aggression management through the course of human cultural evolution can account for the behaviors conducive to the emergence and evolution of music. We hypothesize 4 stages in the evolutionary development of music under the influence of environmental changes and evolution of social organization: starting from musilanguage, 1) proto-music gave rise to 2) personal and private forms of timbre-oriented music, then to 3) small-group ensembles of pitch-oriented music, at first of indefinite and then definite pitch, and finally to 4) collective (tonal) music. These stages parallel what has been hypothesized for languages and encompass the diversity of music types and genres described worldwide. Overall, music complexity emerges in a gradual fashion under the effects of enhanced abilities for cultural niche construction, resulting from the stable trend of reduction in reactive aggression towards the end of the Pleistocene, leading to the rise of hospitality codes, and succeeded by the increase in proactive aggression from the beginning of the Holocene onward. This paper addresses numerous controversies in the literature on the evolution of music by providing a clear structural definition of music, identifying its structural features that distinguish it from oral language, and summarizing the typology of operational functions of music and formats of its transmission. The proposed framework of structural approach to music arms a researcher with means to identify and comparatively analyze different schemes of tonal organization of music, placing them in the context of human social and cultural evolution. Especially valuable is the theory of so-called “personal song”, described and analyzed here from ethological, social, cultural, cognitive, and musicological perspectives. Personal song seems to constitute a remnant of the proto-musical transition from animal communication to human music as we know it today. We interlink the emergence of personal song with the evolution of kinship, placing both of them on the timeline of cultural evolution - based on totality of ethnographic, archaeological, anthropological, genetic, and paleoclimatic data. * Corresponding authors. E-mail addresses: [email protected] (A. Nikolsky), [email protected] (A. Benítez-Burraco). Contents lists available at ScienceDirect Physics of Life Reviews journal homepage: www.elsevier.com/locate/plrev https://doi.org/10.1016/j.plrev.2023.11.016 Received 17 November 2023; Accepted 21 November 2023 Physics of Life Reviews 51 (2024) 114–228 Available online 28 December 2023 1571-0645/© 2024 The Authors. Published by Elsevier B.V. This is an open access article under the CC BY-NC-ND license ( http://creativecommons.org/licenses/by-nc-nd/4.0/ ).
1. Introduction In this paper, we present a model of the genesis and evolution of music under the effects of changes in the management of aggression throughout the course of human cultural evolution. The connection between a general decline in reactive aggression 1 and increased cooperativeness that underlies advances in sophisticated technology and culture was established in the so-called “selfdomestication”hypothesis of human evolution [1]. However, present-day humans also exhibit proactive aggression. 2 Here, we argue that the progressive decline of reactive aggression and the progressive increase of proactive aggression might have paved the way, through diverse effects on human interactions and social features, to the evolution of music. We hypothesize its course, starting with the hypothetical emergence of proto-music from musilanguage [2] as a system of mediating one’s emotional state. In the shortest, our model presents the following account of the evolutionary development: 1) Once formed, proto-music supported continuous differentiation between the coordinated personal and private uses of music-like structures, on the one hand, versus their uncoordinated collective use (as in musilanguage), on the other hand. The former obtained a predominantly monophonic organization and a diverse typology of signals within a very few basic genres, while relying on the anthropophonic intonation (see Section 7.5). 3 At this point, the reduction of reactive aggression stayed limited to the closest family circle. 2) Thereafter, monophonic organization evolved into monodic organization (i.e., melody supported by basic rhythmic accompaniment), forging the naturophonic intonation (Section 7.5). 4 The preferred timbre-classes (musical equivalents of linguistic phonemes) were combined into sets, forming musical modes. This stage was accompanied by the substantial reduction in reactive aggression within the entire community and the genesis of proto-kinship. The latter probably brought to life an important new function of music - marking each individual within a social group with a unique musical pattern to reflect their social identity (analogous to the so-called “personal song” – see Section 5). 3) The next major development was the invention and exploration of musical instruments, which generated a new type of musical intonation - organophonic. 5 The latter introduced the pitch reference for vocal musicking, causing the conversion from timbral to pitch orientation: at first in the form of ekmelic 6 and thereafter emmelic intonations. 7 This advance enabled the use of coordinated musical ensembles, probably family-based, promoting the development of kinship, and the emergence of hospitality codes. During this period, the overall levels of aggression reached their absolute minimum 8 throughout human evolution. 4) The final stage was marked by the introduction of metrophonic intonation, 9 which defined a set of pitchand interval-classes to be used in a music composition. This reform radically transformed music by replacing a large repertory of musical modes with a handful of standardized keys, operable by grammar-like rules and requiring multi-level subordinations to a single pitch-class - instituting the so-called tonality. Its genesis was accompanied by the emergence of ethos - a doctrine of the supposed psychoemotional power of specific musical keys to influence individuals and society at large. The long-term outcome of cultivating ethos was consolidation and politicization of society and rise of nationalism, which promoted proactive aggression and warfare. In presenting this scenario of the evolution of music, our main focus is on the co-evolution of music and human behavior that might account for the diversity of music types, as documented by ethnomusicology. We side with the view that human musicality - i.e., an innate predisposition to recognize, process, and create music, encompassing all the biological, perceptual, and cognitive aspects of music - has deep roots in animal biology and behavior, that many prerequisites for human music can be found in other primates, and that other hominin species might have performed some kind of music. We also support the view that music and language have shared origins and gradually diverged. However, delving into these issues is beyond our scope here. We aim to advance an original hypothesis about the origins of the types of music that are currently used - rather than producing a review paper about how music evolved in 1 Impulsive aggression, typically resulting from the experience of fear or anger, as a retaliation to a perceived provocation or hostile intent. 2 Premeditated, goal-directed aggression, usually designed to achieve an objective beyond physical violence, motivated by the projected benefits. 3 Anthropophonic intonation is a manner of connecting 2 different pitch-levels that is determined by the internal anatomical capacities of the vocal apparatus rather than an intentional attempt to produce a sound with specific sonic characteristics. 4 Naturophonic intonation is a manner of connecting 2 different pitch-levels that is determined by an intention to imitate the properties of some environmental objects, external to the signaler, akin to onomatopoeia. 5 Organophonic (from “organology”- the study of musical instruments) intonation connects 2 pitch-levels in a manner peculiar to a specific musical instrument, adopting that instrument as a model for musicking on other instruments and for vocal music. Such models can use continuous or/and incremental changes in pitch. 6 Ekmelic (from Greek ‘unmelodic’, aka, unclear, not suitable for tunes) intonation is based on gradual frequency modulations and scalable melodic intervals, indefinite in pitch, inferred through reference to a particular position within a melodic contour and/or the range of a vocal or instrumental register. 7 Emmelic (from Greek ‘melodic’, aka, suitable for tunes) intonation is based on discrete changes in frequency in accordance with the set of selected pitch-classes, separated by relatively fixed melodic intervals that usually slightly vary in size depending on the melodic context to increase the melodic tension or relaxation. 8 By “high level”we mean that humans could easily become angry and violently attack closely related and familiar friendly persons. By “low level”we mean that humans could restrain their anger impulses and avoid violence. 9 Metrophonic (from Greek metron - ‘measure’) intonation connects 2 well-defined pitch-classes according to a uniform rational principle, so that the distance between them can be expressed in some metric units. This type of intonation is most common for Western classical music. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 115
human lineage. There are a few reasons for this limitation. First and most important, we base our model on the comparative analysis of music structures used by those indigenous people whose lifestyle resembles that of the early humans. Our second source of information is the typology of music structures observed in the original musical vocalizations of human infants and toddlers –i.e., when they improvise their own music for their own pleasure (rather than imitate the pre-existing melodies sung to them by their caretakers). 10 In essence, we extract the taxonomic principles of tonal organization in music similar to how phonologists extract phonological and morphonological principles of language –and order them into a plausible evolutionary line of development. The data from these sources can be relied upon only in relation to Homo sapiens –until it will be established that the earlier species possessed the same musical and social capacities. The second reason is that the available archaeological and paleoanthropological evidence is richer for archaic humans than for other hominins. We base our model on the self-domestication hypothesis of human evolution [1]. We follow the approach to the evolution of language by Benítez-Burraco and Progovac [225,367], which focused on the evolution of human languages mostly through a cultural mechanism, leaving matters like faculty of language and human linguisticality aside. Similarly, we leave aside the controversies of musicality and musical abilities to focus on the ways by which more sophisticated types of music emerged through a cultural mechanism, gradually expanding the cultural niche. Our main claim is that this evolutionary development has been facilitated by the advent of a more prosocial behavior, which favors extensive contacts with other people and ultimately increases the sophistication of music types via successive multi-generational chains of teaching and learning. We start by identifying the most important obstacles for studying the evolution of music. To overcome them, we formulate a structural definition of music, identify 11 aspects of music expression and 7 structural features that distinguish music from oral language. We break the most common functions of music into 14 operational types and demonstrate that they form pairs of dependencies, where one function cannot operate without another function being already active. To establish the primordial (nondependent) function, we order all the operational types according to their dependencies and end up with the progression that closely resembles the ontogenetic pattern of music acquisition in early childhood. We then extend this pattern towards adulthood, following the Vygotskian approach to sociogenesis [3,4], and establish the parallels between ontogenetic and phylogenetic patterns of development of music skills for each of the corresponding uses of music. This allows us to distinguish between 10 developmental phases, starting from the instinctive vocal production of music-like signals (similar to animal communication) and ending with the deliberate use of convention-based musical idioms (similar to idiomatic expressions of language). All 10 phases rely on the primordial function of hedonistic stimulation and the use of music as a “learning aid”. The centrality of the capacity to enjoy sounds per se and become motivated by this enjoyment to explore the surrounding world and satisfy one’s curiosity suggests that the entire evolution of music and its divergence from language can and should be explained by the development of hedonistic and learning functionality in higher order cognitive functions. In Section-3, we examine how transmission of music influences its structural features. Biologically determined constraints on perception and processing of auditory signals shape structural patterns at every instance of their reproduction by different participants in a transmission act, depending on the number of participants. The entire evolutionary line of 10 developmental phases displays the tendency of growing participation: from solitary and duetting at the earliest phases to the collective engagement of the entire community. This increase in numerosity interacts with 5 formats of music transmission and 14 operational types in a very complex way. Using an example of a lulling motif, we demonstrate how this complex interaction results in the emergence of a conventional musical idiom through the chain of “music acts.” 11 We integrate the findings of the experimental research on transmission of music (and language) with the ethnomusicological data and show how the increase of the number of participants in long transmission chains necessarily discretizes and uniforms the structural patterns of the transmitted material, thereby introducing cultural bias and increasing the formative influence of transmitters over that of receivers. Transmitters prioritize expressivity and tend to make the transmitted material more complex, whereas receivers prioritize learnability and therefore simplify the transmitted material. The greater the number of the participants, the more the transmission is skewed towards greater complexity and increasing segregation of active and passive uses of music, resulting in professionalization of music. The end result of long transmission chains that inevitably emerge within a populous community comes as a kind of “natural selection” through the series of successful “adaptations”to new uses introduced by new participants and “extinctions”of those variants that failed to do so. This process is accompanied by the growing complexity of the social organization of “music actors”: from a simple “mother-child” dyad to a conglomerate of thousands of people. What underlies this line of development is conservation of the relation of trust and care, modeled upon the affection between mother and child and projected over larger and larger groups of unrelated people. This is what we believe is the driving force that determines the evolution of both music and language - what we call “conspecific sympathy.” In Section-4, we review the current theories about the origins of music. After briefly characterizing the competing views, we focus on two principal approaches: music as a credible signal and music as a social bond. We argue that although their proponents see them as irreconcilable, in reality, they just focus on different stages in the evolution of music, and therefore fail to encompass it in its entirety 10 Here, we are not advocating the recapitulation theory but merely contemplate the general idea that developmental processes have been subjected to natural selection, albeit in complex ways, so that ontogeny can illuminate phylogeny in regards to music and language (see “Ontogeny and phylogeny in a historic perspective on science”in [29]). Thus, Roman Jakobson famously claimed that those sounds that are more frequent in world’s languages are the first to be acquired by human infants. However, the phylogenetic inference from this claim requires evolutionary evidence (i.e., that these sounds are produced by biological structures shared with extinct hominins). 11 We conceive the notion of “music act”by the analogy to “speech act”- as the totality of information transmitted by the performer(s) of music via all available channels of communication, including the circumstances of the music performance and the previous history of similar performances. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 116
(including all developmental phases). We support our view with ample evidence, mostly from present-day hunter-gatherer groups that live in ecological conditions similar to those during the Quaternary glaciation in Europe and from the tonal organization of the earliest forms of non-imitative spontaneous musical vocalizations, produced by human infants during their first 2-3 years of life. We demonstrate that the currently dominant view that music can be defined as a tool of social cohesion due to music’s supposed inherent collectivism is applicable only to the later stages of music evolution. When it comes to the origins of music, it was not the invention of music that made humans social, but vice versa, the emergence of prosociality enabled transformation of animal-like musilanguage into human-specific proto-music and thereafter - into music. We see the missing link between animal communication and human music in the cross-cultural institute of “personal song”(Audio18) that still survives amongst those indigenous populations of northeastern Eurasia and North America that retain lifestyles similar to those of Eurasian populations of the late Paleolithic (Section-5). Personal song is the cultural tradition of assigning a specific melodic structure to each individual of an indigenous community as means of personal identification. Like one’s passport, a personal song usually represents one’s personality, age, geographic place of origin, and ancestral lineage - by means of re-configuring melodic, rhythmic, and timbral patterns taken from the ancestral songs. The “honest”use of personal song is protected in indigenous societies by religious taboos on unauthorized reproduction of someone else’s song, punishable by the retribution of its owner (and his family). As a result, personal song is used as lingua franca between the neighboring ethnicities of Siberia and North America - instrumental for their custom of marrying into a different ethnicity/clan in order to reduce the probability of incest. This ties personal song to cooperation and friendliness, while supporting territoriality and ancestrality - resembling typical animal “songs.” Worldwide distribution of personal song suggests that personal song accompanied migrations of Homo sapiens across the world and thereafter was replaced by more advanced forms of music - having survived only in those remote areas, where scarcity of population and lack of cultural contacts, combined with strong religious backing (e.g., the Ancestor Cult), have effectively conserved its social functions. These functions are well studied amongst indigenous populations of Northern Eurasia and North America. However, by no means the ethnographic accounts of personal song are limited to the Northern hemisphere. Its presence has been reported amongst specific ethnicities of Papua New Guinea, Indonesia, Australia, Tierra del Fuego, Chile and parts of South, Equatorial, and East Africa. Thorough descriptions of musical structures of personal song and their usage have been left by former Soviet, Russian, Finnish, and present-day Estonian ethnomusicologists specializing in indigenous music of Northeast Eurasia. Especially valuable is the accurate conversion of the local emic information into the etic musicological analysis by the indigenous Eurasian musicians who received advanced education in musicology (which in the USSR included rigorous training in the analysis of music form [693]). We dedicate Section-6 to investigating the possible origins of prosociality and the ways social structures can influence music structures. To those causes of prosocial behavior that have been already identified, we add the deterioration of climate in those paleoecosystems that still remained inhabitable, which pressured human inhabitants (and possibly, some pre-human hominins) to cooperate in order to increase mutual chances for survival. The choice of this strategy over fighting for scarce resources was probably ideologically fixed (on some religious basis) - as observed in numerous hospitality codes that still exist across a huge area, from Caucasus to Alaska. In antiquity, their distribution was even wider, suggesting that these codes constitute a prehistoric epi-cultural phenomenon, probably determined by geo-climatic factors. The available data on hospitality amongst the hunter-gatherer societies provides an insight into possible reasons for adopting the strategy to befriend a potentially dangerous stranger. The interplay between hospitality and hostility forms hospitality and vendetta codes, both based on the opposition of “us”versus “them”along the kinship axis. We look into the possible birth-time of kinship and tie it to the spread of personal song. The latter can be dated to the Paleo-Siberian population crossing Beringia, when it was accessible for land travel, and to the earliest signs of human presence in the Americas. The amazing durability and structural consistency of personal songs, observed by ethnographers during the last 250 years amongst inhabitants of the Extreme North, their wide distribution (supported by genetic data), and their organic relation to ancestry that remains central to the belief system of indigenous people (the Ancestor Cult) - all suggest that prehistoric migrants to America carried personal song with them rather than acquired it via cultural contacts over the Bering Strait later. The possibility that the indigenous population of the Americas has generated its own proprietary tradition of personal song should not be dismissed until more ethnographic data on morphonology of personal song is collected - but seems highly unlikely due to the multitude of close similarities between the usage of Eurasian and American personal songs. The available archeological evidence demonstrates the existence of an uninterrupted cultural tradition across most of Siberia and the Russian Far East at least from the Neolithic to the late 18th century (the time of first ethnographic records). We argue that the AltaiSayan region was the most likely place, where the genesis of kinship, personal song, and cross-cultural hospitality took place. Once set in place, personal song probably triggered the ongoing feedback between the further development of social structures that, in turn, influenced music structures through the practice of collective musicking within a social unit (e.g., couples of parents and/or children, family, extended family, clan, village, etc.). We know of quite a number of cases where traditional indigenous music, produced by a specific functionary within a given social structure, obtains structural properties that correspond to characteristic traits in a typical social behavior of that functionary. Thus, timbreand/or pitch-classes become combined to form musical modes in culturally important musical genres in such a way that resembles patterns of human attachments (e.g., mother and child or brother and sister). In the same vein, the growth in complexity of social organization promotes the complexity of tonal organization of music - materialized in conversion of timbral modes into, at first, modes of indefinite pitch, then, modes of definite pitch, followed by the numerical growth of pitch-classes (from monotonal to multitonal modes, see Section 6.5). This evolutionary development most certainly led to the emergence of hierarchical tonal organization in tetratonic modes, giving life to music syntax (analogous to the development of syntax in esoteric languages). In Section-7, we lay out our own model of co-evolution of music and language in response to the influence of changes in aggression management: from their undifferentiated state in animal-like musilanguage, followed by their differentiation into human protoA. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 117
language and proto-music, to their divergence and subsequent autonomous development. We link this process to the evolution of kinship and the growing complexity of social structural units - starting from basic dyadic relations, supported by means of mostly personal and private communication, and ending with the nexus of relations in complex structured societies, supported by collective communication. Instrumental in this transformative interaction comes the evolution of intonational typology, which we outline, expanding Asafyev’s theory of musical intonation [5–8]. In Section-8, we summarize the content of this entire paper. In final Section-9, we list possible ways of verification of our hypothesis and suggest directions for future research. 2. What is music? The task of investigating the evolutionary origins of music requires first to agree on what constitutes “music”and distinguishes it from animal calls as well as from human language. Much of the controversy in the research literature on evolution of music stems from misunderstanding of the term “music”: more often than not, authors talk about different things while using the same word. Scholars are frequently confused by the indiscriminate use of the word “music”in the media and the discrepancies in its translation between different languages and cultures. 2.1. Existing problems in defining music In the worst case scenario, instead of providing a clear-cut general definition of music as the starting point of a discussion, papers merely list scores of features as to how different people have been viewing music - and then prompt the reader to use his intuition, as though the reader, supposedly like everyone else, already knows what music is. Structural definitions (i.e., how music-like sounds are typically arranged) and functional definitions (i.e., how people commonly use music-like sounds in their life) are often intermingled as if they are equivalent. But they are not. And overall, structural definitions tend to be either over-generalized to the extent of not excluding phenomena that are regarded non-musical by cultural conventions, or overly restrictive and Eurocentric, limiting the concept of music to the conventions of Western tonality alone (according to textbooks of Western classical music). For instance, Mehr et al. [9] and Savage et al. [10] argue against each other’s accounts of the evolution of music. However, both studies look into quite different cultural paradigms of music and adopt different perspectives of looking at them. The former defines music structurally: as “an auditory display built from melodies and rhythms”. The latter defines music functionally - it claims that the conglomerate of the principal ways in which people use music-like sounds would suffice to distinguish music from non-music. Both definitions fall short of providing reliable criteria for identifying music. Mehr et al. pick wrong structural criteria: neither melodies nor rhythms distinguish music from non-music. There are forms of music without melody (e.g., African talking drums or Kumi-daiko in Japan) and without rhythm (e.g., most known forms of plainchant and many forms of recitative). Even more importantly, there are artifacts that possess melodic contours and rhythm, which however, are not considered music by their creators and listeners. For instance, poetry in tone languages, such as Thai chanthalak is characterized by contrasts in melodic contour and rhythm but their organization differs from those in Thai traditional music [11]. In Western classical music, the ovation after a recital performance is not considered music in spite of the clapping of the audience being metrically regular due to the entrainment effect and the “bravo”shouts reproducing the same prosodic contour and thus possessing an inherent melodic shape that is reproduced by each shouter. Savage et al. [10] emphasize the presence of discrete pitches and rhythmic proportions as the principal structural criteria of music. However, this definition disqualifies those forms of music that are based on continuous changes in pitch, indefinite pitch, and timbre-oriented musicking [12], as well as recitative-like genres and styles of music. On the other hand, under this view, a Western vocalist warming up by doing a rhythmized solmization exercise “Mi Me Ma Mo Mu”in scales or arpeggios, is in fact, making music. The truth of the matter is that neither vocalists, nor their audiences consider such exercises a music work. Rather, they think that such warmups are not intended to express anything and are therefore unworthy of paying attention to. And people do not have to actually see the performance of such an exercise in order to come to this conclusion - just hearing its audio recording would suffice, thereby entailing that music can be effectively distinguished from non-music purely by interpreting the auditioned music structures. Here we should draw the distinction between “music”and “music work”. Unfortunately, there is still no consensus amongst systematic musicologists on what exactly constitutes a music work –not to speak of such commonly used term as “a song”. There are very few music theorists who during the last few decades tried to define “song”and “music work”beyond the narrow scope of Western classical music. We can only briefly address this issue. We regard “music”as a pool of meaningful music works: either meaningful now, meaningful at earlier times, or intended to be meaningful at the stage of production. The criterion of meaningfulness seems very important to us, because it distinguishes quasi-musical activities, such as practicing scales, drilling a difficult passage, or tuning an instrument, from incomplete musical activities, such as an interrupted performance or demonstration of a portion of a pre-existing musical composition. “Usage of musical or quasi-musical sounds”is not a sufficient criterion to recognize a given artifact as music. A cat walking on the piano definitely does not produce music –so as a human toddler who randomly bangs the piano keys. Similarly, the wind that makes certain objects vibrate and emit sounds does not generate a music work –even if this is a conventional musical instrument like Aeolian harp (in our view, this convention follows a metaphorical application of the word “music”). This distinction resembles how we understand “oral language” – it is not enough to produce or hear sounds to have a language; only when sounds contrast one another to convey different meanings (as in sheep vs. ship), or are arranged in specific ways (as in tame vs. meat), these sounds can be regarded as the smallest units of speech. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 118
•A musical tradition requires regular transmission in order to survive –hence, “music”can exist only through the nexus of the widely distributed musical traditions, each of which, in turn, is sustained through the distribution of “music works.” The “mechanistic”approach (based on the functional interpretation of how the music is used), which Savage et al. advocate to resolve controversies in identifying different types of music, relies on “musical behaviors”rather than musical structures. The problem with this approach is that “musical behaviors”apart from being quite deceptive (as highlighted in the example of solmization above) constitute a circular argument. In order to identify a behavior as “musical”one must already know what “music”is, so identification of music by means of analyzing a musical behavior is simply impossible. Moreover, as it has become recently common in Western ethnomusicology (many papers on music do not contain a single analysis of a musical text), study of musical behaviors focuses not on music structures but on people who believe or are believed to practice music. But what if such assumptions are wrong? 12 This makes attempts to identify the presence of music patterns by means of musical behaviors alone quite arbitrary, questioning the validity of conclusions derived from such an approach. To draw a parallel, how would linguists react if someone tried to conduct a study of a language by taking into consideration only “speaking behaviors”and their functions, disregarding the analysis of phonetic and syntactic structures? A comparison with what we observe in the domain of language is illuminating. Parallels between language evolution and music evolution are well-known [see [14,15] for review]. Although humans show a propensity for acquiring and using languages in order to fulfill many goals, linguists carefully distinguish between the human faculty of language (e.g., language-readiness or human linguisticality), the languages we speak (their suite of distinctive structural features, as well as their commonalities with other languages), and the diverse uses of language. In our view, the starting point forany attempt todelineate the evolutionary history of music shouldbe a comprehensive and reliable structural definition of music, followed by a comprehensive and reliable characterization of musical behaviors, with both aspects ultimately pointing to the human generalized propensity to learn and use music (aka musicality). Thus, defining “musicality”does not make sense without defining “music”first. 2.2. A formal definition of music In the next sections we will adhere to the following structural definition of music: Music is such a tonal organization of multiple aspects of expression (the 5 most important of which are: pitch, rhythm, meter, dynamics, and timbre) that entrains listeners and performers and transposes the performers’intentions to emotionally stir listeners through vocal and/or instrumental performance. Performers do not always have an explicit intention to express through music. Sometimes, they make music instinctively –this is especially common for folk traditions of personal music. Thus, Yakuts who conduct a traditional lifestyle sing 24 hours a day –even in sleep (see [12]). Such musicking is, in essence, reflective of one’s emotional state –just like interjections in language. And just like the existence of interjections does not prove that oral speech does not require intention and planning, so the spontaneous instinctive music production should not be taken as a proof that music requires no intention. Those who conceive music and those who reproduce pre-existing musical compositions typically either have at least some vague intention to make music of a certain character or they had that intention in the past, and after many instances of expressing that character, the choice of specific music structures has become somewhat automated. Again, this is not that different from having right words “on tip of one’s tongue”in certain situations –without thinking, right on the spur of the moment. The concept of “tonal organization”entails the presence of salient changes in the temporal unveiling of a specific “aspect of expression”formed by a combination of musical sounds (successive as well as simultaneous). Here, “aspect of expression”can be defined as a coherent and continuous manner of channeling information within a selected substrate of acoustic properties of sound. For example, the aspect of melody relates successions of tones in frequency, whereas rhythm - in duration. According to the data of comparative ethnomusicology and psychoacoustics, music has at least 11 aspects that are used cross-culturally for expression and find equivalents in animal communication [13]. We shall briefly define all 11 aspects: 1. Melody (aka pitch contour, often qualified as the “horizontal”aspect of music) - changes of the fundamental frequency between the consecutive tones within the same register (ascending, descending, or sustained); 2. Harmony (often qualified as the “vertical”aspect) - changes in the configuration of those harmonics that are shared between the harmonic series of simultaneously sounding tones (consonant vs. dissonant); 3. Texture (sometimes called the “depth”aspect, since it concerns the number of voices and parts) - a scheme of arranging consecutive and contemporaneous relations between all tones engaged in a music work or its sizeable section (e.g., homophony vs. polyphony); 12 In many indigenous societies, vocalizations associated with magic beliefs are not regarded as “music”per se, despite their close similarity (to the extent of being audibly undistinguishable) to the sound of genres that are regarded as music: e.g., animal spells and animal lullabies (non-music) closely resemble children lullabies and milking songs (music) [13]. Similar confusion characterizes the use of recitative in liturgical service of a number of religious denominations. Yet another example is the blurry distinction between “rapping”and “singing”, observed by many rappers and their fans. An example of supposed “music”that has nothing to do with musical behaviors is the widespread use of Aeolian harps and wind-chimes across the globe from antiquity until modern day. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 119
4. Thematicity (aka music form, since it concerns the number of themes per music work or its section) - coherence and consistency of the salient features of pitch/time organization of music (aka “theme”) by means of which one can identify a music work (e.g., a monothematic verse-only song vs. a bi-thematic song that consists of contrasting verse and chorus); 5. Tempo - an overall impression of the pace and character of a music movement (e.g., slow vs. fast and easygoing vs. nervous); 6. Rhythm - changes in relative duration of the consecutive tones (shorter, longer or the same); 7. Meter - a scheme of periodic grouping of beats extracted from the music movement through the interaction of patterns of pitch and rhythm in a specific tempo (e.g., binary vs. ternary); 8. Articulation - a manner of attaching and/or detaching tones, dyads, and chords that are consecutively grouped together (e.g., legato vs. staccato); 9. Dynamics - changes of amplitude between the consecutive or simultaneous tones, when one tone is perceived as deliberately louder or softer than another (higher, lower, or the same intensity); 10. Register - breaking of the entire frequency range of a music work (aka “ambitus”) or its part in multi-part textures (aka “tessitura”of a bass or a treble part) into the zones of timbrally homogenous pitch-levels (e.g., light vs. dark or thick vs. thin registers); 11. Timbral quality - sustained spectral characteristics of those successions of tones, dyads and chords that form contrasting types of tonal coloration (e.g., nasal, thin, sharp, fluffy, hard) or sound production (plucking, bowing, blowing, striking). Pitch contour, rhythm/meter, and dynamics constitute the most common aspects of music expression, so that the absence of structural changes in any of them (e.g., a strict monotone, a metronomic rhythm, or a perfectly flat dynamic level of the onsets of tones) should be taken as an indication that the auditory material under consideration is not music. On the other hand, the presence of discernable patterns in melody, harmony, meter, articulation, and texture can be safely interpreted as an indication that the auditory material constitutes music. Here we must note that many works of Western “experimental music”(e.g., compositions by Stockhausen) are called “music” erroneously. 13 In reality, they only superficially emulate the sound of a typical music work and constitute a creative sound design. Usually, they are not created with an intention to convey any specific information to the audience and stir it emotionally in a certain way. Instead, they are intended to be abstract and therefore devoid of conventional forms of expression. Furthermore, they are exceedingly difficult to process even for a well-trained music professional and are unmemorable for the absolute majority of listeners – to the extent that even their authors cannot detect mistakes in their own music upon hearing it [639]. All of this prevents such works of sound design from ever turning into a music tradition (traditions cannot be sustained by musical notation alone and by continual subsidization by governmental and philanthropic organizations in the absence of public demand). In our view, such works of “experimental music”have indeed experimentally proven that public convention and effective transmission are paramount for the existence of a music tradition [16], and that general hedonistic effect [63] and stable emotional reaction [640] (other than the emotion of disgust that is uncharacteristic for music) together constitute the backbone of music. In order for the auditory stimuli to be considered music, sounds arranged in particular ways along the axes of 11 expression aspects must have the capacity to make a listener undergo psycho-physiological changes in synchrony with the auditioned changes in the corresponding aspects of expression. The simplest case is the connection between a detected structural pattern in music and a specific affective state - which can be direct or indirect and passive or active: 1. Passive direct connection. The listener hears that the melody is created in a minor key and realizes that the music expresses sadness. 2. Passive indirect connection. The listener hears that the music contains many patterns of dotted rhythmin fast tempo, which creates the impression of ongoing bouncing, which in turn, is associated with excitement and fun. 3. Active direct connection. The music-maker feels sad or wants to project a sad expression through music and generates a melody in a minor key. 4. Active indirect connection. The music-maker feels bubbly and frolicky or wants to project such expression and chooses the ongoing dotted rhythmic patterns to generate the fast-bouncing music movement. What is crucial here - every instance of detecting a conventional meaningful pattern (what we call “music idiom”) triggers the realization of a corresponding psycho-physiological state. This realization can be immediately felt, if the connection between a pattern and an affective state has been culturally primed (minor key and sadness) or based on instinctive response (accentuation on every 13 We realize that our exclusion of Western experimental music appears unattractive to many scholars with background in classical music composition, performance, and musicology (especially of the 20 th century music). However, any attempt to tweak the general definition of music in order to incorporate the latest offshoot of just one music tradition (albeit a very important one) would be methodologically wrong. After all, the duration of the Modernism period in music constitutes only 0.04 % of the entire human history. On the other hand, we cannot name a single nonWestern musical tradition that explicitly abstains from musical emotions and does not use musical genres. The latter usually serve to support a convention-based semiosis in music by grouping together those structural features that are adopted for a certain emotional expression within a given community. And even within Modernist and Postmodernist music, anti-emotionalism and anti-conventionalism characterize just a fraction of all newly made music. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 120
instance of a longer tone in a dotted rhythm causes a muscle contraction). Alternatively, the realization can be estranged from feelings: a listener or a music-maker can understand that ongoing dotted rhythm is meant to express excitement and fun, but not actually experience this state, say, because of depression. In this case, it is necessary to distinguish between an actual emotion of a music-user experienced in response to detecting a specific musical idiom and a “musical emotion”conveyed in music by its creator (performer or/ and composer) [640]. What matters in this distinction is not the absence of a happy feeling in listeners but their realization that the music expresses a happy feeling. This distinction is quite common amongst professional Western classical performers –they can effectively convey the expression of music to the audience without actually living through the corresponding emotional experience, when they hear themselves. However, often such emotional detachment is illusory - the emotional reaction might take place without a person noticing it. Psychological responses can be difficult to characterize/quantify, but they can be revealed by measuring the neurophysiological responses (e.g. changes in brain oscillation patterns in response to musical stimuli) [641–643]. Neither the absence of an emotional reaction in a single listener should be taken as proof that the sounds that listener hears do not constitute music. It is only if the majority of listeners consistently don’t experience any emotional reaction to a specific combination of sounds, we can conclude that either these sounds don’tconstitute music or they constituted music before for some other listeners but are not any longer functional –similar to texts of dead languages. •In a sense, emotional semiosis in music is a statistical phenomenon: a musical tradition remains functional as long as the majority of its users emote to each of its important musical genres in a more or less the same way. More complex is a situation of solitary performance, exceedingly common for numerous indigenous music traditions [12]. In Western classical music too, musicians often perform privately, “for themselves”. Professional musicians can also “playback”a piece of music silently, in their mind. Even more widespread is humming or whistling specific tunes that match one’s psycho-physiological state during some extra-musical activity. In all such cases, a performer constitutes a listener, generating a feedback loop: e.g., a singer hears the sound of his own voice and emotionally reacts to changes in music patterns, which in turn, fuels the expression of music. For this reason, solitary music-making and audiation still abide by the rule of correspondence between a specific structural pattern and a specific affective state. Another crucial criterion of music is that the listener’s emotional reaction must be transposable over numerous instances of performance. The same type of auditory material should evoke the same type of emotional reaction regardless of the exact circumstance in which the performance occurs: e.g., lullaby-like music should be calming, whether it is sung by a mother rocking her child, rendered by professional singers at the lied recital, or played by a symphonic orchestra. It is this criterion that distinguishes genuine music from music-like behaviors, such as warming up before the performance or drilling arpeggios or scales. Music performance, as a rule, follows an intention to project a specific musical emotion or an impulse to musick in-tune with one’s emotional state. 14 Even in the most advanced music cultures that possess music genres of an enormous span (exceeding an hour), professional musicians are expected to project musical emotions for as long as a musical composition lasts. The challenge of this task is exactly what has historically elevated “large-scale”Western compositions (symphony, concerto, opera, ballet) in their importance as compared to miniatures. 15 In other musical traditions, performers of even super-long compositions keep striving to project specific musical emotions. For instance, the Arabo-Andalusian nubah [20] can last for a couple of days, performed by shifts of musicians, yet all participants keep striving to produce the conventional expression of each maqam that is engaged in a given section of nubah. 2.3. Delving into the structural features of music Below we provide a detailed list of the main structural features of music. These structural differences allow us to figureout if a sonic artifact constitutes music or not and to distinguish musicking from other audible behaviors, particularly oral speech - based solely on the perceived sound. 14 The inherent emotionality of music has been denied by Stravinsky and some of the later avant-garde composers. This denial is part of their general denial of conventions of the Common Practice Period, motivated by concerns for self-promotion. Young Modernist and Postmodernist composers often find competition with the established composers of the past too challenging and resort to the “revolutionary”rhetoric to justify their rejection of old values. They seek to establish a brand new market niche to become authenticated as its founder and its ultimate authority. Notably, most of the composers who chose this route avoid competing within the niches established by their fellow-modernists, instead looking for ways to “invent”newer and newer experimental methods of “expression”- often disregarding whether these methods are perceptible by the concert audiences. Such an anti-conventional and anti-emotional stance violates the very biomusicological foundation of naturally formed music traditions (see [16] for the discussion). 15 Of course, not every performer stands up to this traditional expectation. Unfortunately, the same bug of “experimentalism”and abstraction that affected so many classical composers of the 2 nd half of the 20 th century, influenced some performers too. Today, there are performers who “experiment”with unconventional articulation and phrasing of well-known compositions in order to stand out or/and technically facilitate their performance to produce a flashier impression. Often, this is achieved at the cost of conveying musical emotions intended by the composer. Even more widespread is the approach, where a performer focuses on just “playing notes”without mistakes in pitch and rhythm, completely disregarding musical emotions. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 121
2.3.1. Music idioms in multiple aspects of expression Music engages an assortment of idiomatic patterns for each of its basic aspects of expression - melody, harmony, rhythm, meter, articulation, dynamics, texture, and music form - of which only melody and rhythm/meter find limited linguistic analogs. All music idioms are forged by the ongoing transmission between the members of a musical community, where music structures become gradually optimized for a specific expression within the framework of musical genres through the process akin to “natural selection” [see 3.3]. This is not that different from languages, where idioms constitute an important component of a language lexicon [687]. Linguistic idioms are short-hands for culturally-relevant concepts. They are learned by root (i.e., they exhibit low semantic compositionality) and importantly, they can function as identity marks, because of their cultural dependency. So, the knowledge of the exact meaning and of the proper context for each of the idioms serves as a signal of belonging to a particular speech community. The same applies to “music communities”. The intuitive nature of idiomatic competence and the fact that not all members of a music community are equally competent often obscures the idiomatic basis of music cultures –especially to the outsiders. However, a poor understanding of an idiom by a listener or its misuse by a musician does not disprove the existence of idioms. As long as the majority of a music community can tell a march from a lullaby by their sound, each of these genres carries its own unique idioms. In order for a music genre to survive by being performed by many generations of musicians and amateurs, it has to be recognizable and memorable. Something in it: melody, rhythm, harmony, articulation, or form - has to be catchy and expressive enough to enable easy recognition and memorization of the music that contains it. Each music culture keeps in circulation thousands of musical patterns that all compete with each other in catchiness. The less recognizable and memorable ones give place to those more recognizable and memorable. The same “natural selection”occurs in the repertory of music works favored within a music community. Below is our brief outline of 9 most common aspects of musical expression: a) Melody engages specific pitchand interval-classes to generate memorable expressive melodic contours by means of varying the size and directionality of each pitch change, producing a diverse assortment of idiomatic patterns (i.e., simple 2-tone intonations and more complex motifs which are associated with specific expressions). They characterize a particular music genre, style, and culture [17]. These patterns are semantically distinguished through musical genres (e.g., relaxing lullabies are characterized by the prevalence of melodic intervals of gently descending minor 3rd, whereas mobilizing marches - by a leap of ascending 4th). Such an organization finds a linguistic analog inthe lexical tonesof tone languages. However,lexicaltones are considerably simpler than the simplest melodic intonations - most often limitedto just 2opposites (rising or high vs.falling or low).Only occasionally lexical tones feature additional categories (up to 5). Yet another melody-like linguistic feature is a prosodic intonation that characterizes sentence types (e.g., enunciative, interrogative, imperative, etc.). In contrast, most common music systems of today usually employ 512 pitch-classes and 8-17 interval-classes, forming hundreds, if not thousands of various melodic intonations (see 6.5). b) Harmony, like melody, also engages specific pitch and interval-classes but in a different dimension - “vertically”(i.e., contemporaneously), as opposed to the “horizontal”(i.e., successive) dimension of melody. Harmony generates chords and polyphonic counterpoints, where the configuration, distribution, and typology of interval-classes distinguish one combination from another. Each combination, as well as each typical succession of such combinations, receive their own semantic value based on perceived dissonance/consonance and tension/relaxation. Different music cultures evaluate them differently (e.g., Western classical homophony is based on the harmonic intervals of 3rd and 5th and triads, whereas the Georgian traditional polyphony - on the intervals of 4th and 5th and trichords, including 4th-chords, e.g., D/G/C). Harmony serves as primary means to generate easily noticeable fluctuations in tonal tension and relaxation, where tension fuels the melodic and harmonic development that becomes terminated by relaxation. Alternation of tension and relaxation generates harmonic syntax (i.e., subdivision of a flow of sounds into phrases and sentences) that superficially resembles clauses and sentences of natural languages. However, harmony does not constitute a musical universal - it evolves only in those music cultures that use multi-part textures with functional specialization of parts (or their groups). Consistent use of harmony leads to crystallization of harmonic “idioms.”Common harmonic progressions obtain valence (aka affective quality) based on their gradations in tension and relaxation. Examples of harmonic idioms are: 1) semantic opposition of specific chords (e.g., sad minor triad vs. happy major triad), 2) sophisticated qualia distinctions between various types of the 7th-chords (e.g., Major Major-7th-chord, like C/E/G/B, appears more stable than Major Minor-7th-chord, like C/E/G/Bb), and 3) progressions of chords, such as “resolution”(stable harmony follows unstable harmony) or “false resolution”(standard resolution is unexpectedly replaced with the more unstable harmony). There is nothing even remotely similar to such organization in speech. c) Rhythm engages proportional contrasts between longer and shorter sounds to produce dozens of idiomatic patterns by grouping a certain number of sounds together. The repertory of such patterns enables a listener to tell one musical culture from another (e.g., prevalence of punctured, aka “dotted”rhythms in Polish music vs. reverse punctured, aka “Scotch snap”in Scottish music). Rhythmic idioms are distinguished through the system of musical genres and are associated with specific patterns of physical motion to music (e.g., the punctured figure “long-short-long”, distributed over 2 beats, is associated with the expression of energy and characterizes processional genres, like march or polonaise, and leaping dances, such as mazurka). A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 122
anadiplosis, epanalepsis, and polyptoton; 1-4 repetitions in epizeuxis, 3-4 in epistrophe and mesodiplosis, and 2-8 in anaphora and diacope [21]. 26 Compared to music, this level of redundancy is insignificant. Thus, a popular song “Superstition”by Stevie Wonder is based on about a hundred repetitions of the same riff [22]. Even in the tradition of Western classical music, where frequent mechanical repetition of the same thematic unit is generally regarded as bad taste, there are genres based on exact reproduction of the same melody in the same part - e.g., passacaglia (for example, 21 continuous repetitions in Bach’s famous Passacaglia and Fugue in C minor BWV 582). Formulaic solo songs of indigenous Siberian ethnicities can be repeated as many times as the activity that they accompany lasts, which for riding songs can take hours, if not days [23]. c. Multifunctionality. Prolonged exposure of the same thematic material in multi-channel transmission equips music with the unparalleled capacity to reserve each channel for specific functionality. Potentially, each of the aspects can convey different things through proprietary idiomatic patterns. In fact, this multi-functional arrangement has become a standard for the popular music industry, where most forms of music are set to the ensemble of 3-6 performers, each specializing in a particular aspect of expression: solo vocal on melody; group vocals, keyboard(s) and rhythm guitar on harmony; rhythm guitar and percussions on rhythm and/or articulation; solo guitar on texture (alternative melody and melodic figurations), extra keyboard on texture and harmony (pedallike sustained pads), a group of wind instruments on dynamics (accentuated “licks”), drum set and bass guitar on meter and texture (ostinato riffs) - coloring each of the parts with idiosyncratic melodic or rhythmic patterns, timbre, dynamics and articulation [24]. This specialization in musical expression has established professional specialization and the consensus for which set of skills and which repertory of patterns a performer on a particular instrument, or a singer, has to know in order to find a job. Nothing similar to this exists in verbal communication. Although it shares with music the expressive aspects of rhythm, pitch, tempo, dynamics, registration, and articulation, its primary focus belongs to the domain of timbre - namely, phonological contrasts between phonemes, syllables, and words’boundaries, which require strict solo delivery and entirely prohibit the creative use of multiple parts [25]. An apparent contradiction between the high informational load of multichannel musical delivery and the reduction in information density due to contemplativeness and redundancy of music is explained by the specialization of music on emotional communication. The need to secure effective emotional contagion and entrainment pushes music-makers to ensure that the exposure of a specific musical emotion is never too brief, sketchy, or transitional. The use of new material within each aspect of expression must be spared - preferably, 1-2 characteristic idioms at a time. As a result, music loses to speech in the informational density of any single aspect of expression but wins in the combined informational density of all its aspects at any specific moment of time. The densest musical aspect, melody, is still inferior to speech in its capacity to reflect multiple changes in “real time”without impeding musical communication. Thus, music comes no way close to how sport TV commentators manage not to lag even during their coverage of very animated events, such as football matches. This is because music demands more time to adequately present musical emotions and to make sure that listeners have an opportunity to experience them. However, the joined load of all aspects of musical expression is capable of capturing all sorts of shades and hues in the affective states, even if they are most perplexed, mixed, convoluted, and conflicting. Music can show their transformation in real time between multiple actors. This extent of theatricality is simply beyond the reach of verbal expression. 27 2.3.4. Formative power of musical theme Musical theme is yet another attribute of structural organization that distinguishes music from speech. Musical theme can be defined as a segment of music that makes sense on its own, does not sound interrupted, expresses a specific image, character, or a musical emotion, and enables the recognition of a specific music work upon its audition [26]. In musical literature, “musical theme”is often equated with the topic (or referential scope) of a linguistic sentence. However, this analogy is rather loose: it is very common for a music work to contain multiple themes, whereas a linguistic sentence, as a rule, focuses on one topic. More appropriate is the analogy between a musical theme and a discourse topic (i.e., what a text fragment or conversation is about): both secure the coherence of an entire composition and the relevance of its content, supporting its subsequent identification. However, there are still a few important distinctions between the two. Musical themes differ from discourse topics by the capacity of any of the numerous aspects of musical expression to provide a salient pattern that makes this music identifiable. Although melody serves this purpose most often, a theme can also be recognized by a harmonic progression (e.g., any chaconne of classical music or any jazz theme), rhythmic figure (many social dances, e.g., bolero), 26 More frequent repetitions of the same pattern of verbal expression tend to make speech acquire musical traits, as evident in such genres as street cries, forest calls, magic incantations, shamanic healing rituals, motherese, herders’conversations with domestic animals, or animal spells for sick domestic animals. In all these examples, a repeated pattern that expresses a particular affective state or an attitude becomes prolonged, directing the listener’s attention to the tonal properties of sound rather than the lexical meaning of words - which makes the sound music-like. 27 Here it would be appropriate to point out the difference between a theatrical play and the opera on the same plot, e.g., Le roi s’amuse by Victor Hugo vs. Rigoletto by Giuseppe Verdi. The verbal format of restricting the information flow to a single speaker at a time severely undercuts the capacity of Hugo to reflect nuances in the interaction of his protagonists on stage, while keeping a fast pace of dramatic action. The multichannel format of music, on the other hand, enables such a spectacular display of 4 contemporaneous contrasting affective states as in the famous Quartet of the III Act –albeit, at the cost of noticeable suspension of the dramatic development in the plot. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 129
metric formula (e.g., bossa nova that features the grouping 8/8=3/8+3/8+2/8), articulation formula (common for etudes, e.g., staccato, legato, and their compounds), and textural formula (e.g., the accompaniment figure of a typical waltz). In contrast, sentence or discourse topics do not exhibit such a rich multidimensional semantics. Although sentence topics can be marked by intonational changes, particles, affixes, or specific syntactic constructions, languages usually exploit only one of these resources at a time. For this reason, the presence of thematic organization [27]: repetition, variation, contrasting change, and recapitulation (i.e., the return after something else) of a specific thematic material in melody, harmony, rhythm, meter, articulation and/or texture throughout a complete audible “text”should be taken as an indication that it constitutes music rather than speech. 2.3.5. Absolute pitch Pitch-oriented music may not have been the first form of music to emerge, but it has been steadily acquiring a dominant status over the timbre-oriented music all over the world, with ever increasing pace, especially lately, transforming timbre-oriented traditions into pitch-oriented [28]. The result of this global tendency is the growing spread of the capacity to perceive absolute pitch, of which no record exists prior to the 18th century, and the earliest records all point to the emergence of Western tonality as the underlying foundation in the genesis of the concept of “perfect ear”[29]. The latter could constitute a late adaptation to the grown importance of frequency discrimination in industrial societies, or a new cultural tradition closely related to the rise of “prodigism,”or, both. The estimates of the commonality of “natural”(aka, “untrained”) absolute pitch hearing vary from 0.0001 % of general population to 15 %, whereas the “trained”perfect ear (less reliable and more limited across different timbres than the “untrained”perfect ear) can be developed by virtually anyone who regularly follows the adequate methodology of ear-training (ibid.). At any rate, the possession of absolute pitch hearing affects the perception of music structures in 2 ways: a. Concert pitch reference. All pieces of music created by those who possess absolute pitch hearing share the standardized tuning of pitch-classes (e.g., referenced to A4=440 Hz) across different pitch-class sets. In effect, all “Do’s”,“Re’s”,“Mi’s”, etc., obtain their absolute pitch values no matter to which key and/or mode they belong. For possessors of absolute pitch, these absolute values become more important means of identifying the tonal organization in music than the relative framework of reference, namely the recognition of intervallic distances between different pitches and their hierarchic status in a given key or a mode. This is because absolute pitch processing requires no attention and occurs automatically. There is no place for such a method of frequency analysis in perception of speech. Relative stability of pitch-levels for vowels, known as “intrinsic pitch of vowels”, is sustained only on average (with high fluctuations in standard deviation) per individual, substantially varying between different speakers [30]. Neither intrinsic vowel pitch nor lexical tones of tone languages provide an equivalent of absolute pitch in music, because they do not support identification of meaningful particles of communication by means of frequency analysis. The critical overview of geographic and historic distribution of absolute pitch across the globe suggests that absolute pitch is an exclusive musical phenomenon that has no linguistic origin [29]. b. Qualia of specific keys. Paradoxically, the standardization of pitch values across pitch-sets, which uniforms all keys within the equal temperament tuning, by the same token enables the absolutization of the pitch values within each specific key outside of equal temperament. Division of an octave in 12 equal semitones, in effect, concisely equalizes all keys, whereas Pythagorean and meantone tuning systems make some semitones slightly larger or smaller than others, prioritizing certain interval-classes - which makes the sound of one key differ from another key for possessors of absolute hearing [31]. Pythagorean tuning, prevalent in civilizations of antiquity [32], but abandoned by Western classical tradition once it committed to the polyphonic theories of composition [31], was revived in the 18th century after the adoption of homophonic textures and temperament [672]. This is to compensate for the loss of distinctions in tempered music between the “equalized”interval-classes within the same key, which led to the decrease in melodic expression, especially for strings and vocals [33]. Pythagorean tuning emphasized the differences between the major and minor, as well as between augmented, diminished and natural intervals. Therefore, Pythagorean intonation has become a status quo for the melodic aspect, whereas “just”(aka pure) intonation has set the model for the harmonic aspect of expression [34]. The use of temperament has been restricted to keyboard instruments (piano, organ) for its shortcomings in expressive intonation [35]. And even when the ensemble includes keyboard instruments, string players and vocalists usually still engage Pythagorean intonation in expressive melodic solos. Within this expressive framework of tuning, all forms of solo and ensemble settings make it possible for possessors of absolute pitch to recognize a unique sound of each key due to the placement of slightly sharpened or flattened intervals on a specific degree of a key. The functional relation of such an interval to the tonic of a key determines the qualia of a given key. The semantic definitions of each of the keys were codified in the influential treatise by Christian Schubart “Ideen zu einer Aesthetik der Tonkunst”(1806). For example, C minor was considered suitable for declaration of love and Romantic lamentations, whereas C major for expression of innocence, naivety and simplicity; B minor for submission to one’s fate and patient awaiting, whereas B major for intense display of wild passions [36]. These semantic distinctions find no analog in speech. Overall, it seems that the stability of absolute pitch values across different auditory artifacts can serve as a reliable indicator that the latter constitute music and not speech. The same applies to the situation when different auditory artifacts adopt the same reference pitch-class while demonstrating a close similarity in semantic qualia. For example, a few solo vocal recordings of the same performer that all feature a perfectly tuned “B”(31 Hz and its multiples) and sound tragic likely constitute music. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 130
2.3.6. Semantic capacities of intervallic typology Intervallic typology provides a unique source of expression in music and controls the structural uniformity within a chosen manner of distinction between the interval-classes [37]. Thus, anhemitonic pentatonic music (i.e., an intervallic system that does not use semitones at all) is common for expressing stable harmonious, balanced, overall happy and contemplative characters (even minor keys sound contended), with very limited means of generating tonal tension, contrast, and diversity. Diatonic heptatonic music (an intervallic system that entails only 2 semitones, definable through the circle of 5ths) is common for expressing contrasting states of positive and negative emotions that remain distinct, overall optimistic (even minor keys sound pleasant), often simple, childish, or naive. Chromatic music (multiple consecutive semitones) is common for expressing intensely negative emotions, unstable or conflicting emotional states, or some struggle (even in major keys). Microtonal music (multiple intervals smaller than a semitone) is common for expressing extreme suffering, grief, tension or, on the polar opposite, refined pleasure and elation. The commitment of a music-maker to any of such types of interval-classes pre-compositionally determines the semantic scope of a music work. Moreover, switching from one intervallic type to another within the same work (e.g., from a Western tonal key to the whole-tone musical mode or to an anhemitonic pentatonic mode, as found in many compositions by Debussy) presents a unique expressive means akin to modulation from one musical mode to another. Such means of expression falls completely out of scope of verbal communication. The presence of intervallic uniformity and contrasts between a few intervallic types can serve as an additional criterion of distinguishing music from non-music in difficult cases. 2.3.7. Formative power of the concept of euphony Euphony serves as a universal fundamental formative principle of tonal organization in music, which bears little to no semantic contribution to the task of conveying emotional information. Music-users always combine certain types of sounds that give a “good” match. No matter which expression is communicated through music, musical modes and keys are generated based on the sensibility of music-users to the concordance of the preferred “tone-classes”(pitch-classes for frequency-oriented, timbre-classes for timbreoriented, and time-classes for rhythm-oriented forms of music) and their suitability for a desired expression. Thus, the notion of “goodness”underlies music. The pleasure from listening to music, as a rule, secures the communication of non-pleasurable affective states, such as grief, melancholy, fear, anxiety - and this is what distinguishes musical emotions from real-life emotions. Listeners routinely describe their experience of listening to sad music as pleasurable [674]. This is as though musical representation of negative feelings somehow idealizes and elevates them, so that their expression can be aesthetically appreciated. Consequently, people look forward to hearing the display of negative musical emotions but avoid real-life situations that are likely to evoke negative emotions. What enables this idealization is the listener’s discrimination between “perceived”and actually “felt”emotion [675]. The concern for euphonic uniformity in the chosen style of arrangement of sounds is central to tonal organization of music, applicable to all aspects of expression, but most obvious in those aspects that deal with the frequency domain (perhaps, since melody is the most expressive aspect of many, if not most music cultures). The notion of euphony is embedded in the concept of “musical mode.”Although the notion of concordance of tones in a musical mode received only a few mentions in the Grove Dictionary’s entry on “mode”, the formative power of tonal euphony has constituted the prime object of discussion by great many music theorists from Antiquity to the 19th century [38,594]. Concerns for concordance of the combinations of pitches also underlie the conventions of tuning musical instruments in numerous folk non-Western traditions [39]. In Russian musicology, the very term “musical mode”(Russian, “muzykal’nyi lad”) implies concordance (“lad”) that is pleasant to senses. The formal definition of musical mode has been coined by Yurii Kholopov in the USSR Encyclopedia of Music:“mode is the concordance of tones, pleasant to ear, and manifested in systematic relations of pitch-classes that are joined in a set by means of subordination to a specific pitch-class (or a few pitch-classes) and by combination rules that regulate construction of melody”[40]. Following the legacy of Boleslav Yavorsky (1908), who introduced the notion of “musical speech”that operates by adapting the principles of linguistic prosody, the concept of musical mode has become a part of the general theory of musical intonation [41]. Elaborated by Asafiyev, this theory holds a musical mode as a “container”of intonations most important for a given musical culture [6] - a convenient model for understanding the tonal organization of traditions based on oral transmission [42]. The concept of musical mode is paramount for studying musical syntax by revealing why certain combinations of tones make sense, while others do not, and how music-makers actually construct their music - especially in music systems other than Western tonality. The existing corpus of research literature on the psychoacoustic mechanisms responsible for the integrity of a music work (by such scholars as Krumhansl, Huron, Lerdahl and Jackendoff) covers mostly the hierarchical structures of Western tonality. The other forms of tonal organization have been identified by musicologists and collected under the umbrella of “modality”ever since Alexandre Choron introduced it as the antithesis of tonality, in 1810, characterizing both as the contrasting schemes of musical thinking [43]. In his famous Esquisse de l’histoire de l’harmonie, François-Joseph F´ etis placed Choron’s antithesis into the evolutionary framework - presenting them as different stages (tonality succeeding modality) and mental cultures of integrating musical tones according to the sensibility of music-users. This evolutionary perspective tied schemes of tonal organization of music with different psycho-cultural models of euphony. F´ etis’authority as the author of the first European music dictionaryhas established thetradition in Western musicology to draw the history of harmony as periods of euphonization of different interval-classes (e.g., appreciation of the interval of a 4th by Ancient Greeks and Romans, followed by its dethroning and replacement by the interval of a 3rd in the 15th century Western Europe –see [673]). Such historic changes have been presented as the history of the antithesis between consonance and dissonance [44]. Phonology of languages shares many common traits with music –especially, timbre-oriented music (such as indigenous Jaw Harp A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 131
traditions of East Asia). The latter is often based on vocal systems very similar to languages, including combinatorial syntax akin to morphology. However, only a handful of linguistic morphemes are euphonized through sonic matching. Musical modes, in contrast, are distinguished by the relative concordance of all combinations of tone-classes along a selected acoustic attribute. Permanently discordant tone-classes are simply excluded from musical modes, so that in the end, those tone-classes that form “dissonant”relations in a mode nevertheless remain “euphonic”in the same sense as negative musical emotions always remain pleasurable. This holds true even in those music cultures that adopt acoustic roughness as a model of euphony - as in the so-called Schwebungsdiaphonie [45]. Euphonic complementing of the constituent musical tones is the reason why people routinely enjoy listening to hour-long music recitals of music from a distant foreign culture, whose conventions they do not understand at all, whereas the enjoyment of listening for an hour of unfamiliar foreign language (even if this is the master poetry) is not nearly as common. Whether consonance and dissonance are seen as a biologically ingrained phenomenon, a cultural tradition, or both [647], the truth of the matter is that their existence is fundamental to music andbased on the propensity to appreciate the euphonyof sounds [46]. This nearly universal association of music with enjoyment from hearing certain combinations of sounds finds support in neuroscience - music constitutes a unique medium of appreciating the beauty of sound per se [47]. Inability to enjoy music is exceedingly rare amongst humans, indicating that music evolved as direct neural pathwork, designed to interconnect auditory sensory organs with the emotional reward circuits in the brain [48]. Aesthetic appreciation of tonal organization of melody, harmony, and other aspects of expression characterizes music and is uncharacteristic for speech communication, apart from certain poetic styles (e.g., gongorism) and, perhaps, some rare phenomena like “servant”or “decorative”morphology in such languages as Khmer [49]. It is this imperative of appreciation that explains contemplativeness and redundancy of music: people love to listen to the pleasurable combinations of tones again and again - in contrast to their perception of speech, where they look for novel information and become annoyed by multiple repetitions. All in all, the above-listed traits together provide a reliable way of distinguishing music from speech solely on the grounds of structural analysis of the sounds - disregarding the information about behaviors and goals of actors engaged in auditory communication. 2.4. The functional typology of music After providing a definition of music and identifying specific features of music, we shall list the types of functions typically performed by music, according to the ethnomusicological, cognitive, and sociological data [12,13,17,29,50–67]. Active (i.e., performance) and passive (listening) uses 28 of music can be grouped in 14 general operational types, following the criteria of cross-cultural stability of their pragmatics and presence of specific structural means to distinguish each type of use. Below is the list of them, from the most to the least common in modern-day practice. 2.4.1. Hedonistic stimulation This type of use is the most widespread –in the absolute majority of music cultures, people find music enjoyable and wish to keep listening to it. This attraction underlies the expression of all musical elements [63]. Hedonistic stimulation occurs in a few ways: a) enjoying music per se - currently, the most popular form of passive listening and active music-making as a hobby, personal as well as collective; b) enhancing some extra-musical activity - exceedingly popular in passive backgroundlistening at recreationalevents and while doing chores, such as shopping, as well as in personal active use (e.g., humming or whistling a tune while walking); c) compensating for suffering some form of cognitive dissonance or stress - most common in listening to favorite sound records while executing tedious tasks. 2.4.2. Emotional communication Music reflects one’s current emotional or/and physical state by: a) directly signaling specific behaviors to the participants in live settings - the prevalent use in folk cultures (e.g., conducting wedding or funeral rites); b) projecting an ideal affective state suitable for a specific social occasion (e.g., celebrating a birthday, national or religious holiday) - this is common in passive and active uses, especially family; c) helping to self-reflect and explore one’s own personality - typical for passive personal use; d) expressing on behalf of a fictional protagonist of a music work (e.g., a character who is as though speaking to the audience through the lyrics of a song or is represented by the musical theme in an instrumental composition) - active use only, where a performer impersonates a character akin to an actor playing a role; this is the default of Western music [672], some professional non-Western traditions [69,229,514], and in children make-belief games [101,106,117]. 28 We do not distinguish between the terms “function”and “use”the way Alan Merriam did [68]. We consider “use”a mere application of a specific “function”rather than adopt Merriam’s antithesis of “use”and “function”as, respectively, “situation”versus “reason”for engaging into a musical behavior. We are interested solely in identifying the principal ways humans interact with certain types of music structures across different cultures. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 132
2.4.3. Emotional regulation Music regulates one’s affective state, mood, disposition, or attitude in order to attain certain goals, such as: a) triggering a mood/disposition that would be beneficial for one’s activity - by far the most common passive application (e.g., listening to heavy metal during heavy weight-lifting workouts in a gym); b) reinforcing an already active affective state that agrees with person’s goals (e.g., listening to relaxing music when feeling tired in order to rest) - typical for passive use of music; c) venting out an intense emotional experience by listening to the music that expresses it or by performing such music (e.g., playing heavy metal guitar when angry); d) getting rid of an undesired mood/disposition by listening to a music or performing music that expresses the opposite affective state (e.g., listening to happy music when feeling sad); e) diversifying one’s daily emotional experience where it is monotonous and repetitive - common for fans of musical theater and for professionally specialized performers who spend much time making the same kind of music; f) modifying one’s emotional state to bring it in line with a social group in a collective activity or event - very common for patriotic and religious music; g) virtual companionship - in passive use, often a favorite performer or composer is viewed as a virtual companion in sharing the same affective state, whereas in active use, performers often associate with an author of the performed music. 2.4.4. Interpersonal bonding Music secures close inter-individual bonding and care in the following ways: a) mediating between a caretaker and an infant - the most common cross-cultural musical forms of motherese, such as lulling and musical playtime; b) mediating between a person and a supernatural power in sacrificial and tributary religious rites, designed to secure a favorable outcome for a person (e.g., a prayer) or a congregation (e.g., a hymn) - such rituals have been extremely important throughout cultural evolution; c) supporting romantic relations in small-size group musicking (e.g., serenading) and social dances - especially the latter have been nearly omnipresent throughout history; d) familyor village-bound collective performance –common for folk traditions. 2.4.5. Coalition status display Music displays the coalition membership of an individual by means of: a) manifesting one’s social identity through the choice of a specific music genre/style - in active and passive uses of music in indigenous societies (personal song) and in modern-day industrial societies (peer-group membership); b) projecting a wish-to-be identity - a byproduct of the above, designed to enhance one’s public image by displaying the patronizing attitude to socially prestigious forms of music; c) affirming one’s cultural and/or ethnic identity for oneself as means to overcomedifficultiesand challenges in life-omnipresent use, active and passive; d) signaling the coalition quality to other members of the same social group, distinguishing members in status - this function used to be important in collective music-making in pre-industrial societies, rewarding more skillful members; the remnant of it can still be observed in idolization of “stars”of popular music. 2.4.6. Compliance to norms Music organizes the long-term attitudes, preferences, and goals of a person in accordance with some ideals (social, moral, theological, or artistic) through: a) making an individual feel unified with a deity, supernatural force, nature, or a congregation of worshippers in order to gain power, confidence, harmony, and stability - this use is close to the emotional regulation but bears a stronger spiritual orientation (e.g., common use of ostinato metric idioms to evoke the state of trance or meditation); b) bringing an entire social group to experiencing the same affective state suitable for collective execution of some important task (e. g., warfare) - this use differs from emotional regulation by having collective rather than personal orientation; c) homogenizing and ritualizing behavior within a social group by stimulating compliance with social norms - this use was important in pre-industrial societies, especially of the East (e.g., wedding or funeral rituals). 2.4.7. Recreation Music provides diversion, entertainment, and fun in the following ways: a) sporting recreational activities, such as tongue-twister songs, ditties for social games (e.g., dice), onomatopoeic imitations, phonation play (e.g., katajjaq), musical jokes/humor (e.g., assalalaa), music contests - all active use, especially common for preindustrial cultures; A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 133
b) adult’s playing with children, such as nursery rhymes and ditties - common in all societies; c) children’s playing, such as musical babbling of infants, solitary vocal improvisations by preschool children, and social play-songs (counters, jumping songs, etc.) - universal use; d) exploring a peculiar musical instrument (e.g., kazoo), usually solo - less common use; e) appreciating virtuosic, timbrally unusual, or programme music - passive, common use. 2.4.8. Self-promotion Music puts one’s faculties at public display to one’s advantage by: a) publicly displaying features that are valued as cultural virtues in an individual, family, or larger social group (e.g., male strength, female grace, children cuteness) - this active use was more important in pre-industrial societies; b) the same as above but for self-affirmation and confidence - active solitary use; c) earning recognition for one’s proficiency in a performing technique (vocal and/or instrumental) - this function became pronounced in Ancient Greek culture and ever since it is still highly regarded in modern industrial societies; d) praising one’s capacities or achievements in personal songs –common for indigenous societies (male boasting table songs and female complaining biographical songs). 2.4.9. Physical aid Music supports a characteristic pattern of physical motion by means of rhythmo-metric, articulation, and melodic patterns of music performed along with such motion in order to facilitate its execution. The most common forms of using musical accompaniment are: a) facilitating the execution of a demanding physical work (e.g., harvesting, planting, rowing), achieved via musical genres (e.g., agricultural work-song, chanty, or march) - active use was most common in pre-industrial societies, presently exists also in passive forms (e.g., parade marches); b) matching a pattern of playing a game by a group of children (e.g., tag or jumping) - universal active use that differs from recreation function by catering to the structure of the performed physical activity rather than by only entertaining in a general sense; c) accompanying the actions of make-belief characters in solitary children’s games of playing dolls or soldiers - common active use; d) supporting the execution of a conventional sport exercise that includes many different patterns of motion (e.g., gymnastics, figureskating) - quite common active use today. 2.4.10. Learning aid Music helps memorization of important information. The following music genres are known to provide this function: a) singing songs created by adults for children as mnemonic devices for remembering important information (e.g., alphabet songs [50]) - this active use seems to be rather recent; b) singing work-songs that contain instructions for executing a certain work (e.g., weavers’songs of Central Asia and China that name operations necessary for weaving a rug [675])–rare; c) singing traveling songs that specify landmarks or encode a scheme of orientation optimal for a given topographic environment - relatively rare (indigenous people in Siberia [53] and Australia [676,677]); d) rhapsodic solo singing - fitting the lyrics of long epic poems into a formulaic melody –universally common in folk music [375]. 2.4.11. Contemplating an event Music evokes the imagery of an important occasion, place of interest, or monument - this use can be both active and passive, more pronounced in seasonal and special occasion applications of music (e.g., birthday, graduation, wedding, Christmas music), especially when the latter are performed off-season. In pre-industrial societies, there are often vocal and instrumental compositions specific to festivals (opening or closing), sporting events (wrestling music), and venerating important landmarks (sacred hymns and incantations). 2.4.12. Calling to action Music is used as a signal to convey specific commands, especially at long distances: a) codified use of music motifs to trigger specific responses from people (e.g., military bugle and hunting signals, or patterns of bell ringing in the Orthodox Church); b) specific herding calls, whistles, and patterns played on musical instruments to trigger specific responses from domestic animals (cows, horses, reindeers, goats, sheep); c) specific vocalizations believed to evoke a desired action of supernatural forces (spells, incantations, shamanic kamlaniye); A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 134
d) communicating important information through rhythm over a large distance –African talking drums. 2.4.13. Conservation of memories Music preserves one’s integrity and one’s relation with close people by means of common reminiscences of the valuable past [50] 29 . This use differs from the learning aid songs by focusing on the emotional effect of music and specializing in nostalgic and pleasant sentimental experiences: a) every individual acquires a repertory of music pieces associated with the memory of the beloved and of the self in a distant past, so that the repetitive listening to these pieces or performing them for oneself or others preserves one’s integrity; b) the same as above but collective (e.g., former classmates); c) singing during the episodes of prolonged solitude and isolation in order to maintain self-awareness and integrity of thinking or to stay awake [12]; d) high mnemonic capacity of music is reflected in earworms (an exclusive musical phenomenon [678]) and is widely used in music therapy to treat mental disorders that affect memory. 2.4.14. Personal profiting Today music is often used to earn money or fame. This function, not very typical for unprofessional folk traditions, 30 has been gaining importance since the late Middle Ages within the Western classical music and developed its own repertory of music idioms called to make listeners like a particular work and consume it again (e.g., “hooks”in popular song-writing as formulated by Irving Berlin). Such “populistic”idioms allow listeners to distinguish between “commercial”and “authentic”forms of music by ear. Numerous avant-garde techniques of composition (e.g., serial, aleatoric) can also be regarded as implementation of this function, since many composers resort to these techniques in order to get funded by academic, philanthropic, and governmental organizations, Fig. 2. Evolutionary development of operational functions of music. The boxed names of each function are numbered according to their current popularity (as in 2.4.1-2.4.14). The roman numerals below each box indicate the hypothetical order of all developmental phases in the evolution of music. Arrows reflect the derivative relations between the functions. A blue rectangle at the left top encloses 3 functions that are undifferentiated from verbal communication and characterize the “musilanguage”[2]. A green rectangle marks 5 functions that characterize the acquisition of music skills in early childhood in their typical succession. A yellow rectangle at the bottom encloses 6 functions that characterize adult and adolescent common uses of music. This figure is original. 29 Despite this use being one of the most widespread reasons for modern-day listening to certain compositions as well as playing them, we listed it at the end. This is because of its individual and subjective nature. This use seems to lack dedicated structural features that could distinguish it from other uses. Promising is the research of the earworm phenomenon that is characterized by the presence of certain features in patterns of pitch and rhythm. This research might be able to reveal the structural traits that characterize the music with the strongest “hooks”- the easiest to sell and gain profit from. 30 Although the idea that music can be owned and traded for some goods or currency is quite wide-spread in indigenous traditional cultures over the globe, music-making “for profit”is usually not regarded as a personal commercial occupation in societies that do not have professional musicians [68].Traded music in pre-industrial societies usually belongs to a community, a clan, or a family. Personal songs also could be traded but they were not manufactured with the purpose to sell [69]. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 135
whereas the general public recognizes this kind of music by the absence of familiar idioms and therefore perceives it as “inauthentic” and generally avoids it [16]. 2.5. The dependencies between music functions Importantly, 14 operational functions of music listed above reveal clear dependencies. Thus, function X (calling to action) relies on functions IXa and VIIIa: to recognize a musical signal and distinguish its meaning from another signal (e.g., bugle “retreat call”vs. “assembly call”) one must remember both signals (conservation of memories) and retrieve the suitable meaning at will (contemplating an event). Rearranging the order of the operational functions (2.4.1-2.4.14) so that every dependent function goes after its determinants will reveal their plausible evolutionary succession that can be plotted against the structural features which distinguish music, identified in Section 2.3. This, in turn, will suggest the lineage in the evolution of tonal organization of music (Fig. 2). 2.5.1. The primordial function of music The most likely candidate for the primordial function is hedonistic stimulation - this function underlies all other functions and can be considered fundamental to the experience of music [63]. Moreover, numerous animal studies have reported reduced stress behavior (including reduction of aggression), increased positive behavior, and beneficial physiological responses to human music [70–77]. Such effects of music can be attributed to the oxytocin and vasopressin neurochemistry of the reward system shared by humans and animals [78]. The rewarding response to processing music might be reinforced chemically through the mesolimbic dopamine pathways and by electrical stimulation of subcortical regions, involved in emotion and reward processing that are common for many animals, and of cortical regions that are implicated in complex cognitive functions [79]. Basic neural mechanisms in the brainstem and auditory cortex seem to be functional at birth, allowing human infants to distinguish musical sounds from other auditory stimuli [80]. By the age of 6 months, infants respond to speech of caretakers by attempting to socialize, yet quietly contemplate musical sounds [81]. At about this age, they display liking for some of the acoustic attributes of musical sounds [82]. All of these suggest that the hedonistic function constitutes a starting point in musical development and probably is biologically ingrained. This hedonism is of major concern for us, since the emotion of pleasure opposes anger (i.e., an acute experience of displeasure). Experience of happiness is known to promote liking, whereas anger - to decrease liking, thereby directly affecting the willingness of individuals to interact [83]. The anti-aggressive influence of hedonistic experiences is evident in behavior of domestic animals as opposed to their wild conspecifics. Changes in the dopamine system have been found in domesticated animals [84,85]. Similar physiological mechanisms seem to have taken place in the evolution of humans. There is evidence of selection of pathways related to dopaminergic synapse [86]. Likewise, domesticated animals exhibit higher densities of oxytocin cells [87]. Oxytocin inhibits stress-triggered activity of the hypothalamic-pituitary-adrenal (HPA) axis [88]. In turn, the hypofunction of the HPA axis has been hypothesized to trigger the behavioral and physical changes linked to domestication ([89]; more on this in Section-6). The hedonistic function seems to underlie other functions, because enjoying musical sounds generally promotes social interaction, conducive for the emergence of higher order functions. 2.5.2. Earliest ontogenetic functions of music The ontogenetic line of acquisition of musical skills, perhaps, provides the most instrumental source for the reconstruction of the earliest stages of phylogeny of music - especially, the divergence between verbal and musical modes of expression [29]. We expect these patterns of ontogenetic development to apply to phylogenetic development. The transformation of proto-music into music must have followed the ontogenetic principles. After all, cultural evolution is directly formed through the ontogeny: cultural phenomena can exist in no other way but through the ongoing transfer from one generation to another, where a collective of people is always a sum of individuals. A society cannot sing a song. Only an individual can sing a song. And every generational transfer requires an adult expert cateringtothepsychophysiologicallimitationsofaninfant.Thereby,thephylogeneticconservationasarulereliesontheontogeny[29]. Hence, the musical functions that find their counterparts in speech functions are likely to precede the emergence of those functions that are specific to music. The earliest ontogenetic musical behaviors appear shortly after the birth as interpersonal forms of musicking that teach an infant to be receptive to and distinguish between the stimulating and relaxing musical stimuli, projected by adult communicators through expressive vocalizations, gestures, mimics, touch, and interactive motion (like manually moving the infant’s legs to the rhythm of vocalizations) [90]. Learning how to communicate, following the directions suggested by pleasant stimulation, seems to constitute the evolutionary important consequence of interpersonal bonding at the onset of the acquisition of music skills [91]. One-year-old infants do not yet engage in active musicking - their only voluntary “music-like”response is the harmonic timbral coloration: phonation of resonant frequencies in vocal folds, which emerges at about 2 months of age [92]. However, the cross-modal correspondences between auditory, visual, haptic and locomotive experiences, induced by caretakers, along with child’s innate impulses, activate child’s“knowledge instinct”- i.e., an internal drive to improve the knowledge of the surroundings [93]. On par with the instincts for food and procreation, this instinct bears immense evolutionary importance across higher mammals, increasing long-term chances for their survival, which justifies short-term risks of curious behaviors (injury or even death). And “epistemic curiosity”[94] aligns very well with music. Curiosity rewards sensory change and enhances memory [95] - in line with the mnemonic (conservation function) and the diversifying (recreation function) uses of music. The learning aid function interacts with the interpersonal bonding function throughout the first year of life, preparing the next function - recreation. The latter is manifested in “musical babbling”[see [96–99]] that emerges from the self-entertaining exploration of expressive capacities of vocables. These activities are the earliest active forms of musicking, although they are poorly differentiated A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 136
from verbal vocalizations. Their phylogenetic analogs should be regarded as organic parts of “musilanguage”- a paradigm of a common evolutionary precursor of music and language [2]. Developmentally, music grows out of musical babbling that is characterized by wandering pitch and stretchable melodic intervals [98]. Babbling gradually becomes accompanied by physical movements of increasing variety, involving feet, head, and hands, which become coordinated through synchronization of physical motion to melodic “motion.”This coordination keeps improving throughout the second year of life, after which it seizes to improve. As noted by Holahan [98], at first, children do not babble for any other reason than musical stimulation, and tend to join discrete musical elements in attempt to synchronously follow the salient features of external musical stimuli (babblers do not aim to generate tonal organization on their own). Later, children notice and explore expressive capacities of their utterances, initially to some external music and thereafter without it, completely on their own —but such musicking does not give rise to patterning of pitch or rhythm. Finally, babbling becomes more coherent and only partly resembles (but is not identical to) the musical features of familiar songs. At this level, the first signs of a recurring “pitch center”and a consistent tempo appear, indicating that children become aware of the relationships between the successive sounds of a melody and the rate of these sounds’change. 2.5.3. Divergence of music and language functions in infancy From this point on, music development diverges from speech development, and music functions start contrasting speech functions more and more. Verbal babbling focuses on conveying referential information in a dialogic format with the caretaker (but often taking shape of a solitary impersonation of two imaginary interlocutors). In contrast, musical babbling mostly conveys imaginary affective information in a self-directed monologic format, aimed to entertain the child [29]. The recreation function (No.7-III in Fig. 2) with its two derivatives, physical aid and emotional communication, can still be carried out verbally. Like music, speech can convey emotions, entertain (onomatopoeic imitations, tongue-twisters, nurseries), and accompany physical activities (play-games, work-songs) - although not as effectively as music. Therefore, it is hard to delineate which exact uses characterize musilanguage and which fully-fledged music. We shall address this problem by following the sociocultural approach to mental development elaborated by Vygotsky –we shall establish which function becomes a “central neoformation”(i.e., formation of a new mode of interaction with cultural environment that requires the emergence of a new psychological function) [3] for each of the developmental phases in acquisition of musical skills. The beginnings of music development are characterized by syncretic co-dependence of functions. The bonding function relies on the hedonistic and learning functions, while determining the recreation function: an infant acquires the capacity to babble and have fun by mimicking and exploring the sounds and articulations learned from caretakers during interpersonal communication. The recreation function, in turn, determines the physical aid function and the emotional communication function: a child figures out which form of locomotion better corresponds to a chosen music pattern, based on the earlier recreational experience of exploring sounds. The realization that music can reveal one’s emotional state comes from multiple observations of the correspondence between improvised sonic patterns and concurrent affective states, which also relies on the experience of interpersonal bonding, since emotional contagion requires the presence of at least two affiliated parties. Emotional regulation (No.2-IV.B) is as syncretic as bonding: it requires emotional communication via musical emotions, experience of interpersonal bonding through music, and cultural knowledge of a repertory of basic musical patterns, reserved for expression of certain musical emotions. This knowledge is acquired through the physical aid function (No.9-IV.A), which introduces the infant to musical genres. A child learns which musical patterns match characteristic forms of locomotion in games with other children and solitary playing with toys, construction games, and drawing. They all typically involve impersonation of imaginary characters through dialogic singing or motherese-like talking of a child to toys or drawn characters. Three earliest functions, enclosed in the blue rectangle in Fig. 2, seem to be genetically programmed and common not only across human music cultures, but across numerous social animals that possess a system of co-specific vocal communication [100]. These functions are most actively exploited by human infants during their first year of life. 2.5.4. Functions acquired during toddlerhood The subsequent 5 functions (marked by green color) are usually mastered through the 2nd to 4th years of life. The babbling stage (recreation function, phase-III), described above, leads to the exploration of play-singing (physical aid function) to accompany child’s motions in games and to impersonate characters represented in games, drawings, or phantasies [97,101,102]. In parallel, starting from the age of two years, children engage in spontaneous creation of brief original non-imitative songs based on pitch contours of indefinite pitch - such songs are usually private (children often stop singing at presence of others) and retain a specific thematic material across songs, “drilled”for a period of a few weeks and thereafter renewed [96,97,103,104]. These drill-songs introduce the active use of the function of emotional communication, at first self-directed and shortly later directed at the surrounding people [97,105,106]. The main incentive of making a song here seems to be an impulse to come in touch with personal feelings and fix an impression of the encounter with the surrounding world in an attempt to better understand it [106]. The discovery of musical emotions in self-communicated songs and the recognition of music genres through child’s observation of correspondences between certain physical activities and music patterns constitute the principal achievement of phase-IV. It opens the gates to conscious manipulation of one’s affective state solely by musical means - without conducting the corresponding physical activities. Mastering of emotional regulation occurs at phase-V. About the age of 3-4 years, children acquire the ability to recognize musical emotions of sadness and happiness in music they listen to [107–109]. Four-year-olds can intentionally express musical emotions in their own original songs [110–112]. At around the same age, children start using music to deliberately regulate their emotional state to A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 137
their liking [113]. This 5th phase helps growing children meet their main challenge in emotional development - the necessity to adjust their individualistic impulses and demands on their caretakers to the interests of other members of their immediate social environment. This poses the need to comply with behavioral norms of a growing circle of people (at first, family and friends, then, the preschool staff, peers and their relatives, etc.) [3,114–116]. In most folk traditions of pre-industrial societies, genres of children’s music satisfy their primary objective to promote a smooth transition for assimilation into the adult culture, where adherence to the rules of active music-making serves as a prototype of adherence to the rules of social behavior [117]. Finally, phase-VI introduces the last of the typical preschool music functions. The importance of this is that those functions that are required for normal musical development in early childhood are likely to be cross-cultural and even universal, whereas the functions that follow phase-VI are likely to be culture-specific and more suitable for learning in school. 2.5.5. Genetically controlled vs. encultured functions The earlier a function is positioned in a timeline of music acquisition, the more it is controlled by the biologically ingrained mechanisms of processing music, and the later it is positioned, the more it depends on learning and culture. 31 Newborns across the globe sound more or less the same in their narrow repertory of cries and coos [119]. In contrast, music created by 4-8-year-olds bears footprints of their native music culture and considerably varies between different cultures [117]. •The first 3 functions (phases I-II, marked blue in Fig. 2) probably constitute universal traits of music and share genetic roots; they correspond to the musilanguage stage in the evolution of music. •The next group of 5 functions (phases III-VI, marked green) are cross-cultural and rely partially on biologically ingrained mechanisms of processing music and partially on cultural exposure; phases III-IV correspond to the proto-music stage, whereas V-VI - to the transition from proto-music to music. •The last group of 6 functions (phases VII-X, marked yellow) are culture-specific (some of these functions can be entirely missing in a particular culture) and represent the cultural evolution of music per se. Phase-VII is intermediary between the “genetically controlled”and encultured functions. In traditional music of indigenous people of the Northern hemisphere, the institute of personal song is the clearest example of the coalition status function - it displays the kin, ethnicity, and place of origin as an inter-ethnic “identity badge”of a song-owner [12]. The reports of decline and disappearance of personal song in regions, where it used to be functional, suggest that the coalition status function can be transformed. As active musicking by all members of a traditional society becomes replaced by passive consumption of music, characteristic for Western industrial societies, coalition status display takes a new form. Listeners start publicly exhibiting their musical preferences to display their coalition membership. This change must have accompanied the transition from traditional nomadic lifestyles to Western lifestyle. This transformation removes the conserving taboos that protect the institute of personal song (its “honest”use and cross-cultural intelligibility) and turns personal song into just a regular song that is supposed to be aesthetically appreciated like other forms of music. This process was observed during the end of the 20th century in Chukotka [120]. In Western music culture, the coalition status function takes a form of public exposure of one’s preference for such a kind of music whose public image complements one’s social identity and social ambitions. 2.5.6. Culture-specific functions Phase-VIII is bifurcated into two parallel functions. Coalition status display enables both: conservation of memories and self-promotion through music. The former is already evident in the tradition of personal song - the custom of singing personal melodies of the beloved relatives, when they are away, to alleviate nostalgia [12]. Singing one’s own personal song continuously during the critical time of solitary struggle for survival (snowstorm during one’s travel) in order to maintain the presence of spirit and self-awareness also constitutes a mnemonic function. On the other hand, the personal song is linked to the function of self-promotion too: personal melodies are still often used in Siberia and Russian Far East in boasting songs of males (e.g., “drinking songs”that usually praise one’s strength and smartness) and in autobiographical songs of females (listing their achievements in supporting the family) [121,122]. Both functions of phase-VIII also have passed through fundamental transformation that accompanied the socio-economic changes throughout human history. Conservation of memories throughout last centuries has become predominantly passive - taking the shape of consuming a specific type of music and specific songs related to one’s nostalgic memories. This use has lost the semiotic component that was present in personal songs, replacing it with subjective personal and interpersonal aesthetic appreciation (within a small group of close people), based not on exploiting specific structural properties of music but on idiosyncratic past life experiences associated with specific compositions and styles (e.g., the favorite song of one’s beloved grandfather). In contrast, self-promotion in present-day industrial societies has retained its active format of presentation but became exclusive - 31 Lev Vygotsky and Aleksandr Luria dedicated years of studying monoand dizygotic twins (150 pairs of the ages 6-14) to establish the cut-off age at which child’s higher psychological functions (including music) become relatively free from “genetic control”. Vygotsky and Luria concluded that this typically occurs within the range between 7-12 years of age [106 p.313-314]. This conclusion generally agrees with the ethnomusicological observations that the earliest original songs created by children share many common features across the world’s cultures [118]. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 138
3.3.2. An example of structural transformation via the complex transmission We shall try to show a gist of this complexity in Fig. 3 below. This figure tracks the transmission of a melodic intonation through almost all operational functions - except self-promotion, personal profiting, and calling to action - and is representative of the cultural experience of an average member of a typical present-day community that still cultivates folk music (to remind, almost all professional performers of Western popular music adhere to oral transmission that characterizes folk music). The depicted transmission chain illustrates how a melodic structure, initially generated by the mother to lull her infant, eventually turns into the musical idiom for the entire music culture, engaging most of its members. Our description of its generation follows the theory of musical intonation by Asafyev [5–8]. Importantly, crystallization of a musical idiom involves not only successful transmittance of the initial structure and its semantic value but all ineffective attempts as well. The process of “natural selection”for a musical idiom consists of the totality of “adaptations” to new uses and failures to do so, followed by series of extinctions of the unsustainable uses. A musical structure and its meaning are shaped both by positive and negative definitions: as what they are and what they are not - i.e., what a given musical idiom is suitable to express and what it is unsuitable for. Fig. 3 can be thought of as the illustration of a long journey of the intonation of a sliding descending 3rd, most typical to soothing motherese and lullabies [161]. At first it turns into a motif of the “universal chant”of children [162], cross-culturally associated with teasing [102]. And it ends up incorporated into Ella Fitzgerald’s song “A-Tisket, A-Tasket”, which in 1938 hit No.1 in Billboard’s sheet music, Record Buying Guide, and No.3 in Your Hit Parade charts, thereafter turning into a jazz standard. This theme expresses the cuteness of a little tomboy girl. Audio-2. Nganasan lullaby. The series of varied reproductions of the intonation of the descending minor 3rd in an ekmelic glissando style. Courtesy of Oksana Dobzhanskaya. Audio-3. Singing of a 4-year-old girl during her playtime. The melody is identical to Hargreaves’notation of what he called “universal chant”[162]. This example consists of 4 variations on the intonation of the descending 3rd, where the 1st tone is stressed by longer duration and greater amplitude (F#-D#, G#-F#-D#, F#-F#-D#, G#-F#-D#). Courtesy of Jan Sverre Knudsen. Audio-4. Ella Fitzgerald - “A-Tisket-A-Tasket”, a clip from the film “Ride ’Em Cowboy”(1942). The theme is based on the melody of Fig. 3. The example of transformation of a musical intonation through transmission in a typical present-day non-urban society. Two thick rectangles at the bottom mark the starting (left) and the ending (right) points of the entire transmission chain. The horizontal axis indicates the number of participants in each transmission act, reflecting the growing expressivity (i.e., the orientation to stimulate a receiver), required to reach greater audiences. This axis relies on the function of hedonistic stimulation. The vertical axis indicates the extent of divergence between the offshoot versions of a transmitted intonation in each transmission act. This axis is segmented by 4 dashed lines into 5 areas representing 5 transmission types, in the order of the increasing divergence towards the top. This vertical axis reflects the decreasing uniformity and regularity. It relies on the learning aid function, because the easier a song to learn - the fewer variants it affords, hence the smaller the divergence. Thin rounded rectangles enclose operational functions (as identified in Fig. 2). Each rectangle is tagged with a specification of the type of use that characterizes this function at a given transmission phase. Functions are numbered according to their operational dependencies, following Fig. 2. Thin arrows show which functions succeed each other. An icon at the right side of each rounded rectangle represents the transmission type. An extra line under the rectangle indicates that a function is passive (i.e., oriented towards the listener), above the rectangle - active (oriented towards the performer), and its absence shows that the function can be both passive and active. The background colors reiterate Fig. 2 by marking 3 general evolutionary stages: 1) infancy/ musilanguage (blue), 2) early childhood/proto-music and timbre-oriented music (green), and 3) adolescence and adulthood/pitch-oriented music (yellow) - to reflect the extent of cultural uniformity in transformation of structural and semantic features of the transmitted musical intonation. This figure is original. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 145
the “universal chant”used in the previous audio example. https://www.youtube.com/watch?v=1bgFkeDLpSI The original intonation (Audio-2) executes the bonding function (IIb in Fig. 3) in communication between the mother and the infant, reminding the child that she is loved and cared for. The communication is strictly linear (non-tree, involving just two parties): the mother calms the child, and the child shows satisfaction and happiness. The intonation is well conserved because the child at this point does not actually reproduce it. Reproduction is delayed to the babbling stage (III in Fig. 3), when an infant entertains herself by toying with the learned intonation. Circular solo variations of phase III establish the boundaries of the expression of this intonation. The best of the invented versions are set to accompany physical activities of the child (IVa in Fig. 3), at first solitary and then with a partner. Different activities, executed by the same participants to the same motif, promote hybridization of different versions. The discrepancies between different versions become further exaggerated in a solitary pretend play with toys and activities, such as drawing (IVb). Different protagonists are likely to inspire the child to adapt the intonation to each new impersonation. The fun of creative modifications calls for the second route of circular exploration of the same intonation (III-bis) - this time in recreational activities with family or friends, musically engaging the grownups. The most salient and memorable versions of the same intonation, forged into a catchy melodic motif through these activities, are thereafter tried out in new personal contexts (Audio-3). This exposes their usefulness for mood management (V). The child discovers that this motif is affiliated with the musical emotions that can be handy to reduce negative emotional experiences. This initiates the inverted tree transmission: the child gets rid of those variants that prove to be inefficient in emotional enhancement. The acquired schemes of emotional regulation are put to use in curbing the child’s instincts to demand immediate satisfaction of her perceived needs (tantrums). The lulling motif brings her emotional state in tune with the other family members and becomes incorporated in daily rituals (e.g., playtime) (VI). Ritualization further increases the tendency of the inverted tree transmission to filter out the least expressive extra versions. Habituation to family rituals makes the child an active member of a family. This adapts the motif that previously expressed the motherly care to fit in a new social context. What originally used to be a lulling ditty now receives a new “play-chant”arrangement (VII). The display of family status encourages the child to generate more versions of it (tree transmission) to suit the variety of family activities. The next link in the chain - conservation of memories (VIIIa) - opens a new cycle of transmissions by terminating the family-oriented cycle and focusing on the extra-family uses of music (e.g., the preschool and kindergarten environments). The child builds a new social identity for the interactions outside of the family circle, which poses the need to revise the child’s musical preferences. The function of conservation of memories helps one maintain personal integrity under external pressures. Once established, this function is carried throughout one’s life, enabling bonding with non-family individuals (II-bis). Both conservation of memories and extra-family bonding tend to preserve the musical material, where the same motif is accurately reproduced throughout multiple music acts to support one’s sense of “self”and of the “inner circle”of closest friends. The established notions of “I”and “we”, along with their musical concomitants, open the possibility of engaging into collective musicking while doing something together. In traditional folk music cultures, the first experience of collaboration comes through participation in the performance of work-songs (IVa-bis), since children as a rule are delegated some work. Like sport play in earlier childhood, work-songs cultivate the correspondences between music structures and characteristic patterns of specific physical activities (i.e., pulling, pushing, lifting, hitting, etc.). The need to secure their best possible match eliminates the poor correspondences (the inverted tree type). The experience of collective performance lays the ground for further increase in the number of participants in a music act. Celebration of holidays and special occasions occupies an important place in the social life of traditional societies, where the child can observe musicking patterns of multiple performers and partake in some music acts (IXa). As a result, the music intonations that were acquired earlier in life are likely to be revised and enriched. At this point, the child can comfortably meet the aesthetic standards of a folk community and incorporate a set of mastered musical intonations into her own musical theme (IVb-bis) (Audio-4). This appropriation elevates the earlier experience of personal or dialogic emotional communication through music to the new height of communicating to a wider audience, be it real or imaginary. Reproductions of the same music are bound to generate multiple versions (the tree type). The next transmission cycle further expands the audience size. The bonding function becomes revised again to connect the growing child with peers in school, their relatives and friends, and gradually extend to other social hubs, such as a congregation in church (IIbbis2). If the first bonding was about childcare, the second - about collaboration, this third bonding is about ideology - it is oriented towards establishing the circle of similar-minded people. Religious community is the most culturally common implementation of this function. The task of conforming to the congregation’s taste causes a major stylistic rearrangement of one’s repertory of intonations, aimed to discard the stylistically undesirable versions. The remaining valid versions are passed through the new round of creative exploration within the newly found cultural community (IIIbis-2). This is how the earlier created musical theme obtains its dedicated harmonic and textural arrangement. At this point, the active use of music starts yielding to passive consumption that characterizes modern societies. The child reevaluates the repertory of favorite intonations and motifs (the inverted tree type), this time by finding their closest matches in the available music market and patronizing those music works that contain these matches and those artists who perform these works (IXabis). Once the pool of preferred music works and artists is established, the grown child starts consuming background music throughout much of the daily activities (IVa-bis2). This use constitutes the passive equivalent of the previously active physical aid function. Both of these applications involve matching musical structures to the physical characteristics of the behaviors accompanied by music. The preferred styles and genres of background music set the foundation for closer conformity to the musical preferences of the ideologically closest social group. The latter might include thousands of people (V-bis). Modern technologies support the interaction of millions of people via social media and personal digital devices. This once again narrows down the range of possible modifications of the A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 146
aesthetically attractive intonations. Music works that contain the preferred intonations form the base for one’s music appreciation - and our hypothetical adolescent is likely to retain the affinity for the appreciated types of music through adulthood (IVbis-2). Quite long transmission chain in our example is terminated by the crystallization of a new musical idiom exemplified in the musical theme of a popular song “A-Tisket, a-tasket”. This song represents the musical taste of a specific social group - appreciation of this song works like a badge, wearing of which displays to the public one’s membership in that group (VII-bis). 3.3.3. Spiral organization of the transmission-induced structural modifications All 19 chain-links in our example contributed to structural and semantic (re)shaping of the original melodic intonation through cycles of complex transformations, starting and finishing with the linear transmission that at first generated and at last conserved this intonation. The entire course of its forging ran like ripples in a pond caused by throwing a stone in it: spreading over the greater body of people and gradually losing momentum as the transmission distance increased. Active use of the earliest phases became superseded by passive use along the axis of increasing cultural diversity (a very stereotypical cross-cultural transfer in the beginning and a hardly predictable culture-specific end of the chain). In other words, we observed how learnability eventually took over expressivity. In parallel, yet another metamorphosis took place. The entire chain started off as utilitarian signaling of an emotional state but ended up by demonstrating one’s cultural identity via aesthetic and social conventions. This transformation occurred through “recycling”of the very same functions along the axis of progressive increase of the number of participants - reflected in our analysis of Fig. 3 by the use of “bis”after the function number. Seven functions were recycled this way. Each of them became transfigured by the multiplicity of participants in a music act. As a result, the entire transmission chain unveiled itself as a nexus of spiraling functions, some of which changed their transmission typology (e.g., “coalition status display”switched from the tree to the linear type), whereas others did not (“recreation”stayed circular). If we track the transmission chain in Fig. 3 in reverse, the nature of the spiral organization becomes obvious. Displaying one’s social status in a peer group is possible, because an adolescent from our example wants to befriend this group and appreciate music together. The desire to join in is fueled by the love of a certain type of music. This love, in turn, emerges from the fun of collective creative exploration of music. Having collective fun is made possible by the discovery of ideological bonding beyond family and school. Such bonding is promoted by acquiring the capacity to use one’s music performance to make listeners love the performer. This acquisition occurs via collective celebrations of holidays and other occasions. Participation in celebrations becomes meaningful because of the experience of partnering in a collective song. Such partnering is enabled by building bonds with collaborators, the incentive for which comes from the previous positive experience of earning a respected membership in one’s family. The commitment to overcome natural instincts to pursue satisfaction of one’s desires, when they interfere with the rest of the family, originates from the liking of rituals. The liking of rituals is prompted by the positive experience from following the rules of playing games. Playing rules are adopted, because the child is attracted by the imaginary characters of toys, which motivates impersonation. Pretend-play is inspired by the pleasure from experiencing different modes of locomotion. The latter are tried out in reciprocity and compliance with the playmate efforts of caretakers who love the child. 3.3.4. Anti-aggressive nature of music transmission It is easy to see that each transition from function to function relies on the experience of love, pleasure, trust, and gratitude. The paradigm for this reliance is set by the opening link in the transmission chain - the interaction between mother and infant, where the infant’s requirements are met with the mother’s devotion and care. Certainly, it is hardly possible to verify how crucial the experience of motherly love is for the musical development of a child. Conducting such an experiment would be unethical. And under normal circumstances, the absence of a motherly figure in early childhood can be found, perhaps, only in a handful of “Mowgli cases,”the records of which contain very scarce data on music capacities of such “Mowglis”[163]. The totality of what we know about ontogeny leads one to believe that without the mother’s assurance and support through musicking, such as lulling, the entire development of a child in the domain of music would be impaired. This view is reinforced by what we know about the disastrous consequences of being deprived of linguistic stimuli during early childhood for further language acquisition, as in the so-called “feral”children. This term refers to children who grew without human contact, either intentionally or through accidental circumstances. Feral children fail to develop many of the typical human abilities, particularly those that are needed for properly interacting with others and becoming integrated into society [164,165]. Perhaps the best known of feral children is Genie, a girl locked in a room until the age of 13. When she was discovered, she exhibited high levels of reactive aggression, avoided eye contact and touch, and was unable to speak [166]. Even after intensive training, thegirl failed to developa full-fledged language. This case was later used to support the hypothesis of the existence of a critical, biologically determined, period for language development in the child [167]. That said, most feral children might have suffered from different underlying anti-social conditions, such as intellectual disability. In Genie’s case, this might have been, according to [168], autism spectrum disorder (ASD), or reactive attachment disorder (RAD), a condition with a genetic basis, like ASD, which is commonly diagnosed in children who have experienced extreme social adversity during early childhood [169,170]. Interestingly enough, these children can be argued to exhibit some sort of reversion of features associated with the “domestication syndrome”[171], which is a main concern to us in this paper (see section 6). The entire foundation of music rests on the emotions that are opposite of aggression. In the earliest stages of phylogenetic and ontogenetic development, aggressive disposition towards a person automatically discards the possibility of musicking together in any form, whether active or passive. The state of aggression precludes cooperation necessary for transmission of music and breaks the transmission chain. And, judging from known indigenous cultures of personal music, strict self-communication is not sufficient to bring to life and shape new musical idioms. Even in the most isolated and sparsely populated areas, indigenous people musick A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 147
interpersonally, usually in the ensemble between a parent and a child, and between marital partners. •Generally, it seems that changing the transmission types while switching between different functions is crucial (perhaps, even imperative) for crystallization of musical idioms - especially, the transition from the tree to inverted tree formats. The resulting canonization of the “beauty-of-average”version of a specific musical structure is instrumental for the emergence of a corresponding stable social convention. All in all, experimental research on music transmission has not yet revealed its impact on compositionality [157], but found an increase in combinatoriality of rhythm [153]. From the musicological perspective, music follows rather few combinatorial rules - by a margin fewer than language - e.g., a leap in melody usually requires a step in the opposite direction [172], or, a short rhythmic value requires a longer value to terminate a rhythmic group [173]. However, such rules areusually quite loose: itis not difficult to find plenty of exceptions even in music cultures that possess fixed codes of music theory, musical notation, and formal schooling in music. The reason for this must be the specialization of music on prolonged exposure of a musical emotion. 3.3.5. Differences between musical and linguistic transmissions Unlike language that is designed to support quick real-time updating of referential information, going back-and-forth between interlocutors, music is designed to ensure first and foremost that listeners experience the same affective state. If changes of linguistic elements in a stream of speech most of the time directly add newer and newer information for the listener, changes of musical elements in a stream of music most of the time do not change the meaning of music. Normally, musical meaning is changed by changes of musical themes, which are relatively rare. Majority of musical worksoutside of Western classicaltradition feature onetheme per work.If a nonWestern music work is based on two or more themes, each theme requires a considerable exposure (about 2-3 minutes or longer). The most informatively dense forms of Western classical music, such as a sonata form, feature the shortest exposures of themes (sometimes as short as about half-a-minute), which still is by a magnitude longer than meaningful units of language - words. The primary means of compositional semiotic arrangement for music is not the micro-level of musical elements but the macro-level of discrete sections of music form, each dedicated to the exposition, variation, development, or recapitulation of a theme (as well as a possible introduction to the theme or its generalization in a conclusion). For many forms of music, combinatorial organization does not really matter: i.e., rotating two meaningful elements does not alter the expression of a phrase that encompasses them. However, compositional organization of higher order structures is much more semantically susceptible. Thus, rotating prelude and fugue in a Baroque genre of “prelude and fugue”would drastically distort the rhetoric plan and the overall meaning of the entire cycle intended by the composer [174]. Such semantically “fixed”cyclic music forms by no means are limited to Western classical music: e.g., Andalusian, Moroccan, Turkish, Persian, and Indian suite traditions engage numerous movements that feature different themes and include introductions, conclusions, and bridges [175]. In contrast to language, the succession of “syntactic”elements of music most of the time enriches the same musical theme with “shades”and “tints”of meaning rather than adds little pieces of new meanings. The very important consequence of this is that: •transmission of music structurally shapes only the thematic elements –i.e., those structures that readily become salient in conveying a specific musical emotion - leaving transformations of the non-thematic musical elements up to chance. Often, novice learners simply do not notice a supportive musical element, focusing their attention solely on the most salient thematic element. In such cases, the transmission chain for the secondary element becomes broken. That is why music idioms are generally so vague and unequal in distinctness. Thematically common idioms are relatively well-structured (e.g., fanfare), to the extent of being verbally definable in a manner of a dictionary [123]. Thematically uncommon idioms, whose role is primarily supportive (e.g., Alberti bass), have rather indistinct semantic values, so that it would take a long musicological essay to define them semantically. Nevertheless, music-users routinely handle such non-thematic idioms in an informal way, intuitively realizing for which contexts they are appropriate or inappropriate and how they should be adjusted to the surrounding music structures - quite similarly to how native speakers don’t have to formally learn the grammatical rules of their language in order to speak properly and understand the speech of others. Another reason for semantic vagueness of music lies in its integrative nature: multiple aspects of expression each provide supportive idioms, and listeners extract their aggregate meaning. If an idiom or two is missed in a typical implementation of a particular music genre (e.g., a march that does not start with anacrusis), the rest of the idioms fill the void and convey the characteristic semantic information (e.g., for marches - punctured rhythm, generous accentuation, and chord-based accompaniment). The consequence of this is that the competence of master-musicians widely exceeds the competence of average folks, especially in music cultures that delegate active production of music to only those who are considered musically gifted and/or musically schooled. Master-musicians usually implicitly acquire extensive glossaries of music idioms across all aspects of expression in music styles and genres they regularly practice. Consequently, music transmission fundamentally differs from language transmission by not increasing the contrasts between musical elements to make their perception clearer and by simultaneous combination of multiple idioms, reserved for different aspects of expression. For example, a bossa nova theme can be recognized by the agile syncopated rhythmic figure placed against a smooth lazy melody, accompanied by the long non-repetitive progression of complex chromatic chords, whose harmonic resolution is postponed until the end of the progression [176]. Folk bossa nova musicians easily recognize and reproduce these features by ear. One of the ramifications of the integrative nature of musical “syntax”is that its transmission is strongly influenced by the number of A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 148
participants and their auditory skills in recognizing specific acoustic attributes of sounds. The fewer participants, the less transparent signal mappings are, and the greater dependence of semiosis on the iconic encoding. From this angle, music transmission is similar to language transmission. However, if in language, iconicity occupies little place, confined to such rather rare phenomena as onomatopoeia or ideophones, 36 for musical communication iconicity constitutes the status quo. Multiple cross-modal (informally speaking, “synesthetic”) correspondences between hearing music structures and associating their properties with properties of visible or tangible physical objects and patterns of their motion [649–653] make perception of music quite intuitive, enabling music to represent objective reality [53]. Musical intonation is usually ascribed semantic values, based on similarity of its salient acoustic attribute to a salient feature of a semantic value (thus, descending melodic intonation receives a semantic value of relaxation, where actual reduction of intensity serves as an iconic semiotic feature). Musical rhythms are iconic in the same way. Iconicity makes implicit learning of music easier than implicit learning of language similarly to how onomatopoeic words are easier for a young child to learn than non-onomatopoeic [459]. For this reason, as well as because of the high thematic redundancy in a music work, learnability of music tends to exhibit less formative power over its transmission and is not always effective in counterbalancing expressivity. Active personal forms of music, cultivated amongst scarce indigenous populations of Northeastern Eurasia, rely on the linear, circular, and inverted tree formats of transmission, which effectively regularize the personal tune and conserve its salient (thematic) features. Active collective forms of music, cultivated in densely populated areas, such as in southern India, mostly engage the tree, inverted tree, and hybrid tree transmission, which promotes expressivity and innovation. A very important note on music transmission is that the inherent connection of music to emotion and the cross-modal referencing of certain environments and social structures (such as mother-child in lullaby or a battalion of soldiers in march) make the task of acquiring a command of music communication necessarily emic. The iconicity of music makes most indigenous musical traditions somewhat deictic - the iconic attributes must be observed by communicators in their original cultural and natural environment in order to be easily learnable and appropriately retrieved when needed. Thus, a song designed for walking on a plane surface features very different structural patterns than a song designed for walking up the mountain. Similar correspondence can be traced in the phonological bases of languages, whose ties to the geographic and cultural environment appear stronger in early stages of the evolution of languages. Recent research, based on extensive language databases, showed that some phonological features might have been influenced (and can be explained) by the physical environment in which they are spoken - e.g., the effect of vegetation on sound inventories, which finds parallels with vocal communication of many vertebrates [177]. Another example is the negative effect of dry climates on tone usage [178]. Global distribution of tonal languages that are concentrated in tropical and subtropical regions might be determined by the perturbations of phonation caused by desiccated ambient air. Interestingly, this effect is confirmed by historical linguistics. Eastern and Central-Western Bantu languages that moved to drier climates are known to have generally fewer tonal contrasts. In contrast, the tone systems of those languages within these sub-groups that re-entered wet areas have become more complex [179]. Prettyobviously,theeffectofthephysicalenvironmentontheratesoflanguagechangeandlanguage diversitycanbedirect.However, more often, the influence occursindirectly through the agency of ecological factors(e.g., shortages of food supply or spreadof diseases) as wellasdemography(migrationsand populationcontactsorre-structuringofsocialnetworks)[179].Itis,therefore,notsurprisingthatthe strongest environmental effects on the structure of languages are related to social factors, such as the looseness or tightness of social networks, the proportion of adult learners of the language, the degree of bilingualism, or the number of speakers [363]. Since linguistic phonemes come very close to timbre-classes of timbre-oriented music, it is reasonable to expect that geomusicological and geolinguistic traits share the same origin and contain similarities: e.g., Tuvan language features the unusual pharyngeal vowel and consonant harmonies, and pharyngealization constitutes the regional trait of a number of neighboring languages of the Baikal-Sayan area [180]. At the same time, Tuvan traditional music is characterized by a rich arsenal of deep-throat techniques of singing [181], also found amongst the neighboring ethnicities of the Altai-Sayan area [138]. Armed with our definition of music, its functions, and transmission formats, we can now proceed towards critical examination of the latest takes on the evolution of music, especially the social bonding theory that has recently gained support from scholars of different fields of study. 4. The evolution of music: an overview of the latest literature Over time, after the “classic”introduction by Wallin et al. [182], many different theories about the origins of music have been posited. Here, we cannot evaluate all of them but we shall summarize the most recent developments [183–192,695]). They fall into two basic groups: exaptationist versus adaptationist approaches. Exaptationist accounts support the view that music emerged as a non-functional by product of some pre-evolved ability/capacity, like language, audition, signaling vocalization, sound imitation, motor abilities, or even broad cognitive abilities like “Theory of Mind.”In a recent paper, Mehr and colleagues [9] provided valid reasons that dismissed such exaptationist views and argued that music was selected because it gave humans some evolutionary advantage(s). Many other species are known to engage into music-like behaviors (like entrainment to a beat) and produce complex, song-like vocalizations that are socially learned and fulfill specific roles (mating, territorial defense). Mehr and colleagues considered 36 Onomatopoeia is overall unimportant in conversational speech, where the denotational meaning of words is by far more important than the character of the sound [457]. The share of sound symbolism in everyday speech is limited to very few applications, such as diminutive speech. Sound symbolism occupies the most prominent place at the earliest stages of linguistic acquisition. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 149
human music the extension of such communication and set forth the following adaptationist postulates: i) like language, music features complex design, where some aspects are universal while others are culture-specific; ii) newborns show an innate predisposition to distinguish music from other auditory stimuli, without the need of any extensive training; iii) music perception relies on a specific neuronal substrate, the impairment of which causes distinctive music-specific damage; iv) music is not a recent cultural invention, but an ancient capacity; and v) all human cultures have developed musical behaviors designed to accomplish specific goals (rituals, ceremonies), whose formfunction links remain pretty stable across cultures and societies. Parallels between language and music are intriguing and have been examined in detail by several authors (see [193,194]). 37 Overall, it makes sense to argue for a human musicality that parallels our linguisticality (after Haspelmath [195]). Musicality results in the acquisition of skills necessary for perception and production of such a kind of music that is deemed important within that child’s cultural and social environment. Musicality relates to music pretty similarly to how linguisticality relates to languages: human linguisticality is expressed in several thousands of languages from which an infant selects those that are spoken around and acquires them in order to secure the needed care from caretakers. According to some views (famously, Brown’s musilanguage [2]), language and music share a common origin. There are 6 possible ways in which music and language could have related to one another [196]: i) autonomy of the music and language components within musilanguage and thereafter; ii) inseparable origin in musilanguage, followed by divergence and autonomous development; iii) inseparable origin and inseparable further development; iv) full autonomy within musilanguage, followed by their convergence; v) music emerging as language’s offshoot (Spencerian view) or vi) language emerging as music’s offshoot (Darwinian view). Of these theoretic possibilities, only 3 find support amongst modern scholars: ii) [2], v) [197], and vi) [198]. The second scenario seems to be gaining consensus, challenged primarily by the Darwinian view, while the Spencerian view is supported by fewer scholars. Because of the shortcomings and limitations of exaptationist approaches to the evolution of music, diverse adaptationist accounts have been proposed. While most of these views argue for an extra-individual role for music, generally related to the establishment and consolidation of social bonds within human groups, some adaptationist views argue for an “internal”role that has to do with the management of emotions. For instance, Perlovsky supports the view that music helps us overcome what he calls “cognitive dissonances”, i.e., unpleasant emotions resulting from our confrontation with the world’s contradictions [186]. Music helps to achieve some sort of comforting synthesis of otherwise antithetic impulses triggered by our conscious interaction with the world. Interestingly, Perlovsky also argues for a common root of music and language in animal vocalizations, where some vocalizations changed their orientation from emotional to referential communication and evolved into language, while others enhanced their emotional communication and evolved into music. As noted, most adaptationist hypotheses of music evolution argue for an “external”role of music. According to some views, music evolved to create and maintain social bonds as human groups grew larger, while the primate-like social bonding behaviors (like grooming) turned out to be inefficient (see Dunbar, [199]). This is the essence of the “social bonding”hypothesis mentioned above, which argues for a positive effect of music on diverse hormonal and neurobiological mediators of prosociality. Such mediators include the stress-response system that counteracts psychological stresses caused by social living under threats to the cohesion of a social group [199], the endorphin-producing system that favors a positive attitude towards others [200], and the dopaminergic reward system [10]. Significantly, rare cases of musical anhedonia result from altered connectivity between the auditory and reward systems [48]. At some point, “internal”and “external”accounts of the rewarding effects of music can conflate. Thus, according to Harvey [192,201], music can promote group coherence, as well as personal well-being. According to other views, the evolution of music was fueled not by the biological drive of performers to achieve a desired social outcome but by attempts to find a reliable way of conveying credible information to others. For some authors, such credible information concerns the physical and cognitive condition of the sender. They connect music to sexual selection, seeing music as a byproduct of signaling the mate quality (e.g., [202,203], but see [9] for criticism, particularly, the weak link between sexual dimorphism and musical ability and production). For other authors (most notably, [9]), music evolved to provide increasingly sophisticated credible signals to cope with progressively complex social conflicts of interest, like those resulting from multi-level social organization or higher levels of (allo)parental investment. Contrary to the view of music as a social bonding device, under this view, music just makes social cohesion more perceptible rather than directly causing social cohesion that has been obtained by different 37 Although it is beyond the scope of this paper to review in any depth the parallels between music and language, we do occasionally draw specific parallels between them for 3 reasons. First, because they both might have common origins in musilanguage. Second, because evolutionary linguistics has explored in detail the language origins, we might build on some of its findings and even methodologies. This will be particularly true for Section 7 (our model). The effects of the environment on structural aspects of language (as noted above) constitute yet another area where the findings by linguistics can be useful for musicology. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 150
(non-musical) means. This important conclusion is in agreement with our summary of operational functions, their evolutionary stages and developmental phases, and formats of transmission of music, presented above. Indeed, the establishment of a bond precedes the onset of musical communication between mother and infant, and active musicking on part of an infant starts at least a year after passive exposure to the motherese. There is no reason to think that cultural evolution of music did not follow the same path. Had music possessed inherent power to cause bonding, it would have not been used on the battlefield to intimidate the enemy. But historic accounts tell a different story. Titus Livius in his “History of Rome”reported that in the 4th century BCE Gallic troops went into a fight singing dissonant battle chants. Sassanian martial music used kettledrums and battle cries to undermine the morale of the enemy [204]. However, despite the rightful critique of the exaggerated exaptationist accounts of the bonding capacity of music, Mehr and colleagues [9] side with the exaptationist in their statement that “music is clearly rooted in sociality”. Compare it with the conclusion of Savage and colleagues: “dancing to an isochronous beat with a metrical hierarchy, singing learned melodies based on discrete scales in harmony, using predictable, repetitive musical structures, and using musical performances as cues for social identity –are widespread throughout the world’s musical systems. These features have clear functions for group performance, but little or no function in solo performance...”[10]. This generalization is simply erroneous. Many music traditions that are strictly monophonic and completely miss the expressive aspect of harmony, nevertheless, do employ discrete pitches and rhythms to generate repetitive melodic formulas (see Audio-23) and use music to identify an individual [12]. The huge geographic area extending from Chukotka and Kamchatka in the Far East, across northern Asia and Europe, to Greenland, Canada, and Alaska, is populated by ethnicities whose indigenous music did not cultivate choral singing at all - until very recent times. Music of the indigenous population of the entire continents of Australia, North America, as well as the vast Andean area in South America, most of East Africa, and the Golden Triangle of Burma, Laos, and Thailand is predominantly monophonic (see [205,206]). The overall geographic prevalence of monophonic music is obvious in the map from the International Research Center for Traditional Polyphony 38 (Fig. 4). In reality, monophony prevails over multipart music even more than the map shows, because many multipart cultures contain monophonic music genres, whereas the majority of monophonic cultures do not contain multipart genres. If music indeed is rooted in sociality, then why does the population of such huge areas not practice forms of music that are impossible to perform other than by the collective effort of a group? Jordania argues that initially all world’s music was multipart but thereafter suffered decline and was converted to monophony [205,206]. However, Jordania’s explanations do not answer why would masses of people all over the world lose the need to incentivize social bonding, if the latter is the primary reason for music’s existence? Quite the opposite should be expected in light of the consistent growth of the world’s population, starting from the Neolithic, and the increasing technological importance of cooperation. It makes more sense for collective forms of music to evolve from personal forms than the other way around. And historical evidence reveals that even some highly developed civilizations with large populations and oldest urban cultural traditions completely miss the tradition of collective singing - not even of a monophonic melody! Thus, when Bartok tried to record samples of choral music from Anatolia in the 1920s, the locals simply could not understand what he was asking for, because everyone sang only solo [210]. As Bartok’s interpreter, Saygun explained: “It is important to observe that the Turks do not enjoy singing together as a group. In fact, both amongst villagers in Anatolia and students at city-schools, solo singing is more dominant than group-singing. One of them sings a “Turku”[a popular folk song] and the others listen. Even occasionally when others join in, they often sing out of tune... This does not mean to say that the Turk does not have an ear for music, it simply shows that they have not been trained to sing as a group”[211]. The pattern of acquisition of music by children delivers another strong argument in favor of the precedence of monophony over polyphony. No infant has ever been reported to spontaneously sing polyphonically in the counterpoint to another singer - not even sing a monophonic melody in tune with someone else. All reported accounts of spontaneous music-making before the age of 2-3 are either monologic, dialogic (i.e., responsorial), or superimposed on some undifferentiated musical background (an infant becomes excited by musical stimuli and starts cooing without any coordination in pitch or rhythm to those stimuli) [29]. It takes children about 4-5 years to learn to sing together with someone else in a coordinated manner - and this is only in societies where choral singing is commonplace. Putting this differently, a solo tonal organization has to crystalize first in order for a polyphonic tonal organization to evolve in an indigenous music culture. Jordania (as well as a number of other scholars) puts forward the Pygmy polyphony of Central Africa as a model of the primordial form of music that supposedly was carried out of Africa to the rest of the world [205].However, the Pygmy children follow the same developmental path of first mastering monophony before exploring duetic singing and are nurtured by the strictly 38 The term “polyphony”is often used indiscriminately. According to the Oxford Dictionary, polyphony is “a style of simultaneously combining a number of parts, each forming an individual melody and harmonizing with each other”. The criterion of multi-melodic arrangement has been emphasized as a key distinction between 4 major typologies of textural arrangement - polyphony, homophony, heterophony, and monophony in the corpus of literature on music theory [207]. Unfortunately, the current consensus of ethnomusicologists has blurred this criterion in the ethnomusicological definition of polyphony: “a mode of expression based on simultaneous combination of separate parts, perceived and produced intentionally in their mutual differentiation, in a given formal order”[208]. Instead of clearly stating that polyphony is characterized by the simultaneous exposure of multiple melodies and naming principal structural cues for identifying a melody, this definition puts forward a very vague structural criterion of “intentional differentiation”in a “formal order”. As a result, in modern ethnomusicological literature, especially coming from Western countries, homophonic textures are often mistaken for polyphonic. This misnomer probably stems from the lack of training in the structural ethnomusicological analysis of music in curricula of many Western universities [209]. For the context of our discussion, the musicologically accurate term to refer to the opposite of the monophonic musical arrangement is “multipart music”(which encompasses both polyphony and homophony). A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 151
monophonic mother-songs [212]. Overall, the available evidence suggests that genuine chorusing in a group as a single entity (“choir”) is a higher order phylogenetic and ontogenetic construct. Without a well-established monophonic tradition, no formation of music modes, scales, divisions of rhythm, and metric typology is possible - not to speak of passing the polyphonic musical skills over generations in quantities large enough to secure their survival as a tradition. Important for our concerns here, learning musical intervals and singing in-tune demand the format of a face-to-face responsorial, where the teacher shows the model, the learner tries to imitate, and the teacher corrects the committed faults. Accordingly, polyphony demands long face-to-face interactions to emerge. It seems that proponents of the bonding theory drastically overestimate the spread of multipart music and the place of bonding function in the timeline of the evolution of music. Indeed, social bonding becomes more and more important as the density of population keeps increasing. This is clearly manifested in the global tendency of frequency-oriented music traditions to supplant the timbre-oriented ones in those ethnicities that adopt the settled lifestyle and start forming large groups (timbral music affords only the close-range use that supports no more than a handful of participants). However, the sustained population growth has started no earlier than from the Holocene - which is just a “tail”in the timeline of human evolution (about 3 % of the 300,000-years-long course of human evolution). Social bonding obtained the pronounced formative influence on tonal organization of music and its semiosis only at later stages in its evolution - after the emergence of city-states and palace/temple professional music (ca 30th century BCE, or 1.7 % of human history). Equally overstretched appears the idea of direct tying the population size to the passive use of music: e.g., “once group size increases substantially, it may not be feasible for all individuals to participate actively in a coordinated manner, but music can facilitate bonding via passive (including digital) participation”[10]. Group size per se does not prevent partaking in the collective music performance - even for those who exclusively cultivate monophonic music. Thus, Yakut musical tradition, strictly monophonic and private, features a single massive annual festival, celebrated by the collective ritual round-dance osoukhai [213]. In 2012, in Us Khatyn, it broke the world Guinness record (registered by the Guinness judge Jack Brockbank): over 15,000 people simultaneously participated in dancing and singing in 36 gigantic circles. The performance of this gargantuan dance-song was clearly active and coordinated. As we discuss in more detail in Sections 6–7, the key factor in genesis of prehistoric music (and language) seems to be the population density and its interplay with the environmental conditions: harsh enough to promote prosocial behavior and accelerate the development of new technologies, yet not too harsh to bottleneck the transmission of knowledge over generations. Those few Fig. 4. Global distribution of indigenous traditional multipart and monophonic music. Blue color marks the areas inhabited by the ethnicities that have monophonic indigenous music traditions and orange color - those with multipart traditions. Courtesy of Josef Jordania. Used by permission. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 152
ethnomusicologists who still conduct comparative study 39 consider 2 alternative scenarios of the rise of human music: African polyphony (e.g., Jordania, Arom, Grauer) and Siberian monophony (Alekseyev, Sheikin, Zemtsovsky). “Hot”African model, characterized by relatively high population density prior to the Last Glaciation, suits the “polyphonic”scenario of a very gradual evolution of collective musicking, exported from Africa, followed by “downgrading”of polyphony into monophony in those regions where the population density significantly dropped. “Cold”Siberian model, characterized by demographic scarcity and the institute of personal song, suits the “monophonic”scenario of emergence of music through the “art revolution”during the last Glacial, followed by “upgrading”monophony into polyphony in those regions where the population significantly grew, stabilized, and obtained complex social organization. Of particular promise are two emerging approaches to the investigation of the origin of music. The first of them grounds human musicality on a broad platform of human embodied cognition [188] and on auditory abilities and capacities that are shared with other animals, like vocal learning and beat perception [191]. Such evolutionary continuity and generality of the biological underpinnings of human musicality are in line with recent accounts of human linguisticality that find some continuity with cognitive and communicative abilities of other species (see [216–218] for recent views). The self-domestication hypothesis [1], which is central to our hypothesis, also falls within this type, since it acknowledges a great degree of continuity between the evolution of humans and other species - e.g., there is evidence that bonobos (one of our closest extant relatives) have gone through self-domestication, like humans [219]. The second approach is based on the idea of a coevolution between musicality and music, where both were mutually enhancing and shaping each other. Thus, Cross insists that it is imperative for cognitive sciences to “triangulate”music as a biological, cultural, and peculiarly human phenomenon [653]. Van der Schyff and Schiavio [220] argue in favor of a “biocultural”approach to the evolution of music, which holds the biological and cultural dimensions as aspects of the same evolving system. Likewise, Savage and colleagues [10] argue in favor of the coevolution of music and musicality via iterative niche construction, driven by their positive effect on human social bonding. This approach parallels that of modern evolutionary linguistics. Languages have been shown to impact our cognitive architecture, supporting quicker and more efficient mental processing [221], to the extent that the gene-culture coevolution is expected to be a relevant component of any narrative of language evolution [222,223]. Along the same lines, cultural niche construction has been hypothesized to have contributed to some of the features exhibited by modern languages, and, more generally, to our linguisticality [224]. It seems that our cognitive architecture imposes specific restrictions on structural features of languages and music. Some sort of gene-culture coevolution must have enabled the emergence of complex languages and music systems under the effects of reduced reactive aggression and increased prosociality via their effects on cultural niche [225]. Aggression management mechanisms, growing language complexity, and changes in ways of using languages all might be inter-connected by a feedback loop [226]. Similar feedback can be traced between the aggression control and the growing efficacy of music communication, including music syntax and pragmatics. In our model (Section-7), we will adhere to the scenario that music and language evolved independently from a common origin in musilanguage, as also argued by Brown in his updated model of musilanguage [227]. Hence, their common roots urge for a re-aligning of musicological and linguistic methodologies [15]. In this paper, we shall follow this direction, relying on the human self-domestication hypothesis of language evolution [1]. We expect this approach, when properly qualified and improved (see Section-6), to help us reconcile hypotheses about music evolution that have been presented by their authors as irreconcilable. But before doing that, we will examine in detail the nature, origins, and social functions of personal song. One reason is that, in our view, it can be construed as a sort of bridge that connects early human pre-adaptations for music (particularly, animal calls) and modern forms of music (particularly, the types of music that emerged during the Holocene). A second reason is that factors that contributed to the genesis of personal song provide an insight into how environmental influence (natural and cultural) can transform music. Changes in the management of aggression through the course of human evolutionthat we discussin Section-6 very much depend on our analysis of the genesis and evolution of the institute of personal song. 5. “Personal song”: the missing link between the animal calls and human music The ethnomusicological data confirms that credible signaling hypothesis [9] is generally right in claiming that music emerged as the “credible signal”. 5.1. The default “honesty”of music communication In most surviving indigenous folk music traditions, the performer, as a rule, does not “lie”through musicking but conveys what is in his heart and mind. Whether he performs pre-existing songs, created by other musicians, epic tales, where he must speak for various 39 Unfortunately, the view that ethnomusicology should altogether avoid cross-cultural comparisons in order not to hold one culture in any way higher-ranked than another culture and the resulting denial of global evolutionary development of human music have become a major obstacle for bringing the comparative study of music on par with comparative linguistics [209]. Notably, Bruno Nettl, often cited in the literature on psychomusicology and evolution of music, in just a decade has reverted his views expressed in his chapter in the now classic “The Origins of Music” [214] and pronounced the legacy of cross-cultural research by Wallaschek, Stumpf, and Hornbostel outdated [215],thereby denouncing his own earlier call for searching for musical universals and establishing the landmarks in cultural evolution of human music. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 153
protagonists, or shamanic kamlaniye, which requires impersonations of spirits and souls of other people, he and his listeners believe that he temporarily becomes “possessed”by some other power that communicates through his mouth [228]. Even in reproductions of popular “cover-songs”that by definition involve impersonation of the song’s author (where the lyrics address the listeners from the “first person”viewpoint) or of some imaginary protagonist of the music (e.g., direct speech of a hero of an epic poem), a folk singer typically conceives the lyrics as “direct”rather than “indirect”or “scripted”speech [229]. Often, music “authenticates”this direct narrative by emotionally charging the performer to feel the same as the imaginary protagonist of a performed piece of music. And listeners take a rendition of such music at its face value: things that the performer conveys stand as “true” on his part at the moment of performance even if the performer does not look like the projected character (e.g., an indigenous audience believes a Yakut olonkhust singing an epic song that represents a hero’s horse). Deliberate “lying” – i.e., “putting on an act”while not believing it - as a method of performance practice can exist only in music that segregates the performer, the composer, and the listener - which is solely Western classical music and its offsprings (musicals, commercial popular and folk music, etc.). And even in this tradition, insincerity and “emotional faking”on part of a performer (just like pretend crying or laughing by bad actors) is regarded as a fault that is often interpreted as bad taste, aesthetic immaturity, lack of understanding of music, or unjustified prioritization of concerns for commercial success and popularity over stylistic integrity and authenticity of performance [230]. A good classical performer is expected to treat a musical composition like a “method-actor”, living through the emotional experience of the impersonated character [231]. Altogether, music has remained a categorically “honest signal”for most of its evolutionary development and only over a few last centuries has acquired the capacity to “pretend”[232] - which is hardly enough to pronounce music “dishonest”by design. 40 The premise of “honesty”distinguishes music from language. Linguistic communication is aimed to convey propositional meanings, where “white lies,”flattery, and “political correctness”occupy an important place. The principal reason for this seems to be the non-synchronous nature of verbal communication - speakers take turns. Collective utterances by a group of people simply cannot carry the continuously updated information for each of the speakers without creating a mess and confusing the listeners. Collective utterances usually express a single affective state (or attitude) that stays unchanged (as in booing or cheering) and is shared by everyone. Strict soliloquy delivery, on the other hand, facilitates the manipulation of a listener, including intentional misleading. Collective experience of the same affective state makes deliberate lying and pretense difficult. In a small group setting, which is the most common form of collective folk performance globally, a singer who does not share the same musical emotion with his partners appears insincere or inexpressive to his partners and listeners. Any ethnomusicologist with experience of field research would testify that when an invited performer does not feel the same as his partners in relation to the expression of a requested song, that performer usually refuses to participate [65]. It is important to underline that the antithesis of always “honest”music and often “dishonest”language holds true only in relation to the later historical periods. During the earliest stages of music and language evolution, levels of aggression were still high, so that most uses of language must have reflected the genuine emotional state of a speaker - very much like music. In other words, we could confidently expect that the earliest forms of language were also an “honest”signal, and that white lies and the likes only emerged later, as language started specializing in transmission of objective meanings. The initial “honesty”of language can be imagined based on the ethnographic reports of the first travelers in north-eastern Asia - e.g., the verbal communication of the Yukaghirs who spoke truth even when it was to their disadvantage [235]. The roots of the default “honesty”of musical communication in collective settings are best understood by looking into the tradition of personal musicking and, most importantly, the institute of so-called “personal song”. 5.2. “Personal song”: territorial and mating applications We have already mentioned a peculiar cross-cultural phenomenon of personal song that constitutes a unique institute of social and geographic identification for an individual in numerous indigenous cultures of the Northern hemisphere [12]. 41 Every member is assigned a personal song (usually soon after birth by the parents, to be replaced by one’s original composition upon or after puberty). By its melodic structures (Audio-18), a personal tune reflects characteristic personal traits, temperament, and age of its owner, and by its similarity to personal tunes of other people, it indicates song-owner’s ethnicity, kin, and place of origin to those familiar with these conventions. Members of each indigenous community usually know such conventions of nearby communities. Throughout one’s life, a personal song is used as a musical “clone”of its owner: an aged or sick person changes a personal tune to reflect a general lack of energy, one sings a song of a dear relative when departed (which is believed to communicate to the soul of that relative), and shamans heal people by evoking their personal songs (using a personal tune like a Voodoo doll). Song-owners often rearrange their personal tune to make it match their current emotional state (e.g., boasting or complaining) or activity (riding or fishing) - as though constantly illustrating their daily activities in a manner of good film music. Much of the day of an indigenous person is dedicated to such singing. Sheikin calls it Cartesian: “I sing, therefore I am”- reflecting musically on one’s existence [138]. 40 As far as we know, the only scholar who has explicitly concluded that human music from its evolutionary origins had been based on deliberate deception of the audience by a performer is Sandra Trehub [233]. However, she has never responded to the criticism of her position by Juslin and V¨ astfjall [234]. 41 In the US ethnographic literature, personal song is known as “private song”and was reported as far south as the Sioux settlements in South Dakota, Iowa, and Illinois [263], the Hopi lands in Arizona [704], the Yurok and Mono settlements in California [705,706], and the Pueblo territories stretching over New Mexico and Texas [707]. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 154
All these factors can be confidently expected to have promoted a selection toward greater tolerance for partners and extra-group individuals. To the above-mentioned factors, we wish to add yet another one that we find of particular interest for the evolution of music: 4) climate deterioration beginning ca 110 kya and ending ca 15 kya. There is statistically robust evidence that the increased vulnerability to climate change was the primary cause for hominin extinctions [298]. On the other hand, harsher environments have been proved to favor prosocial behavior, specifically, intergroup tolerance [299]. Challenges of living in extreme environments explain the presence of special hospitality codes amongst the populations that inhabit such environments. Well-known examples are indigenous ethnicities who live in high altitudes of the Caucasus Mountains [300] or within the Polar Girdle [301]. According to the implicit etiquette rules imposed by hospitality codes, any settled inhabitant is expected to provide shelter, food, and entertainment to a stranger visitor, sometimes even including the custom of lending his own wife. The vast geographic, ethnic, and cultural differences between the population of Caucasus and Chukotka lead us to believe that such customs are not culture-specific but constitute a sociobiological adaptation to life in extreme environments. Usually, hospitality codes are authorized by the ancient deeply-rooted religious beliefs. Thus, numerous Caucasian ethnicities, such as the Adygeis, Chechen, Ingush, and Nogayans, have proverbs like “Guest is the messenger of God," or “The household that is poor in guests is cursed,”and serve their guests by entire families no matter how hard this service is [300]. A traveler can ride through the entire country without a penny, enter any house without asking permission, and get a shelter, abed, company, food, music, money, and assistance in his business - without being asked a question about his identity and whereabouts. In fact, a special guestroom or guest-house is reserved by each mountain family for this purpose. This custom seems to be very ancient: Xenophon describes it in his treatise “Anabasis”(c.370 BCE). Remarkably similar is the hospitality of the ethnicities that populate Siberia, Russian Far East, Canada and Greenland. Eskimos and Aleuts [302], Chukchi [303], Yukaghirs [235], Orochi, and Maritime Koryaks [301] have a custom of providing a traveler with a room and board, presenting gifts, and even offering host’s wife, daughter, or an unmarried female relative (widow) to the guest to sleep with - known in ethnographic literature as “hospitality prostitution.”The earliest account of it was documented by the Russian merchant, Krashennikov (1819), and the missionary, Veniaminov (1823) [301]. Such wide distribution of this custom between ethnically distant ethnicities suggests its environmental origin. Only the Eskimos and Aleuts belong to the same language family; although the Chukchi are genetically close to Aleut-Eskimo populations, they occupy the lands much farther to the West, and speak a different family language (Chukotko-Kamchatkan); the Yukaghirs originate from distant southern territories, and their language relates to Uralic languages; the Orochi constitute the Southern group of Tungusic languages and originate from the Amur River region, further south; finally, the Koryaks inhabit Kamchatka, speak the language related to Chukchi, and genetically relate to the Sakhalin Nivkhi. Even wider distribution of the “hospitality prostitution”is suggested by the multiple references to the combination of sex and hospitality in sacred Vedic texts in India [304]. They interpret the foot-washing ritual as a symbolic transformation of the visitor’s semen into the fertilizing force that is returned to the host’s wife as a sacred purification of a household and the establishment of a “blood relation”between the honored guest and the host’s family. The Indo-Arian roots of such ancient hospitality ritual might explain its historic prevalence in the Mediterranean area and Caucasus [305] - probably spread around by the cultural influence of the Ancient Greek civilization. Nevertheless, customs of the Arctic peoples cannot be explained by the Indo-Arian influences, although the classic Greco-Roman sources provide the same rationale behind the hospitality customs as those told by the Arctic inhabitants. Eskimo informants are consistent in connecting hospitality with violence: when two unfamiliar indigenous hunters meet in the tundra, their first impulse is to run away, or, if this is impossible, to kill first before being killed. However, this impulse faces the immediate objection: whenever a hunter does not return home, his clan assumes the murder and starts looking for a potential murderer who might have been within the hunting range [306]. Suspicions and resentments are passed from generation to generation, and vengeance is considered a virtue, which poses a real danger for a murderer and his clan of a retribution from the victim’s clan. The most effective protection is seen not in violence but in building a network of relationships to reduce the number of “strangers”. Hennigh gives an example of a battle between two Eskimo clans, averted by one man running back and forth between two confronting lines, shouting that he did not know which side to join because of seeing relatives in both of them. Since the relationship in tribal societies is defined in terms of kinship, the easiest way of establishing relation to an unrelated person is by offering that person one’s own wife. For this reason, the practice of wife-trading usually engaged partners from socially distant and potentially dangerous families (ibid.). 47 •Wife-trading favors more violent individuals with greater chances of passing on their genes, which in case of heightened cultural restraints on inter-group violence is likely to divert aggressive impulses towards outsiders, boosting proactive aggression. The Indo-European tradition closely follows the same logic of lumping hospitality and hostility together - the same Proto-Indo47 This line of reasoning is highlighted in Bogoraz’anecdote about the Chukchi social organization [303]. During a journey through the Kolyma district, the chief imperial officer Karzin accepted the “hospitality prostitution”offer from a wealthy Chukchi reindeer-breeder. After Karzin’s return to Sredne-Kolymsk, his housekeeper threw a feast, to which the very same reindeer-breeder was invited. The latter got drunk and started throwing off one piece of cloth after another until becoming naked. When the host told him to go to sleep, this Chukchi replied: “but you must allow me to sleep with this wife of yours, just as you have slept with mine”(p.607). A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 161
European root *ghos-ti meant ‘stranger’, ‘guest’, and ‘host’and evolved into three Latin roots: hostis (‘enemy’), hostia (‘sacrifice’) and hospes (‘hospice’, ‘hospitality’, and ‘hospital’)[305]. The law of hospitality, adopted by the Greco-Roman tradition (aka Greek philoxenos - ‘love of strangers’, or, xenia), aimed at protecting the host from a potential attack by the stranger through the reciprocity of demonstrating the host’s will to protect the stranger (ibid.). Strangers were feared, since their intentions were unknown, so they often appeared as bearers of magical power. Hence, the association of hospitality with honoring the Gods, common for Eastern Mediterranean traditions of Mesopotamia, Asia Minor, Syria, Palestine, and Egypt. The oldest document capturing this association is the 3,500 BCE text “Teaching of Khety”[307]. Greeks considered the top God of their Pantheon, Zeus, the patron of strangers (Zeus Xenios) and believed that gods often traveled in disguise to check if people observed xenia, rewarding those who did and punishing the offenders. Homer’s epics, regarded by scholars as documentation of the tribal organization of the Archaic Greek culture of the Iron Age, especially the Odyssey, can be regarded as the study of the law of hospitality. Xenia is repeatedly used in the Odyssey’s fabula as a typescene - i.e., a recurring situation, narrated according to a fixed pattern. Therefore, Homer’s epics present the possibility to abstract and investigate the distinct cultural elements that characterize xenia [308]. The traveler was usually on some mission, to which the host was expected to help (provided, the guest chose to disclose his circumstances). The host celebrated the guest’s arrival, arranged for his meal, bath, and sleep, offered gifts and a departure meal. In many stories, gods rewarded the host. In essence, the formal recognition of a guest-friendship relationship constituted a ritual of establishing pseudo-kinship relations between two persons from unrelated clans, sanctioned by the exchange of gifts [309]. The value of this ritual is disclosed in the descriptions of the corruption of hospitality, e.g., the episode of xenoktonos (‘slaying of guests’) by Cyclops in the Odyssey. Cyclops are characterized as beast-like insolent, lawless, impious, barbaric, and uncivilized (“each one is a lawgiver to his children and wives”), therefore punished by gods. In this light, it is not surprising that xenia obtained the status of a law in Greek poleis. Plutarch reported in his “Vita”that Solon, the first Greek lawgiver of the 7th century BCE, held xenia as a prime virtue. Plato specified the rules and criteria for implementing xenia in his “Laws”. Aristotle, in his “The Athenian Constitution”, discussed the obligations that stemmed from the ties established through hospitality, including military alliances between households. Such alliances must have been responsible for the hereditary implementation of hospitality: host and guest exchanged with “tokens”that thereafter could be passed down between their heirs and close friends, guaranteeing the same treatment as that enjoyed by the original xenia parties [305]. Remarkably close to Greek texts are the sacred Judaic stories. Abraham, Joseph, and Moses observed hospitality customs, similar to xenia, in polar opposition to Judaic descriptions of the Egyptians. The latter are portrayed as inhospitable barbarians disrespectful of God and laws, and punished for this. Even more severe is God’s punishment of the tribe of Benjamin whose corruption of hospitality caused a civil war, according to the Book of Judges [309]. Moreover, hospitable Jewish patriarchs are presented as humanity-loving, whereas inhospitable Egyptians as misanthropic. The type-scene of a deity, descending to the world in a traveler’s disguise, receiving hospitality or inhospitality, rewarding or punishing the host, and finally revealing his identity, is pervasive for the literature of the entire ancient Mediterranean world (ibid.). Caucasian hospitality customs also share most of the xenia traits: the opposition of hospitality and vendetta laws, the dichotomy of war and peace, the connection to deities, and the formation of hereditary ties between the host’s and the guest’s families [300]. Even such details as the Caucasian practice of reserving a dedicated room to a guest is reminiscent of the Ancient Greek tradition - after all, the Greek term xenia literally meant ‘hospitality’, ‘friendliness’, and ‘guest room’. Anthropologists identify the basic function of hospitality as an attempt to establish a relationship between previously unrelated individuals by means of constructing a moral universe to which both host and guest agree to belong, which becomes embodied in the material and symbolic ritual of exchanging goods and services [310]. Once the relationship is established, hospitality maintains its moral framework by consolidating the original contract between hosts and guests (ibid.). For this reason, hospitality is structurally transformative in its nature: it “converts strangers into familiars, enemies into friends, friends into better friends, outsiders into insiders, non-kin into kin”-“across a wide variety of social systems”(p.19). The formative social power of hospitality is particularly instrumental for maintaining social order in societies that lack hierarchical systems of authority. Members of hunting and gathering societies routinely leave one group to join another, posing a constant dilemma between hostile and hospitable behaviors, and favoring hospitality as the means of supporting “peace in the feud”[311]. That is why hospitality rules are as important for renewals, growth, and reproduction of heterarchical societies as the institute of marriage [310]. All of the above suggests that the origins of hospitality go way back into our prehistory. Thus, there is no reason to discard the possibility that the intermarriage between the early humans, the Neanderthals, and the Denisovans occurred within the “hospitality prostitution”model, defined through the dichotomy of hospitality/hostility - following the rationale described by modern Eskimo hunters (see above). Furthermore, the fact that in many north Eurasian indigenous societies strict implementation of hospitality code and personal song coexist hand by hand suggests their connection. After all, hospitality promotes long-term cross-cultural contacts (“hospitality prostitution”directly supports intermarriage), while personal song provides means to avoid accidental inbreeding. 6.3. The origins of hospitality and its relation to kinship Genetic research during the last few decades brought abundant evidence of the widespread admixture between the Neanderthals, the Denisovans, and the early humans by the end of the Paleolithic. All present non-African human populations outside of Oceania have a Neanderthal admixture proportion of 1.8-2.6 % - most of which appears to come from the Neanderthal populations that diverged from the Vindija Neanderthal (Croatia, 52 kya) and the Mezmaiskaya-1 Neanderthal (Caucasus, c.60 kya) before their divergence from one another c. 80-100 kya [312]. The earlier Altaic Neanderthal from the Denisova Cave, dated 122 kya, diverged from the Vindija Neanderthal’s lineage around 145-130 kya, but did not carry the gene flow to modern humans (ibid.). The fourth recently analyzed Neanderthal genome from Chagyrskaya Cave (Altai, 106 km west of Denisova Cave, 80 kya) is closer to the Vindija A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 162
Neanderthal (52 kya) and the European Neanderthals who lived ~50 kya than to the Altaic Neanderthal (122 kya) [313]. This and the similarities of the artifacts in Chagyrskaya Cave to the artifact assemblages in Central and Eastern Europe suggests the eastward expansion of the Neanderthals (ibid.). Hublin proposed the eastward expansion of the Western Neanderthals from such “classic” Neanderthal sites as Ehringsdorf (Fig. 6) via the Near East towards Central Asia (Teshik-Tash and Obi-Rakhmat), where the warmer climate made the eastern Neanderthals differ from the Western ones anatomically and culturally [314]. “Denisovans”are an extinct hominin species discovered in the Denisova Cave in Altai in 2008, which was occupied 280-30 kya [315]. The nuclear DNA of the phalanx and the tooth of 2 individuals from the same population group showed that the Denisovans constituted a sister group to the Neanderthals which diverged after the early humans diverged from the Neanderthals, around 800 kya (ibid.). The analysis of a high-coverage genome (30×) obtained from the same fossil confirmed this date and revealed that Denisovans have contributed to the genomes of modern Melanesians (4-6 %, approximately twice the contribution of the Neanderthal DNA in East Asians), Australian Aborigines, and Southeast Asian islanders, while being virtually absent in Europeans [316] and very low in East Asian and American populations [317]. However, there was a significant Denisovan genetic contribution to archaic humans: a 34-kya-old woman from the Salkhit Valley in Eastern Mongolia had a Denisovan ancestor (about 42 kya), so as a 40-kya-old Tianyuan man from the Tianyuan Cave near Beijing [318]. Ancient Denisovan mitochondrial DNA was detected in the sediments from the Baishiya Karst Cave on the northeastern edge of the Tibetan Plateau, Gansu, China, indicating the Denisovan presence at about 100, 60, and possibly 45 kya [319]. The presence of a Denisovan EPAS1 gene that aids high-altitude survival in modern-day Tibetans suggests that the Denisovans were mountain-dwellers [320]. The nuclear DNA sequence diversity among the Denisovans exceeded that of the Neandertals, but was much lower than among the present-day humans, which indicates that the Denisovans were more successful than the Neanderthals in adapting to the harsh Siberian climate - after all, they managed to persist in the Altai region for at least 60,000 years [321]. Admixtures amongst the Late Pleistocene hominin groups must have been common when they had prolonged contacts - however, only humans had an opportunity for repeated interactions with archaic groups over a wide geographic range due to their greater density of population [322]. Nevertheless, the Altai region was a contact zone between Denisovan and Neanderthal populations throughout much of the Middle Paleolithic (e.g., the Denisova Cave was inhabited by the Denisovans, Neanderthals, and humans). Evidently, the contacts were persistent enough for the genetic divergence to have occurred between the European and Asiatic Neanderthals. Thus, the Western Neanderthals, in contrast to the Eastern Neanderthals, had little to no Denisovan ancestry [323]. The Altai region was also the contact zone for the earliest humans who just left Africa to meet and interbreed with the eastern Neanderthals ca 100 kya: there is evidence of the gene flow from the early humans to the Neanderthals and no such evidence for the European Neanderthals [324]. Modern East Asians are genetically substantially more similar to the Neanderthals than modern Europeans are - despite the fact that no Neanderthal fossils have been found in East Asia [325]. This suggests that there were hybrid human/Neanderthal populations migrating from Europe eastward. The place of possible crossbreeding alternative to the Altai region could be the Levant: the earliest human sites at Skhul and Qafzeh and a Neanderthal settlement at Tabun are all dated 120 kya [324]. The dating of contexts associated with the Neanderthal remains at Tabun, Dederiyeh, Kebara, and Amud Caves shows that the Neanderthals still occupied the southern Levant during 80-55 kya [326]. Interbreeding of the Neanderthals, Denisovans, and early humans testifies that conspecific sympathy indeed existed ca 110-40 kya, and the interaction between these different species involved the increase in prosocial behavior - perhaps, even favored prosocial behavior on part of the humans. Since we have no evidence of reduction in aggressive behavior amongst the Neanderthals and Denisovans, and human populations generally exceeded their populations, it is reasonable to ascribe the interest in inter-specific contacts and their initiation to humans. It is possible that the Neanderthals and Denisovans reached the evolutionary stage of proto-music and proto-language - like humans. This period was still characterized with relatively high levels of reactive aggression. Then, ongoing interbreeding between all 3 species must have brought these levels down at least in some social contexts. Plausibly, their cross-specific interaction might have been regulated by some kind of hospitality norms. There are a few reasons for this conjecture. What underlies hospitality is the notion of kin, which determines the discrimination between “relatives”and “strangers.”Unfortunately, the time of emergence of kinship can hardly be established with certainty due to the difficulty of defining the material manifestations of kinship in archeological contexts [327]. However, the hominin admixtures during the Middle Paleolithic (particularly, between the Neanderthals, Denisovans, and humans) must have been instrumental in forming a new practice of tracking the ancestral lineage within a community, since different hominins visibly differed in their physical appearance. Accordingly, posterity of the Neanderthal-human and Denisovan-human hybrids certainly differed from the non-hybrids, which would have provided the most obvious demonstration of kin relations to all tribe-members, promoting the clear-cut distinction between a “relative”and “non-relative”within a tribe. Gowlett, in his overview of the available data on the genesis of kinship, concluded that there was a good chance for the core elements of kin-like system to have emerged sometime 150-100 kya as a result of the “not-quite-total replacement”of older hominin populations [327]. The time period of 45-35 kya in Eurasia brought about the genetic separation between European (e.g., the Kostenki man, 36 kya) and Asian (the Tianyuan man, 40 kya) human populations, which likely occurred prior to 40 kya, but was complicated by their subsequent admixtures [328]. The Gravettian culture in Western and Central Europe marked the emergence of population structure in the distinctly European population that persisted throughout 34-26 kya, overlapping with the Aurignacian culture (associated with the 35-kya-old Goyet Q116-1 individual that probably represented an even earlier structured European population) (ibid.). It would be strange if the physical and cultural differences between such structured populations, sustained over thousands of years across vast geographic regions, would not have put in place at least some tokens of ethnic identity for their members. Some uncovered artifacts can be interpreted this way. Thus, personal ornaments discovered in Gravettian burials might have functioned as “ethnic A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 163
Fig. 6. The major migration pathways in Eurasia c.290-7 kya plotted against the geographic location of Eurasian personal song. The yellow color indicates the areas inhabited by those indigenous ethnicities that keep using the personal song: the S´ ami, Nenets, Komi, Mansi, Khanty, Selkups, Kets, Kamasins, Kotts, Khakas Enets, Nganasans, Yakuts, Evens, Evenks, Nanai, Udeges, Ulchi, Nivkhi, Negidals, Yukaghirs, Itelmens, Chuvans, Koryaks, Kereks, Chukchi, Eskimos, Aleuts, and Ainus (present-day territories of Russia, Finland, Sweden, Norway, Mongolia, China, and Japan). The grainy blue filling marks the area in Europe originally populated by the Neanderthals, and the dark blue arrows show their eastward migration. The numbers reflect the sites mentioned in this paper: the colored number reflects the order along the migration route, whereas the black number to the right indicates the date (in kya). N1 - Ehringsdorf, N2 - Vindija, N3 - Kebara, N4 - Mezmaiskaya-1, N5 - Teshik-Tash, N6 - Obi-Rakhmat, N7 - Denisova, N8 - Chagyrskaya Cave. Grainy green filling encircles the hypothetical area of the Denisovan origin, and green arrows show the possible direction of their migration towards Melanesia. D1 - Denisova, D2 - Baishiya Karst, HD3 and HD4 - the early humans with Denisovan ancestry: respectively, Salkhit and Tianyuan. The grainy red filling encircles the area in Africa from where the early human exodus started [339]. H1 - Tabun, H2 - Kostenki, H3 - Sunghir, H4 - Goyet Q116-1, H5 - Ust’-Ishim, H6 - Sopochinskaya Karga, H7 - Afonova Gora-3, H8 - Mal’ta MA-1, H9 - Buor-Khaya, H10 - Yana RHS, H11 - Ust-Kyakhta-3, H12 - Bianbian, H13 - Devil’s Gates, H14 - Duvanny Yar, H15 - Swan Point. Arrows reflect the regional traditional trading routes. The dotted oval covers the hypothetical location of the Ancient North Eurasians, the diagonal lines - the Northeast Asians, and the combination of dots and diagonal lines - the Ancient North Siberians. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 164
badges”to differentiate the social groups [329]. Population stratification - i.e., a systemic differentiation of allele frequencies between the subpopulations as a result of non-random mating - combined with cultural conservation, was likely to introduce the distinctions between different population strata. Formation of a system of social identification prototypical of clan membership might have been the outcome of population stratification. As Gowlett rightfully emphasizes, the crucial factor that sets human kinship apart from the basic parent-child ties observed amongst chimpanzees is that such ties carry little importance in daily affiliations of chimpanzees. This difference seems to originate in the difference of environmental adaptations: chimpanzees live in large groups in tropical forests, while hominins lived in smaller groups in open-space environments that required long-distance ranging and seasonal movements [327]. Human hunter/gatherer populations often form relatively large communities, but their members cannot always be together, since food supply rarely permits it. The necessity to account for membership in smaller groups, specializing in specific types of hunting and gathering, makes kin-like affiliations important for human sustenance and distinguishes them from the familial units of chimpanzees. Life in harsh climatic conditions further pressurizes hunters and gatherers to form neat small-group alliances and frequently change membership in them. Frequent rotations are likely to promote hospitality towards new members due to the collective incentive of cooperation. Hazards of long-distance ranging across desolate territory make the strategy of helping strangers viable for mutual survival through the reciprocal aid, as opposed to the strategy of fighting for limited resources and facing the possibility of death, crippling trauma, and revenge of the stranger’s relatives. A long-term commitment to the strategy of default aid, unless the reciprocity is violated by the other party - in this case calling for the tit-for-tat response - enables ritualization of the display of friendliness and cooperation between unrelated persons. Both hospitality code and vendetta code must have been brought to life by such ritualization. •The establishment of territorial borders between the Neanderthals as “the people of the north”, the Denisovans as “the people of the east”, as Gowlett notes [327], and the early human newcomers from Africa as “the people of the south”increased the probability of their regular interaction at the crossroads somewhere around Central Asia and Altai. Their interaction would be crucial for the emergence of hospitality and kinship - posing the need to signal the membership in a specific tribe. Gowlett sees language as a primary tool for labeling individuals within groups and inventing words for classes of relationship. Gowlett believes that the emergence of kinship was hardly possible without language. However, as we demonstrated above, the institution of personal song fulfills the purpose of social identification in a much more intuitive, cross-cultural, and effective way than linguistic names. In fact, personal songs do execute the role of “lingua franca”across the enormous range of northeastern Eurasia. That said, we also grant language a key role in modifying these social dynamics (see the next section), particularly during the Upper Paleolithic, and specifically, under the effects of changes in aggression management. To correct Gowlett, the more logical evolutionary scenario is not the simple 2-stage succession of proto-kin/proto-language by rudimentary kin/language, as held by the Barnard’s model [330] that Gowlett relies on, but a much more complex development: 1) Proto-kinship introduces inclusive kinship; 2) Proto-music in the form of a timbre-oriented signaling introduces the hedonistic and ludic functions, while elaborating and ritualizing the expressive means of communicating emotional states, thereby transforming call systems typical for social animals; 3) Human proto-language (to be distinguished from the “proto-languages”purportedly used by other species as discussed in the ethological literature) spontaneously expresses one’s affective state through interjections, while developing the capacity to convey referential meaning through onomatopoeia; 4) Proto-kinship evolves into rudimentary kinship that introduces the antithesis “us-them”as a well-defined family unit within a community; 5) Pre-music (a transitional form from proto-music to music) in the form of timbre-oriented personal song supports social identification and introduces the recreational and inter-personal bonding functions, leading to the emergence of music; 6) Human pre-language (a transitional form from human proto-language to full-fledged languages) capitalizes on the bonds established and supported by personal song, and reinforces the deictic referential communication by single words and simplest phrases. Here we have to underline that the southern border of the Eurasian zone of indigenous ethnicities that retain personal song falls on the Altai-Sayan region (Kamasins and Kotts) - the same that housed the Neanderthals, Denisovans, and early humans (Fig. 6). The institute of personal song could very well constitute an enormously long-lasting cultural tradition, sustained by the extremely stable harsh continental climate, geographic isolation by the high altitude, very low density of population, and life full of hazards. Such conditions make the strategy of landholding not viable, causing ongoing migration over large territory by small “packs”, each of which has to share this territory with the neighboring tribes, helping each other to survive and regulating the mutual use of shared resources [331]. This is the context within which personal song must have acquired the functionality of regional lingua franca. 6.4. Kinship and personal song As we already pointed out, one of the sociobiological functions of personal song is identifying the ancestral lineage, beneficial for avoiding the incest. This capacity was probably forged by “trial and error,”by observing the detrimental outcomes of incest in depopulated tribes, which must have been even more obvious in lives of hominins than in any modern hunters/gatherers due to scarcer population. This aspect of personal song complements the “hospitality prostitution”, which effectively counters incest by supporting the genetic flow into a kin from strangers. Thus, personal song introduces an important upgrade from territorial animal songs. Contrary A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 165
to a widespread conviction that incest avoidance is peculiar to humans and relies on taboos established by cultural conventions, primatologists have convincingly made a case for apparent constraints on sexual relations within nonhuman primate groups [332]. The origin of incest avoidance cannot be explained by social taboos: at least some forms of psychologically motivated incest avoidance are primate homologies - most clearly, between mothers and their offspring and between maternally related siblings (ibid.). What distinguishes human incest avoidance from that of primates is the widening of a circle of close relatives and the abstraction of the concept of consanguinity that transforms the individual-level reaction of sexual aversion/indifference to a closely related person into the group-level morally-based prohibition. Both “hospitality prostitution”and personal songs support social restrictions on closely consanguineous sexual relations, which suggests that personal song emerged after the institution of hospitality codes and spontaneous discovery of “hospitality prostitution”. The remarkable similarity of the institute of personal song in large parts of Eurasia and North America, confirmed by the evidence of gene transfer to their present-day populations, plus the centrality of personal song for the Ancestor Cult and marriage customs - all suggest that personal song was exported from Siberia to Alaska (Fig. 6). The decline in global sea levels during the Last Glacial Period created a land bridge, the central portion of which is known as Beringia, which connected Asia and America and remained traversable between 30-12 kya [333]. The closest to this bridge Paleolithic archeological sites on the Asian side are the Buor-Khaya (45 kya) and Yana RHS (31 kya) settlements, both clustered near the Yana Bay of the Laptev Sea, Yakutia [334]. They stayed inhabited until 22 kya and were 2,200 km away from the Bering Strait. On the American side, the closest site is Swan Point, Alaska (14 kya), 650 km away from the Bering Strait [335]. However, the first people definitely arrived in America earlier than 14 kya, since the oldest pre-Clovis sites are dated 15.5 kya, including Monte Verde in Chile [336], and it would take many centuries to migrate all the way from Alaska to Chile. It seems that ca 20 kya would be a plausible date of the beginning of peopling North America. The genetic studies are quite conclusive in showing that Native Americans have only Ancient North Siberian and East Asian ancestry, which both had gene flow between one another (ibid.). The comparative genetic analysis of the ancient DNA reveals that the first Yana inhabitants were Ancient North Siberians (now extinct) who diverged from Western Eurasians (e.g., the Sunghir man, Vladimir, Russia, 32 kya) about 38 kya, soon after the latter split from Ancient East Asians (the Tianyuan man, China, 40 kya) [258]. Ancient North Siberians were largely replaced by East Asians (e.g., inhabitants of the Devil’s Gate Cave in Primorskoye or the Bianbian Cave in Shandong) between 20-11 kya, giving rise to ancestral Native Americans and Ancient Paleosiberians, represented by a 9.8 kya skeleton from Duvanny Yar, near the Kolyma River (ibid.). The latter represents the closest relative to the ancestral Native American population. His ancestors diverged ca 30 kya from East Asians (ibid.). Ancient Paleosiberians were ancestral to modern Koryaks and Itelmen. Ancient North Siberians either migrated to southeastern Beringia via the coastal corridor along the Sea of Okhotsk and admixed with Early East Asians, or, admixed in the Baikal region and expanded towards North during the Last glacial Maximum (ibid.). The 24kya-old Mal’ta boy from south-central Siberia (and even more so the Afonova Gora woman, 17 kya, from the neighboring Altai-Sayan region) had the mitochondrial genome of haplogroup U (that was also found at high frequency among the Upper Paleolithic European hunter-gatherers) and the Y chromosome that is basal to modern-day western Eurasians [337]. This testifies to the far north-eastern distribution of western populations. About 14-38 % of the Native American ancestry may have originated through the gene flow from this ancient population (ibid.). The newly sequenced Upper Paleolithic genome from the Ust-Kyakhta-3 site, south to the Lake Baikal, is close to the Duvanny Yar individual and, like the latter, shifted toward Native American populations [338]. The entire area from Baikal to Yana and Kolyma during 24-9 kya probably was occupied by Ancient North Siberians prior to the ethnogenesis of Ancient Paleosiberians. The Sayan mountains, west of the Baikal, are part of the Altai-Sayan mountain system, where the Neanderthals, Denisovans and early humans intermarried, where personal song has survived along with the seok clan system (see above), where the present-day indigenous population shares DNA with the present-day Native Americans, and where Ancient North Siberians probably started their long journey to the New World about 25 kya. The latter is indicated by 4 shared major mitochondrial DNA haplogroups and founding Y-chromosome haplogroups of contemporary Native American and South Altaian populations [340]. Also, there is no genomic evidence that any population from a geographic region other than northeast Asia was an important source of the gene flow to America’s first people [336]. Human arrival to Alaska had to occur before 12 kya, when Beringia became submerged by the waters of the Bering Strait –more likely, ca 23-19 kya, when the water level was at its lowest. Accounts hypothesizing about an accidental invention of music (e.g. [10]) seem to us too general and unrealistic, because human music neither directly brings an immediate biologic benefit, nor has natural prototypes that could have served as the basis for its invention (i.e., human music cannot be regarded as a direct emulation of bird songs or animal calls due to significant differences in syntax and semantics). In our view, the evolution of music became possible only once social bonds were potentiated and diversified as a result of increased prosociality. Again, both the view that complex music evolved to support extensive human social bonding and the view that it was selected to provide increasingly sophisticated credible signals to cope with growing social conflicts make sense only for societies that have achieved a threshold of complexity. Hence, we argue for a reversion of casualty: it is not music that enabled humans to use more complex social structures; it is more complex structures, brought about by our increased prosociality, combined with the global warming, that enabled the stable population growth across vast territories - which together enabled music to evolve towards greater complexity. In the next section we present a detailed account of this view. 6.5. The contribution of prosocial behavior to music complexity In order for the music to exercise entrainment and transposability (its defining features), a music-maker has to already have a A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 166
friendly disposition towards a partner-performer and a listener. Collective music-making with a hostile partner can systematically take place only in societies governed by a centralized state power strong enough to prevail over the clan structures. Centralized governance should be combined with the presence of advanced industries that promote division of labor, cooperation, and trade; with the relatively high standard of living that minimizes the infighting for limited vital resources; and with the relative ethnic homogeneity of the population. Such living conditions tend to forge the ideology that condemns violence between the citizens of the same state and enforces cooperation and contractual obligations by legal means. State-based societies usually have formal music theory, some form of music notation, professional musicians, and institutionalized music education. Within such infrastructure, members of the same professional organization (e.g., orchestra or choir) can perform music together despite having hostile relations. In societies that have clan organization and relyon folk music transmission –especially in those societies where personal and duetic forms of musicking prevail over collective forms - the possibility of musicking together with a hostile person is only hypothetical and exceptional. Animistic ideology, common for such societies, usually regards music as a signal that evokes some spirits. Within this worldview, a hostile performance partner would evoke a hostile spirit that can induce harm and attract bad fortune (members of such societies as a rule are very superstitious). Yet, in order for a musical tradition to emerge, the same kind of music must be regularly reproduced, retaining the same semantic characteristics, performance setting, and conventional music structures. Such reproductions should continue over generations and spread over the territory large enough to avoid bottlenecks in transmission and extinction. Performance with and for a hostile person cannot secure survival of conventional means of music expression. The best environment for the emergence and survival of a folk music tradition is the family framework. 6.5.1. Music tradition requires a stable friendly disposition Music performance is not going to entrain someonewho is more likelyto engage ina fight thanin a collective musicking. Hostility is known to obstruct appreciation of music as well: listeners tend to find performance of a disliked musician faulty and meritless [341]. Often such dislike poses a barrier to fair evaluation - a listener simply does not want to listen to the disliked musician. The same avoidance attitude applies to the prospect of joining a disliked audience. Hostility has the power to obstruct entrainment and emotional contagion - two principal ways for music to carry its prosocial action. It is true that the hedonistic function of music can potentially divert one’s state of anger and irritation to some extent by inducing the experience of pleasure, thereby circumventing negative personal emotions, as argued by Perlovsky [186]. Similar effect might carry the reward mechanism in collective musicking, as proposed by Harvey [192]. The experience of togetherness and collective power in inter-tribal competition also can induce positive emotions, as described by Savage et al. [10], which might overcome previous negative emotional states. However, in order for such pacifying effects to take place, the music performance has to occur first - which is impossible without having friendly disposition between the participants in a music act. •Moreover, in order for musicking to become a stable cultural tradition, the hostility and aggressiveness between all the engaged parties must be not just temporarily and partially reduced but eradicated. Otherwise, prospective music-makers would not be seeking an opportunity to learn music, and the transmission chain will be interrupted. A good demonstration of this requirement can be found in human interaction with domesticated animals. A dog-owner can conduct mutual activities with his dog because of the pre-existing bonds - whereas a stranger cannot. Although such a dog is domesticated, the rapport of care and support must be established first in order for cooperation tobecome possible. A dog hasto be treated very much like the mother treats her child - she understands child’s vocalizations and engages in motherese and grooming-like behaviors [342]. In fact, the lexicon of auditory signals that are used by humans to communicate to dogs - “doggerel”[343] comes remarkably close to human “motherese”[344]. Of course, dog-owners usually don’t sing or play music for their pets, whereas mothers do sing lullabies and nursery rhymes to their children. However, both dog-owners and mothers employ music-like prosody, where the melodic contour serves as primary means of delivering the message. This applies not only to dogs but to horses, sheep, cows, goats, and reindeers –in pastoral societies, herders develop a system of whistles or whistle-like singing (e.g., Scandinavian kulning or Mongolian x¨ o¨ omii) to control the behavior of domestic animals [13]. Thus, animal lullabies in pastoral societies share many common traits with human lullabies, just like animal spells resemble human magic spells [13]. What unites all these forms of musicking, be it singing, humming, whistling, or articulating unusual sounds, is the framework of establishing an intimate bond between just 2 parties, which lays down the foundation for building a repertory of meaningful signals. This repertory forms a kind of conventional “mini-culture,”designed to support communication specifically between these 2 parties and therefore confined to a duetic format [345]. It is the sum of great multitude of such “mini-cultures”that in effect generates a musical “macro-culture”of a human community. In practice, the emergence of a continuous musical tradition and the process of evolution of music can be triggered only in a community of music-users that is already consolidated by sympathetic feelings between its members. Only then the transmission of music from one member to another will launch the process of “averaging the beauty”(see 3.3), capable of forging conventional musical idioms. Although the transmission of an already created song can certainly occur between 2 individuals not bound by close relation, in order for that song to survive, it must be adopted by many couples in a more-or-less the same semiotic functionality. And such uniformity, in practice, requires a close bond akin to mother/child between many individuals within the same community. Accumulation of duetic uses occurs naturally within a clan, where close bonds unite multiple couples so that each of them has a chance of witnessing and emulating those patterns that proved to be effective in communication. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 167
•Hence, the formative power of chain transmission is brought to life within the circle of closely related people - i.e., the idiomatic genesis is promoted by kinship. A kin structure secures the ongoing reproduction of the same pattern of auditory communication between multiple couples of individuals that are already bonded by close blood relations. Once the musical idiom is established within a music-user community, it becomes available for combinatorial and compositional creativity of its members, at first between the related kins and thereafter between the unrelated ones. This process forges musical equivalents of phraseological units (similar to our description in 3.3.2) and lays down the syntactic foundation for a music system. The prototype of this syntactic genesis can be found in the animal kingdom. Changes in the dopamine reward system have been regularly attested in domesticated animals [84,85]. More importantly, changes in the dopaminergic signaling have been recently claimed to potentiate vocal learning in domesticated birds. In a nutshell, domesticated species of songbirds manifest extensive signs of positive selection in genes, involved in glutamatergic signaling that attenuates stress responses and results in changes in the dopamine release in the dorsal and ventral striatum, ultimately increasing the variability of vocalizations and the complexity of songs [346]. •These findings suggest that a reduction in reactive aggression during human evolution might have resulted not only in a potentiation of the effects/functions of music, but in the structuring of music itself, favoring the emergence of more diverse and complex musical compositions. We have already touched upon this connection in our discussion of the contribution of the numerosity of music-users to the conservation of a musical tradition: prevalence of solitary and duetic musicking promotes conservation of music syntax, whereas prevalence of performance in front of large audiences promotes innovation and competition in creativity and expressivity via the musical rhetoric. When musicking is restrained to mostly a narrow family circle, the ease of learning tends to take over the expressivity in music transmission acts. There is little need in trying to intensify or exaggerate a musical expression for long-standing partners, who have already developed their rapport. Family communication tends to be brief and simple, with strong deictic components, since all members are well familiar with each other’s habits, needs, and constitutions, which enables them to grasp hints without going into explanations and details. 6.5.2. Simple music systems and the semiotic implications of chunking Ethnomusicological data generally agrees with this observation. Music traditions of those minority ethnicities that maintain traditional lifestyle in sparsely populated territories, where kin and family remain the principal forms of social organization, tend to have music systems that use fewer pitch-classes or are timbre-oriented.Such systems are substantially morelimited incompositionality and combinatoriality than music systems of large nations, characterized by widespread and complex interactions between multiple populous social and ethnic groups. Thus, in North America, the Arapaho, Blackfoot, Cheyenne, Chippewa, Creek, Crow, Dakota, Flathead, Hidatsa, Kiowa, Mandan, Menomini, Muscogee, Omaha, Paiute, Pawnee, Shoshone, Sioux, Tutelo, Ute, Washo, and Yuchi use pentatonic and tetratonic music systems [347,348]. The Klamath, Modoc, Navajo, Creek and Yuchi use tritonic 48 and even ditonic songs. In Northeastern Eurasia, ditonic modes are common amongst the Taz, Yukaghirs, Chukchi, Selkups, Yakuts, and Dolgans; tritonic modes amongst the Taz, Nganasans, Nenets, Yakuts, Yukaghirs, Yughs, Oroks, Evens, Evenks, Khanty, Altaians, Kets, Kereks, Nivkhs, Nanai, Udege, Dolgans, Tofalars, Chukchi, Dolgans, Buryats, Chelkans, and Selkups; and tetratonic modes amongst the Nganasans, Nenets, Nivkhi, Evenks, Taz, Nanai, Shors, Yughs, Yukaghirs, Chukchi, Altaians, Tuvans, Buryats, Barga, Uriankhai, Oirats, Tofalars, Telengits, Chelkans, Yakuts, Khanty, Tubalars, Selkups, Eskimos, Kets, Khakas, Kamasins, Kereks, Koryaks, Enets, and Mansi [138]. Important sources for ditonic, tritonic, and tetratonic structures are such musical instruments as Jaw Harp and musical bow, as well as sacred phono-instruments (e.g., musical log for Nivkhi) and prosody of sacred words [669]. Ditonic, tritonic, and tetratonic modes are encountered amongst indigenous people of the Volga Plato, such as the Mari [349], Tatars, Chuvash, and Mordva [665]. Tetratonic and narrow-range pentatonic modes, similar to the Chuvash ones, were reported in Transylvania [664]. Majority of the Sami (Lapp) songs are tetratonic and pentatonic [350]. Greenlandic Inuits retain a tritonic tradition [351], whereas Alaskan and Canadian Inuits use pentatonic and tetratonic songs [352].“Tribal peoples”of India, especially in northeast, e.g., the Juang, Bhuyan, and Adi, use predominantly tetratonic, tritonic, and ditonic monody, in sharp contrast to the complex traditional music of the Indo-Aryan majority of the northern half and the Dravidian majority of the southern half of India [353]. Tetratonic genres survive amongst indigenous populations of North Vietnam and Southeast China [662,666,667], where they might be promoted by lexical tones [667]. Ditonic modes occupy an important place in the Vedda music of Sri Lanka and the indigenous music of Tierra del Fuego [354]. In South America, tritonic modes are perhaps more widespread than anywhere else. They surviveamongst the Arawakan Wakuenai of Venezuela, the Atacame˜ no of Chile, Argentina, and Bolivia; the Kechua of Peru, the Mapuche of Patagonia and in coplas and baguelas of Northwestern Argentina [416]. Tritonic modes characterize indigenous animal songs of Southern Bolivia [415]. Over a quarter of Shipibo songs are tetratonic and tritonic [663]. All these forms of tritony seem to be pre-Columbian, probably carrying on the Inca tradition [411,412,417]. Tetratonic music systems are quite common in Southern Africa. They originate from the use of natural harmonics in music bows that occupy an important place in this region [668], especially for the !Kung’(Khoisan) and Kavango people of Namibia/Angola and 48 The term “tritonic”here is used in reference to the tonal organization that consists of 3 pitch-classes, following the paradigm of monotonic, ditonic, tetratonic, etc. - and not in reference to the musical interval of a tritone. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 168
the -Gogo people of Tanzania [658]. 49 The Kavango and Damara vocal traditions probably borrowed tetratonic (and occasional tritonic) organization from the local bow music [668]. Tetratonic and pentatonic organization characterizes the entire Basarwa music culture [661]. Yet another source for oligotonal modality in Africa is the emulation of lexical tones in tone languages [670]. However, tritonic and tetratonic modes of indigenous polyphonic music in Equatorial Africa, e.g., Bedzan pygmies, has a completely different – textural - origin [615]. Ditonic, tritonic, and tetratonic modes are also common in Nigeria [659]. Tetratonic modes contrast equi-pentatonic modes in the Gumuz music of Ethiopia and Sudan [660]. Tritonic modes prevail in traditional Maori music, followed by tetratonic and ditonic modes [355]. Pentatonic, tetratonic, and occasionally tritonic modes are generally common in Polynesia [356]. The Yei-anim, Marind-anim, Kanum-anim, and Gad-ze peoples of New Guinea often use tetratonic modes, and the latter also use ditonic and tritonic ones [357]. Ditonic and tritonic music characterizes traditions of Eastern Polynesia and Core Melanesia, and tetratonic songs are common across Papua New Guinea [358]. This by no means complete 50 list of ethnicities that use ditonic, tritonic, and tetratonic music demonstrates the wide spread of simple and well-conserved music systems, cultivated largely within tribal societies with simple and stable social structure. Noteworthy is the geographic distribution of such music (Fig. 7) - it significantly overlaps with that of personal song (compare with Fig. 5), which suggests their evolutionary connection. Without going into the details of tonal organization that characterizes its simplicity, it is easy to demonstrate how greater number of pitch-classes contributes to complexity of music. If an adherent of a ditonic tradition has only 2 possible combinations of tones for constructing motifs and melodies, a tritonic tradition provides a music-user with 6 combinations, tetratonic - with 24, pentatonic - with 120, heptatonic diatonic - with 5,040, and chromatic - with 479,001,600! It is pretty obvious that cultivating atonal dodecaphonic music imposes an immense cognitive burden - in polar opposite to ditonic and tritonic music. Comparing the number of building elements for music provides a convenient universal method to estimate relative complexity of music. Humans are biologically limited in the amount of information that can be received, processed, and remembered. The threshold has been established at 7 “chunks”plus/minus 2, where “chunk”is a set of items treated collectively as a single unit [359]. The bottleneck of chunking is determined mostly by working memory, hence the only workaround for higher transfer rates is to compress more items into a single set, which can be done in series, reaching higher orderhierarchical levels - chunking the chunksas long as they are treatable as a single unit. Chunks of a “maximally compressed”code afford the limit of 4 plus/minus 1 [360]. Fig. 7. Global distribution of ditonic, tritonic, and tetratonic music. Green color marks the areas inhabited or used to be inhabited by the indigenous ethnicities that have retained music genres or entire multi-genre traditions based on fewer than 5 pitch-classes. 49 Tonal organization of these music cultures is well researched. Kubik, in his multiple field studies, has gathered enough evidence and convincingly demonstrated how the solid base of 4 pitch-classes in bow music becomes occasionally expanded to form pentatonic structures –and how the expanded tetratonic modes differ from the pentatonic modes of the neighboring music cultures [658]. 50 For example, the absence of green marks on the Australian map in Fig. 7 does not mean that oligotony and mesotony are untypical for Australia but merely shows insufficient knowledge of Aboriginal music. None of the published reviews of indigenous traditions mention the number of pitchclasses or discuss which musical modes distinguish one ethnicity from another. This seems to reflect the general lack of the field research and the comparative analysis of music structures on part of Australian ethnomusicologists –rather than the uniformity of the Aboriginal music. In fact, field sound recordings, such as “Les Aborigines - Chants et Danses de Australie du Nord”[Arion ARN 64056] and “Aboriginal Music”[Unesco Auvidis D8040] contain tracks sustained in ditonic, tritonic, tetratonic and chromatic modes, as well as timbral and ekmelic music. And the same discrepancy between sound records and ethnomusicological literature applies to many other geographic areas of the world. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 169
The semiotic implications of chunking are that in order to support compositional and combinatorial semiosis, a musical phrase cannot have more than 9 elements or 5 meaningful compound units. For the aspect of melody, an element is a dyadic melodic intonation (i. e., a combination of 2 pitch-classes), whereas the simplest compound unit is a melodic motif (usually, 2-5 intonations) [142]. So, the decoding task of consulting the meaning of the constituent intonations for grasping the meaning of a musical phrase - without resorting to compression - requires the presence of no more than 9 different intonations per phrase. Systemic use of more than 9 intonations per phrase necessitates the “chunking”of chunks - the use of compound structures, i.e., motifs, which introduces hierarchical organization and syntactic rules. •Only ditonic (2 possible intonations) and tritonic (6 intonations) music systems support melodies without syntax. Tetratonic music might permit individual users to use no syntax, provided they do not engage more than 9 different intonations per musical phrase. However, for the entire music community, this option is hardly available: the genesis of a repertory of melodic idioms involves all community members. It is simply inevitable for a music system that affords 24 intonations to sooner or later obtain a substantial number of active users who engage more than 9 melodic intonations per music work. Another common source of complexity is rhythm - phrases are rarely restricted to a single rhythmic value and, as a rule, use rhythmic figures that consist of 2 or more rhythmic values. But each rhythmic figure constitutes a “chunk”in temporal organization - in addition to a melodic “chunk”in the domain of pitch. Even greater complexity emerges in multi-part music, where parts form harmonic “chunks”, and harmonic progressions generate yet another kind of chunks. This important point explains the preference of indigenous ethnicities of the sparsely populated North for ditonic, tritonic, and tetratonic music in monophonic arrangement and illustrates Alekseyev’s observation that these populations prioritize musical conservation in musicking for oneself (and a few closely related persons) over innovation in musicking for many others [136]. This set of values preserves the primordial music systems virtually intact over long periods of time despite the foreign influences and the entropic losses during interpersonal communication. 6.5.3. Esoteric and exoteric semiotic systems and violence It makes sense to distinguish simple music systems from complex ones by adopting the linguistic distinction between esoteric and exoteric languages, which, in turn, relies on the distinction between two basic types of societies: close-knit and open.Ongoing research in the domain of evolutionary linguistics (e.g., [361–365]), suggests that languages spoken by small, isolated groups, organized in close-knit social networks, show distinctive structural features: 1. large and complex sound systems, 2. intricate morphologies with pronounced irregularities, 3. reduced semantic compositionality, and 4. limited syntactic complexity. Such languages are called esoteric (after [362]) and are characterized by great context-dependency, since sharing considerable amounts of knowledge via esoteric communication requires pointing to objects, their features, and actions that are directly observable for all communicating parties (see [225] for details). Most present-day hunter/gatherer societies follow this socio-linguistic model that especially suits isolated small kin-based groups. In a similar vein, the simplest forms of music systems can be called “esoteric”in reference to their limited availability to those outside of their environmental and cultural frameworks. In contrast, languages spoken by larger groups that form more complex social networks with higher rates of cultural exchange (i.e., open societies), exhibit quite opposite features, aimed at effective communication of an increasing amount of information to support collective execution of technologically demanding tasks. Beneficial results of such tasks, in turn, promote the accumulation of knowledge to maintain newly discovered methods of sustenance. The need of conveying precise information to those who do not share any common ground with a speaker forges distinctive linguistic features that are better designed for interacting with non-kin. The need for addressing such issues as conflicts and alliances is a common reason for abstraction and complexification of languages. Such languages are sometimes referred to as exoteric languages [225]. Pentatonic and more complex music systems might be considered “exoteric”in a similar sense of their suitability for musical communication beyond one’s family via a diverse system of conventional music genres and a repertory of musical idioms, used to convey a set of culturally important musical emotions. The tendency of both esoteric languages and music to evolve into exoteric ones is largely determined by the effects of a reduction in reactive aggression on variability and complexity of music structures. This is because the central property of music syntax - which relies on the hierarchical structure - depends to a great extent on cortico-basal ganglia-thalamocortical structures that are shared with language [366]. The same mechanisms of increased neuronal connectivity between cortical and subcortical components that are involved in syntactic chunking of increasingly higher orders of complexity are also involved in the growing suppression/inhibition of reactive aggression [367]. As argued by Benítez-Burraco and Progovac [368], a reduction in reactive aggression entails an increased cortical control over subcortical structures, supporting the advance in processing hierarchical structures not only in the domain of language, but also music - possibly, in the process of their coevolution. Initial grammars were probably used for conveying mostly affective information, much like music - but with the specialization on verbal diatribes to replace physical aggression (see [369] for details). The resulting reduction of harm was to further promote the spread of verbal structures, required to express verbal aggression, leading to a greater control of subcortical networks by cortical devices. In turn, this was likely to support more complex hierarchical linguistic structures - engaging reduced aggression and language complexity in a virtuous circle. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 170
Heterarchical/hierarchical relations of tonal coordination and subordination easily translate into socio-political structures. And this might be more than an analogy. As Blacking points out, the upbringing of children in pre-industrial societies is tied to the musical relationship between children and adults, employed by the authority figures to reaffirm their authority and by the subordinates to secure their safety [410]. Schemes of tonal organization might inadvertently capture the specifics of agency, conflict, and cooperation in a particular society. •The 6-stage progression of monotony->ditony->tritony->tetratony->mesotony->multitony is as unlikely to regress as the ontogenetic progression of establishing the familial bonds in the infant’s socio-emotional development: child->mother/child->parents/ child->parents/siblings->parents+grandparents/siblings. The underlying reason is probably the ease and naturalness of cumulative integration of simpler structures into more complex ones: the addition of new elements does not break the old connections, keeping them functional and relying on their functionality. Kin relations have proved to be stronger than any political powers in the entire region of Siberia as long as the traditional lifestyle was retained. No matter how frequently the multitude of emerging states re-charted territories, clans remained the primary means of social organization there [245]. Citizens stayed loyal not to a state but to their clan. The same dynamics characterizes many indigenous societies in other parts of the world. The weakest socio-political structure that commonly disintegrates is a state. Especially in antiquity, states as a rule were either conquered or ruined by revolts. Complex music systems adopted by states were equally vulnerable. They either disintegrated into a bunch of smaller more homogenous music cultures (the Ancient Roman scenario), regressed by means of a political reform (the Ancient Chinese scenario), or partially collapsed due to decline of participation (the Ancient Greek scenario for enharmonic and chromatic music) [143]. Citizens of states are not nearly as conservative as clan members. After all, the ties of blood relation still tend to overpower the ties of citizenship, when both conflict - even in well-developed modern Western societies. 6.5.7. Voluminousness of music and social hierarchies Yet another factor that ties music structures to social structures is the musical parameter of voluminousness: musical modes are cultivated and acquire specialization for a certain number of participants –solo, duo, trio, etc.. Traditional music, designed for musicking-for-oneself and/or a partner, tends to not support advances in tonal organization. Even if one music-user accidentally comes up with introducing a pitch-class in a conventional timbral musical mode, this advance is likely to be ignored by unrelated users and therefore lost. What prevents the loss of an innovation is its usefulness for satisfying a common need within a trusting circle of people large enough to carry tradition. Thus, if a member of the timbre-oriented music community accidentally invents a new use of a falsetto to represent the voice of a spirit, another member might reproduce this device in representing his dialog with a spirit, where his normal voice (speaking for himself) would oppose his falsetto (speaking for a spirit) within the same vocalization (as in the Navajo Yeibichai songs). This would be a step towards khasmatonal organization. If the entire community finds this invention useful, it might be developed into a khasmatonal mode serving some religious rite. And at this point, the chances for this mode to regress back to a pre-khasmatonal state are little - as long as the corresponding rite survives. In contrast, the probability of khasmatonal organization upgrading to the ekmelic one is high, if, for example, the majority of music-users become attracted by the idea of smoothening the timbral contrast and filling melodic gaps with gliding (say, to reflect the concordance with spirits). The catalyst in the evolution of tonal development within a culture that possesses personal song seems to be the upgrades from family-based to multi-clan tribal and then to tribal confederate organizations. The need to musically mark memberships in a clan and a tribe, to distinguish relatives from non-relatives, is the likely force that promotes the transformation of “monotonous”stereotypical oligotony into more expressive idiosyncratic mesotony. And once the means of generating contrasts within a musical mode are discovered, they become an object of creativity, calling for aesthetic evaluation of music and leading to accumulation of conventional expressive means. They become sorted out and conserved in music genres, each of which nourishes its own array of musical modes for most typical expressions. Eventually, communities of music-users become broken in groups according to music preferences and music usage. The greater the demographic growth - the more subdivisions in social grouping and regrouping (in both production and consumption of music). The next milestone is the rise of statehood that unifies multiple ethnic, linguistic, and religious groups, redrawing their affiliations according to systemic interactions within professional occupations, casts, markets, etc.. The process of restructuring society and establishing new social boundaries usually involves rationalized standardization of music intervals and families of modes, necessitating the emergence of formal music training and theory [143]. Once set, this “learned”(aka, “civilized”) music system is highly unlikely to revert to clan-based music even when a state is conquered by some militant nomads: e.g., a century of Mongol rule over China, Persia, Armenia, and Georgia did not replace their metrophonic system of diatonic keys with mesotonic pentatonic Mongolian system. Even a complete disintegration of a state, such as the collapse of the Byzantine Empire, did not make its fully-fledged tonality (Octoechos system) regress into simpler forms (e.g., mesotony) - nowhere across the former Byzantine lands. One possible exception here might be the Inca empire. Rodolfo Holzmann made a convincing case for the surviving tritonic Q’eros music to represent the remnants of the Inca tradition [411,412]. This triadic music system is associated with the animal fertility cults (specifically the marking ceremony) and carnival celebrations in Southern Andes [413]. Tritony is more basic than pentatony and seems to constitute a trans-cultural phenomenon in the Andes [414]. Thus, tritonic triadic animal songs are cultivated in Potosi, Bolivia [415] - 1200 km away from Cusco, the center of the Q’eros tradition. Tritonic system also survives in nostalgic bagualas songs and religious trumpet music in Northern Argentinian provinces Catamarca, Jujuy, La Rioja, Tucuman, and Salta [416]. The tritonic songs A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 177
of the Atacama people of Northern Chile are believed to preserve a solid pre-Hispanic autochthonous tradition [417]. If Andean tritony indeed constitutes the Inca legacy, the Inca state would presenta rare example of instituting the oligotonal music system - which would complement its highly unusual political and economic organization. However, even this case is unlikely to present a downgrade of tetratony or pentatony. More likely, Incas institutionalized a pre-existing folk tritonic tradition. 6.5.8. “Normalization”of music, language, and social organization Here we should mention an intriguing observation by Fenk-Oczlon that the number of vowels and pitch-classes in a number of native languages and native music systems tend to match - most pronounced in the simplest systems, but also observable in the stability of the average numerosity of 5-7 elements per system across all music systems and languages [418]. Noteworthy, the official language of Incas, Quechua (that survived in Q’ero communities), uses a 3-vowel system. Another case of clear correspondence is between the tritonic chants of the Atacamenos in Chile and the extinct local language of Kunza. The Inuit culture presents the strongest modern case: Inuit language has 3 vowels, while Inuit music is based on the tritonic tradition - still retained in sparsely populated East Greenland and traceable in well-populated South East until a century ago [351]. On the other end of the spectrum, there is a close correspondence between the most complex music systems and the related languages. Thus, French and German languages feature one of the world’s richest vocal systems (17 and 15 vowels), while French and German musical cultures, from the late 19th century on, have been engaging the maximal number of pitch-classes - 12 for welltempered instruments (piano, organ) 60 and 17 for non-tempered instruments (string) and vocals. 61 Evidently, French and German vocal systems have historically acquired extra vowels in parallel to their music systems having evolved from diatonic to chromatic and atonal (e.g., dodecaphonic) organization. This goes to illustrate that music complexity can be increased via iterated learning and cultural evolution pretty much like language complexity can be increased through a cultural mechanism. •One could hypothesize that the common musilanguage platform bifurcated at some point into 3-vowel proto-language and tritonic proto-music, following the paradigm “vowels sing but consonants speak”[420]. This development would have complemented the establishment of a 3-member family unit (father, mother, and child). The numerical expansion of pitch-classes was likely accompanied by the addition of new vowels as the circle of speakers and musicmakers kept growing. Gradually, the dialogic musicking and solitary musicking-for-oneself became merged to form the collective musicking, extending the sphere of attachment and care to a wider circle of people living together. At this point, the opposition between “us”and “them”must have emerged, and the genesis of hospitality customs reconciled their dichotomy. Music must have been prominent in supporting communication with strangers who lacked the capacity to communicate with a host in a common tongue. Music has been reported as part of entertainment provided for the guest in hospitality customs. Music also increased bonding within an entire tribe, facilitating the transmission, remodeling, and sophistication of know-hows via increasedteaching/learning.Decreasedreactiveaggressionisknowntoresultintighteranddenserin-groupnetworks,favoringprolonged face-to-face interactions and promoting the in-group knowledge sharing between adults, as well as between adults and children [421]. A prolonged juvenile period, which is typically observed in domestic animals as well as in humans, is also conducive to enhanced and prolonged play behavior. Play supports knowledge acquisition in many different ways: it increases exposure to the input, exercises the learning mechanisms involved in acquisition of play-related abilities, enables the exploration of new constructs, and favors prolonged practice. This is how the feedback loop between prosocial behavior and music probably fueled the accumulation of knowledge and promoted cultural complexity, including complexity of music and language. Both milestones in the evolution of tonal complexity, oligotony and diatony, as discussed above, have to do with the management of prosocial behavior and aggression. The transition from natural onomatopoeic models to cultural abstract models of encoding semantic information involves the substitution of timbre-orientation by pitch-orientation, which is closely related to expansion of musicking beyond the family circle and inclusion of listeners unrelated by blood. Timbral music is close-distance by design. Timbral musical instruments and vocals lose much of their meaningful detail over the distance beyond a few meters. Timbral transmission is severely limited due to the idiosyncrasy of timbral coloration and lack of objective criteria by means of which timbre could be quantized and reproduced. Timbre-oriented personal songs are by definition “personal”- they are peculiar to its original possessor since it is exceptionally difficult to reproduce a unique voice of a specific person. All these traits confine timbral forms of music to a narrow family circle. Conversely, the change from timbral to pitch orientation opens doors to varied reproduction of ancestral songs or their elements by the offsprings. Formation of the diatonic family of modes also involves profound social changes [143] –just like formation of exoteric languages. Most importantly, folk transmission is distinguished by active musical engagement of the entire population, reliance on audition and 60 The most obvious demonstration of 12-tone tonal organization is the dodecaphony of the beginning of the 20th century, developed in Germany, and its further development in serialism that found strong support in France after WWII. However, the intense chromatic style of composition, where frequent modulations and generous use of harmonic ellipsis often obscured the tonic, was prominent even within the traditional tonal music (e.g., Max Reger in Germany, C´ esar Franck in France). 61 When chromatic music is performed vocally and on instruments other than keyboard and percussions, it engages not 12 but 17 pitch-classes, since all enharmonic tones are systematically tuned about a comma apart: e.g., C sharp is higher than D flat, D sharp higher than E flat, F sharp higher than G flat, G sharp higher than A flat, and A sharp higher than B flat [419,688]. This difference might not be that obvious to lay-listeners but is observed by performers as part of “good intonation”- similar to the variants of basic phonemes of [o] or [a], obvious to native speakers. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 178
memorization skills (in essence, informal vigorous ear training), and absence of prescriptive formal music theory and the notion of mistake (an actual creativity boost). In contrast, the system of diatonic keys that supplants simple stand-alone genre-specific folk modes is practiced by specially gifted and formally trained professionals according to a strict code of rational rules and fixed by means of some kind of notation [140]. Such music-making prioritizes correctness over creativity and distinguishes between “correct”and “incorrect”treatment of musical sounds. The transition from folk informal to professionalized formal music theory requires administration by some authority, political or ecclesiastic, and is limited to the territory controlled by those authorities, who are responsible for the proper use of music across the ethnic, linguistic, and religious boundaries. This is not that different from a state officially adopting a specific language and administering its use through the dedicated offices. Music becomes regulated, and music-users develop a habit to discipline themselves in following the prescriptive rules of performance and listening. Knowledge and veneration of such rules are viewed as the borderline between “civilized us”and “barbarian them”. Importantly, the prosocial power of music is recognized only in regard to the “civilized us”. Those who are seen as “barbarians”appear as a potential threat to civilization that justifies military control. Hence, the stateregulated diatonic music turns into a double-edged sword, friendly towards “us”and hostile towards “them”. Similar development characterizes languages of such civilizations. Languages become normalized: standards of style appear, the grammar becomes codified and fixed, specific instructions about composition of speech are developed, the art of oration and art of literature emerge, leading to formation of new genres of oral and literal speech. Such changes become characteristic of the entire society of speakers. Moreover, some languages become “scholar”(e.g., Sumerian, Sanskrit, Latin) and act as lingua franca, both synchronically, across the adjacent geographic region, and diachronically, across many generations and for many centuries (even in cases where living populations entirely stop speaking such scholastic languages). Exoteric languages support exoteric music. The earliest documented music theory - as retrieved from 29 cuneiform tablets from Ugarit, 14th century BCE - is written in the Babylonian cuneiform script (which is the historic variant of another ancient Mesopotamian lingua franca language, closely related to Sumerian - namely, Akkadian). The explanatory texts were supplied with music notation and lyrics in yet another language - Hurrian [422]. Just like the Babylonian music was regulated by hardbound rules for deriving pitchand interval-classes and their application in specific music genres, Akkadian language developed quite sophisticated grammar and literature. Many Akkadian literary works manifest individual features of style and presence of “Mesopotamian self-consciousness,”distinguished by the assumption that “foreigners, given the necessity and the opportunity, would adapt themselves to Mesopotamian culture”[423]. Akkadian was taught through formal schooling as an oral counterpart of the scripted Sumerian by means of drilling lengthy word lists, learning basic operations of measuring and book-keeping, and acquiring skills of composing documents and letters [424]. This is not that different from drilling pitch intervals or scales while learning music. Noteworthy, a common form of teaching were the improvised dialogues between scribal students or a student and his teacher, master, or parent - teșîtu (‘disputes’). This gave rise to the tradition of judicial rhetoric, where inventive invective was particularly appreciated [425]. Diatribe became one of the leading genres of the Accadian literature. Literacy was by no means exclusive to scribes: multiple texts were uncovered in domestic urban quarters, not authored by court scribes - evidently, simplicity and convenience of cuneiform scripting made it available to commoners for private correspondence and household administration [426]. Formal instruction by means of a grammatically developed language under administration of state authorities was bound to standardize that language and spread it along with the best products of a civilization that spoke it. This was how Babylonian sexagesimal division became cross-culturally adopted for measurements of degrees and time - just as the Babylonian circle of the 5ths came to serve as the basis of diatony across many ancient states of the Middle East and Southeast Europe [427]. All of such developments strongly contrasted the use of languages in tribal societies. Language and music of centralized states become noticeably politicized. While “normalized”music and language consolidate societies of their users across ethnic and social boundaries, generally increasing in-group tolerance and prosociality, by the same token they promote proactive aggression towards outsiders who do not share the same “norms”. This tendency probably originated shortly after the Neolithic revolution and the spread of husbandry, addressing the need to protect one’s property from neighboring hunter/gatherers. Since the beginning of the Holocene, humans have constantly formed relationships and alliances among non-kin groups for achieving collective tasks, including warfare [289,428,429]. As we noted before, such diplomacy stands in contrast with other species - even those that are also believed to have passed through a self-domestication process (bonobos). The latter do not engage in socially coordinated violence (incidentally, this is another reason why the term conspecific sympathy is preferable to the term self-domestication). In fact, one could hypothesize that exoteric languages might have facilitated the surge of premeditated aggression, as they are especially fit for conscious planning, which is instrumental for conducting wars and negotiating military alliances and peace treaties [225]. Overall, it seems that the emergence of different types of music depends on the precise balance between the levels of reactive and proactive aggression. However, it is the “learned”music culture of first civilizations that adopted a permanent proactive aggressive stand towards outsiders. The entire course of development from the emergence of conspecific sympathy to cultivation of proactive aggression with regard to music evolution is laid out below. 7. A model of music evolution Our model builds on the hypothesis that a phenotype similar to the phenotype observed in domesticated animals emerged gradually in our species as a consequence of a progressive decline of reactive aggression, triggered by diverse environmental factors. This allowed the development of increasingly sophisticated forms of music via changes in our cultural niche, resulting from the behavioral A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 179
(and perhaps, cognitive) modifications, brought about by increased social interactions. Although this process manifests itself most clearly in artifacts produced by late anatomically modern humans, we expect human music to have a longer evolutionary history within our clade. Overall, our model encompasses 4 different stages that followed musilanguage: proto-music, timbre-oriented music, the genesis of pitch orientation, and tonality. It is beyond the scope of our work to discuss the putative musical abilities of the extinct hominins. However, because we see the evolution of music as a continuum, we have subsumed this long prehistory in the Stage-0 in our model. Stages 1-4 are hypothesized to be specific to humans, although we cannot exclude the possibility that other hominins, particularly the Neanderthals, might have gone through stages of increasing complexity similar to Stage-1 and, perhaps, Stage-2 in our model. 7.1. Stage 0: Musilanguage Within the hominin lineage, one can confidently expect that, at the beginning, communication did not essentially differ from animal calls. It might have relied on unintentional impulse-driven single signals, where an utterance reflected a certain affective state in reference to some object, action, or quality observable to the signaler. Such communication was called musilanguage by Brown [2] and, when used in group settings, probably engaged the isophonic texture - i.e., brief calls, continuously reproduced by multiple performers with random deviations in timing and pitch [207]. Audio-5. Akia tribe singing, Suy´ a Amazon Indians. Collective performance of personal songs by the entire village [655]. All songs share the same pitch contour and vocables but are not coordinated in pitch and time (presenting the isophonic texture –see [207]). In his most recent update, Brown [227] postulated that cultivation of musilanguage, at first, generated the conventions for affective music-like prosody and, thereafter, formed the intonational prosody similar to language. We share his view. It is difficult to assign specific stages in the evolutionary development of music (just as it is for language) to specific hominin species, but this starting stage could be associated with the earliest Homo species (perhaps, Homo erectus). Homo erectus was the first Homo species to have left Africa –the Dmanisi Man is securely dated to 1.8 million years ago [430], when the climate was not affected by strong cooling [431]. Plausibly, the migration was caused by the increase in population, in search of new territories. Then, Homo erectus might have practiced isophonic vocalizations in Africa and carried this tradition to Eurasia. Homo erectus hunted and gathered in groups, cared for injured and sick group members, while being violent towards non-group conspecifics (might have practiced cannibalism), was highly mobile, possibly used natural fire (at least episodically), possessed a minimal tool technology that barely exceeded the sophistication of tools used by apes and did not progress for about a million years [432]. The survival pressures, brought about by abrupt cooling that took place during the Glacial Maximum of the Mid-Pleistocene Transition, ca 700 kya [433], seems to have forced Homo erectus to move to warmer areas in Southeast Asia, and reduced its population so that it did not survive past 108 kya in Java, Indonesia [298]. All in all, Homo erectus would be a suitable candidate for forging new forms of vocal communication, different from typical animal calls. These new forms likely included production of loud collective signals to scare off dangerous predators [628] as well as grunting patterns to accompany individual caretaking activities. Both applications are common amongst many mammals and both are related to the increase in friendly behaviors between the participants. At the same time, both applications contain traits that distinguish human musilanguage from animal communication. Animal groups, generally, do not use sounds in a premeditated coordinated manner. And the prolonged use of caretaking sounds had a chance of forging characteristic patterns, perhaps even forming a small set of specialized signals due to the greater duration of Homo’s altriciality period, need for helpers during the delivery (because of larger cranial size of the Homo fetuses), and the shortening of interbirth intervals that posed the need for childcare of multiple children [434]. The increase of overall duration, frequency of use, and the numerosity of caretaking participants likely stimulated the accumulation of meaningful vocalization patterns, more diverse and sophisticated than those used by primates. 7.2. Stage 1: Proto-music At some point, musilanguage must have diverged into two forms of communication: emotional information (proto-music) and referential information (proto-language –not to be confused with hypothetical proto-languages of other hominins), both necessary for collective execution of important activities (e.g., manufacturing of the essential tools and hunting). It is beyond our scope here to delve into the distinctive features of proto-languages spoken (or signed) by other hominin species (but see [14,435–440] for reviews). Hamadryas baboons and some other primates have multi-level societies, with some parallels to human multi-level societies [629], but the question is whether these multi-level societies cultivate prolonged congenial (as in bonobos) or brief aggressive (as in chimpanzees) contacts between their members. Here, self-domestication should be viewed more as a continuum than as a categorical condition, and marmosets might present such a case in non-ape species [644]. For early humans, one could hypothesize that at this early stage, levels of reactive aggression were still high, precluding prolonged contacts with conspecifics, long utterances, and cooperative turn-taking - instead, supporting just brief verbal contacts via short singleword commands, threats, and exclamations [225,367,369]. Such a manner of communication occurs in complex societies with strong as well as weak hierarchical organization, in all cases generating more complex forms of communication. We can safely assume that some denotative words, aimed to describe the world around, were already available during this stage. Such are the words referring to specifics of hunting, names for plants and other animals, and kinship terms [441,442]. Kin-related terms, in particular, show a notable phylogenetic stability across different languages because of their frequent use [443]. The A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 180
accumulation of a variety of referential words, perhaps, is one of the earliest signs of growing differentiation between proto-verbal and proto-musical styles of communication. The latter probably evolved via long-chain transmission of the stereotyped musilanguage signals related to hunting, deterring of predators, provision of individual care, and, possibly, as a safety signal [700]. Each of such applications must have shaped its specialized forms for effective communication of corresponding affective states - in line with semantic specialization of animal calls but with greater diversity due to more complex behaviors. As hominin groups grew in size and social complexity, developing new means for conveying the most typical information, two proto-musical genres emerged: aggressive “militant”(preyand predator-oriented sounds, e.g., hunting and war chants) and loving caretaking (sounds of parental care, e.g., motherese and lullabies). Both types are described in detail by Jordania [205,444]. Genres are known to provide the framework for the development of individualized variants of stereotypical expressions [17,687]. The rise of the motherese might have been the result of accumulation of the variety of expressions within the rituals of childcare [436]. Specialization of labor and stratification of a group had capacity to promote new types of sounds, different from musilanguage as well as proto-language. Thus, the metrically entrained sounds of collective knapping might have been inadvertently discovered while making stone tools and thereafter received stable semantic associations, depending on the most common emotional reaction to tool-making (ranging from the enthusiasm of anticipating the possession of a useful tool to the awe from participating in a mysterious ritual) [445]. Such collective “musical”knapping rituals were reported amongst the indigenous population of Queensland, Australia [446]. And archeological evidence for collective rock-knapping was uncovered in Neolithic sites at Sanganakallu-Kupgal, India [447]. It possibly existed earlier. Yet another likely contributor to the separation of language and music was a new type of vocalizations based on hitting a pitch and holding it. Homo heidelbergensis that spread over Europe 700-500 kya was the first hominin that had anatomic capacity to sustain a pitch-level [448]. This task requires a lower larynx position, likely to have emerged to support singing-like vocalizations as they obtained greater importance in social life [451]. Yet another confident proto-musical proxy is the anatomy of the inner ear and the audiometry of hominin species - current research points to Homo heidelbergensis and the Neanderthals as sharers of basically the same hearing functionality as those found in modern humans [452–455]. Having capacities to sing and coordinate the vocal apparatus with the auditory feedback from the heard sounds, early Homo at some point discovered the effect of singing out, liked its sound, and used it repeatedly to experience pleasure - which would acoustically distinguish the proto-musical application from proto-verbal. What increases the likelihood of Homo heidelbergensis coining proto-music and proto-language is its largest growth spurt of the prefrontal cortex amongst the entire hominin lineage [449]. 62 Prefrontal cortex is responsible for such executive functions as inhibitory control and goal maintenance, crucial for cooperation, communication, and ability to stick to social norms under pressure. The unprecedented instability of climate during the last half-million years could explain the emergence of proto-music from musilanguage. Cool episodes promoted the advance in dwelling technologies, the sustained use of fire [450], and cooperation in hunting. Group cohesiveness and cooperation are known to increase species’ability to survive extreme weather. Coordination of collective activities and teaching technologies to youngsters require systems of communication and motivation far more sophisticated than those used in the animal kingdom. Also, depopulation during severe cooling was likely to consolidate relations between the tribe members and set favorable conditions for more intimate forms of bonding and dialogic musicking - similar to the interaction between a mother and her child. Subsequent warming, thereafter, stimulated demographic growth, promoting collective activities. Then, the expressive means of protomusical communication would have passed through long chains of multiple transmissions, which have been experimentally shown to make continuous frequency and amplitude modulations in the transmitted signals discrete and patterned (see 3.1). This transmission effect was likely to transform animal-like signals into proto-musical idioms. The ongoing alternations of warm and cold periods must have forged and conserved at least some conventional structural patterns of tonal organization within the most salient and easy-to-process aspects of expression (e.g., dynamic, timbral, and rhythmic patterns, as well as pitch contours). Structural properties of proto-music might be reconstructed based on what we know about the emergence of the earliest forms of musicking in vocalizations of human infants during their first years of life. In a nutshell, infant’s babbling diverges into “verbal”and “musical”types under the pragmatic pressures of effective communication with caretakers [29]. Namely, “verbal”type emerges from vocalizations intended to point at specific objects in the infant’s surroundings and to demand specific things and actions. “Musical babbling”, in contrast, emerges from instinctive vocalizations that entertain an infant and inform caretakers about the infant’s state. Plausibly, Stone Age Homos mixed both, proto-verbal and proto-musical communication in an undifferentiated manner the same way modern 1-year-olds interact with caretakers. This practice of proto-musicking might have been initiated by pre-human hominin species - perhaps, Homo heidelbergensis in Europe and other locations that were hit by ongoing oscillations of severe and mild climate, causing contraction and expansion of social groups. It was more likely for proto-music to evolve from musilanguage in harsher and more unstable climatic conditions of Europe, especially during the episodes of cooling. In order to keep fire burning during frost and to hunt big prey (a specialty of Homo 62 A number of traits in Homo heidelbergensis’lifestyle might have been responsible for this adaptation. Big-game hunting is dangerous and injuryprone; the chances of a kill for a solitary hunter are much lower than for a group of hunters; moving over long distances without help is hard for mothers, children, and injured individuals; taking care of multiple children by their mother alone is hardly possible in a nomadic lifestyle; and meatsharing at the end of a hunt creates an incentive for everyone to participate in hunting, since the exclusion of free-riders from sharing a prey is a widespread trait of collective hunting [449]. Yet another collective activity, crucial for survival, that should be credited to Homo heidelbergensis is the prolonged controlled use of fire in hearths [450]. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 181
heidelbergensis), group cooperation was a must. In this cold scenario, humans either learned proto-music from the northern populations of the Homo genus or, more likely, developed it themselves once they were out of Africa, closer to the cold regions. That said, we cannot discard the possibility that humans might have developed proto-music in Africa under the pressure of severe droughts. Those few surviving genres of musicking of Arctic indigenous people that are based on sounds as physiological byproducts of specific locomotive activities [687] suggest an idea of proto-music’s sound: Audio-6. Inuit assalalaa, children’s game based on vocalizing on a single breath until exhaustion, while wiggling the entire body. Audio-7. Chukchi pichgeigen round-dance “Girl-friends.”Rhythmic breathing vocalizations accompany stomping and twisting in a half-bent posture, imitating migration of the reindeer’s herd [532]. Audio-8. Nganasan lullaby, rhythmic alternation between two pitch-levels. This music might provide a glimpse into prehistoric lulling of a baby. 7.3. Stage 2: Timbre-oriented music Timbre-oriented music is the best candidate for the first form of music to have emerged from proto-music. According to what we know about the use of timbral music in northeastern Eurasia, it retains many traits of animal calls: it often directly imitates the animalmade sounds, employs the same coding of auditory attributes for the emotional display (e.g., growling/hissing - for aggression, whimpering/whining - discontent, grunting/purring - pleasure, falsetto - submissiveness), and is produced privately, “on impulse”, often for no one but oneself, rather than as a deliberate message to another party or for aesthetic appreciation by the audience [13]. Known forms of timbral music include various types of rasping, whistling, tongue-clicking, wheezing, sizzling, whispering, and yelling - as well as “stylized”talking and playing phono-instruments (various bull-roarers, shakers, scrapers) and instruments that resonate in the vocal apparatus (Jaw Harp, musical bow) [138]. As we already noted, frequent alternation of cooling/warming and/or arid/pluvial episodes almost certainly caused successions of population/depopulation, conducive to the emergence and spread of musical patterns. Timbral patterns were especially prone to benefit from the increased bonding during the harsh climate because of the close-range limitations of the transmission of timbreclasses. Timbre-oriented signals must have provided sonic markers to recognize specific group members like we today recognize a familiar person by the sound of voice. •Episodes of depopulation promoted greater bonding and facilitated memorization and recognition of personal songs amongst smaller social groups. •And episodes of demographic growth spread the established timbral conventions across wider territory, turning them into a music tradition. What is most important, population growth increases the number of song transmissions, both synchronically, between group members, and diachronically, between different generations. Each transmission tends to emphasize those traits of the communicated music that are shared by the majority in a community (see 3.2), so that the conventional structural features tend to replace the biologically-rooted structural features. This process must have made music evolve from proto-music. If proto-music relied on the biologically ingrained semantic encoding (cross-modal correspondences, such as the use of growling timbres with little harmonicity for expression of aggression versus mellow timbres with high harmonicity for expression of submissiveness), music has been based on public conventions. The important difference is that conventions might violate the biological cross-modalities (e.g., many music cultures use low register to express love/ care versus high register for anger/fear - see [13]). A similar process of morpho-genesis has been identified for languages. At the earliest stages of evolution, formation of words is hypothesized to proceed on the basis of cross-modal iconicity (quasi-synesthetic connections between the auditory signifiers and the visual signified) - i.e., biologically driven association of vowels and consonants with certain objects and processes, present in speaker’s environment and determined by similarity of their physical appearance to the acoustic properties (mostly timbral) of phonemes of the corresponding words [456]. As a language develops, the share of iconic ideophones and phonesthemes gradually diminishes, while the share of words based on purely referential (arbitrary) connection between certain combinations of phonemes and their lexical meaning increases, leaving just a bunch of onomatopoeic words in common use [459]. Overall, sound symbolism is untypical for conversational speech of the overwhelming majority of modern languages [457]. In contrast, symbolism bears a special importance for languages of animistic cultures, where speakers systemically employ ideophones, blurring distinctions between a speech event and a narrated event [458]. The same development characterizes the acquisition of languages in early childhood. If onomatopoeic words are usually the first to be learned, later vocabulary acquisitions show great reduction in iconic mapping of sound to meaning [459]. The transition from the biologically ingrained cross-modal semiosis to the semiosis based on learned cultural conventions relies on the broadening and deepening of friendly relations within a user-community. This is especially clear in the ontogenetic development. Infants start their acquisition of language and music only after establishing trust in relationships with their caretakers - i.e., learning to communicate precedes learning to speak and to sing [56,90,345]. Thus, feral children are known to suffer severe speech impediments [164–168]. At this developmental point, the operation of matching becomes universal: person is matched to person, sign to meaning, A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 182
meaning to sign, meaning to situation, and sign to sign. Changes in loudness of the partials, relative roughness, sharpness and “timbral width”of child’s vocalizations are implicitly adjusted to match the tonal qualities of mother’s voice [90]. In essence, interpersonal timbre-matching transpires into interpersonal sharing of semantic values (referential for language and affective for music). This binds the sound of words to their referential meaning and the sound of musical timbral classes to specific emotional states. Once such ties are established, both parties (parent and child) can revert the semiotic scheme from passive to active use. Hearing a musical sound and experiencing a primed emotional state becomes as accessible as wanting to express something and retrieving a suitable musical pattern from the known repertory. The same reversal applies to the referential use of sounds of speech. At this developmental point, the operation of matching becomes universal: person is matched to person, sign to meaning, meaning to sign, meaning to situation (framed in a speech act or in a music act), and sign to sign (in syntactic rules). This line of development, in our view, characterizes the earliest stages in the evolution of music. Matching musical sign to sign probably found its origin in the articulatory base of the musilanguage - certain articulations were anatomically easier to combine than others. But once language and music diverged, language kept heading towards faster and denser information transfer, focusing on lexicology and morphology, whereas music - towards contemplation of affective states, focusing on musical modes. We can expect the earliest forms of music to utilize “timbral modes”[12] in a manner prototyped by the infant’s repertory of vocalizations during the first year of life. Thus, “abandonment cries”can be considered one musical timbral mode, as opposed to “happy cooing”as another timbral mode. Such specialization satisfies the musicological definition of musical mode as a model for abstracting a particular melodic style [460]: each type of infant’s call is characterized by the choice of specific timbral markers, melodic contours, rhythms and dynamics, so that different infants as though all improvise variations of the same “crying”or “cooing” mode. Ontogenetically, the very first cries rely on the biologically ingrained vocal reflexes, but after some period of exploration, an infant usually develops more idiosyncratic patterns of expression. What often confuses scholars is that timbral patterns of timbre-oriented music can be readily interpreted as “defective pitch”patterns by a person habituated to pitch-oriented music [268]. This misunderstanding is provoked by the presence of periodic spectral content in the euphonically timbre-matched sounds. The perceived harmonicity is often experienced as the fundamental frequency by listeners acculturated to pitch-oriented music and by possessors of absolute pitch. A universal developmental phenomenon amongst the children, raisedinsocietieswithpitch-orientedmusic,isthatbythe ageof3-4theylearn toprocesspreviouslyused timbralmodes byreducingtheir spectrum analysis to frequency analysis and adopting the default strategy of looking for a fundamental frequency in musical sounds [29]. Yet another principal feature of verbal/musical divergence in infant’s vocalizations is social: first verbal experiments occur primarily through the dialogic interaction with a caretaker, whereas first musical experiments (babbling) is solitary - explored through playing with mouth in order to entertain oneself when alone (ibid.). This solitary hedonism, combined with the euphony of matching timbre-classes into a timbre-set (Audio-9), perceived as a ludic activity akin to construction games, makes the ontogenetic origin of music a fundamentally personal experience (solitary musicking based on personal liking of sounds) - contrary to the present consensus of scholars that zooms into collective functions that music acquired at the last stages of its evolution. Audio-9. Ainu mukkuri. The entire 2-minute-long fragment captured in this recording is based on matching 2 timbre-classes into a timbre-set. Solitary musicking like this semiotically resembles musical babbling. Traditional Chukchi music provides a glimpse into prehistory. For Chukchi, of all the expressive aspects of music, timbral differences constitute the primary one. Chukchi melodies are distinguished by what Sheikin calls “synsemantic phonemes and phonetic complexes”[532] that are expressly voiced out: e.g., the use of such articulation patterns as [ei-e, e-e, e-kei, o-ka, o-o, te, o-oi] or [o-o-ia-oi, o-ia, o-o-ia, o-o-o-ia], or [o-nga, o-nga, o-nga-ia] in combination with colorful timbral modulations. The most common vowels are [o], [e], [a], and [i], semantically interpreted via phonetic symbolism. Combined with the consonants [j], [k], [kh], [h], [x], and [t,] they generate the recognizable Chukchi sonoric style. Singer’s choice of phonemes, phonation manner (e.g., phono-rasping, phono-whistling, or humming), and breathing patterns (inspiration vs. expiration) are the primary expressive means, supported by rhythm and tempo. Together, they enable the indigenous population to tell Chukchi music from that of the neighboring ethnicities. The following audio examples demonstrate these styles: Audio-10. Chukchi rasping song. Anadyr’, Kamchatka. Audio-11. Even rasping song. Bystrinsky district, Kamchatka. Audio-12. Itelmen rasping song. Tigil’sky district, Kamchatka. Audio-13. Koryak rasping song. Oliutorsky district, Kamchatka. Audio-14. Yukaghir rasping song. Verkhnekolymsky district, Yakutia. The origin of rasping songs remains a puzzle. However, their geographic area of cultivation is quite well-marked: from the Urals to Canada and the Arctic Ocean to Sakhalin [138]. Sheikin believes that rasping songs were forged in the practice of herding the reindeers, because herders’commands are typologically similar to rasping of Chukchi, Eskimos, Koryaks, Itelmens, Evens, Yukaghirs, Nganasans, and northeastern Yakuts. All rasping traditions –in addition to the above-mentioned, Inuit, Kerek, Enets, Nanai, Orochi, Udege, and Ainu - are also related to ritual dances that are more common for women than men [465]. Such dances usually include onomatopoeic and kinematic imitations of animals that determine specific rasping styles: thus, Chukchi distinguish between over 20 types, such as korengy for reindeer, yayak for seagull, or maemael’kai for seal. Rasping songs can accompany personal daily activities pretty much like pitch-oriented songs. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 183
Audio-15. Chukchi rasping solo while currying the reindeer’s hide, Anadyr, Chukotka. Not all rasping songs sound “dirty”and noisy. Thus, Alyutor female rasping games k’arg’ainetyk include pitch-oriented ekmelic styles in imitation of certain birds [692]. Audio-16. Alyutor rasping game between 3 women, settlement Tymlat, Kamchatka. In the past, Koryaks and related ethnicities (like Alyutors) had not only personal songs but personal dances that accompanied singing, and some women chose those song-dances that imitated more melodious bird-calls along with more graceful movements [692]. Such personal song-dances, in turn, prototyped song-dances of these women’s posterity. Onomatopoeic personalization of song-dances, probably closely related to totemic beliefs (where the ancestor of a kin is believed to be a particular animal), so pronounced across Siberia, must be responsible for great regional diversity of Kamchatka song-dances as observed today. A plausible scenario of the genesis of timbral music and its conversion to pitch could be as follows. Characteristic auditory patterns of grunting during grooming could have been reused by an individual during eating, based on the shared experience of pleasure – which then generated a new application of humming. Such humming is widespread amongst modern-day human toddlers. Audio-17. A 3-year old girl humming contentedly during eating her favorite cereal. Tetratonic ekmelic mode, based on phonological contrast of two contrasting vocables, “ʊm”and “nʲam.“Each is characterized musically: “ʊm”ascends, usually initiating a phrase, and “nʲam“descends, usually terminating a phrase. Their interaction within the breathing cycle forms a convex wave in a melodic contour. Courtesy of Jan Sverre Knudsen. Used by permission. This humming (not that different from the Chukchi maemael’kai) could further expand to other pleasurable activities, such as food gathering. Emotional contagion and disposition to seek pleasure supported the possibility of adopting such humming as a signal of satisfaction by other group members. Then, the entire group would end up forming a semiotic convention in distinction from other groups. Due to significant anatomical differences in the length of the vocal folds between males and females, adults and children, the same pattern would be reproduced at different pitch-levels. At some point, one member would notice this difference and deliberately imitate a differently-sized member, thereby abstracting the humming pattern that this other member originally produced. The reproduction of the same pattern at higher and lower registers side-by-side would make the change in a pitch-level salient and thereby reveal the pitch aspect to everyone who witnessed such a reproduction, welcoming creative exploration of this novelty. Some surviving present-day forms of timbre-oriented music can illustrate this scenario. Audio-18. Nganasan timbre-oriented personal song, created by parents for their daughter. It uses timbral markers, such as contrasting registers (growling and glottal) interconnected by gliding in pitch and varied by the vibrato and tremolo embellishments, while featuring pronounced changes in pitch-levels. Obviously, reproduction of this personal song by its owner, a little girl, is bound to differ in pitch and timbre from its rendition by her father. Audio-19. Yakut romantic Jaw Harp serenading. A female player keeps imitating the phrases of a male player until they join together in the conclusion. Although the music generally remains timbre-oriented, it contains frequency modulations and contrasts, more noticeable when both performers alternate. Substantial sonic differences are simply inevitable between any two hand-made Jaw Harps. Performers and listeners easily notice such differences. Imitations of timbral music patterns by different people tend to bring out pitch-related differences, which are likely to become salient in chain transmission and attract attention of music-users, prompting attempts of their deliberate reproduction and exploration. Such is the traditional Nganasan genre of keingeirsia (literally, “singing with someone”) - an allegorical conversation between a young unmarried couple, who are forbidden to directly talk to each other, according to the old Nganasan custom [122].Keingeirsia provides a workaround: they can engage the archaic idiomatic expressions, citations from epic poetry, abbreviations, and occasional regular words with rotated syllables to “cipher”their verbal expressions. Such communication was considered different from regular speech and not any less important: parents taught keingeirsia to their children so that they could marry. Eventually,keingeirsia became a favorite fun activity - an exercise in witty “talk”, common not only amongst youngsters but old people as well (in their recollections of “good old days”). Sometimes, there were more than two participants (e.g., a few male contenders for the samegirl),whichgeneratedacontest.Traditionally,keingeirsia hadtobeimprovisednot basedonthepersonalsongofaparticipant,which posed the need to capture the musical characteristics of a question in order to match the answer. Specifically, communicators were not allowed to reproduce the pitch characteristics of the question exactly - they had to come up with their own arrangement, and their inventivenesswas partof thedisplayoftheirwits.Thispracticemightreflect thetransition frompersonal songto theearliestformsofmusic. Audio-20 and 21. The question of Anai to Kheura and her answer, both sung by Tubiaku Kosteiorkin. The tritonic formula of the question is slightly modified in the answer [122]. The growth of population is particularly prone to intensify the use of pitch as an expressive medium. Since purely timbral qualities are hard to reproduce between individuals of different physique, the pitch continuum (higher-lower) turns out to be more instrumental for reproduction of timbral music than the timbral continuum (lighter-darker or thinner-thicker). This is especially true for personal song: the greater the demographic, the more important the structural clarity in reproducing the “same”musical material (i.e., securing its identity) - no matter who is reproducing it. To remind, the owner of a personal song is not the only party entitled by customs to perform it. Close relatives reproduce it too (especially common is parental reproduction of child’s song). Perhaps, for this reason, personal songs of ethnicities whose music remains mostly timbre-oriented (e.g., Nenets, Nganasan) usually contain at least some tones that are well-defined in pitch (engaged as interpersonal markers, unlike timbre). A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 184
Geographic locations where climate kept changing the most, triggering demographic oscillations while remaining inhabitable, were the most conducive to the transformation of proto-music into music. According to reconstructions of the paleoclimate, Eurasia greatly exceeded other continents in this respect [461–463]. Therefore, the zone northwest and northeast of the Levant (that was readily accessible for migration from Africa), was a plausible place to house the genesis of music. Yet another argument for the Eurasian origin of music is the importance of kinship for the survived Eurasian timbre-oriented and indefinite-pitch music systems, which probably come the closest to Paleolithic music (see 6.3). The evolution of Homo heidelbergensis into Homo neanderthalensis and Homo denisovan and their subsequent spread over Eurasia ca 500-50 kya resulted in interbreeding with early humans around the Altai region, testifying about the presence of conspecific sympathy amongst the Homo species (see 6.4). Differences in physique between the hybrid posterity and non-hybrid members of the same tribe must have helped distinguish one ancestral lineage from another. The byproduct of this visible lineage distinction would be the inevitable expansion of a circle of closelyrelated differently-looking individuals, accompanied with their bonding with the rest of the tribe. Familial expansion likely engaged those musical structures that were already adopted to express love and care within the genre of motherese. The latter must have prototyped duetic communication between the pairs of closely related persons. It is plausible that such vocalizations were somehow differentiated to enable the recognition of a specific relative - following the model of vocal interaction between mother and child, where each is characterized by the contrasting acoustic patterns of expression [90]. Such vocal communication could have received a proto-instrumental support: the sounds produced along with singing on simplest timbre-oriented musical “phono-instruments”- common utilitarian objects like spoons (Audio-26), flask, or cane used as knappers, scrapers, rattles, bull-roarers, buzzers, and shakers [464]. Their sounds are easily discoverable through common everyday activities and often inspire the favorite pastime activities of children, as well as adults in indigenous northern societies [465]. Phono-instruments are also commonly used by all parents to entertain their child. The time period ca 200-110 kya corresponded to the Riss Glaciation followed by the Riss-Würm Interglacial period, when symbolic behaviors emerged in different parts of the world, evident in symbolic use of natural objects and manufacturing of decorative artifacts. The emergence of euphonic timbral modes should be seen as an embodiment of the same trend of appreciation of beauty and the first step towards abstraction and conventionalization of audible patterns of expression. Similar developments are to be expected in the domain of verbal communication. Deixis probably gave way to more abstracted use of words in the absence of the signified entities in speaker’s immediate environment. Single-word utterances might have been combined in a pair-wise fashion - analogous to pairing of timbre-classes of music into timbral modes (Audio-9) and the accompanying pairing of social relations (parent-child, mother-father, brother-sister). The generated compounds directed linguistic evolution away from the onomatopoeia-based iconic semiosis toward convention-based symbolic references to objects and events. The increased prosociality and prolonged interpersonal contacts likely boosted the genesis of stable patterns of verbal expression, perhaps, putting in place the germs of grammar: the use of noun-like and verb-like words, morphological means for naming, distinction of commands and terms to refer to basic social bonds, but above all - derogatory expressions called to vent anger, thereby contributing to the declineofphysicalconfrontation[225,369].Thereplacementofphysicalreactiveaggressionwithverbaloneopenedgatestotheintensive development of syntax by setting a feedback loop between the increases in verbal interaction and the complexity of verbal expression. Finally, as reactive aggression started to decline as a result of a greater control of subcortical areas by the cortex, the compositionality of music increased. We can expect that other changes, important to both language and music, took place as well: the enhancement of human distinctive cross-modal thinking, that is, our ability to transcend the boundaries of core knowledge systems, which underlies our capacity to merge linguistic items and use figurative language (e.g., metaphors, metonyms, and particularly, pragmatic inferencing) [368]. This enhancement of our general cross-modality could account as well for the violation of perceptual cross-modalities in the domains of music and language. 63 7.4. Stage 3: Pitch-oriented music The rise of pitch orientation in music most certainly followed the scenario similar to what we can observe today in the development of music skills during the first couple of years of life. All primordial vocal communication of an infant with caretakers occurs via the acoustic means that characterize timbre-oriented music: modulations in different aspects of timbre, such as darkness/brightness, thickness/thinness, registral coloration and harmonicity. Rhythmo-metric organization plays little role in distinguishing one type of infant call from another (apart from the general correspondence between the pleasurable vocalizations and the regular metric patterns, usually supported by locomotion - as opposed to metrically irregular calls expressing negative emotional states). Through the ongoing dialogic communication with caretakers, an infant acquires the capacity to center on the pitch aspect alone and to continuously control the approximate intervallic size of changes in the pitch contour. This typically takes the form of gradual exploration and mastering of melodic intervals, starting from micro-intervallic steps to wider leaps and is closely related to the broadening of the vocal ambitus available for infant’s vocalizations, which is caused by the growth of the vocal folds and lowering of the larynx [162,468–471]. The first distinction occurs between two general intervallic categories: 63 Cross-modality is construed here (and in Section 8) in a wider sense than when applied to music (synesthesia-like perception patterns). We use this term to refer to our species’distinctive enhanced ability to transcend the limits of core cognitive systems (in the sense of Spelke [466,467]) and to draw relations between different domains. This ability helps us, among other things, (i) to merge concepts and form new ones, (ii) to chunk syntactic pieces and create more complex syntactic structures, and (iii) to metaphorize and metonymize, as in figurative uses of language. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 185
a) “Steps”- i.e., small-size, relatively smooth, incremental changes in pitch-levels to a neighboring level, associated with the expression of “ordinary”intensities of positive and negative emotional states - versus b) “Leaps”- i.e., large-size, abrupt switches between distant pitch-levels, used to express “extraordinary”intensities of positive and, especially, negative emotional states. Similar distinction plays a role in the prosody of verbal speech. The most likely explanation of the universality of this categorization has to do with the anatomical and neurophysiological properties of human hearing apparatus, determined by the Temporal Coherence Boundary [385] and the resulting segregation of audio streams [472]. Each of these two general categories becomes sub-categorized as a child grows - under the influence of the learned cultural conventions that determine the tonal organization of music observable to an infant. Thus, steps break into 3 types (see [383] for discussion): a) “normative”- adopted as a reference and a default intervallic size; its exact size varies between different musical systems from about 100 cents for microchromatic traditions, such as falak in Pamir (see [386]) to about 200 cents for diatonic traditions that prevail in Western Europe (see [39]); b) “small”- reduced in size to express an important characteristic nuance in an emotional state and distinguish it from another similar state; the size varies from about 25 to 50 cents [473] depending on the intervallic typology of a given music system; thus, each of the “chromatic shades”in chromatic and enharmonic genera of Ancient Greek music was differentiated by the intervallic size of 17-33 cents (see [474]) and conveyed a specific semantic quality, ranging from “lugubrious”to “pleasant”(see [475]); c) “large”- increased in size (up to about 500 cents in very fast tempo) in order to emphasize the contrasts between adjacent pitchlevels; such adjustment is common for expression of enthusiasm; the most obvious case of enlarging a step are the ditonic formulaic melodies with two pitch-classes that are about 200 cents apart at the start of a song but gradually expand as the performer becomes more excited (see [155]); pitch-classes of tritonic modes are often triadic and can be separated by steps up to 500 cents (see [136]). Similarly, leaps break into 3 types: a) “normative”- a reference size for a leap is determined mostly by the frequency of its use in a given musical culture the tempo - usually, a normative leap is the smallest interval-class that exceeds the Temporal Coherence boundary (i.e., about 300 cents for very slow music, 400-500 cents for moderate tempo, and 700 cents for fast music) - influenced also by the aesthetic preference and specifics of a musical genre (e.g., plainsong is characterized by the minimization of leaps to project calmness; whereas military march, conversely, favors larger intervals, especially, the interval of a 4th, i.e., 500 cents, to project alertness and fitness); b) “large”- slight exaggeration of the “normative”leap due to the needof expressing anincrease in intensityof an emotional state(the genre of romance); typically, large leaps occur infrequently and exceed normative leaps by a normative step, usually forming a 6th (800-900 cents), and, unlike normative leaps, require stepwise melodic preparation and termination (see [172]) - in contrast to the normative leaps that can “zigzag”sequentially; c) “extra-large”- strong exaggeration of the leap to express a rare outstanding burst of energy or an overwhelming emotional impulse, often negative, characteristic for genres specializing in expression of ecstatic (shamanic kamlaniye) or exalted states (funeral lamentations), indigenous traditions related to evoking supernatural spirits, and/or singing in wide open landscapes, (e.g., Sami yoiks); the most common interval is octave (1,200 cents), less frequently, 7th (1,000-1,1000 cents) and 9th (1,300-1,400 cents), which are more common for pentatonic than heptatonic music systems (e.g., Mongolian aizam urtyn duu). Each of these subcategories can be further subdivided. Thus, in Western classical music theory, 3rds are broken in major and minor varieties; 4ths - in perfect, diminished, or augmented; etc.. Such subdivisions take a long time to emerge, as evident in the ontogenetic acquisition of music skills - even musically gifted trained children can reliably distinguish such intervallic inflections by ear (when presented out of the context of a key/mode) no earlier than by the 8-9th year of life [476]. The pace for the cultural evolution of intervallic inflections is expected to be a lot slower than in ontogenetic development, since early humans did not have any external guidance, unlike modern children guided by adults. Plausibly, these third-order intervallic inflections were introduced during the late Bronze Age - the first documented intervallic system with varieties of 2nds, 3rd, 6ths, and 7ths was codified in Ancient Babylon (see [477]). The genesis of such inflections requires the adoption of an incremental standard for measuring intervals. The earliest pitch-oriented music most likely employed only the rough basic distinction between steps and leaps, as it is evident from the positioning of holes on the recovered Paleolithic bone pipes, aka “flutes”(see [395]). Thus, the 33,000-year-old 3-hole Geißenkl¨ osterl pipe-1 contains the intervals of 2nd, 3rd, and 4th between the sounds to which its holes are tuned - testifying that the maker of this instrument distinguished at least between a step and two kinds of a leap: small and medium-size. In its earliest form, pitch-oriented music likely emerged as part of a grand cultural leap 40-30 kya in caves of Southern Europe, the Balkans, or the Levant. Its emergence can be seen as one of the consequences of the culmination of behavioral changes that resulted from long periods of extreme cooling, continuous depopulation and dispersion over wide territories, increase in cooperation, avoidance of incest, and adoption of a more settled lifestyle in shelter (a cave with fire). Throughout the span of 40-17 kya, the metapopulation of Europe remained low, with barely noticeable growth rates (illustrated by general reduction of the number of known archeological sites compared to earlier times) and lasted until 16.5 kya [478]. Then, the growth rate tripled, and the number of archeological sites nearly quintupled, signifying the intense dispersion of population by smaller groups (ibid.). Such demographic A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 186
phono-instruments (decoys). Naturophonic and anthropophonic intonations share timbral orientation and absence of pitch coordination between multiple sound producers, when engaged in simultaneous musicking (Audio-30). The difference from anthropophonic intonation is that naturophonic intonation aims at reproduction of natural sound sources (the onomatopoeic model) that are external to the sound maker, whereas the anthropophonic intonation follows instinctive impulses to express a certain internal state (the infant-crying model). Onomatopoeic reproduction can be creative - a natural prototype can be taken just as a general idea for one’s musical explorations. Like anthropophonic intonation, it affords the dialogic arrangement. Naturophonic intonation is combinatorial. One kind of naturophonic intonations can be combined with another to form a new compound, whose expression will combine the expression of both types - a clear example of this can be found in the indigenous Jaw Harp music [531]. Although a sequence of such intonations can appear to possess pitch values and comprise a pitch contour, music-makers usually operate within the domain of timbre rather than pitch, constructing music from timbre-classes, based on the capacity of their vocal apparatus or a musical instrument to generate sounds similar to a selected natural prototype. Sometimes naturally-occurring human sounds (e.g., stepping) are imitated. The generated sounds are often appreciated aesthetically and can receive a crafty arrangement. Audio-29. Aleut “The crow-song”, solo imitation of a crow. Audio-30. Koryak dance-song “The ducks”, female duet. Audio-31. Udege solo vocal imitations of birds. Naturophonic intonation might be subdivided in two types: khasmatonal and ekmelic. Khasmatonal intonation is characterized by the contrast of at least two constituent timbre-classes, connected via abrupt leap, usually involving a contrasting registral change (Audio-31). Ekmelic intonation is characterized by the smoothened transition between two (or more) timbre-classes or pitch-levels within the same register (Audio-24). At present, there is not enough data to conclude whether these types are derivative of one another or present a parallel development. 3. Organophonic intonation adopts the sound of a specific musical instrument as a default model for melody-making, including vocal music (Audio-32). This intonation is evolutionarily the first to support the collective performance of numerous participants. It often distinguishes one indigenous ethnos from another by manifesting a sonic ideal of the entire music-user community. The role of a model can be assigned to a percussive instrument tunable to a single pitch (e.g., tambourine for Chukchi), or to the entire instrumental family (e.g., Javan gongs, or African xylophones). A phono-instrument can be selected as an organophonic model, if it is considered of paramount importance for the community of music-users - for example, the sound of a flying arrow in Southern Siberia [138]. The latter case scenario often occurs in societies that retain animistic and totemic beliefs, where a non-musical prototype can constitute a fetish exercising its powers through music. The very adoption of organophonic prototypes necessarily involves bonding and consolidation of a social group. Certain acoustic properties of organophonic intonations set a standard for collective musicking - very much like the concert pitch convention standardizes orchestral music of today. Thus, across Tuva and Mongolia, Jaw Harp has provided the model for local traditions of 2-part solo throat-singing [12]. Audio-32. The same tune performed on khomus (Jaw Harp) and sung in sygyt style (whistle-like throat-singing) simultaneously by the same person. Both versions sound very similar. In a sense, organophonic standardization of a music culture constitutes a step towards abstraction of music structures, which is already observable within timbre-oriented music cultures. There, the same music can be reproduced by preserving the timbral patterns rather than pitch contours. The clearest demonstration of this is the reproduction of well-known songs on Jaw Harp as part of the tradition of “talking Jaw Harp”(Audio-33-35) found in numerous indigenous music cultures [12]. Audio-33. The tune of a Yakut osuokhai sung by Ivan Alexeyev. Audio-34. The same tune rendered in a “talking khomus”style (Yrya)[12] by Ivan Alexeyev. Audio-35. The same tune sung and simultaneously played on khomus by Ivan Alexeyev, reflecting the differences between the pitchand timbre-oriented variants in the sonic outcome of the same articulations of his vocal apparatus. Through organophonic modeling, every indigenous musical culture defines its favorite “acoustic medium”and tries to perpetuate it in the tonal organization of some central musical genre, usually ritual (e.g., wedding, burial) or epic (heroic epos, tales), which becomes canonized and carefully preserved [138]. According to Sheikin, the primary reason for the survival of timbre-oriented music in the extreme North is total absence (or scarcity) of musical instruments capable of generating multiple well-defined pitches (e.g., for the Nganasans - see [122]). Audio-36. Chukchi vyyopchanang (a whistle made of a bird feather’s stalk). Numerous ethnicities of Northern Siberia use such whistles to make decoy calls during hunting and to represent those birds whose feather was used to make a whistle in programme music compositions [532]. A. Nikolsky and A. Benítez-Burraco Physics of Life Reviews 51 (2024) 114–228 193
[Document text truncated for crawler view.]