Composing Music by Selection: Content-based Algorithmic-Assisted Audio Composition
Full text
COMPOSING MUSIC BY SELECTION CONTENT-BASED ALGORITHMIC-ASSISTED AUDIO COMPOSITION Gilberto Bernardes de Almeida A dissertation submitted in partial fulfillment of the requirements for the degree of Doctor of Philosophy in Digital Media — Audiovisual and Interactive Content Creation Dr. Carlos Guedes, advisor Dr. Bruce Pennycook, co-advisor July 2014
! ! ii! Copyright © 2014 by Gilberto Bernardes
! ! iii! Abstract Musicians today have access to a vast array of multimedia content on personal and public databases, which consequently expands the possibilities for audio processing. Yet, composing with audio samples is still a very time-consuming task. A major reason for the disconnect between the state-of-the-art technology and current compositional practice is the lack of effective browsing methods for selecting, manipulating, and assembling samples for a particular application context. My dissertation addresses the aforementioned mismatch by proposing an analysissynthesis framework for assisting musicians in the daunting tasks of selecting and assembling audio signals, namely by incorporating algorithmic music strategies in the processing chain. I address problems raised by the implementation of audio signals in algorithmic composition by creating higher-level descriptions of sound objects, which drastically reduces their audio data representation and yet provides meaningful and highly flexible information. The proposed strategies for audio stream segmentation and description rely on musicological theories, psychoacoustic dissonant models, and contentbased audio processing techniques. Using these frameworks, I finally present algorithmic strategies for style imitation and genuine composition that recombine collections of annotated sound objects for a variety of musical contexts from installations to concert
! ! iv! music. EarGram, the proof-of-concept software I developed as a result of this study, integrates the proposed analysis-synthesis framework in a concatenative sound synthesis system. I also critically analyze some common issues in concatenative sound synthesis and propose the following three solutions that increase user flexibility and control, in particular for creative purposes: (1) meaningful visualizations of the corpus in relation to specific sound typologies; (2) prioritization strategies and/or weights in the unit selection adapted to particular application contexts; and (3) minimization of concatenation discontinuities between audio units by audio processing techniques. In sum, this dissertation proposes a description scheme for representing sound objects that offers relevant information for the sound-based composer as well as suitable descriptions for automatically modeling the temporal evolution of musical structure. In addition, the sound objects’ descriptions are highly flexible and allow the manipulation of audio signals in known computer-aided algorithmic composition strategies linked to symbolic music representations. Ultimately, earGram helps musicians to easily manipulate audio signals in creative contexts—particularly by assisting in and automating a sound mosaic, which allows greater focus on the creative aspects of music making. ! ! Keywords: algorithmic composition, music analysis, recombination, audio synthesis, concatenative sound synthesis. !
! ! v! Resumo Atualmente os músicos têm acesso a um vasta gama de conteúdo multimédia em bases de dados pessoais e públicas que, consequentemente, expande as possibilidades para o processamento de áudio. No entanto, compor com amostras de áudio é ainda uma tarefa bastante morosa. A razão fundamental para a discrepância entre o estado-da-arte em tecnologia e a prática atual da composição é a falta de métodos de pesquisa eficazes para selecionar, manipular e montar amostras de áudio num contexto de aplicação específico. A minha tese aborda a divergência anteriormente referida ao propor um sistema de análise e síntese para assistir o compositor nas exigentes tarefas de seleção e montagem de sinais de áudio, nomeadamente por incorporar estratégias de música algorítmica na cadeia de processamento. Eu abordo problemas que resultam da adopção de sinais de áudio na composição algorítmica através da criação de descrições de objetos sonoros de um nível mais alto que a sua codificação digital. Desta forma, reduz-se drasticamente a sua representação de áudio e providencia-se, no entanto, informações relevantes e altamente flexíveis. As estratégias propostas para a segmentação de áudio em objetos sonoros e a sua consequente descrição baseiam-se em teorias musicológicas, em modelos psicoacústicos de dissonância e em técnicas de processamento baseadas no conteúdo de sinais de áudio. Finalmente, apoiando-me em descrições de áudio, apresento estratégias algorítmicas para imitação de estilo musical e composição genuína que recombinam
! ! vi! coleções de objetos sonoros previamente anotados, e que se adaptam a uma série de contextos musicais desde instalações até música de concerto. EarGram, o software que implementa o modelo que desenvolvi e que valida os conceitos apresentados, integra o sistema proposto para análise e síntese num algoritmo de síntese concatenativa de som. Aqui, também analiso criticamente algumas questões pertinentes e conhecidas da síntese sonora concatenativa e propus as três soluções seguintes, as quais aumentam a flexibilidade de controlo do utilizador em contextos criativos: (1) visualizações representativas do corpus em relação a tipologias de som específicas; (2) estratégias de priorização e/ou factores de ponderação na seleção de unidades adaptados a contextos de aplicação específicos; (3) minimização das descontinuidades resultantes da concatenação de unidades de áudio através de técnicas de processamento de áudio. Em suma, esta tese propõe um esquema de descritores para representar objetos sonoros, que oferece informação relevante ao compositor de música baseada em som, assim como descrições apropriadas para a criação automática de modelos da evolução temporal da estrutura musical. O esquema analítico desenvolvido demonstra também uma grande flexibilidade e permite a manipulação de sinais de áudio em estratégias de composição algorítmica assistida por computador altamente ligadas a representações musicais simbólicas. Em última analise, earGram auxilia os músicos na manipulação de sinais de áudio em contextos criativos—particularmente por assistir e automatizar o processo de montagem de mosaicos sonoros, permitindo-lhes um maior foco nos aspectos criativos da composição musical. Palavras-chave: composição algorítmica, análise musical, recombinação, síntese de áudio, síntese concatenative de som.
! ! vii! Acknowledgments It goes without saying that this dissertation is a collaborative work. My initial words of gratitude go to my supervisors: Carlos Guedes and Bruce Pennycook, whose orientation, endless revisions, criticism, and support were seminal for completing this dissertation. I am deeply indebted to them for their rigor, patience, and encouragement that allowed me to foresee ways to connect computational research to the world of contemporary music. I must also thank both my supervisors for their outstanding hospitality during my visits to New York University Abu Dhabi and University of Texas at Austin. I want to acknowledge my colleagues and friends at FEUP: George Sioros, Rui Dias, Filipe Lopes, Rui Penha, Gustavo Costa e Rodrigo Carvalho for their continuous support, endless comments, and corrections of several stages of this document. I would like to express my gratitude to the composers Nuno Peixoto, Rui Dias, and Ricardo Ribeiro, which showed interest and took the time to learn the software I developed—earGram—in order to use it in their compositions. Their comments, questions, and suggestions were quite useful for improving the usability of the system. I would like to thank Jessica Pabón, Steven Snowden, and June Snowden for their English corrections on the use of terms and grammar. I want to acknowledge Adriano Monteiro, Adriano Torres Porres, and William Brent for
! ! viii! their more or less direct help in some components of earGram. I would like to thank Daniela Coimbra for her exemplary modesty, willingness, patient help, and never ending generosity, which provided me the necessary mental force and courage to withstand deadline pressures. I must also thank David Silva for his kindness and total availability to help me with the formatting and editing of the document. I must acknowledge Fabien Gouyon for the heavy task of proofreading some chapters of the manuscript and his many helpful remarks. I had the honor of being supported by a research grant from Fundação para a Ciência e Tecnologia (grant SFRH / BD / 46205 / 2008), without which it would be almost impossible for me to complete this doctoral program. I also have the privilege of having an excellent group of friends that not only provided me great support in several stages of the research, but also kept an eye on me in order to not lose the motivation to be on stage. During the research period, I would like to highlight some really important artists/friends with whom I had the pleasure to share the stage: Henk van Twillert, Hugo Marinheiro, Fernando Ramos, Isabel Anjo, and Cláudia Marisa. I would like to thank my sister Marisa Bernardes de Almeida for her support and endless explanations of mathematical concepts. My greatest gratitude goes to my parents: Joaquim da Silva Almeida and Maria Rosa Moreira Bernardes, for their long-term journey of love and sacrifice for me.
! ! ix! Table of Contents Abstract Resumo Acknowledgments Table of Contents List of Figures List of Tables List of Abbreviations List of Sound Examples Contents of the CD-ROM Notes on Terminology CHAPTER 1 – Introduction 1.1 - Motivation 1.2 – Approach 1.3 - Concatenative Sound Synthesis 1.3.1 - Technical Overview 1.3.1.1 - Analysis 1.3.1.2 - Database 1.3.1.3 – Unit Selection 1.3.1.4 – Synthesis 1.4 - Time Scales of Music iii v vii ix xiii xvi xvii xix xxii xxx 1 4 6 10 11 13 13 14 15 15
! ! xvi! List of Tables 2.1 - Comparison between criteria of music perception of three representative sound-based theories by Pierre Schaeffer, Denis Smalley, and Lasse Thoresen. 3.1 - Comparison between computational schemes for the description of perceptual attributes of sound inspired by Schaeffer’s typo-morphology. 3.2 - Description scheme used to characterize the audio content of sound objects in earGram. 3.3 - Perceptual attributes of musical dissonance according to Terhardt. 4.1 - Flowchart of the algorithm that reduces the Bark spectrum representation to a single value. 38 45 52 59 70
! ! xvii! List of Abbreviations List of de abbreviations (ordered alphabetically). ASP Aural Sonology Project CAAC Computer aided algorithmic Composition CSS Concatenative sound synthesis DBSCAN Density-based clustering GRM Groupe des Recherches Musicales HPCP Harmonic pitch class profile IRCAM Institut de Recherche et Coordination Acoustique/Musique K-S Krumhansl-Schmuckler K-K Krumhansl-Kesseler MDS Multidimensional scaling MIR Music information retrieval MIDI Musical Instruments Digital Interface MPEG Motion Pictures Expert Group PCA Principal Component Analysis
! ! xviii! PROGREMU Program of music research PSOLA Pitch synchronous overlap add QT-clustering Quality-threshold clustering TOM Traité des Objects Musicaux (Treaty of Musical Objects) TTS Text-to-speech
! ! xix! List of Sound Examples 1 - Three musical phrases synthesized in spaceMap by navigating the visualization of the corpus with the same trajectory and utilizing the following pairs of audio features: (1) fundamental bass and spectral variability; (2) width and sensory dissonance; and (3) noisiness and loudness, respectively. 2 - Sound generated in soundscapeMap by recombining audio units from a corpus resultant from the onset segmentation of a tropical forest recording. The target was defined by navigating the soundscapeMap’s interface, which largely corresponds to a trajectory that goes from sparser and smoother (lower-left corner of the interface) to denser and sharper (top-right corner of the interface). 3 - Sound generated in soundscapeMap by recombining audio units from a corpus resultant from the onset segmentation of a storm recording. The target was defined by navigating the soundscapeMap’s interface, which largely corresponds to a trajectory that goes from sparser and smoother (lower-left corner of the interface) to denser and sharper (top-right corner of the interface). 4 - Three phrases generated in sounscapeMap that synthesize targets that change gradually from the most consonant to the most dissonant—all remaining parameters, density-sparsity and smoothness-sharpness, remain unchanged.
! ! xx! 5 - Sound example generated in shuffMeter. It exposes the utilization of different instrumental sound clusters extracted from a collection of both drum and bass samples. 6 – Sound generated by recombining and layering different sound clusters extracted from Bob Marley’s Don’t Worry, Be Happy, sung by Bobby McFerrin, utilizing a predefined meter of 4/4. 7 – Sound generated by recombining and layering different sound clusters extracted from Bob Marley’s Don’t Worry, Be Happy, sung by Bobby McFerrin, utilizing a predefined meter of 3/4. 8 - Sound generated in infiniteMode’s structSeq that recombines and extends the initial 28 seconds of Jean-Baptiste Lully’s Les Folies d’Espagne (1672, LWV 48). 9 - Sound generated in infiniteMode’s chordSeq that recombines and extends the initial 28 seconds of Jean-Baptiste Lully’s Les Folies d’Espagne (1672, LWV 48). 10 - Random recombination of audio segments extracted from Jean-Baptiste Lully’s Les Folies d’Espagne (1672, LWV 48), which have also been used in sound examples 8 and 9. 11 - Electronic composition Schizophonics (2012) by Rui Dias. 12 – Scale produced by the organization of sound samples according to the brightness of the audio units from the Porto Sonoro sound bank utilized in Urban Sonic Impression (2013). 13 – Scale produced by the organization of sound samples according to the brightness of the audio units from the Porto Sonoro sound bank utilized in Urban Sonic Impression (2013). 14 - Excerpt of the sound installation Urban Sonic Impression (2013) by Rui Dias and Gilberto Bernardes.
! ! xxi! 15 - Live recording of Nuno Peixoto’s song Your Feet (2012), performed by Rita Redshoes (voice) and Nuno Aroso (percussion). 16 - EarGram’s “sonic transcription” of the composition Your Feet (2012) by Nuno Peixoto synthesized with piano and clarinet sounds (sound example 19). 17 - EarGram’s “sonic transcription” of the composition Your Feet (2012) by Nuno Peixoto synthesized with piano and clarinet sounds (sound example 19). 18 - EarGram’s “sonic transcription” of the composition Your Feet (2012) by Nuno Peixoto synthesized with piano and clarinet sounds (sound example 19). 19 - Synthesized MIDI version of Nuno Peixoto’s song Your Feet (2012) with piano and clarinet. 3. !
! ! xxii! Contents of the CD-ROM software ReadMeFirst.txt earGramv.0.18.pd absOverview.pd COPYING.txt dependencies abs 2d.tabread-help.pd 2d.tabread.pd 2d.tabwrite-help.pd 2d.tabwrite.pd 2dPlot-help.pd 2dPlot.pd average-help.pd average.pd bagFifo-help.pd bagFifo.pd barGraph-help.pd barGraph.pd beatInduction-help.pd beatInduction.pd bus.input.pd bus.output.pd
! ! xxiii! bus.pd bus0.output.pd collection.pd colorGrid-help.pd colorGrid.pd correlate-help.pd correlate.pd covariance-help.pd covariance.pd data K-K-profiles.txt psi-functions.txt Temperley-profiles.txt vera.ttf dbscan-help.pd dbscan.pd depot-help.pd depot.pd featureDisplay.pd first-help.pd first.pd grain.pd hpcp-help.pd hpcp.pd indispenser-help.pd indispenser.pd k-means-help.pd k-means.pd keyInduction-help.pd keyInduction.pd list-bounds-help.pd list-bounds.pd list-combine-help.pd list-combine.pd list-datasort-help.pd list-datasort.pd
! ! xxiv! list-euclid-help.pd list-euclid.pd list-maximum-help.pd list-maximum.pd list-minimum-help.pd list-minimum.pd list-permut-help.pd list-permute-help.pd list-permute.pd list-sliding-help.pd list-sliding.pd list-stats-help.pd list-stats.pd markovRetrieve-help.pd markovRetrieve.pd markovStore-help.pd markovStore.pd menu.window.pd meterInduction-help.pd meterInduction.pd minmax-help.pd minmax.pd mtx-stats-help.pd mtx-stats.pd mtxOptimization-help.pd mtxOptimization.pd num2rgb-help.pd num2rgb.pd onsetDetect-help.pd onsetDetect.pd parallelCoordinatesDef.pd pitchCommonality-help.pd pitchCommonality.pd pitchSalience-help.pd pitchSalience.pd pitchShift~-help.pd
! ! xxv! pitchShift~.pd qt-clustering-help.pd qt-clustering.pd quantize-help.pd quantize.pd randomWalk-help.pd randomWalk.pd rgb2pd-help.pd rgb2pd.pd rgbAverage.pd roughness-help.pd roughness.pd rSlider-help.pd rslider.pd rt.buttons.pd runningAverage-help.pd runningAverage.pd runningCollAverage.pd runningCovariance-help.pd runningCovariance.pd runningStats-help.pd runningStats.pd scale-help.pd scale.pd similMtxDist-help.pd similMtxDist.pd smooth_triple.pd spaceMapDefH.pd spaceMapDefV.pd starCoordinates-help.pd starCoordinates.pd starPlot-help.pd starPlot.pd stats-help.pd stats.pd store.param.f.pd
! ! 1! Chapter 1 Introduction Music and technology have been closely linked since ancient times. It is even unthinkable to speak and discuss music and its history without considering the technological developments associated with it. Musical instruments like the piano and violin, for instance, are a remarkable result of the collaboration between music and technology. Musical instruments not only constitute major pieces of technological mastery, but are also seminal for the development of musical expression. As Curtis Roads notes, “the evolution of musical expression intertwines with the development of musical instruments” (Roads, 2001, p. 2). Given the close link between music and technology, it does not seem surprising that the rapid expansion of electronic technology in the late 19th century had a tremendous impact on musical practice a few decades later. In the beginning of the 20th century, the ability to record, amplify, reproduce, and generate sound by electronic means tremendously affected the way we perceive, interpret, and compose music. In the late 1970s, the advent of affordable personal computers offered another avenue for the production of music by electronic means. Computers have become a fundamental tool.
! ! 2! However the music community was, and is still to a certain extent, reluctant to use computers as “creative machines” under the assumption that they are not capable of producing relevant artistic results. The early days of computer music systems relied almost exclusively in symbolic music representations, in particular the Musical Instrument Digital Interface (MIDI) standard. Symbolic music representations encode physical actions rather than an acoustic reality (Rowe, 2001), and model closely the behavior of a piano keyboard, as well as traditional concepts of music notation (Rowe, 2009). Despite its clean, robust, and discrete representation of musical events, symbolic music codes have many drawbacks. For instance, the MIDI standard, one of the most common symbolic music codes, was recognized since its inception to be slow and very limited in its scope of representation (Moore, 1988).1 Audio signals, and in particular digital audio signals, are the most common music representations used today. Contrary to symbolic representations, audio signals encode the music experience, or, in other words, the physical expression or performance. Even if it is a very precise, flexible, and rich representation of the auditory experience and opens up possibilities others than the MIDI or any other symbolic music representation, audio signals also pose crucial problems. Audio signals’ low-level representation reclaim the use of algorithmic strategies, importantly including the field of sound and music computing and music information retrieval (MIR), to extract information from the content of the signal. The field of research concerned with the extraction of information from audio signals is commonly addressed as content-based audio processing, which gained increasing attention in recent years given the large expansion of multimedia content over personal and public databases. Due to the considerable increase of audiovisual contents, it became crucial to develop algorithms for browsing, mining, and retrieving these huge collections !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 1 For a comprehensive discussion of symbolic music representation, particularly its limitations please refer to Loy (1985) and Moore (1988).
! ! 3! of multimedia data (Grachten et al., 2009). A substantial body of knowledge has been presented over the last few years, which offers various solutions to help users deal with audio signals in the era of digital mass media production. The widespread availability of multimedia databases not only affected how users access, search, and retrieve audio, but also enacted critical transformations in how creative industries produce, distribute, and promote music. Research on multimedia information retrieval has also been gradually incorporated in creative work, despite the gap between state-of-the-art research in multimedia information retrieval and usability. From a creative standpoint, processing audio data is still a very elaborate and timeconsuming task. Currently, to create electronic music one usually needs to use software that emulates old analog-tape production means (e.g. audio and MIDI sequencers). These software workstations demand a considerable amount of time to select, segment, and assemble a collection of samples. Despite the large and ever-increasing amount of audio databases, sound-based composers must manage tremendous difficulties in order to actually retrieve the material made available in the databases. One of the most evident and prominent barriers for retrieving audio samples is the lack of appropriate and universal labels for sound description adapted to particular application contexts and user preferences. In this study, I aim to improve music analysis and composition by devising an analytical framework that describes the audio content of sound objects by minimal, yet meaningful, information for users with a traditional musical education background. Consequently, the audio descriptions will be tested as possible representations of sound objects in computeraided algorithmic composition strategies (CAAC) greatly attached to symbolic music representations. The ultimate goal is to devise CAAC strategies that deal almost exclusively with audio signals in order to ease the manipulation of audio samples in creative contexts. In addition to the reformulation of known CAAC to process audio
! ! 4! signals, I study new strategies for composing based on the idiosyncrasies of computer music and the description scheme. The framework proposed will be integrated into an algorithm for concatenative sound synthesis (CSS) and implemented as software (earGram) to test and verify several strategies to analyze and reassemble audio (a detailed description of CSS can be found in section 1.3). 1.1 - Motivation After completing a Master of Music degree at the Conservatory of Amsterdam, which opened possibilities for aesthetic experimentation with interactive music systems, I had the chance to enroll in a new Doctoral program between two renowned Portuguese Universities—University of Porto and the New University of Lisbon—under the auspices of the University of Texas at Austin. At first, I was integrated into a project coordinated by my supervisors: “Gestural controller-driven, adaptive, and dynamic music composition systems” (project reference UTAustin/CD/0052/2008). My involvement with the project gave me a solid theoretical and applied knowledge of generative music, which became seminal for fulfilling the objective of this dissertation. By the time I enrolled in the PhD program, I was mainly concerned with the compositional possibilities of using audio signals as the primary music representation in interactive music systems, in particular the use of large collections of audio samples as raw material for musical processing. One of the major reasons motivating my research was the poor sound and expressive qualities of MIDI synthesizers. A major influence is the work of Tristan Jehan, namely his PhD dissertation Creating Music by Listening (Jehan, 2005), and soon it became clear that I would work at the intersection of many fields including sound synthesis (namely CSS), algorithmic composition, CAAC, and interactive music (see Figure 1.1).
! ! 5! The ultimate goal of this dissertation is twofold: (1) to reshape the compositional experience of working with audio samples, and (2) to devise an intuitive and intelligible guided search and navigation through large collections of sound-objects oriented towards music composition. Figure 1.1 - Overlapping fields addressed by my research, inspired by Ariza (2005). The model I propose aims at reformulating the audio-content description of CSS system audio units through a musical theory and practice standpoint, and targets an audience more familiarized with music theory than with music technology. While my intent is to minimize the usage of computer science terminology, some is unavoidable—particularly concepts related to music information retrieval.2 In addition, earGram will allow the fast exploration of compositional practices by incorporating several CAAC techniques related to symbolic music representation as unit selection strategies in a CSS system, thus proposing new approaches to explore creatively large collections of audio segments. !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 2 Music information retrieval is “a multidisciplinary research endeavor that strives to develop innovative content-based searching schemes, novel interfaces, and evolving networked delivery mechanisms in an effort to make the world’s vast store of music accessible to all” (Downie, 2004, p. 12).
! ! 6! 1.2 - Approach In this dissertation I claim the following hypothesis: Sharing the same constitutive elements manipulated through reciprocal operations, morphological and structural analyses of musical audio signals convey a suitable representation for computer-aided algorithmic composition. In other words, I suggest that analysis3 and composition share the same structural elements and can thus be (computationally) seen as complementary operations of a close musical activity cycle. While analysis fragments the sound continuum into constituent elements according to a bottom-up approach in order to reveal and abstract representations of the various hierarchical layers of musical structure, composition elaborates these same elements in an opposite fashion by organizing musical elements from the macrostructure down to the lowest level of musical structure (top-down approach). The interaction between analysis and composition cannot be discussed without considering music theory. Music analysis and composition not only depart from music theory, but also the constant dialogue between the two fields contributes to music theory with new principles and compositional systems (see Figure 1.2 for an abstract representation the interaction between several agents of the cycle). Any analysis-synthesis computational approach must describe musical structure. Music theorists have recognized and identified in the temporal span of the music continuum several hierarchical levels (Roads, 2001). The composer’s task is undoubtedly to elaborate !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 3 Analysis refers to the general process of separating something into its constituent elements and to a certain extent to the examination of the elements or structure of something, typically as a basis for discussion or interpretation. However, it does not imply music analysis, which focuses essentially on the interpretation and elaboration of the elements provided by the analysis carried here.
! ! 7! the several levels of creating a sonic work. Analysis often examines a compositional impulse, while composition often elaborates an analytical impulse. In order to pursue the aim of this dissertation, I intend to computationally model the music cycle present in Figure 1.2. Specifically, I aim to design a computational system that learns from given musical examples, and/or relies on music theory knowledge, in order to generate meaningful musical results with minimal user interference. The analysis agent encompasses two operations: listening (perception) and learning (cognition), while its complementary agent is composition (action). These two agents are in a constant and reciprocal dialogue with music theory, a repository of knowledge constantly populated with new knowledge generated by the two aforementioned agents. Figure 1.2 - Basic building blocks of the musical life cycle computationally modeled in this dissertation. My analysis of audio signal content aims at providing representations and revealing patterns of the musical surface higher than the sample temporal unit. In order to do so, I will devise a bottom-up or data-driven computational model for the automatic segmentation and description of sound objects and musical patterns according to criteria of musical perception grounded in sound-based theories by Pierre Schaeffer (1966), Denis Smalley (1986, 1997, 1999), and Lasse Thoresen (2007a, 2007b). Alongside a critical ANALYSIS listening & learning (perception and cognition) COMPOSITION (action) MUSIC THEORY AND PRACTICE
! ! 8! discussion of the criteria of musical perception proposed in the cited theories, I will devise a set of descriptors for characterizing sound objects; the description scheme is adapted to the idiosyncrasies of a CSS system. Relying on the sound objects’ descriptions, I then identify and model higher structural levels of the audio data by grouping sound objects into recognizable patterns up to the macro-temporal level. Outlined from a music theory and practice standpoint, my model is adapted for music analysis and composition. The outcome of the model intends to provide a rich representation of the audio content in a compact and meaningful representation. However, it does not provide a successful answer to the ultimate goal of the analyst, which is to explain the organization of several events and to reveal how meaning derives from those organizations. Instead, the model provides information that can either allow a different view over the sound material or establish comparisons between vast amounts of material that are not traceable by human senses. A human interference is mandatory in order to determine the causal linkages between the sonic objects and to determine the relationships between patterns (if this level of syntax exists). Figure 1.3 – Hierarchical organization of the music time scales considered in the analytical module. patterns sound objects samples
! ! 9! Segmenting the several layers in an audio continuum, along with the description of its constituent units (sound objects), not only provides a groundwork for the analyst, but also for the sound-based composer. In other words, the outcome of my analytical model is suitable for guiding the composition process by reciprocating the analytical operations (i.e. through a top-down or knowledge driven approach). The outcome of my analysis offers the composer a good representation of the audio source’s structure and allows a fast and intuitive reorganization of the segments from the macrostructure to the basic element of the musical surface (sound object). One can compose the macrostructure in earGram by selecting sub-spaces of the corpus that can be assigned to a particular piece, performance, or even to different sections of a work. The process is manual, but guided by several visualizations that expose the structural organization of the corpus, such as similarity matrices and 2D-plots. Some patterns of the audio source(s) structure may also be revealed through the use of clustering techniques in combination with the visualization strategies aforementioned. The recombination of the sound segments in earGram is automatic and it is mostly done by adapting CAAC algorithms related to symbolic music representations to function as selection procedures in CSS. The CAAC strategies can be guided by music theory knowledge or models created during analysis from user-given examples. As the name implies, CSS deals with the concatenation or juxtaposition of sound segments, that is, the horizontal dimension of musical structure (e.g. melody, metrical accents, dynamics and properties relating to timbre). However, it is also my intention to expand the CSS scope of action to handle the recombination of units in the vertical dimensions of musical structure (units’ simultaneity) as a cause of timbre creation and variance, control of the event density, and (psychoacoustic) dissonance. Finally, in earGram I will explore the idea that all sonic parameters, such as brightness and sensory dissonance, can be as important as parameters like pitch and duration, which are commonly seen as primary elements of musical structure. I envision all criteria for
! ! 10! sound description as fundamental “building blocks” for compositional systems. This is not to say that every piece composed by these means must use equally all sonic parameters, but that all sonic parameters may be taken into careful consideration when designing a musical work and seen as primary elements of musical structure. 1.3 - Concatenative Sound Synthesis CSS is “a new approach to creating musical streams by selecting and concatenating source segments from a large audio database using methods from music information retrieval” (Casey, 2009). Briefly, CSS uses a large “corpus” of segmented and descriptoranalyzed sounds snippets, called “units”, and a “unit selection” algorithm that finds the best matching units from the corpus to assemble a “target” phrase according to a similarity measure in the descriptor space. The first CSS software appeared in 2000 (Schwarz, 2000; Zils & Pachet, 2001) and their technical basis strongly relied on concatenative text-to-speech (TTS) synthesis software—a technique presented in the late 1980s (Schwarz, 2004). CSS began to find its way into musical composition and performance beginning in 2004, in particular through the work of Bob Sturm (2004, 2006b) and Diemo Schwarz (Schwarz, Britton, Cahen, & Goepfer, 2007; this paper documents the first musical compositions and installations exclusively produced by CataRT, a real-time CSS software developed by Schwarz). Currently CSS is considered state-of-the-art in terms of sample-based techniques and content-based audio processing. The technique is at an interesting phase of development and attracts a broad audience of users, researchers, and developers from the scientific to the artistic community. CSS shows great potential for high-level instrument synthesis, resynthesis of audio, interactive explorations of large databases of audio samples, and procedural audio—especially in the context of interactive applications, such as video games. Despite its mature development at engineering and technological levels, CSS is rather undeveloped in terms of aesthetic
! ! 17! all time scales. Even if these values offer a comprehensible suggestion of the duration of each scale of music, they may also provide erroneous information because in a concrete musical work all time scales are identified and definable based on their function, rather than their effective duration. 1.5 - Outline of the Dissertation I have divided this dissertation into two large parts. Part I (chapters 2, 3, and 4) provides a review of musicological theories for sound description and then offers an algorithmic description scheme for describing and modeling the content of audio signals at various hierarchical levels. Part II (chapters 5, 6, and 7) presents algorithmic strategies to automatically recombine segmented-analyzed audio units, then summarizes the original contribution of this study, and finally provides guidelines for future work. Chapter 2 presents an overview of three major musicological theories for sound description by Pierre Schaeffer, Denis Smalley, and Lasse Thoresen. In addition, I critically compare the three aforementioned theories with a particular focus on their criteria for the morphological description of sound objects. Chapter 3 discusses the conceptual and technical considerations that assisted the creation of a description scheme adapted to the automatic characterization of sound objects. Chapter 4 examines the higher layers of musical structure from an analytical standpoint, focusing on how visualization strategies and statistical analysis help reveal and model musical structure. The visualization strategies are supported by two topics, which are extensively discussed: (1) the computation of similarity between sound objects, and (2) clustering algorithms that help group sound objects that expose similar features and reveal temporal patterns of musical structure. Chapter 5 provides an overview of the technical and conceptual background of the
! ! 18! framework’s generative strategies, particularly an historical perspective of sample-based techniques and compositional systems that contributed to earGram’s design. In addition, it is details the articulation between the two major modules of earGram—analysis and composition—and how their interaction establishes a compositional system. Chapter 6 describes generative music strategies implemented in earGram, from the organization of the macrostructure down to the lowest level of the generated music. Finally, Chapter 7 discusses my conclusions and original contributions to then provide guidelines for further study.
! ! 19! PART I: ANALYSIS Ce n'est point avec des idées, mon cher Dégas, que l'on fait des vers. C'est avec des mots. — Stéphane Mallarmé
! ! 20! Chapter 2 Sound Morphologies: From Pierre Schaeffer and Beyond This chapter presents an overview of three representative analytical theories of soundbased works by Pierre Schaeffer (1966), Denis Smalley (1986, 1997) and Lasse Thoresen (2007a, 2007b). Each theory is largely presented according to three topics: (1) methodological premises; (2) conceptual framework; (3) and morphological criteria of sound perception. The first two topics acquaint the reader with select information regarding the foundations and guidelines of the analytical theories necessary in order to then focus special attention on the third topic, laying the groundwork for the development of a computational description scheme presented in the next chapter.
! ! 21! 2.1 - Describing Sound Sound description is an essential task in many disciplines from phonetics and psychoacoustics to musicology and audio processing, which address it for a variety of purposes and through very distinct perspectives. The two disciplinary approaches to sound description most relevant to this dissertation are computational and musicological. Computational refers to content-based audio processing strategies, namely the use of audio descriptors to provide an automatic characterization of an audio signal’s content. For example, a computer can easily describe how bright, loud, and stable a sound is by inspecting characteristic present in its digital signal representation. Content-based audio processing systems that extensively use audio descriptors tend to exclude the analytical operations of the system from the interface. The computational descriptions in content-based audio processing systems like Shazam4 and Moodagent5 take place during the implementation, or training, phase of the algorithm and are hidden from the system’s interface, thus preventing the user from accessing them. Contrarily, creative applications like Echo Nest Remix API6 and CataRT (Schwarz, 2006a) give access to the generated audio descriptions and even allow their manipulation, because it is an inherent process of music creation. However, most audio descriptors extracted computationally, like spectral kurtosis and spectral flatness, are not adapted to the terminology of musical practice and are meaningless to musicians. The gap between computational descriptions of sound and music practice/theory is essentially a problem of terminology, because even if many audio descriptors measure musical or perceptual properties of sound, they are always addressed according to the mathematical operation involved. Developing a set of descriptors adapted to current music practice will increase the usability for those musicians more familiar with music !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 4 http://www.shazam.com. 5 http://www.moodagent.com. 6 http://echonest.github.io/remix.
! ! 22! theory and practice than with music technology. By unpacking the language, the usability of content-based audio systems would increase considerably, and appeal to a larger audience, most-importantly including musicians. Indebted to previous research by Ricard (2004), Peteers and Deruty (2008), and Schnell, Cifuentes, and Lambert (2010) (each of whom was inspired by Pierre Schaeffer’s typo-morphology), the strategy I apply here will offer a description scheme adapted to the needs and knowledge of musicians. I will now discuss musicological approaches to sound description for two reasons: (1) to present a theoretical basis of the mechanisms behind sound description; and (2) to provide a succinct set of descriptors adapted to music imperatives, particularly composition. In addition, utilizing perceptual criteria like mass, harmonic timbre, and dynamics, based on musicological literature describes abstract sounds independent of their sources because they rely on perceptual characteristics of the audio signal disregarding causal relationships. 2.1.1 - A Western Musicology Viewpoint: From Note to Noise The emergence of electroacoustic music in the 1940s extended significantly the practice of music creation with new instruments, a myriad of tools resultant from the possibility of recording and diffusing audio, raw material that has been unexplored so far in music composition, etc. Until then, music composition was confined to acoustic instrumental and vocal models and uniquely focused on the manipulation of the following four elements: pitch structures (melody, harmony, aggregates, etc.), rhythm (meter), timbre (restricted almost exclusively to orchestration), and form (theme, motives, macroform, etc.) (Thoresen, 2007b). These musical elements convey a clear understanding of musical structures that are highly tied to the concept of musical note. The musical note, as the basic unit of composition, favors pitch and duration as primary musical elements over timbre or other attributes of sound.
! ! 23! The appearance of new electronic instruments and sound manipulation strategies broke the paradigm linking sound to the physical object producing it, and allowed composers to work with dimensions that were previously inaccessible or totally disregarded in music composition, particularly the use of all sonic phenomena as raw material for a composition or expanding the act of composing to the sound material itself. In electroacoustic music, the basic structural unit of the composition is no longer the musical note. Instead, the concept of sound object comes to the fore, significantly extending the spectrum of possibilities (from note to noise) without indicating a priori sources or known causes. Electroacoustic music opened the exploration of timbre and reformulated the notion of spectrum as a compositional strategy. As a result, much electroacoustic music was particularly resistant to traditional analysis and categorization. In addition, the new dimensions explored in electroacoustic music existed for some decades without any theoretical ground or formal definition that could articulate the relevant shift within musical composition. Clearly, a unique set of terms and concepts was needed to discuss, analyze, and interpret electroacoustic music (Smalley, 1986). In the early years of electroacoustic music theory, the discourse was largely monopolized by engineering terminology, consequently lacking theoretical and aesthetic reflection. In 1966, Pierre Schaeffer presented Traité des Objets Musicaux (TOM)—the first substantial treatise on the subject, which addresses the correlation between the world of acoustics and engineering with that of the listener and musical practice. While the technology used by Schaeffer is now outdated, and his argument far from the model presented here, his overall perspective in TOM is valid because of the approach taken to listening and the new concepts and taxonomies of timbre and sound description. 2.2 - A Schaefferian Approach to Sound Based Theory TOM was the first major essay that attempted to understand and devise an analytical
! ! 24! theory for sound-based works. In TOM, Schaeffer outlines a Program of Music Research (PROGREMU) that provides several stages of action like the definition of different types of sounds along with their morphological description, characterization and organization. These stages aim to abstract musical value from audio signals for particular musical contexts (Landy, 2007). Although it provides a solid foundation for musical composition, TOM is “situated rather in the area of hearing than making, it is descriptive rather than being operational” (Chion, 1983, p. 98). Schaeffer reframes the act of listening to sound by articulating a phenomenological theory that is primarily concerned with the abstracted characteristics of sounds, rather than their sources and causes (Chion, 1983). The theory articulates modes of listening to sound that ultimately establish the basis of a solfeggio for sound-based works. In order to acquaint the reader the basis of TOM’s methodology, four concepts coined by Schaeffer—(1) concrete music (musique concrète), (2) listening functions, (3) reduced listening, and (4) sound object— will be examined next. Their order reflects a top-bottom organization of music practice/theory and human perception principles. According to Schaeffer, concrete music denotes the music created by a group of composers working at the French Radio, which later became the Groupe des Recherches Musicales (GRM). Schaeffer provides an explanation of this term, as many others, by referring to binary and antonym concepts (Landy, 2007). The term “concrete” is used to represent a musical reality, a new creation paradigm that opposes abstract music, which was the prevailing composition model for vocal or acoustic instruments at the time. It emphasizes that the raw material for a composition is based on pre-recorded sounds. In contrast, traditional Western music composers start with an abstract idea of the work, which only later achieves its concrete form when performed. Schaeffer’s theory derived from particular active listening functions. He describes four listening functions related to different ways of perceiving and understanding sound: écouter, comprendre, entendre and ouir. John Dack and Christine North translated them
! ! 25! to English as: to listen, to comprehend, to hear, and to perceive, respectively (Chion, 1983). To listen refers to the identification of the sound-producing event through the sound. In this case, sound is seen as an index of an event. To comprehend implies the identification of a message transmitted by the sound. This listening function is well illustrated by speech, in which sound is only a “vehicle” carrying meaning for words. To hear is to perceive the intrinsic properties of a sound. These qualities allow us to distinguish between different instruments, for instance. To perceive refers to the discernment of the raw-sound data with no intention of interpreting or qualifying it. It is the lowest level of our auditory perception and can be seen as a kind of passive listening (Chion, 1983). Another key concept in TOM is reduced listening, which is a listening attitude that focuses on the morphological qualities of sound rather than its causes or meaning. Reduced listening neglects the phenomenon of source identification that is highly linked to vocal and instrumental music and describes Schaeffer’s methodological approach adopted in TOM for analyzing the qualities expressed by sound events. As Jean-Claude Risset notes “in the first instance, Schaeffer placed the accent on the primacy of the listening experience and on the necessity to develop a solfeggio of effects as opposed to causes” (as cited in Thomas, 1999, p. 37). It also would be difficult to approach sound with reduced listening as a strategy without available technology for sound recording and diffusion, because we must be exposed to a sound many times to fully grasp its morphology. The repetition of the physical signal prompts more awareness of its perceptual attributes and relegates the listener’s attention on the sound source to a secondary level (Schaeffer, 1966). Schaeffer’s reduced listening denotes an attitude toward listening that is characteristic of its time: to listen to sounds whose cause is invisible, such as radio broadcasts, telephone conversations, or recorded sounds. The focus on the sonic matter is a guiding principle for the compositional approach of the composers associated with the GRM.
! ! 26! Lastly, let us address the concept of sound object. The sound object is defined as the basic unit of musical structure, which resembles the concept of note in traditional Western music. In other words, the basic unit of composition and analysis is the concept of sound object that encompasses sound events that are perceived as an entity. A sound object can be identified by its particular and intrinsic perceptual qualities that unify it as a sound event on its own and distinguishes it from all other sound events (Chion, 1983). After enlightening seminal concepts and the methodology present in TOM, I will delve into the core of Schaeffer’s treatise and present first a brief overview of the basic organization of PROGREMU, and then focus on the second stage of this program— morphology—which aims at outlining perceptual criteria for describing sound. 2.2.1 - Program of Music Research The core of Schaeffer’s TOM is the PROGREMU, which guides the user through “the art of practicing better listening” (Chion, 1983, p. 38) in relation to musical activity. PROGREMU is divided into five stages: (1) typology, (2) morphology, (3) characterology, (4) analysis, and (5) synthesis. The first two stages of the PROGREMU are commonly addressed together as “typomorphology,” and as the most detailed stages undertake three tasks in relation to sound objects: (1) to identify; (2) to classify; and (3) to describe. The first task aims at identifying sound objects from an audio stream. The resulting segments are further classified into distinctive types, and, finally, exhaustively detailed according to their morphological characteristics. Typology takes care of the first two operations and morphology the third. In sum, sound objects are categorized into a typology based on perceptual attributes. While the ultimate aim of Schaeffer’s typology is to assign “value” to sound object and derive their suitability for musical activity, morphology offers a refined and precise description of the sound objects and their inner structure. As
! ! 33! Figure 2.2 – Smalley’s (1986) attack-effluvium continuum. (Copyright 1986 by Palgrave Macmillan. Reproduced with permission.) Similar to the spectral typology’s pitch-effluvium continuum, in morphology Smalley also presents an equivalent concept named “attack-effluvium continuum” that describes the range of possibilities offered by the rate and compression of the iteration between attack-impulses. Figure 2.2 depicts all possibilities within the attack-effluvium continuum. While the first two categories describe an iterative behavior with different time scales, the last two categories can be seen as a description of the sound’s granularity—of which the ultimate stage (effluvial) is perceivable as a sustained sound. 2.3.2 - Lasse Thoresen and the Aural Sonology Project The Aural Sonology Project (ASP) is a research program initiated in the 1970s at the Norwegian Academy of Music in Oslo by Lasse Thoresen with the assistance of Andreas Hedman and Olav Anton Thommessen. A major contribution of this ongoing project is an analytical framework for music for which no score is available, or music in which no simple one-to-one correspondence between score and the aural phenomenon exists (Thoresen, 2007a).
! ! 34! The project claims two main influences: the first is Sonology as taught at the Utrecht Institute of Sonology in The Netherlands, and the second, the typo-morphological point of view articulated by Pierre Schaeffer in TOM. ASP draws its fundamental principles on the primacy of the listening experience through reduced listening, a concept borrowed from Schaeffer (1966). However, ASP extends Schaeffer’s theory towards “a pragmatic use of selected structuralist techniques” (Thoresen, 2007a). The musical object is apprehended not only as an objective fact but is partly formed by the listener’s intentions. The ASP developed an analytical approach to sound-based compositions with the following three levels: (1) sound objects, (2) elementary patterns, and (3) patterns of patterns. The first two levels—sound objects and elementary patterns—encompass constituent units or simultaneous layers of the sound continuum, and the third level characterizes the functional relationships between the several units (Thoresen 2007b). The following parallels can be established between Thoresen’s concepts and the time scales defined by Roads (2001) and defined earlier: (1) sound objects is a common concept by both taxonomies; (2) elementary patterns and the meso time scale; and (3) patterns of patterns with the macro time scale. One of the most valuable contributions of the project for the scope of this dissertation is the redefinition of Schaeffer’s typo-morphology into a terminology suitable for describing the musical phenomenon in empirical terms and adapted to practical analysis. It relies on philosophical jargon, everyday language, and terminology from musicology and acoustics, employing terms that are not coined within a consistent phenomenological point of view. Bellow, I present the morphology proposed by Thoresen that is divided in four main criteria, each of which further subdivide into smaller classes: (1) sound spectrum, (2) dynamic profile, (3) gait, and (4) granularity. Sound spectrum is characterized by spectral width and spectral brightness. The width of the spectrum characterizes the magnitude of the spectral components of a signal according to a continuum of possibilities, whose boundaries range from pure tones to
! ! 35! white noise. Within these limits an endless set of possibilities can be described such as monophonic and polyphonic pitched sounds with harmonic or inharmonic spectrum, or even any further saturations of the sound spectrum from unrecognizably pitched structures to a fully saturated spectrum (white noise). Spectral brightness indicates the spectrum’s “center of mass” and has a strong connotation with the perception of the “color” of the sound; it is commonly described by adjectives such as “dark” or “bright.” As Thoresen (2007b) notes, spectral brightness is a well-known phenomenon in linguistics to discern and organize the vowels and consonants, and in the music domain is a crucial feature for distinguishing between several traditional musical instruments, for instance. Dynamic profile “expresses the energy articulation of a sound object” (Thoresen 2007b). Thoresen proposes the following seven profiles, based on Schaeffer’s typology, to characterize the dynamic shape of sound objects: (1) no dynamic profile (dynamique nulle); (2) weak dynamic profile (dynamique faible); (3) formed dynamic profile (dynamique formée); (4) impulse-like dynamic profile (dynamique-impulsion); (5) cyclic dynamic profile (cyclique); (6) vacillating dynamic profile (rëiteré); (7) accumulation-like dynamic profile (accumulé). He (2007b) further characterizes the dynamic profiles by providing a typological description of two distinct phases of each profile: the onset and the termination. There are seven onset types and six termination types, which will not be detailed here, because the extrapolation of a sound typology is not relevant for the scope of this dissertation. The onset phase can be even further characterized by adding an indication of the spectral brightness of the opening transient.7 Gait is Thoresen’s attempted translation of the French word allure, one of the criteria of Schaeffer’s typo-morphology. Gait is closely related to the idea of vibrato, and defined by Thoresen (2007b) as “the undulating movement or characteristic fluctuation that often can be found in the sustained part of sound objects” (p. 139). Thoresen further divides the gait criterion in three categories according to nature of the undulation, which can be !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 7 For a detailed description of typology of the onsets and terminations please refer to Thoresen (2007b).
! ! 36! traced in the pitch, dynamic, or spectrum dimensions. Each category can be addressed by their nature, namely, pitch gait, dynamic gait, and spectral gait. Gait can be also characterized according to the degree of the undulation’s deviation from its mean value (small, moderate, and large) and the pulse velocity of the undulation (slow, middle, and fast). Granularity describes the microstructure of a sound object, that is, the perceptual irregularities. It is analogous to the abrasiveness one can feel when touching a piece of cloth or material, or the visible granularity of a photograph. Granularity is intrinsically related to the notion of iteration, and it is hard to differentiate the two concepts. “Generally, grains are a micro feature of the object in question, whereas iterations are of a coarser kind; thus grains would tend to be smaller, quicker, and be inseparable from the main body of the sound” (Thoresen, 2007b). Thoresen distinguishes nine types of sound objects’ granularity, which result from the combination of two characteristics of the grains: coarseness (small, moderate, and large) and velocity (slow, middle, and fast). 2.4 – Critical Review and Summary This chapter reviewed three major analytical theories by Pierre Schaeffer, Denis Smalley and Lasse Thoresen for sound-based compositions. Special attention is given to the typological and morphological criteria to describe sound objects because they support a major contribution of this dissertation, which is a computational description scheme for sound objects. Smalley’s and Thoresen’s theories are rooted in the seminal work of Pierre Schaeffer, whose typo-morphology has considerably inspired several approaches in music analysis and composition—like Smalley’s spectralmorphology and Thoresen’s ASP. The three aforementioned theories—typo-morphology, spectralmorphology, and aural sonology—were developed for the analysis of electroacoustic music. However, as Smalley (1986) points out in regards to his own theories, they are easily extensible to other music
! ! 37! genres. A GMR researcher named François Delalande asserts a position that may help us further understand the application of such theories to the analysis of music. While referring to Schaeffer’s solfeggio, Delalande (1998) notes that: The morphological analysis of electroacoustic music (based on a resolution into sound objects) is a ‘syllabic’ analysis, which does not provide the means of highlighting pertinent configurations either poietically (a ‘trace’ of compositional strategies) or aesthesically (contributing to explaining the behaviours and representations of listeners). Thus, we do not consider a morphological analysis to be a music analysis. (p. 20) As descriptive properties, the information provided by the analytical theories should not be treated as ends unto themselves, but rather as intermediary characteristics of sound objects. Moreover, the theories, and particularly the criteria for sound description, are broad enough to not be restricted to any music genre or style. In music practice, the application of Schaeffer’s typo-morphology has been nearly inexistent. Among the existent theories, Denis Smalley’s spectromorphology has received the most attention from the music community. Lasse Thoresen has taken the inaccessibility of Schaeffer’s work and presented a simpler, yet systematic, model that synthesizes Schaeffer’s major classes and enhanced applicability, providing a symbolic notation for each descriptor. The following paragraphs provide a critical review of the description schemes of the three aforementioned authors and establish a comparison between criteria used by them. A comparison between the top-level criteria of the three sound-based theories is provided in Table 2.1. In the following paragraphs, I adopt Schaeffer’s typo-morphology as the basis for the discussion.
! ! 38! SOUND MATTER SOUND SHAPE/ FORM VARIATION CRITERIA Pierre Schaeffer Mass Harmonictimbre Grain Pace (alure) Dynamic criteria Melodic profile Mass profile Denis Smalley Spectral typology Morphology (attack-) effluvium continuum Morphology Lasse Thoresen Sound spectrum Granularity Gait Dynamic profile Spectral width Spectral brightness Table 2.1 - Comparison between criteria of music perception of three representative sound-based theories by Pierre Schaeffer, Denis Smalley, and Lasse Thoresen. Schaeffer’s criterion of mass is present in both description schemes by Smalley and Thoresen, under the designations spectral typology and sound spectrum (more precisely spectral width), respectively. The three criteria are very similar, in particular the categorization of sound according to discrete types, whose limits are pure tones and fullsaturated spectra. Alongside the description of the sound objects’ mass by types of sounds, Smalley also adopts a description of the sound objects’ mass in a continuum of possibilities, whose limits are note and noise. Harmonic timbre is probably the most ambiguous criterion presented in Schaeffer’s morphology. Its definition is very vague and its close relation with the mass criterion is somehow misleading. Smalley avoids this criterion altogether and Thoresen presents a sound descriptor that clearly belongs to the harmonic timbre criterion within the mass criteria (sound spectrum according to Thoresen’s terminology). All theories examine the grain or granularity of sound objects. However, while Schaeffer and Thoresen consider it as criterion on its own, Smalley describes this dimension within the morphology criterion, particularly in the attack-effluvium
! ! 39! continuum. Half of the interval of the attack-effluvium continuum, between grain and effluvial states, can be seen as a description of the granularity of the sound. Another ambiguous concept presented in Schaeffer’s TOM is the notion of pace (allure). Similarly to harmonic timbre, Smalley avoids this criterion. Thoresen adopts the criterion and enlightens its definition by providing simpler, yet reliable categories for describing both the nature (pitch, dynamic, and spectral), and the quality of the phenomenon (velocity and amplitude of the undulation). Still, I find Thoresen’s definition of pace unsystematic and inconsistent, namely having in mind its algorithmic implementation, since it does not offer a concise description of the limits of the criteria. The dynamic criterion is transversal to the three frameworks, even if some nuances may distinguish them. All authors give priority to the description of the amplitude attack. Schaeffer and Smalley focus on the description of the overall stability of the sound objects’ dynamic, as well as the type of attack. Thoresen further explores some harmonic timbre characteristics of the attack phase. A final comment should be addressed to the simplifying approaches of the frameworks that follow and acknowledge Schaeffer’s theory. Even though it is seminal to understand the roots of Smalley’s and Thoresen’s theories (i.e., Pierre Schaeffer’s TOM), their contributions provide a much better adapted framework for the ultimate goal of this discussion, which is the formulation of a computational scheme for the description of sound objects.
! ! 40! Chapter 3 Computational Segmentation and Description of Sound Objects The current chapter aims at presenting strategies for segmenting and describing sound objects by computational means. It starts by providing an overview of techniques and tools used in MIR for the computational description of audio signals (§ 3.2), and ends by proposing algorithmic strategies for segmenting an audio stream into sound objects (§ 3.3) along with a musician-friendly description scheme that intends to characterize sound objects according to perceptual criteria (§ 3.4). The description scheme is particularly adapted to musical imperatives and targets a musicians’ audience by relying on the interaction between MIR and musicological literature. 3.1 – Introduction The ever-increasing amount of digital audio made available through public and private databases has demanded a deeper understanding of audio signals, in particular the
! ! 41! formulation of algorithms that can automatically extract information from audio data. Content-based audio processing is a recent technology designed to address the problem of sound indexing—offering new functionality for browsing, interacting, rendering, personalizing and editing musical material—by automating the task of manually annotating large sound databases. Most content-based audio processing research focuses on the recognition of sound sources (Martin, 1999; Eronen, 2001; Herrera, Dehamel, & Gouyon, 2003; Wold, Blum, Keislar, & Wheatom, 1996; Misdariis, Smith, Pressnitzer, Susini, & McAdams, 1998), music classification (Lu, Jiang, & Zhang, 2001; Tzanetakis & Cook, 2001; Ellis, Whitman, Berenzweig, & Lawrence, 2002), and music recommendation (Cano et al., 2005). Research in, and attention to, this field increased significantly when the Moving Picture Experts Group (MPEG), started working around 1996 on MPEG-7, a standard for describing multimedia content. Unlike their previous standards, which were mostly codecs for multimedia content, MPEG-7 targeted the creation of standardized descriptions for multimedia data, along with ways for structuring them (Herrera, Serra, & Peeters, 1999; Kim, Moreau, & Sikora, 2005). The primary purpose of MPEG-7 is to easily allow users or agents to search, identify, filter, and browse audiovisual content. MEPG-7 represents audio signals using audio descriptors—a research topic that has captured a lot of attention from the MIR community and consequently extended how computers manage audio. 3.2 - Audio Descriptors A critical feature of systems that deal with content-based audio processing, at any level, is the selection of audio data representations. The output quality of these systems is commonly dependent on such representations. The most common approach to represent audio in such systems is the adoption of audio descriptors, which measure properties of audio signal content and wrap audio features to sets of values. For example, the
! ! 42! brightness of a sound can be extracted by the audio descriptor spectral centroid, which measures the center of mass of the time-domain representation of an audio signal and expresses the brightness of a sound in a single value. Despite the numerous developments in this area, even state-of-the-art technology cannot compare with the accuracy, fastness, and detail of human perception and cognition. The computation of audio descriptors involves the use of various and sometimes overlapping approaches. Not even in the context of the MPEG-7 is there a standard way of obtaining these descriptions, or a customary approach on how to use them (Herrera et. al, 1999). Some of the most common techniques for extracting relevant features from audio data are through signal processing, computational auditory scene analysis, and statistics. (Herrera et. al, 1999). Despite the idiosyncrasies of the various audio description approaches, there are common taxonomies applied. Descriptors can be classified according to the representation of their output as well as their level of abstraction. According to Schwarz (2000), audio descriptors can be organized into three different classes: (1) categorical (class membership); (2) static (a single value); and (3) dynamic (temporal evolution). Specific to the level of abstraction, audio descriptors can also be organized according to the following three categories: (1) lowlevel, (2) mid-level, and (3) high-level. Low-level descriptors are computed from the digitized audio data by simple means and with very little computational effort in a straight or derivative fashion. Literature in signal processing and speech processing documents an enormous amount of different low-level features that can be computed from the audio signal representation, either on the time domain (e.g. amplitude, zero-crossing rate, and autocorrelation coefficients), or on the frequency domain (e.g. spectral centroid, spectral skewness, and spectral flatness) (Schwarz, 2000). Most low-level descriptors make little sense to humans, especially if one does not master statistical analysis and signal processing techniques, because the
! ! 49! Dixon (2005) point that the computation of short-term timing deviations is particularly relevant when attempting to find the location of beats. I had to implement a new algorithm for offline audio beat tracking in earGram because there are no available tools in Pure Data (earGram’s programming environment) to compute such task. Initially, my algorithm infers the tempo (beats per minute) of audio data stored in a buffer by finding the highest value of the accumulated spectral flux autocorrelation function. Then, in order to find the beat location, my algorithm starts by selecting the ten highest peak values of the spectral flux function (i.e., the ten onsets with higher growth values), and, relying on my hypothesis that one of these ten onsets corresponds to a beat location, the algorithm inspects for each selected onset the location of the beats according to the induced tempo. The computation of the beat locations allows short-term timing deviations, only if a local maximum is found within 2048 tolerance samples from the predicted location. For each of the ten onsets a score is computed by accumulating the spectral flux values from each prediction. Finally, the beat locations with the highest score are reported. After the segmentation of user-assigned audio tracks into sound objects, earGram extracts meaningful information from the sound objects’ audio signal representations and provides feature vectors that exposes their most prominent characteristics. The audio descriptors used to extract features of the audio will be detailed in the remaining sections of this chapter. 3.4 - A Musician-Friendly Audio Description Scheme In the creation of the description scheme that I will detail in this section, I relied on eight premises (formulated before its creation) to guide, unify, and regulate the set of perceptual criteria devised. In order to clarify the guidelines that assisted the creation of the description scheme utilized in earGram, the following premises are presented to the
! ! 50! reader. Some of the guiding principles of the description scheme were particularly devised to convey its primary use, the characterization of audio units of a CSS system (earGram). However, even if the scheme addresses idiosyncratic features of CSS, its application context is not restricted to this synthesis technique. The scheme encompasses dimensions that can be easily adapted to application contexts that require sound descriptions regardless of the relation between the sonic phenomenon and its source. Premises one to five address general considerations of the scheme, and premises six to eight address the idiosyncratic aspects of CSS. 1) The applied terminology in the scheme should rely on concepts from music theory and practice, in order to offer a more user-friendly experience for people with a music education background. 2) It should promote musical activity, specifically by providing representations of audio signals that can be easily manipulated in CAAC strategies. 3) It must rely solely on the abstract perceptual characteristics of sound—the morphology of sounds—disregarding their source, means of production, or stylistic features. 4) The descriptors’ computation should be definable by a mathematical function. 5) It should consider the emergence of higher-level descriptions of audio signals by associating or manipulating the basic criteria proposed in the scheme. 6) It should cover a continuum of possibilities and avoid the lattice-based organization of sound units (Wishart, 1994). Every criterion should be defined in a linear continuum with limited typological categories of sounds. This feature is appropriated from Smalley’s spectromorphology, namely its pitchand attackeffluvium continuums. 7) All descriptors must have the same range.
! ! 51! 8) The descriptions should be invariable to the units’ duration. In other words, the descriptions should allow meaningful comparisons between units of different durations within the same time scale. Relying on the eight premises listed above, I started to devise the top-level organization of the description scheme, which relies on two concepts borrowed from Schaeffer: matter and form. While the criteria related to matter describe the units’ sound spectrum as a static phenomenon, the form criteria expose the temporal evolution of the matter. The matter criteria express features of the audio in numeric values in a linear continuum interval, whose limits correspond to typological categories; the form criteria are expressed as vectors. In other words, the matter criteria represent each sound object with a numerical value, which is meaningful in relation to a finite space whose limits represent particular types of sounds. The form criteria follow the same approach but provide a contour of the audio features’ evolution. For example, noisiness, a criterion of matter, describes sound objects in relation to two typological limits (pure tone and white noise), and within these limits, sound objects are defined by a numerical value according to its characteristics. Sound typologies (as defined by Schaeffer) are only used here to define the limits of the interval. The dynamic profile, in turn, exposes the evolution of the amplitude of a sound object. Matter is further divided in two other categories: main and complementary. While the criteria under the main category provide meaningful descriptions for the totality of sounds that are audible to humans, the criteria under the complementary category provides meaningful results for limited types of sounds. For example, pitch—a complementary criterion of mass—only provides meaningful results for pitched sounds, thus excluding all sounds that do not fall in this category.
! ! 52! MATTER FORM MAIN COMPLEMENTARY Mass Noisiness Pitch Fundamental bass Spectral variability Harmonic Timbre Brightness Width Sensory dissonance Harmonic pitch class profile Dynamics Loudness Dynamic profile Table 3.2 - Description scheme used to characterize the audio content of sound objects in earGram. In choosing the descriptors that constitute the scheme, I relied on three musicological theories presented earlier—Schaeffer’s typo-morphology, Smalley’s spectromorphology, and Thoresen’s aural sonology—but I did not fully incorporate them into the scheme because of simplicity, usability, and/or technical reasons. Instead, I selected the ones that are more adapted to the practice of music composition, and whose technical implementations are feasible. A major concern behind the description scheme was the use of terminology from music theory and practice. Therefore, without disregarding the use of concise concepts, the terms used in the description scheme attempt to facilitate the usability for musicians with a traditional Western music education. While the conceptual basis of the scheme is entirely mine, the computation of each descriptor relies on algorithms from others—in particular William Brent’s timbreID library (Brent, 2009)—to extract low-level audio features from the audio. I chose Brent’s library for its robustness, efficiency, and ability to work in both real-time and non-real time. The detailed conceptual basis and technical implementation of each description in my scheme follows.
! ! 53! 3.4.1 - Criteria of Mass The mass criteria examine the spectral distribution of a sound object in order to characterize the organization of its components. It not only attempts to detect spectral patterns (e.g. pitch, fundamental bass) but also to provide general consideration of the spectral distribution (e.g. noisiness). The criteria of mass encompass four descriptors: (1) noisiness, (2) pitch, (3) fundamental bass, and (4) noisiness profile. The first is a main descriptor of matter, the second and third are ancillary descriptors of matter, and the last descriptor falls into the form category. 3.4.1.1 - Noisiness The noisiness descriptor estimates the amount of noisy components in the signal as opposed to pitched components. The measure of noisiness is present in all theories described in Chapter 2. However, each theoretician adopts a different term for this descriptor: Pierre Schaeffer designates it as mass; Denis Smalley uses spectral typology; and Lasse Thoresen names it spectral width. I adopted the term noisiness in the detailed scheme, because not only is it an easily understandable concept by both experts and nonexperts, but also because it is related to the algorithmic nature of the descriptor (detailed below). Another feature of the noisiness descriptor inspired by Smalley’s musicological theory is the adoption of a linear continuum to characterize the sound objects. The noisiness descriptor is calculated as a weighted sum of the following four low-level descriptors: (1) spectral flatness, (2) spectral kurtosis, (3) spectral irregularity, and (4) tonalness. Spectral flatness is the most significant descriptor, while spectral kurtosis, spectral irregularity, and tonalness are primarily useful to provide a better distinction between pitched sounds and noisy sounds. Spectral flatness is a very robust indicator of the noisy components of a signal, and
! ! 54! provides reliable descriptions of all sounds. However, it poses a major and pertinent problem: its characterization of pitched sounds is extremely poor. In other words, spectral flatness has a very good resolution for noisy-like signals; however, it is quite crude in relation to pitched sounds. In some experiments I carried out with a corpus of heterogeneous sounds, I immediately noticed a discrepancy between the interval of pitched sounds, which fall roughly in the interval ]0, 0.1[, and the interval that comprises noisy sounds, which inhabit the rest of the scale. As mentioned, in order to reduce this problem, I merged the results from spectral flatness with three other descriptors—(1) spectral kurtosis, (2) spectral irregularity, and (3) tonalness. A brief definition of each of the aforementioned descriptors, along with their contribution to enhance the representation of pitched sounds in the noisiness criterion follows. Spectral kurtosis gives a measure of the flatness or “peakedness” of the spectral distribution around its mean value (Peeters, 2004). The kurtosis of a single sinusoid will be extremely high, while that of noise will be extremely low. Spectral kurtosis is particularly good at distinguishing between pitched sounds that range from pure tones to heavy frequency modulations. Spectral irregularity enhances the difference between jagged and smooth spectra by looking at the spectrum from low to high frequencies and denoting how each bin compares to its immediate neighbors. Spectral irregularity has two common definitions: one by Jensen (1999) and other by Krimphoff, McAdams, and Wimsberg (1994). For practical reasons, I use Jensen’s measure since it defines the irregularity of a spectrum by values between zero and one, avoiding further processing to convey the same numeric interval used in the other descriptors. For jagged spectra (e.g. tone with harmonic spectra), irregularity will be high, and for smooth contoured spectra (e.g. filtered noise) it will be low. Spectral irregularity enhances the distinction between sounds from tones with harmonic or inharmonic spectra to spectral distributions formed of several “bands” (e.g. bell sounds) to spectral distributions formed of an array of sounds, which is non-locatable
! ! 55! in pitch (e.g. sea sounds). Tonalness measures the “perceptual clarity of the pitch or pitches evoked by a sonority” (Parncutt & Strasburger, 1994, p. 93). Sounds with high tonalness values evoke a clear perception of pitch. I use the tonalness measure by Parncutt, which defines the (pure) tonalness as the quadratic sum of the spectral pitch weights (so that its maximum value is equal to one). The tonalness descriptor I use is a slightly altered version of code provided in the Dissonance Model Toolbox by Alexandre Porres (2011). The combination of descriptors detailed above enhances the quality of the noisiness criterion by providing a better definition of pitched sounds and a better distinction between pitched and noisy sounds. Still, I applied some additional processing to each descriptor individually to further enhance the balance between pitched and noisy sounds. Spectral flatness and spectral kurtosis were scaled by exponential and logarithmic functions, respectively. No post-processing was applied to the spectral irregularity and tonalness descriptors. The noisiness descriptor ranges between zero and one. Zero represents a full saturated (noisy) spectrum and one represents a pure sinusoidal without partials. Within these two extremes the descriptor covers the totality of audible sounds including instrumental, vocal, or environmental sounds. 3.4.1.2 - Pitch The name of the second descriptor of mass is self-explanatory; it reports the pitch or fundamental frequency of the units. Pitch is a secondary criterion of mass, since it only conveys meaningful results for pitched sounds, and thus may reduce the corpus to a smaller collection of units. This descriptor is not contemplated in any theory discussed in Chapter 2 because it is highly attached to the concept of musical note and does not provide meaningful descriptions for the totality of perceivable sounds. However, the pitch
! ! 56! descriptor is adopted here since it may constitute an extremely important element in the composition process when dealing with pitched audio signals. There are several robust algorithms to estimate the pitch or fundamental frequency of monophonic audio signals. State-of-the-art algorithms for polyphonic pitch detection are not very reliable. The Pure Data’s built-in object sigmund~ by Miller Puckette is the pitch detection algorithm used to compute the fundamental frequency of (monophonic) sounds. The output of the descriptor is twofold: (1) in MIDI note numbers and (2) pitch classes. I additionally scale the resulting values to the interval [0, 1] to convey the general range of all descriptors. 3.4.1.3 - Fundamental Bass The fundamental bass descriptor reports the probable fundamental frequency or chord root of a sonority. Similar to the pitch criterion, it is a secondary criterion of mass, because it may reduce the corpus to a smaller number of units. I utilize this descriptor as a strategy to overcome the limitation of the pitch descriptor when analyzing polyphonic audio signals. The fundamental bass is computed by an altered version of a Pure Data object from the Dissonance Model Toolbox by Alexandre Porres (2011). The fundamental bass corresponds to the highest value of the pitch salience profile of the spectrum. The pitch salience of a particular frequency is the probability of perceiving it or the clarity and strength of tone sensation (Porres, 2012). The fundamental bass is expressed in (1) MIDI note numbers and (2) pitch classes. The output of the descriptor is further scaled to the interval [0, 1].
! ! 57! 3.4.1.4 - Spectral Variability Spectral variability provides a measure of the amount of change in the spectrum of an audio signal. It is computed by the low-level audio descriptor spectral flux (Peeters, 2004), which compares adjacent frames by calculating the Euclidean distance between two non-normalized spectra. The use of non-normalized spectra not only accounts for spectral differences, but also denotes sudden amplitude changes. Spectral variability is a form descriptor since it describes a temporal phenomenon. The output of this descriptor is threefold: (1) a curve denoting the spectral variability of the unit, (2) basic statistical values that express characteristics of the curve (such as maximum and minimum values, mean, standard deviation and variance), and finally (3) a single value that expresses the overall spectral variability throughout the unit duration. The curve depicts the evolution of the spectrum at regular intervals of 1024 samples, and each analysis window encompasses 2048 samples. Relying on the computed curve some basic statistical properties are then extracted, such as minimum, maximum, mean, standard deviation, and variation. These statistical properties provide a characterization of the curve by a vector with a reduced dimensionality. A single value depicting the overall variability of the overall unit’s spectrum is computed in four steps: (1) dividing the units in two equal halves, (2) computing the spectrum of each half, (3) calculating the distance between the two spectral representations, (4) summing all values resulting from point 3. The output of the descriptor is further divided by the number of reported bins (resulting from the spectral difference computed in point 3 to scale the output to the interval [0, 1].
! ! 58! 3.4.2 - Criteria of Harmonic Timbre The three musicological theories presented earlier provide little guidance for the formulation of algorithmic strategies to describe the harmonic timbre content of a signal. Schaeffer’s criteria of harmonic timbre are very misleading and too inconsistent to be encoded algorithmically. Smalley (1986, 1997) does not provide a specific set of criteria for harmonic timbre; even if he considers this dimension while describing the mass of sound objects under spectral typology. Thoresen’s sound spectrum criteria, in particular the spectral brightness, are the most adapted to a computational definition of harmonic timbre. His criteria also points towards the possibility of including psychoacoustic models as harmonic timbre descriptors, which Schaeffer rejected because (in his opinion) the in vitro psychoacoustic experiments do not fully apprehend the multidimensionality qualities of the timbre (Chion, 1983). Still, Thoresen’s suggestion led me to further investigate psychoacoustic literature, most notably models that examine the sensory dissonance phenomenon, which provide a good description of spectra distributions. The main source for investigating possible usages of psychoacoustic models in my description scheme was Alexandre Porres’ PhD dissertation (Porres, 2012). Porres not only explores several creative applications of psychoacoustic models in signal processing, but also points out the underexplored possibilities of psychoacoustic dissonance models for the automatic description of audio signals’ content. This has not been subject to any study and could contribute significantly to applications such as CSS and alike (Porres, 2012). While concatenative sound synthesis and similar techniques are common and are at an interesting development stage, processes of the same order with the higher-level descriptors, such as the attributes of dissonance here exposed, have not been fully explored, with the exception of some computer-assisted composition works by Sean [Ferguson] (2000) and [Clarence] Barlow (1980). (Porres, 2012, p. 86) Therefore, I will use here the psychoacoustic dissonance models presented by Porres
! ! 65! Figure 3.1 - Dynamic profile of a sound object and the values extracted from the profile. 3.5 - Summary In this chapter, I proposed strategies for segmenting an audio continuum into sound objects using onset detection and beat tracking methods, along with a morphological scheme for describing their most prominent perceptual characteristics. The description scheme results from the interaction between musicological and psychoacoustic theories and MIR research, in particular the literature related with audio descriptors. In addition, the terminology adopted relies on empirical terms borrowed from musical theory and practice in order to increase usability. The description scheme is divided in three major perceptual criteria—mass, harmonic timbre, and dynamics—that unfold in a set of ten descriptors. The descriptors may also be categorized according to their output representation according to two concepts borrowed from Schaeffer: matter and form. Audio descriptors under matter are defined in a linear continuum and adopt the same range: the interval [0, 1]. While the interval limits of each descriptor corresponds to specific types of sounds, there is no strict one-to-one correspondence between regions of the interval and specific sound typologies. To achieve such uniformity in the descriptors’ range some scaling is applied to the descriptors output. However, the scaling factor is not relative to the maximum, minimum, or mean values of the descriptors functions. Instead, the scaling I use in some descriptors is relative to fixed values determined by specific
! ! 66! perceptual characteristics of the features in question. Therefore, not only is a “normalized” range guaranteed, but also meaningful information concerning the audio signal’s content. Despite the recent tendency to adopt large numbers of audio features in content-based audio processing systems in order to enhance their results, earGram purposefully encompasses a very limited number of descriptors. The adoption of a limited number of descriptors relies on recent studies (Mitrovic, Zeppelzauer, & Eidenberger, 2006; Peeters, Giordano, Susini, Misdariis, & McAdams, 2011), which argue that the information expressed by the totality of audio descriptors developed so far expose a high degree of redundancy. Therefore, I can conclude that my description scheme provides a rich representation of audio signals, since it covers the most significant classes of audio descriptors.
! ! 67! Chapter 4 Musical Patterns In this chapter, I extend the analytical tools presented in the previous chapter by inspecting higher layers of musical structure. In other words, while the analytical strategies presented in Chapter 3 departed from the most basic representation of audio signals (i.e., the sample), the algorithms detailed here rely on sound objects’ descriptions in order to extrapolate representations of the higher temporal scales of musical structure. Ultimately, the aim of this chapter is threefold: (1) to create models of the temporal dynamics of the music (§ 4.1); (2) to discuss and propose strategies to compare and group sound objects (§§ 4.2-4.4); and (3) to provide mid-level descriptions of the corpus (§ 4.5). The following sections will introduce the reader to algorithms that may ultimately answer the following questions: how common is a particular characteristic throughout the audio source? Which features are more relevant? Which units recur, and in which order? Are there any outliers? How similar are the units in the corpus? Are any representative groups of units within the corpus? How are they organized in time? Does the original
! ! 68! temporal sequence of the units denote the use of repeating patterns? 4.1 - Probabilistic Models of Musical Structure Given the temporal nature of music, the ability to represent the dynamics of musical structure is at the core of any analysis-synthesis system. Accordingly, earGram adopts strategies to model transitions between sound objects in order to map the dynamics of musical structure. In other words, earGram uses algorithms to learn and encode the temporal evolution of particular audio features of musical structure. In order to do that, I adopt the state-space models named n-grams, which encode sequences of discrete events using statistical properties (Jurafsky, Martin, Kehler, Vander Linden, & Ward, 2000). Ngrams are amongst the most used strategies to encode musical structures computationally (Cont, 2008). They provide a representation of time-indexed sequence of graphs (nodes and edges) where each node refers to a state of the system over time. When dealing with musical elements, the states may represent musical events of different time scales, such as sound objects (e.g. notes, chords, silences), meso structures (e.g. rhythmic patterns, melodic arcs), and macro structures (e.g. sections). I adopted n–grams because they embed a property that is seminal for my framework: they provide the basis for a Markov chain algorithm, which is an algorithm utilized in earGram for generating musical sequences. While the creation of the n-gram representations will be examined in the following section, its application for the generation of musical structures will only be addressed in the second part of this dissertation. The models that will be presented not only learn and encode the dynamics of three elements of the audio source’s structure—noisiness, timbre, and harmony—but also “artificially” establish optimal transitions and overlaps between sound objects based on psychoacoustic theory principles. It is important to highlight that the modeling strategies
! ! 69! implemented in earGram only encode singular features of the original audio data because the goal is not to provide a comprehensive representation of all dimensions of musical structure and their inter-relationships, as it is attempted in many style imitation approaches to music (cf. Cope, 1996, 2001). Instead, I adopt models that provide a basis to assist and ease the process of music creation through sampling techniques by automating some of the parameters of a composition. A detailed explanation of the n-grams creation will be presented in the following sections. Section 4.1.1 details models that learn and encode particular elements retrieved from the structure of the audio source(s), and sections 4.1.2 and 4.1.3 detail psychoacoustic-based models for transitioning and superimposing audio objects. 4.1.1 - Modeling Elements of Musical Structure EarGram creates n-grams that encode the temporal dynamics of the following three elements of the audio source(s) structure: noisiness, timbre, and harmony. My software starts by learning the probability of transitioning between discrete elements of musical structure for each of the aforementioned characteristics, and, consequently, stores all probabilities in a matrix. The modeled events need to be discrete features that are extracted from the sound objects, and the temporal dimension of the models encodes the original sequence of units. The elaboration of transition probability tables is fairly straightforward to compute. However, when dealing with audio signals, to obtain a finite-state space for each modeled element may pose some problems. If the states were directly observable, as in symbolic music representations, no pre-processing would be necessary. However, this is hardly the case when dealing with audio data. Thus, I applied a different strategy in order to create a finite-state space for each of the three musical characteristics. The noisiness descriptor characterizes the units in a linear continuum, whose limits are
! ! 70! zero and one. Given the need to have a finite number of classes to create a transition probability matrix, the range of the descriptor was arbitrarily divided in ten equal parts. Each class is represented by a numerical value from zero to nine, sequentially distributed in the interval from the lower to the upper limits. Timbre is expressed by a single integer that represents the three highest bark spectral peaks. The algorithm to find the compound value is shown in Table 4.1. Finally, the pitch class of the fundamental bass represents the audio units’ pitch/harmonic content. After obtaining the finite-state space for each characteristic, I computed the creation of a transition probability matrix in the following steps: (1) accumulating the number of observations from the n previous states to the following state and, after the totality of the sequence is considered, (2) divide each element of the matrix by the total number of observations in its row. The resulting matrix expresses the probabilities of transitioning between all events. Operation Number Operation Description Example 1 Sort in an ascending order the three peaks with highest magnitude 5, 14, and 15 2 Convert the integers to binary 101, 1110, 1111 3 Shift the 2nd and 3rd numbers by 5 and 10 cases to the left10 101, 111000000, 11110000000000 4 Convert the result to decimal 5, 448, and 15360 5 Sum the resulting values 15813 Table 4.1 - Flowchart of the algorithm that reduces the Bark spectrum representation to a single value. A final note should be addressed to the order of n-grams used. By default, it is adopted a third order n-gram, that is, the algorithm encodes the probability of transitioning between the three last events and the next one. However, the user can easily change this !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 10 This bitwise operation allows the codification of the three values in non-overlapping ranges, which makes the sum of their decimal representation a unique value for any possible combinations of three values that the algorithm can adopt (0-23).
! ! 71! parameter. If the corpus has a considerable number of units, increasing the order of the ngram may enhance the resemblance of the generated output to the original audio. The inverse procedure should be applied to corpora with a very small number of units. 4.1.2 - Establishing Musical Progressions Based on Pitch Commonality All models exposed in the previous section rely on the structure of the original sequence of the units to formulate the probability of transitioning between musical events. In this section, I present a different strategy to determine the probabilities of transitioning between sound objects, which does not rely on the structure of the audio source(s). Instead of modeling a particular characteristic of the audio source(s) by learning its internal organization, the method presented here defines the probabilities “artificially” by applying a psychoacoustic dissonance model, in particular by computing the pitch commonality between all units. Pitch commonality provides a link between psychoacoustics and music theory and it is defined as the degree to which two sequential sounds have pitches in common. It measures the “pleasantness”11 of the transition between two sounds, and can be seen as an oversimplification of harmonic relationships (Porres, 2011). For instance, the pitch commonality of musical intervals is quite pronounced for perfect octaves, less pronounced for perfect fifths and fourths, and more or less negligible for any other intervals. The computation of pitch commonality depends on the amount of overlapping pitch saliences between two sounds. The pitch salience is defined as the probability of consciously perceiving (or noticing) a given pitch (please refer to Parncutt (1989) for a detailed description of its computation). Pitch commonality is calculated by the Pearson !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 11 The concept of pleasantness is understood here as sounds that express a low degree of sensory dissonance (see § 3.4.2.3 for a definition of sensory dissonance).
! ! 72! correlation coefficient12 of the pitch salience profiles across the frequency spectrum of two sonorities (Porres, 2011). It is equal to one in the case of equal spectra and hypothetically minus one for perfect complementary sonorities. For a complete mathematical description of the model please refer to Parncutt (1989) and Parncutt and Strasburger (1994). Initially, earGram creates a matrix that stores the results of the pitch commonality calculation between all pairs of units in the corpus. Consequently, all elements of the matrix are converted into probabilities. The last step is done by dividing the absolute value of each element in the matrix by the sum of all absolute values in its respective row. The resulting matrix is the transition probability table of a first-order Markov chain algorithm. 4.1.3 – Vertical Aggregates of Sound Objects Based on Sensory Dissonance CSS deals primarily with the horizontal dimension of the music, that is, the generation of musical sequences. However, it is current practice to expand the technique to address the synthesis of overlapping units (Schwarz, 2012; Schwarz & Hackbarth, 2012). Despite the popularity of this new approach, the resulting sound quality of the vertical superposition of audio units has been overlooked. So far, there is no consistent method to define the sonic quality of target phrases that encompass vertical aggregates of audio units. The vertical dimension of music is related to the relationship between simultaneous events, or the sonic matter and its constituent components. According to Thoresen (2007b), the primary structural element of the vertical dimension in Western music is harmony. Timbre can be considered a secondary element. The description scheme !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 12 The Pearson correlation coefficient is often used to determine the relationship between two variables by measuring the linear correlation between them. It is calculated by the covariance of the two variables divided by the product of their standard deviations. The Pearson correlation coefficient may adopt values between minus one and one. Zero expresses no association between the two variables, minus one indicates total negative correlation, and one indicates total positive correlation (Taylor, 1990).
! ! 73! presented earlier (§ 3.4) allows the characterization of the vertical dimension of the sound objects, such as the width or degree of sensory dissonance of sound. However, from a creative standpoint, the use of any of these descriptors is confined to the horizontal organization of music. The sensory dissonance descriptor does not express much about the sonic result of simultaneous layers of audio units. I adopted the sensory dissonance descriptor in order to characterize and organize vertical aggregates of audio units, but in a different manner as used in the description scheme. To measure the “pleasantness” of two simultaneous units, I computed the degree of sensory dissonance between the combination of the spectral representations of the two units (see § 3.4.2.3 for a detailed explanation of the computation of sensory dissonance). A matrix stores the results of the sensory dissonance measures between all pairs of units in the corpus (see Figure 4.1). The resulting matrix will be utilized later to guide the generation of vertical aggregates in earGram. Unit number 1 2 3 … 1 1 0.1 0.2 2 0.1 1 0.5 3 0.2 0.5 1 … Figure 4.1 – Example of a matrix that exposes the sensory dissonance between all pairs of sound objects in the corpus. Above, I have detailed the creation of five n-grams that encode optimal transitions and the superposition of sound objects. The creation of the models relies on descriptions of sound objects, whose computation was presented in Chapter 3. The following sections will continue to examine how musical structure can be apprehended and/or extrapolated, but the focus will shift towards higher layers of musical structure. In order to provide
! ! 74! strategies that ultimately expose the higher layers of musical structure, I will first introduce how sound objects can be consistently compared. 4.2 – Audio Similarity Sounds can be compared to other sounds according to numerous properties. Tristan Jehan (2005) summarizes the criteria with which we can estimate the similarity between two songs to the following five categories: (1) editorial (title, artist, country), (2) cultural (genre, subjective qualifiers), (3) symbolic (melody, harmony, structure), (4) perceptual (energy, texture, beat), and (5) cognitive (experience, reference). A definite measure of similarity between two songs or audio samples does not exist (Jehan, 2005). Music similarity is an ambiguous task, not only because it is a very complex multi-dimensional problem, with varied subjective dimensions, but also because it is context-dependent. In other words, the evaluation of the similarity between audio signals is highly dependent on the application context and the user. A systematic computational model of music similarity poses even more problems, which can be roughly summarized in three topics: (1) the complexity of the task, (2) the subjectivity of criteria, and finally, (3) the difficulty of algorithmically considering and/or encoding application contexts and users’ preferences. While addressing computational models of similarity, it should be noted that I do not refer to any work that operates on metadata supplied by humans, or even high-level music representations, such as MIDI. Of interest here is the measurement of similarity between non-uniform audio units, generally ranging from a fraction of a second to a few seconds, which are represented by an audio feature vector. Ultimately, the aim of discussing audio similarity measures in the context of this dissertation is to provide a reliable method for comparing and depicting audio units according to their similarity. In order to compute the similarity between audio samples one usually calculates the
! ! 81! stored cluster from further consideration, and finally (5) repeat the operations from the first point with the reduced set of points until no more clusters can be formed. QT-clustering considers all possible clusters. The candidate clusters are generated with respect to every data point and tested in order of size against the quality criteria. Major advantages of this clustering technique are the detection of outliers that can be treated differently at runtime (for instance, excluded from the recombination) and the precise control over the similarity of grouped units. The major disadvantage of this clustering strategy is its heavy computational cost. 4.3.3 - Density-Based Clustering DBSCAN is a well-suited algorithm to discover clusters of arbitrary shapes in spatial databases. Clusters are defined according to two parameters: (1) the distance threshold or neighborhood proximity between data points and (2) their density (the minimum number of points within the radius of each unit to form a cluster). In other words, each unit in the cluster must not exceed an assigned distance threshold from another unit in the cluster and each cluster has to contain at least the specified minimum number of units within the distance threshold. Therefore, the formed clusters have a typical density of points, which is considerably higher than outside of the cluster. My implementation of the DBSCAN algorithm follows Ester, Kriegel, Sander, and Xu (1996). EarGram starts by inspecting an arbitrary unit that has not been visited by the algorithm. This unit’s neighborhood is examined, and if it includes enough points within the threshold distance, a cluster is defined. Otherwise, the point is labeled as “noise.” However, this unit might later be found in the neighborhood proximity of a different point. If a unit is found to be a dense part of a cluster, all of its neighborhood units (within the distance threshold) are also part of that cluster. This process continues until the density-connected cluster is found and repeated until all units have been visited.
! ! 82! Similarly to the QT-clustering algorithm, DBSCAN avoids defining a priori the number of clusters. However, the algorithm finds arbitrarily shaped clusters that are very different from the ones found by the QT-clustering. It can even find clusters surrounded by (but not connected to) a different cluster. DBSCAN helps define action zones or large groups of interconnected data points according to a proximity distance and density, which may sort and group sound objects by “scales.” The resulting clusters will encompass sound units that express some continuity, that is, each cluster exposes perceptual trajectories of particular audio features. Similarly to QT-clustering, DBSCAN detects outliers, which may help restrict the corpus to a more unified collection of units during performance. The audio similarity and clustering algorithms detailed in the last sections are better understood through visuals. Corpus visualizations were adopted in earGram not only to expose the hidden results concerning audio similarity and clustering, but also to provide tools in which one can navigate, explore, and interact with the corpus. 4.4 – Visualizations EarGram adopts two visualization strategies: 2D-plot and similarity matrix. Both allow the navigation, exploration, and interaction with the corpus, and also aim at depicting various (hidden) analytical stages of the system. Ultimately, the implemented visualization strategies reveal some intrinsic characteristics of the audio source(s), in particular its macrostructure by depicting the similarity between the sound objects that compose the corpus. The visualizations may assist in the decision-making processes during performance. In addition, they also allow interactive and guided explorations of the corpus. A detailed description of the two visualization strategies, along with their technical implementation, follows.
! ! 83! 4.4.1 – Sound-Space Sound-space offers a visual representation of the sound objects’ collection in a 2Dplot. It provides an intuitive visualization of the similarity between sound objects and is particularly suitable for browsing and exploring a corpus of audio units by navigating through its representation. 2D-plots are one of the most common visualizations adopted in CSS, and frequently allow physical interaction with the corpus of audio units (Martin, 2011; Schwarz, 2012; Schwarz & Hackbarth, 2012). Figure 4.3 - Visual representation of a corpus of audio units comprising a single audio source—4 by Aphex Twin—in a 2D-plot whose axes were assigned to the following descriptors: noisiness (x-axis) and spectral variability (y-axis). The units’ color is defined by sensory dissonance, loudness, and duration, by assigning each descriptor to the values of R, G, and B, respectively, and using an additive color model. The following sections detail two different approaches to constructing the sound-space visualization in earGram. The first assigns individual audio features to the axes of the 2Dplot, and the second adopts multiple audio descriptions in the representation, which can
! ! 84! be depicted in two dimensions with the help of multidimensional reduction algorithms. A common element to both approaches is the layer of information that is offered by the units’ color. The color of each unit in sound-space is defined by a list with three elements that correspond to the red, green, and blue values of an additive RGB color model. The R, G, and B values represent audio features from the available set of descriptors. 4.4.1.1 – Two-Dimensional Visualizations of the Corpus Using Binary Sets of Descriptors In sound-space, the use of different binary sets of audio features to depict a corpus of audio units provides valuable information about the corpus, in particular the similarity between its constituent units. In order to create such visualization of the corpus, the user must first assign single audio features to each axis of sound-space. Then, earGram collects the analyzed information of both selected features for all audio units in the corpus, and depicts the corpus according to those values. In other words, the coordinates of each audio unit (single dot) in sound-space are two values (x and y) that correspond to audio features of the units. Therefore, the sound-space visualization provides as many visualizations of the corpus as the number of possible combinations between all pairs of audio features provided by earGram’s description scheme. In addition to the depicted information, the sound-space visualization also allows the definition of target phrases to be synthesized with a high level of precision by navigating its representation. In fact, sound-space functions in a similar way as traditional acoustic instruments. Its response is predictable, direct, and controlled, because the audio units are represented in a scale whose limits are specific types of sound. If the same conditions are met, in particular the same pair of descriptors, the same performance gestures result in the same sonic response. An important distinction is the mutability of the instrument according to the feature space of the 2D-plot. Changing the feature space in sound-space can be seen as changing a preset in a synthesizer. The adoption of different sets of
! ! 85! descriptors to depict the corpus of audio units imposes significant changes in the sonic feedback. If one repeats the same trajectory in sound-space with different sets of descriptors the sonic feedback of the gestures can be regarded as variations. Each compound set of descriptors has its own identity. A two-dimensional representation of the sound objects according to the descriptors pitch and loudness is a clear example of how the sound-space visualization may emulate a piano-keyboard behavior. However, earGram was not designed to primarily emulate this behavior. Instead, its purpose focuses rather on the exploration of all aspects of sound outside the pitch-duration primacy. For example, the combinations of pitch or fundamental bass descriptors with any of the harmonic timbre descriptors (e.g. brightness, width, and sensory dissonance) provide an extended control over the harmonic quality of the pitch/chords. One thus may “modulate” the timbre of particular pitches. The use of noisiness, brightness, width, and sensory dissonance is particularly effective for visualizing a corpus of electronic-generated sound units. This is due to the fact that these types of sounds commonly expose a rich variety of colors (timbre). Finally, any combination of the following descriptors: noisiness, loudness, width, and spectral variability is particularly interesting to visualize and control the synthesis of environmental sounds. I should also remind the reader that the descriptors pitch and fundamental bass do not provide meaningful information to all types of audio units. For this reason they were presented as complementary descriptors (see Table 3.2). To conclude, I just would like to add that the creation of sound events/structures organized by parameters other than pitch, duration, and loudness are not very common in Western music and their use should undergo an experimental phase that goes beyond the temporal scope of this investigation.
! ! 86! 4.4.1.2 - Multidimensional Reduction of the Descriptor Space The most common way to describe and represent units in content-based audio processing applications is to include a large number of audio features in a multidimensional vector. In recent years, a large number of new descriptors have been presented and adopted despite the higher computational costs. We should also keep in mind that the incorporation of a larger number of descriptors does not always represent an improvement in the characterization of the audio signals. High-dimensional vectors are difficult to visualize and are not at all suitable for physical navigation, which is commonly performed in twoor three-dimensional spaces. Hence, the most common solution to this problem is either assigning a single descriptor to each of the axis of the two or three-dimensions representation (as discussed in the previous section), or employing dimensionality reduction algorithms to decrease the number of dimensions to twoor three-dimensions while retaining most of the information provided by the vectors. In this section, I propose a dimensionality reduction algorithm to decrease the number of dimensions of the feature vectors, and examine its implication in creative practices, such as musical composition. Multidimensional reduction techniques are commonly used in content-based audio processing applications. These techniques not only reduce computational costs associated with the matching process, but they also convey the visualization of the corpus in twoor three-dimensions. Hence, the topic could be examined in the current section with respect to the visualization of the corpus or in a different section where I address audio similarity computation. The decision to address multidimensional reduction techniques here is for a very simple reason: while in earGram the visualization of high-dimensional feature vectors in a 2D-plot is a very a pertinent problem, the description scheme does not encompass a very large number of descriptors and therefore does not pose tremendous problems in
! ! 87! terms of computational cost. I noted in the direct contact with a few composers that utilized earGram in their compositions that dimensionality reduction methods raised curiosity, even if the outcome of the algorithm is slightly misleading. The visual representations gathered after applying multidimensional reduction algorithms lack clarity. The axes of the plane are hardly related to any particular feature and what remains is a general and “blind” representation of similarity between the sound objects. The precise manipulation of individual dimensions of the sound matter is the most common approach in musical composition. The same level of accuracy is achieved if single features are assigned to each of the axis of a plane. However, dimensionality reduction methods might be helpful in cases where it is unclear which features the user wants to control, or when it appears that no pair of features will provide satisfactory results. The latter case might be true for applications using sounds of very different natures. Two of the most popular algorithms for dimensionality reduction are principal component analysis (PCA) (Shlens, 2005; Skočaj, Leonardis, & Bischof, 2007) and multidimensional scaling (MDS) (Mikula, 2008; Schwarz & Schnell, 2009). Both methods can be used to obtain smaller representations of high-dimensional feature spaces. EarGram adopts the algorithm star coordinates (Kandogan, 2000) for dimensionality reduction. The algorithm is substantially less known and applied than PCA or MDS, especially to address audio feature vectors. However, star coordinates offers two major advantages over the aforementioned algorithms: (1) the understandability of the axes after processing takes place, and (2) its suitability for both online and offline processing. A clear disadvantage of star coordinates is the need to explore the representation by weighing the variables and assigning different angles to the axes to find interesting patterns.
! ! 88! Figure 4.4 - Mapping of an eight-dimensional point to two dimensions. Axes are named as Cx, each dimension of the point as djx, and P is the final point position (Kandogan, 2000). (Copyright 2000 by Eser Kandogan. Reproduced with permission.) Star coordinates map a high-dimensional point linearly to two dimensions by summing the vectors resulting from the point coordinates arranged on a circle on a two-dimensional plane with equal (initial) angles between the axes with an origin at the center of the circle (see Figure 4.4 for a demonstration of the algorithm). 4.4.2 - Self-Similarity Matrix By depicting pairwise similarity between the same original sequence of sound objects assigned to both vertical and horizontal axes of a square matrix, it is possible to reveal patterns of the audio source(s) that ultimately expose the macro structure of the data. The graphical representation is called a self-similarity matrix. The technique was first introduced by Foote (1999) with the aim of visualizing musical structures. The method consists of building a square matrix where time runs from left to right, as well as from bottom to top, and the cells’ color denotes the similarity between audio units. The similarity between sound objects is computed by the Euclidian distance between their
! ! 89! representative feature vectors. The feature vectors may include a variable number of features from the available set of descriptors, as well as variable weights. The standard approach to audio similarity matrices consists of computing and depicting the similarity between short windows of fixed duration. In earGram, the compared units correspond to sound objects with non-uniform duration. I adopt similarity matrices in earGram mainly to guide the user through the selection of sub-spaces of the corpus that can be used differently during performance. Figure 4.5 – Visualizations of a corpus comprising a single audio track—4 by Aphex Twin—by a self-similarity matrix (left image) and a related visualization whose color of each cell results from two found clusters in the corpus (rightmost image). The middle image is a detail of the self-similarity matrix, which exposes with detail the color of each cell. EarGram provides a different visualization strategy of the corpus that follows the same principles behind similarity matrices. The major difference resides in the color of each matrix cell, which instead of resulting from the distance between feature vectors is attributed to the juxtaposition of the audio units’ colors (resulting from the clustering strategies). The visualization provides very similar information as the traditional approaches to similarity matrices and enhances its clarity in a similar fashion as reducing time time
! ! 90! the noisiness of a signal by a smoothing function. It is important to note that the matrix configurations are highly dependent on the features used to compute the similarity or the clusters. Therefore, in order to find interesting patterns the user may need to explore different collections of features or assign different weights to each descriptor. In addition, both visualizations can provide interesting feedback to the user in understanding how different audio features’ weights and constraints, and/or the use of different feature spaces, can alter the notion of similarity between sound objects and their consequent grouping. Now that I have addressed how sound objects can be consistently compared and grouped to expose characteristics of the higher layers of musical structure, I will conclude the current chapter by providing two descriptions—key and meter—of the corpus/audio source(s) that will be preponderant for some decisions during the composition phase. 4.5 - Mid-Level Description of the Corpus In this section, I will detail the adoption of two mid-level descriptors for meter and key induction. Unlike the audio descriptors proposed in Chapter 3, the two following descriptors do not target individual sound objects. Instead, they characterize collections of sound objects and may be applied as a strategy to constrain the corpus to sub-spaces of units, or simply to provide information that can be used at later stages of the system for the generation of new unit sequences. The conceptual and technical considerations that assisted the implementation of both descriptors follow. 4.5.1 - Meter Induction In music, meter refers to the hierarchical organization of time based on perceived temporal regularities (Lerdhal & Jakendorff, 1983). It consists of a periodic system of
! ! 97! PART II: COMPOSITION Any text is constructed as a mosaic of quotations; any text is the absorption and transformation of another. — The Kristeva Reader, Julia Kristeva (1986)
! ! 98! The aim of Part II is to explore CAAC strategies that automatically recombine audio units by manipulating descriptions of sound objects as well as to suggest methods for incorporating the generative algorithms in a composition workflow. Part II will adopt a similar, but inverse, structure as Part I. In other words, while the first part of this dissertation adopts a bottom-up strategy for analyzing audio data, the second part adopts a top-down approach to algorithmic composition. I will adapt well-known CAAC strategies attached to symbolic music representations to address audio signals and function as unit selection algorithms in CSS. My generative methods were implemented and tested in earGram and are able to build arbitrarily long structures in a way that the synthesized musical output reflects some of the elements that constitute the audio source(s). Yet, due to the particularities of my generative methods, the created music is new and different from the raw material that supports its creation— and any other existing music. EarGram demands little guidance from the user to achieve coherent musical results and it is suitable for a variety of music situations spanning from installations to concert music.
! ! 99! Chapter 5 Organizing Sound Using sound as raw material for a composition is a central concern in electroacoustic music. The simplest approach to compose with sounds in order to create a new composition is by manually manipulating and assembling pre-recorded audio samples. I embrace this method through the recombination of sound objects. However, the recombination process is semi or fully automated by organizing prominent features inferred from the sound objects. The following subsections provide an overview of the technical and conceptual background of the framework’s generative component proposed in this dissertation, in order to place it in a particular historical context and justify its pertinence. The chapter concludes by explaining the articulation between the two major modules of the framework: analysis and composition. More specifically, this chapter provides an historical perspective of sample-based synthesis techniques—sampling, micromontage, and granular synthesis—which contributed to the emergence of CSS. Next, I provide an overview of musical applications of CSS over the last decade. Then, I examine the technical aspects of the framework by asking how
! ! 100! they influence the practice of music composition. The following three compositional approaches will be addressed: (1) the use of sound structure, namely its timbral qualities as the primary material for structuring musical processes; (2) music as a consequence of pre-devised processes; and (3) the notion of “appropriation” as a musical concept. In addition, I will detail the contribution of each topic to earGram’s design, in particular how they influenced the articulation between the analysis and composition modules. 5.1 – From Sound to Music: Technical and Conceptual Considerations 5.1.1 - Sampling In electronic music, sampling (also known as audio collage) is the act of taking a portion of a particular recording and reusing it in a different piece. Apart from previous isolated experiments, musicians began exploring the technique in the late 1940s. The very first sampling experiments were carried almost exclusively in radio broadcast stations, because they had the necessary technology. The most prominent pioneers of sampling are the French composers Pierre Schaeffer and Pierre Henry; they began to explore experimental radiophonic techniques with the sound technology available in the 1940s at the French Radio in Paris—where the current GRM still resides (Palombini, 1993). The advent and widespread use of magnetic tape in the early 1950s opened new possibilities to sampling techniques, in particular the exploration of large amounts of audio samples. It is interesting to note that the use of a large corpus of sounds, a crucial feature of earGram, appealed to composers from the very first moment the technology allowed its manipulation. Karlheinz Stockhausen, John Cage, and Iannis Xenakis are three representative composers of the electronic music of this period. Stockhausen used in Étude des 1000 collants (1952), known simply as Étude, a corpus of millimeter-sized tape
! ! 101! pieces of pre-recorded hammered piano strings, transposed and cropped to their sustained part to assemble a previously devised score that defined a series of pitches, durations, dynamics, and timbres (Manion, 1992). John Cages’ Williams Mix (1951–1953), a composition for eight magnetic tapes, is another piece from this period that explores the idea of using a large pre-rearranged corpus of sounds as the basis of a composition. Williams Mix’s corpus comprised approximately 600 recordings organized in six categories: city sounds, country sounds, electronic sounds, manually produced sounds, wind sounds, and "small" sounds, which need to be amplified (Cage, 1962). Xenakis’ compositions Analogique A and B (1958-1959) and Bohor (1962) are also worth mention, not only for its exploration of a large corpus of short sound fragments but also for the assembling process, which was driven by stochastic principles (Di Scipio, 2005). From the mid 1960s until the 1990s, we witnessed a rapid proliferation of sampling techniques, mainly because of the growing interest of popular music producers and musicians. Sampling featured prominently in renowned bands such as The Beatles, for example in Tomorrow Never Knows (1966) and Revolution 9 (1968), and The Residents, whose song Swastikas on Parade (1976) appropriates and samples James Brown extensively. Later, from the mid 1980s onwards, most electronic dance music has significantly explored samplings techniques. Sampling CDs, a new commercial product that contains rhythmic loops and short bass or melodic phrases, became quite popular among this group of musicians. Commonly, loops featured in these CDs were labeled and distributed by genre, tempo, instrumentation, and mood. Most well known uses of this practice occur in popular music, such as hip-hop, which has immediate roots in the 1960s reggae and dub music of Jamaica, and ancient roots in the oral traditions of Africa. Sampling techniques have been expanded since the 1940s, importantly including the use of various samples sizes, as explored in micromontage and granular synthesis, two techniques that will be further detailed in the following sections.
! ! 102! 5.1.2 – Micromontage Micromontage defines the process of composing musical works by assembling short audio samples, usually known as microsounds. All sounds between the sample and sound object time scales can be defined as microsounds, roughly equivalent to the range between 10 and 100 milliseconds (Roads, 2001). Micromontage treats sound as streams of acoustic particles in time and frequency domains. Curtis Roads offers a systematic survey of the history and origins of microsound as well as its application in music composition in his seminal book Microsound (2001). Roads not only exposes the history and roots of microsound from the atomistic Greek philosophers of the 5th century BC until the modern concept of sound particles by Einstein and Gabor, but also provides a comprehensive overview of the artistic work done in this domain, including his own compositions. Iannis Xenakis was the first composer to develop compositional systems that explored microsounds extensively (Roads, 2001)—“grains of sounds” in Xenakis’ terminology. For Xenakis, all sounds can be seen as the “integration of grains, of elementary sonic particles, of sonic quanta” (Xenakis, 1971, p. 43). Xenakis developed a taxonomy for grains of sounds and sound-particles assemblages, such as “sound masses,” “clouds of sound,” and “screens of sound” (Xenakis, 1971). The Argentinian composer Horacio Vaggione has worked extensively with micromontage techniques and is recognized as a pioneer of using sampling techniques in the digital domain (Sturm, 2006b). Vaggione’s first experiments with micromontage date back to 1982 when he started composing Octuor. All the sound material used in Octuor derives from a set of five audio files that were previously synthesized by the composer. The files were initially segmented into small fragments and later edited and mixed into medium to large-scale structures. Thema (1985) for bass saxophone and tape and Schall (1995) for tape are two other major works from Vaggione that continue to explore
! ! 103! micromontage. In Schall, the composer transforms and arranges thousands of segments of piano sounds to create a variety of textures and themes (Roads, 2001). The initial experiments of these two composers—Xenakis and Vaggione—constitute the most important impulses in both the theory and practice of micromontage. Their works guided most future developments of the technique, which many composers have continued and extended towards different aesthetic approaches and technology, such as Karlheinz Stockhausen, Gottfried Michael Koenig, and Noah Creshevsky. Of notice the work of the Portuguese composer Carlos Caires, in particular his software IRIN (Caires, 2004), which combines graphic and script editing with algorithmic generation and manipulation of sound sequences to ease the creation of compositions through micromontage. Caires’s work points toward interesting directions in regards to how to obtain and compose with very short audio snippets, in particular how to organize and manipulate meso structures. 5.1.3 - Granular Synthesis Granular synthesis is a technique that assembles very short segments of audio to build sonic textures, which can be understood as an extension of micromontage towards a higher degree of automation (namely in the selection procedures). In fact, a pioneer of granular synthesis—Curtis Roads—was under the supervision of Horacio Vaggione—a micromontage pioneer—while experimenting with the technique. Barry Truax, a Canadian composer and researcher, is another pioneer of granular synthesis in its extension towards real-time uses (Truax, 1988). Granular synthesis uses short snippets of sound, called grains, to create larger acoustic events. Grains are signals with a Gaussian amplitude envelope that can be constructed from scratch, like different types of sound waves, or short audio segments obtained by segmenting an audio sample. The duration of a grain typically falls into the range of 1-50 milliseconds (Roads, 1998). Most granulators synthesize multiple grains simultaneously at
! ! 104! different density rates, speed, phase, loudness, frequency, and spatial position. Of note is how Barry Truax’s soundscape compositions demonstrated that granular synthesis is particularly efficient at generating natural acoustic environmental sounds, such as rain, waterfalls, or animal vocalizations. Above, I presented an overview of three representative sample-based synthesis techniques in order to introduce the reader to the state-of-the-art technology and artistic practices before the emergence of CSS. Despite having already discussed CSS in several sections of this dissertation,16 I will address once more this synthesis technique to describe its application in musical composition during the last years. 5.1.4 – Musical Applications of Concatenative Sound Synthesis In 2006, while referring to Bob Sturm’s compositions17 and to his own compositions using real-time CSS, Diemo Schwarz claimed that “the musical applications of CSS are just starting to become convincing” (Schwarz, 2006, p. 13). Regarding the application of CSS to high-level instrument synthesis, Schwarz (2006) added that “we stand at the same position speech synthesis stood 10 years ago, with yet too small databases, and many open research questions” (p. 13). Schwarz furthers his remarks with a prediction that in a few year’s time, CSS will be where speech synthesis is at the time. “After 15 years of research, [concatenative TTS synthesis] now become a technology mature to the extent that all recent commercial speech synthesis systems are concatenative” (Schwarz, 2006, p. 14). Schwarz’s prediction became true regarding the application of CSS to high-level instrument synthesis. The Vienna Symphonic Library,18 and Synful (Lindemann, 2001) are two remarkable examples of state-of-the-art CSS software for instrumental synthesis. !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 16 An overview of the technical components of CSS has been presented in section 1.3 and various aspects of CSS have been discussed in Chapters 3 and 4. 17 Diemo Schwarz was referring to Bob Sturm’s compositions: Dedication to George Crumb (2004) and Gates of Heaven and Hell: Concatenative Variations of a Passage by Mahler (2005). 18 http://www.vsl.co.at
! ! 105! Vienna Symphonic Library had several updates for the last years and increased significantly its database.19 On the contrary, Synful does not rely on its database’s size to provide better results, but in additional processing—using transformations of pitch, loudness, and duration. Nonetheless, Synful fulfills the application of high-level instrument synthesis strikingly well. The application of CSS to instrumental synthesis if of utmost importance for composition, but, even if it improves the quality of the results in comparison to other instrumental synthesis techniques, it does not provide tools that expand a compositional thinking towards new musical ideas. However, these ideas have been explored by different CSS software, such as MATConcat (Sturm, 2004), CataRT (Schwarz, 2006a), and AudioGuide (Hackbarth et al., 2010), and I can summarize them in three major compositional strategies: (1) re-arranging units from the corpus by other rules than the temporal order of their original recordings; (2) composition by navigating through a liveor pre-assembled corpus; and (3) cross-selection and interpolation, which allow to extract and apply the morphology of one corpus to another. Hitherto the above-mentioned compositional ideas have been mostly applied in musical composition by the CSS systems’ developers.20 A significant exception is Schwarz’s CataRT, which has been utilized in many creative projects, even if most of them result from a direct collaboration with Schwarz or from people working at the Institut de Recherche et Coordination Acoustique/Musique (IRCAM), where Schwarz currently works. Matthew Burtner, Sebastien Roux, Hector Parra, Luca Francesconi, Stefano Gervasoni, and Dai Fujikura are contemporary music composers that have worked at IRCAM and employed CataRT in their compositions (Schwarz, 2007). Schwarz has also been performing with CataRT for several years, either as a solo performer or in improvisation sessions with live performers. He is a regular presence in the music sessions of many international !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 19The latest Pro Edition Vienna Symphonic Library comprises 235 GB of instrumental sound samples—an increase of 135 GB since its first release in 2002. 20 Note that my comment may also suffer from a lack of documentation about music composed by CSS. Composers are certainly less concerned with the documentation of the techniques they apply in their practice than researchers that work in the academia.
! ! 106! conferences related to computer music, such as Sound and Music Computing, Live Algorithms, International Computer Music Conference, and New Interfaces for Musical Expression. He has been performing with renowned musicians such as the trombonist George Lewis, the saxophonist Evan Parker, and the clarinetist Etienne Brunet.21 The last application of CataRT that I would like to highlight is the interactive exploration of sound corpora in combination with new interfaces for music expression. For example, the Plumage project explores sound corpora by navigating in three-dimensional visualizations of the corpora (Schwarz et al., 2007), and the project Dirty Tangible Interfaces (DIRTI) uses CataRT to sonify and interact with tangible interfaces such as granular or liquid material placed in a glass dish (Savary, Schwarz, & Pellerin, 2012). Norbert Schnell, another IRCAM researcher and head of the IRCAM Real-Time Musical Interactions team, has recently presented the MuBu library for Max/MSP, which is a set of externals for interactive real-time synthesis of analyzed and annotated audio segments. Similarly to CataRT, the MuBu library was already applied as a CSS system in musical composition, notably to assist composers in residence at IRCAM such as Marco Antonio Suárez-Cifuentes in Caméleon Kaléidoscope (2010) and Plis (2010), and Mari Kimura in Clone Barcarolle (2009). MuBu has also been used in projects dealing with new interfaces for music expression such as Mogees,22 which applies “realtime audio mosaicing” to augment everyday objects and transform them into musical instruments. Apart from the exception of the work developed at IRCAM and the aforementioned commercial CSS software for instrumental synthesis, Schwarz’s prediction about the dissemination of CSS is not yet apparent in contemporary music practice. Most of the remaining compositions or sound examples were mostly produced by the system’s developers, as is true of Tristan Jehan, William Brent, Michael Casey, and!Ben Hackbarth. I believe that many musicians are interested in the technique, but most CSS software are !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 21 The improvisation with George Lewis, Evan Parker and Diemo Schwarz took place during the Live Algorithms for Music conference in 2006, and was later released on CD (Schwarz, 2007). The performance with the clarinetist Etienne Brunet, along with many other examples, is available in Schwarz’s website: http://diemo.free.fr. 22 http://www.brunozamborlin.com/mogees/.
! ! 113! composers of Bach’s family legacy (Geiringer, 1950).24 Until the 20th century, composers who integrated pre-existing music into their pieces adapted the material to their idiom, and their compositions maintained a sense of stylistic unity. Contrarily, appropriation in the 20th century shifted towards the use of “readymade” musical material that “clashes with the prevailing style of the original piece, rather than conforming to it” (Leung, 2008). The neoclassical works of Igor Stravinsky, such as Pulcinella (1920) and The Fairy’s Kiss (1928), are remarkable examples of compositions in which Stravinsky reworked upon a borrowed material. Stravinsky does not appropriate for increasing his own expressivity, but rather for expressing his view of the past (Leung, 2008). The idea of “ready-made” or collage is even more present in the works of Charles Ives and George Crumb. In Central Park in the Dark (1906) and The Fourth of July, the third movement of A Symphony: New England Holidays (1897-1913), Ives presents to the listener a complex interaction between his “imaginary present” and “memorable past.” Ives commonly refers to the past by quoting his childhood tunes (Leung, 2008). Crumb appropriates musical material from others by literally quoting the material in his compositions. In Crumb’s compositions appropriated musical materials cohabit independently, integrating and overlaying uneven aesthetics. A remarkable example of Crumb’s use of appropriation can be found in Night Spell I, the sixth piece in Makrokosmos (1972-1973).25 Another notable example of music appropriation in the 20th century, which cannot remain unmentioned is the third movement of Berio’s Symphony (1969) for eight singers and orchestra, which was entirely conceived as a tapestry of quotes from various works by the following composers: Bach, Beethoven, Brahms, Mahler, Debussy, Ravel, Strauss, Stravinsky, Schoenberg, Berg, Stockhausen, Boulez, and even early works by Berio himself !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 24 Please refer to Burkholder (1983) and Leung (2008) for a comprehensive review of appropriation techniques in early Western music. 25 For a deeper review on appropriation techniques used by 20th century composers please refer to J. Peter Burkholder (1983, 1994), who systematically outlines a large set of “borrowing” techniques found in music with a particular emphasis on the musical pieces of Charles Ives.
! ! 114! (Altmann, 1977). From the 1940s onwards, the practice of appropriation became popular due to technological advances that allowed musicians to record, manipulate and playback audio by electronic means. The gradual massification of music technology tools—in particular the sampler—since the 1940s, provoked an aesthetic shift from an early historical phase designated as acousmatic to a later stage addressed commonly as sampling culture (Waters, 2000). While the first relies mostly in self-referential matter and on the listening experience, the second relies on musical and cultural referential contexts, notably by incorporating and reutilizing pre-existing music recordings data to convey new means of expression (Waters, 2000). As I mentioned earlier, the sampling technique relies on existing recordings and is therefore related to the concept of appropriation as a compositional principle. In fact, it is only in the second half of the 20th century that the term appropriation became a musical concept (Landy, 2007). The first example of an electronic music composition entirely based of borrowed audio material is James Tenney’s 1961 composition Collage #1 (Blue Suede) (Cutler, 2004). In this composition, Tenney recombines and manipulates sound material from Elvis Presley’s song Blue Suede Shoes. Two additional early examples of compositions that explicitly expose the technique of appropriation are Bernard Parmegiani’s Pop'eclectic (1968) and Du pop à l'âne (1969). These tracks were created as tapestries of mostly late 1960s pop records, and assembled with unique and significant relationships between sonorities, genres, and cultural contexts by transitioning seamlessly between small samples. Tenney’s and Parmegiani’s works also question the distinction between low art and high art sometimes also referred to as popular music and art music. Since then the differences between these categories have become less prominent (Emmerson, 2001; Landy, 2007). Another proponent of appropriation in electronic music who explores this overlap between low art and high art is John Oswald. His 1988 CD named Plunderphonic (Oswald, 2001), demonstrates an unusually broad eclecticism by plundering, recombining, and
! ! 115! decontextualizing music from Ludwig van Beethoven to the Beatles.26 The practice of appropriation is even more evident in popular music, namely after the emergence of affordable technology such as the sampler, which was and still is a huge catalyst of the technique. Many concepts are associated with appropriation and expose similar or overlapping approaches, such as sampling, remix, collage, mashup, cutups, cut & paste, blend, crossover, plunderphonics, etc. All of these terms are highly associated with popular music, and in particular with practices and styles such as Hip-hop, Rap, and DJing. The idea of appropriation has been explored in many other fields, which to a certain extent have also influenced many contemporary composers. The idea of appropriation is particularly present in the visual arts. The collages of George Braque and Pablo Picasso, and the ready-mades from the artists associated with the Dada movement are clear examples of such. In literature, an exponent of the cut-up technique, that is, a literary technique in which a text is cut up and rearranged to create a new text is the American writer William Burroughs. In philosophy, I may cite Mikhail Bakhtin, in particular his concept of dialogisms, which has been acknowledge and followed by Julia Kristeva in her intertextual theory (Kristeva, 1969). The system developed here embraces the idea of appropriation by recombining userassigned sounds. In comparison with most CSS systems, earGram uses relatively larger sound segments, whose source is easily recognizable after recombination. Therefore, the resulting music can be seen to a certain extent as a remix or variation of the audio source(s). In addition, if one uses a corpus that comprises sound objects from audio sources with distinct styles, origins, or aesthetics, one may not only recombine sound objects according to morphological features, but also drawing upon the cultural associations of the original pieces. A final note should be paid to the relation between copyright laws and the practice of !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 26 Please refer to Oswald (1986), Holm-Hudson (1997), and Cutler (2004) for an historical and conceptual overview of Oswald’s work.
! ! 116! appropriation. As Simon Waters (2000) points, sampling embeds an ambiguous relation between ownership and authorship. The practice of appropriation raises many problems concerning copyright infringements. I will not unpack the topic here, because it is not of primary importance to my dissertation. However, the reader may refer to Bob Sturm (2006a) for a legal discussion on the subject within the scope of sound synthesis, and Lawrence Lessig (2008, 2004, 2001) for a general take on the subject. Having situated earGram historically and aesthetically, I will narrow my perspective to the practical implications of the various technical and conceptual issues raised in this chapter. In order to do so, I will first discuss design strategies for musical composition (§ 5.2), which will then be examined from an algorithmic perspective (§ 5.3) and more precisely in the devised framework (§ 5.4). 5.2 – Design Strategies for Musical Composition As Gottfried Koenig (1978) points out, it is interesting to note that the concept of musical composition relates to both the act of producing a score or a fixed media work, and to the result of that process. While the concept can be seen as definite in terms of the resulting product, it says nothing with regard to the creative process. It is important to understand the creative process, however, in order to be able to encode it algorithmically (or at least partially) and ultimately generate some coherent musical results. A crucial feature of any computational system that intends to automate the processes of music creation is the need to algorithmically encode some creative features that are inherent to human activity. However, creativity is an extremely difficult concept to circumscribe in a strict definition, in particular because there is a lack of understanding of how our creative mechanisms fully work (Taylor, 1988; Csikszentmihalyi, 2009). Even though computer programs seem to oppose to the idea of limitless originality, they also
! ! 117! offer potentialities that are hardly achieved by humans. Computers can actively contribute to the development of new creative practices and promote interesting discussions concerning artificial creativity. Composing can be seen as a decision-making process. Many choices have to be made during the creation of a musical piece, from high-level attributes such as instrumentation to low-level elements, such as pitches and durations. Musical composition design and practice commonly require one of three distinct approaches: (1) top-down, (2) bottom-up, or (3) the combination of both (Roads, 2001). A top-down approach to musical composition starts by developing and formulating the macrostructure of the work as a predetermined plan or template, whose details or lowerlevel formulation are elaborated at later stages of the composition process. All time scales below the macrostructure are considered and refined in greater detail according to the initial plan, until the most basic elements of the structure are elaborated. In Western music, this compositional strategy has been extensively adopted from the 17th to the late 19th centuries, especially because during this historical period the form or macrostructure of the works was mainly confined to a limited number of options (Apel, 1972), such as the sonata form, the rondo, and the fugue. Many music theory textbooks catalog the generic classical forms (Bennett, 1981), whose widespread use enters into a decadent phase at the turn of the 20th century. By contrast, a bottom-up approach conceives the musical form as the result of a process. The macrostructure is the consequence of the development of small-scale ideas or provoked by the interaction of the lower levels of musical structure. Roads (2001) mentions serialism as a paradigmatic example of a bottom-up musical compositional technique, in particular the permutations resulting from applying the inversion or retrograde operations. Bottom-up compositional strategies may also be found in electronic music in processes such as time-expanding a sound fragment into evolving “sound masses.” These examples create an apparent line between different historical periods.
! ! 118! Top-down approaches were assigned to musical compositional practices before the 20th century and bottom-up strategies from the 20th century onwards. Although some generalization may be made in this regard, the implied distinction is not entirely true. Not only has the musical form evolved continuously from the 17th to the 19th centuries, but also in contemporary music the older forms remain present. This is not to say that the use of preconceived forms has died. The practice of top-down strategies in contemporary music still subsists (Roads, 2001), even if in most cases it does not apply to known forms. The compositional process may also incorporate both top-down and bottom-up approaches. In the case of what I call the hybrid approach, the composition is the result of a constant negotiation between its lowand high-hierarchical layers, which are drawn simultaneously. In electronic music the creative process is not dissimilar from traditional instrumental composition concerning the various levels of decision-making. However, it is possible to point to a clear difference between the two practices, which are related to the nature and idiosyncrasies of the raw material used. While instrumental music departs from an abstract to a concrete realization, electronic music commonly starts from concrete sounds (or synthesis methods) to an abstract level. Therefore, in electronic music the act of composing involves the need to define the elementary units of the composition, that is, the sounds themselves. As Koenig notes “electronic sounds or graphic symbols are not always additions to composition; they are often ‘composed’ themselves, i.e., put together according to aspects which are valid for actual composing” (Koenig, 1978). Koenig’s statement articulates a fundamental aspect of electronic music compositional processes, which must also be taken into consideration in CSS during the choice of the audio source(s) and segmentation strategies. The synthesis quality of earGram is not only dependent on the characteristics of its database, but also on the algorithmic composition strategies for unit selection.
! ! 119! 5.3 - Algorithmic Composition Algorithmic composition is the term used to describe “a sequence (set) of rules (instructions, operations) for solving (accomplishing) a [particular] problem (task) [in a finite number of steps] of combining musical parts (things, elements) into a whole (composition)” (Cope, 1993). David Cope’s definition of algorithmic composition is one of the broadest and most concise descriptions of the field, especially because it does not imply any means of production. It not only encompasses the various historic periods that presented work in this domain, but also restricts its modus operandi to a set of specific and clear procedures. The definition comprises two parts. The first part addresses the general definition of algorithm (Knuth, 1968; Stone, 1972) and the second part restricts the target object of the algorithm problem-solving strategy to the music domain. For Cope, music is defined as an activity that groups musical elements into a whole composition. When designing an algorithmic work, the role of the composer is significantly different from the attitude undertaken in traditional Western compositional approaches. Heinrich Taube (2004) refers to this role as a new compositional paradigm. While creating an algorithmic composition, the composer works on a meta-level because instead of outlining a piece by stating musical events notation or soundwise, he/she designs a model which in turn generates the work. An algorithm, within this domain, constitutes a well-defined set of instructions that define and control particular aspects of the composition. The algorithm must effectively provide a finite number of states and their interaction. However, it does not necessarily convey deterministic results. Algorithms for music composition are commonly initialized by data that alters its behavior, and consequently its outcome.
! ! 120! 5.3.1 – Algorithmic Composition Approaches The earliest experiments on algorithmic composition are commonly traced back to the 11th century (Roads, 1996). Apart from their historic importance, algorithmic music composed before the mid 20th century constitutes isolated experiments with minor significance to the music field. Algorithmic composition establishes itself as a field in its own right in the late 1950s by integrating the power of digital computers in the design of algorithms to assist the generation of musical works.27 Early approaches to CAAC—by Lejaren Hiller, Leonard Isaacson, Isacson, Iannis Xenakis, and Gottfried Koenig—have established the basis of the practice according to the following two major approaches: (1) generative models of music for style imitation and (2) generative models of music for genuine composition.28 Several composers and researchers have further research in CAAC according to these two major trends. The first line of research, whereby generative models of music for style imitation, follows the early experiments of Hiller, Isacson, and Koenig—in particular the formalization of principles from music theory or the emulation of a particular style, composer, or body of works. There are two approaches to generative models for style imitation: (1) knowledge engineering, in which the generation is guided by rules and constraints encoded in some logic or grammar, and (2) empirical induction, in which the generation relies on statistical models resultant from the analysis of existing compositions (Conklin & Witten, 1995). Some of the topics that have been continuously revisited within the knowledge engineering approach to generative models of music for style imitation are: the generation of species counterpoint (Ovans & Davidson, 1992; Farbood & Schoner, 2001); functional !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 27 For a comprehensive review of the history of CAAC please refer to Nierhaus (2009) and Ariza (2005). 28 The concept of genuine compositions establishes a distinction between approaches to music composition, whose starting point for a work focus rather on idiomatic approaches that a clear desire to imitate the style of a particular composer, work(s), period, etc. The concept does not intend to raise aesthetic questions related to the originality and/or validity of a work, or even its definition as art. I am only concerned with distinguishing an attitude towards the act of composing.
! ! 121! harmony as used in Western music from the 17th to 19th centuries (Pachet & Roy, 2001); the automatic generation of rhythmic events, namely in the context of interactive music systems (Eigenfeldt, 2009; Bernardes, Guedes, & Pennycook, 2010; Sioros & Guedes, 2011); and the exploration of serial music operations (Essl, 1995; Ariza, 2004). Concerning the empirical induction methods to generative models of music for style imitation one may highlight the work of David Cope (1993, 1996, 2001). Cope (1996) extensively used transition networks to create representations of musical data extracted from one or more compositions. These representations allow the automatic generation of new pieces that resemble the style of analyzed compositions. The second line of research in algorithmic composition—genuine compositions— applies techniques that are commonly inspired by models outside of music, or formulated from scratch. This research line is grounded in the early algorithmic music experiments carried out by Xenakis, Hiller, and Isacson. The adoption of terminology, concepts, and algorithms from disciplines outside the music domain, in particular from biology, became popular in this domain from the 1980s onwards. Typical examples of such strategies are cellular automata (Beyls, 1989; Miranda, 2001), chaos attractors (Pressing, 1988; Bidlack, 1992; Leach & Fitch, 1995), Lindenmayer systems (Prusinkiewicz, 1986; DuBois, 2003), and artificial neural networks (Hild, Feulner, & Menzel, 1992; Mozer, 1994). Current commercial digital audio workstations, such as (Magic) Garageband29, Ableton Live30, and in particular the programming environments for interactive music creation, such as Open Music,31 Max/MSP,32 Pure Data,33 and PWGL34 provide a large set of tools for the exploration of CAAC techniques. !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!! 29 http://www.apple.com/ilife/garageband/. 30 https://www.ableton.com/en/live/. 31 http://repmus.ircam.fr/openmusic/home. 32 http://cycling74.com/products/max/. 33 http://puredata.info/. 34 http://www2.siba.fi/PWGL/.
! ! 122! 5.4 – Computational Life Cycle of Music: An Analysis-Synthesis Approach In this section, I establish a link between the previous sections of this chapter, which provide theoretical and practical foundations of composition, and the generative music strategies developed to recombine sound objects. I will start by describing the architecture of the framework developed here, and conclude by detailing the interaction between the two major modules of the framework (i.e., analysis and composition) and how they determine compositional strategies. The proposed framework can be summarized and characterized by a compound word: analysis-synthesis. Analysis-synthesis defines two complementary procedures that were addressed in two distinct parts of this dissertation. The first is rather analytical, and its aim is threefold: (1) decomposing the audio continuum into elementary units, (2) describing the content of the audio units, and (3) modeling and depicting the higher structural levels of the audio source(s). The counterparts of these operations are the generative aspects of the framework. Both modules examine the same time scales, but assume an inverse or complementary path. While analysis deconstructs the audio continuum from lower to higher elements of musical structure, the composition module recombines the units using algorithmic strategies, from the higher to the lower levels of musical structure. The proposed framework can be placed along the axis between automatic music generation and assisted-algorithmic composition, merged in the concept of CAAC. Additionally, the generative strategies implemented in the framework are not exclusively focused on style imitation or genuine algorithmic strategies, but rather on a hybrid combination of these. As a matter of fact, in addition to a consistent description of all audio units, the analytical module goes as far as modeling individual elements of the audio source(s) structure, which can be used to automate parameters of the generated music.