Proceedings of the 17th International Symposium on Computer Music Multidisciplinary Research
Full text
P R O C E E D I N G S
CMMR 20 25
The 17th Interna tional Symposium on
Computer Music Multidis ciplinar y Res ear ch
S ound, Music : Space , P lace
Hos ted b y Univ er sity Col leg e London
3- 7 No v ember 2025
© To m H o l b e r t o n , W i l l i a m T i n d a l
Proceedings o f the
17 th International Sym posium on
Computer Music Multidisciplinary Research
3 – 7 th Novem ber , 20 25
London, UK
Organized by
Universi ty Col lege London
London, UK
a nd
The Laboratory PRISM
“Perception, Repre sentations, Image, Soun d, Music”
Marsei lle, Fra nce
Published by
The Laboratory PRISM
“Perception, Repre sentations, Image, Sound , Music”
31, chemin Josep h Aiguier
CS 70071
13402 Marseille Ced ex 09 - France
November , 20 25
All c opyrig hts remai n with the a uthors .
ISBN 979 - 10 - 97498 - 06 -1
Les éditions d e PRISM
Organization
The 17th In ternational Symp osium on Computer Music Multidisciplinary Re-
searc h CMMR2025 “Sound, Music : Space, Place” is co-organized b y Univ ersit y
College London (London, UK) and the lab oratory PRISM (CNRS-AMU, UMR
7061, F rance).
General Chair
Angela McArth ur (UCL An throp ology , UK)
General Co-c hairs
Emma-Kate Matthews (UCL Bartlett Sc ho ol of Arc hitecture, UK)
T om Holb erton (UCL Bartlett Sc ho ol of Arc hitecture, UK)
Sp onsor Chair
T om Holb erton (UCL Bartlett Sc ho ol of Arc hitecture, UK)
P ap er and Programme Chairs
Mathieu Barthet (PRISM CNRS-AMU, F rance)
Mitsuk o Aramaki (PRISM CNRS-AMU, F rance)
Ric hard Kronland-Martinet (PRISM CNRS-AMU, F rance)
Sølvi Y stad (PRISM CNRS-AMU, F rance)
P oster Chairs
A v a F atah (UCL Bartlett Sc ho ol of Arc hitecture, UK)
Ch unling W u (UCL Bartlett Sc ho ol of Arc hitecture, UK)
W orkshop Chair
Tin Ob erman (UCL Bartlett Sc ho ol En v, Energy & Resources, UK)
W ebsite Chair
Emma-Kate Matthews (UCL Bartlett Sc ho ol of Arc hitecture, UK)
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
i
Comm unications Chairs
Bleiz Macsen Del Sette (Queen Mary Univ ersit y of London, UK)
Sara Cardinale (Queen Mary Univ ersit y of London, UK)
Sound & Music Chairs
An ton Spice (Goldsmiths, Univ ersit y of London, UK)
Cecilia Anindia Pradita (Indep enden t)
Demo Chairs
Chris P ettit-Mee (UCL An throp ology , UK)
Chris Rho des (UCL Institute of Education, UK)
T utorial Chairs
Ross Purv es (UCL Institute of Education, UK)
Brian Castriota (UCL Art History , UK)
Conference Steering Committee
Mitsuk o Aramaki (PRISM CNRS-AMU, F rance)
Mathieu Barthet (PRISM CNRS-AMU, F rance)
Keiji Hirata (F uture Universit y Hak o date, Japan)
Ric hard Kronland-Martinet (PRISM CNRS-AMU, F rance)
Sølvi Y stad (PRISM CNRS-AMU, F rance)
P ap er Programme Committee
Mitsuk o Aramaki (PRISM CNRS-AMU, F rance)
Ana M. Barbanc ho (Univ ersidad de Málaga, Spain)
Mathieu Barthet (PRISM CNRS-AMU, F rance)
P aul Ba vister (Univ ersity College London, UK)
Jonathan Berger (Stanford Univ ersit y , USA)
Floren t Berthaut (Univ ersité de Lille, F rance)
Tifanie Bouc hara (Univ ersité P aris-Saclay , F rance)
Marco Buongiorno Nardelli (Univ ersit y of North T exas, USA)
F. Amílcar Cardoso (Univ ersit y Coim bra, Portugal)
Isab el Castro (PRISM CNRS-AMU, F rance)
Greg J. Corness (Illinois State Univ ersit y , USA)
Rup ert Co x (Univ ersit y of Manchester, UK)
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
ii
Natac ha Cyrulnik (PRISM CNRS-AMU, F rance)
Kevin Dahan (Univ ersité Gusta v e Eiffel, F rance)
Charalamp os A Dimoulas (Aristotle Univ ersit y of Thessaloniki, Greece)
Georg Essl (Univ ersit y of Wisconsin, Milw auk ee, USA)
Ic hiro F ujinaga (McGill Univ ersit y , Canada)
Keiji Hirata (F uture Univ ersit y Hak o date, Japan)
T om Holb erton (Univ ersit y College London, UK)
Akinori Ito (T ohoku Univ ersit y , Japan)
T ak a yuki Itoh (Oc hanomizu Univ ersit y , Japan)
Katsutoshi Ito y ama (Institute of Science T oky o, Japan)
Haruhiro Kata y ose (K w ansei Gakuin Univ ersit y , Japan)
Alexandros K on togeorgak op oulos (Univ ersit y of A thens, Greece)
Ric hard Kronland-Martinet (PRISM CNRS-AMU, F rance)
Shengc hen Li (Xi’an Jiaotong-Liv erp o ol Univ ersit y , China)
Vincen t Lostanlen (LS2N, CNRS, F rance)
Luca Andrea Ludo vico (Univ ersità degli Studi di Milano, Italy)
Olivier Mac herey (LMA, AMU-CNRS, F rance)
Akira Maeza w a (Y amaha Corp oration, Japan)
Masaki Matsubara (Univ ersit y of T sukuba, Japan)
Emma Kate Matthews (Univ ersit y College London, UK)
Ben McDonnell (Norwic h Univ ersit y of the Arts, UK)
Julie Mey er (King’s College London, UK)
T omo y asu Nak ano (AIST, Japan)
Eleni-Ira P anourgia (Gusta v e Eiffel Univ ersit y , F rance)
Johan P au w els (Queen Mary Univ ersit y of London, UK)
Ross M. Purv es (Univ ersit y College London, UK)
Matthew Ro dger (Sonic Arts Researc h Cen tre, Queen’s Univ ersit y Belfast, UK)
Charalamp os Saitis (Queen Mary Univ ersit y of London, UK)
Sridharan Sank aran (Financial Rob otic A dvisory Services, India)
Dominik Sc hlienger (Uniarts Helsinki, Finland)
Diemo Sc h w arz (Ircam - CNRS STMS, F rance)
Ro d Selfridge (Edin burgh Napier Univ ersit y , UK)
T akuro Shiba y ama (T oky o Denki Univ ersit y , Japan)
Reb ecca Stew art (Imp erial College London, UK)
Ariane Stolfi (Univ ersidade F ederal do Sul da Bahia, Brazil)
Salomé Sudre (LPCPP , F rance & PRISM CNRS-AMU, F rance)
Shinnosuk e T ak amic hi (Keio Univ ersit y , Japan)
Etienne Thoret (INT, CNRS, F rance)
Keisuk e T o y ama (Son y Group Corp oration, Japan)
Isotta T rastev ere (PRISM CNRS-AMU, F rance)
Pierre Alexandre T rem bla y (Conserv atorio della Svizzera italiana, Italy)
Y udai T sujino (Meiji Univ ersit y , Japan)
Jonathan T yrrell (Univ ersit y College London, UK)
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
iii
A drien Vidal (PRISM CNRS-AMU, F rance)
Dan W ang (Univ ersit y College London, UK)
Gerain t A. Wiggins (V rije Univ ersiteit Brussel, Belgium)
Anna Xam b ó (Queen Mary Univ ersit y of London, UK)
Ry osuk e Y amanishi (Kansai Univ ersit y , Japan)
Sølvi Y stad (PRISM CNRS-AMU, F rance)
Asterios Zac harakis (Aristotle Univ ersit y of Thessaloniki, Greece)
Massimiliano Zanoni (P olitecnico di Milano, Italy)
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
iv
T able of Con ten ts
P ap er Oral Session #1 - T ec hnology-Augmen ted
Music Learning and P erformance
Educational Profiling of Guitar T ablature: T o ols to F oster Self-Guided
Learning ......................................................... 11
Mar c elo R o dríguez and A nssi Klapuri
Prop osal of a Singing Instruction Support System with an Inhalation
Restriction F unction to Demonstrate the Difference in Ph ysical Abilit y
b et w een a T eacher and Studen t in Online Lessons ..................... 33
Hir oya Miur a, A kiyo Ob ar a, Hir oki W atanab e, Y oshinari T ake gawa,
and Keiji Hir ata
GuitarFlo w: Realistic Electric Guitar Synthesis from T ablatures via
Flo w Matching and St yle T ransfer ................................... 46
Jackson L oth, Pe dr o Sarmento, Mark Sand ler, and Mathieu Barthet
The Impact of Nonlinear Distortion on the P erceived Brigh tness and
Roughness of Electric Guitar Chords ................................ 59
R ichar d Berr ebi, A gathe Mangialomini, Etienne Thor et, and
R ichar d Kr onland-Martinet
P ap er Oral Session #2 - Auditory and Multimo dal
P erception
Exploring Individualization of Sound Stim uli ......................... 72
Salomé Sudr e, Sølvi Ystad, R ichar d Kr onland-Martinet, L aetitia
Petit, and Mitsuko A r amaki
Bouba, Kiki, and the Mo vemen t Betw een: T o ward a Sensorimotor
Hyp othesis of Sound-Shape Cross-mo dal Corresp ondences .............. 81
L ou Couston, Mitsuko Ar amaki, Sølvi Ystad, Richar d Kr onland-
Martinet, and Etienne Thor et
Bet ween Bouba and Kiki Lies Grougi: Modelling the Rough Middle
Ground of Sound-Shap e Correspondences ............................ 93
Asterios Zachar akis, Konstantinos V elenis, and Emilios
Camb our op oulos
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
1
Cross-Mo dal Coupling b et w een Pitc h and the P erception of Space, and
its Mo dulation b y Con textual Priors ................................. 105
Zofia A nna Hołub owska, Marlen Herkenb er g, and Mar c
Schönwiesner
Exploring the Impact of Nonlinear Distortion on P erceived Musical
Gestures and In ten tions ............................................ 114
A gathe Mangialomini, R ichar d Berr ebi, Etienne Thor et, and
R ichar d Kr onland-Martinet
P ap er Oral Session #3 - Net w ork ed, Situated, and
Collab orativ e Sound Practices
Connectivit y and A daptabilit y: Ev aluation of Signal Flo w Scenarios in
Shared Outdo or Electronic Practices on Mobile Devices ................ 123
Otso A avanr anta, Oleksandr a Nenko, A lejandr o Montes de Oc a,
and Dominik Schlienger
In teractiv e Theatre Sound Design: A Collab orativ e W orkshop .......... 137
R o d Selfridge
Immersiv e Soundscap es with Net w ork ed Vibrating P anels: Bridging
Place and Space .................................................. 149
Mathieu Barthet, Jér émy Perr ouin, Emily Stifter, Vinc ent Debut,
and Christine Esclap ez
P ap er Oral Session #4 (Sp ecial Session) - Human
Auditory Ecology: Minds, Milieux, Mac hines &
Musics
Solidian Soundscap es: Artistic Researc h on Spatial Qualit y of Sound in
Soil ............................................................. 163
Paul Goutmann and Diane Schuh
Sound, Music, Language, and Human Comm unication: A coustic Nudges
in Sound-Mediated In teraction ...................................... 177
Nami A rimitsu
P ap er Oral Session #5 - Computational Musicology:
Mo deling Div erse and Sp ecialised T raditions
Lev eraging Melo dic Con text for Impro v ed Sv ara Represen tation ......... 185
Thomas Nuttal l, Vivek Vijayan, Xavier Serr a, and L ar a Pe arson
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, Nov. 3-7, 2025
2
Graph Con v olutional Neural Net w orks Approac hes for Melo dic P attern
Analysis in Arab-Andalusian Music ................................. 197
A lessandr o Sel lani, Nic c olò Pr etto, and Ivan Donadel lo
CNN P erformance in Genre Classification on P opular Genres from
Amazonian Region: An Analysis in Relation to Human Ev aluation and
its Implications ................................................... 209
Claudio Gomes and Tsang Ing R en
T o w ard Musicologically-Informed Retriev al: Enhancing MEI with
Computational Metadata .......................................... 221
Nádia Carvalho and Gilb erto Bernar des
P ap er Oral Session #6 - Inclusiv e Sound Design and
Musical In terfaces
Sonic W eather: A study on the Sonification of Daily W eather P atterns
and their Correlation with the Sonic Iden tit y of Lo cations .............. 234
A b e er Mohame d and T ony Sto ckman
A Protot yp e of a MIDI 2.0 Con troller for Inclusiv e Music In teraction . . . . 246
V anessa F aschi, Guido Orsomaria Basi, F e deric o Avanzini, Luc a
A ndr e a Ludovic o, and Emanuele Parr avicini
Sp eaK: a Metho dology to Comm unicate ab out Sounds ................. 259
Patrick Susini, Olivier Houix, Nic olas Misdariis, and R omain
Barthélémy
P ap er Oral Session #7 - F ron tiers in Music
Information Retriev al: Analysis, Classification,
and Separation
AnalysisGNN: Unified Music Analysis with Graph Neural Net w orks ..... 271
Emmanouil Karystinaios, Johannes Hentschel, Markus Neuwirth,
and Gerhar d Widmer
Pitc h Estimation in Real Time: Revisiting SWIPE with Causal Windo wi ng 285
Peter Meier, Seb astian Str ahl, Simon Schwär, Meinar d Mül ler, and
Stefan Balke
Exploring the Classification of Pip e Organ Sounds: Recognizing
Instrumen ts, Organ T yp es, and Registrations ......................... 298
Vinc e During, Pablo Dumenil, and Peter van Kr anenbur g
A Study on the Data Distribution Gap in Music Emotion Recognition ... 307
Jo ann Ching and Gerhar d Widmer
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
3
A Conditioned UNet for Music Source Separation ..................... 322
Ken O’Hanlon, Basil W o o ds, Lin W ang, and Mark Sand ler
Pitc h Sp elling Jazz Lead Sheets and Solo T ranscriptions ............... 334
A ugustin Bouquil lar d and Flor ent Jac quemar d
P ap er Oral Session #8 - Designing, Comp osing, and
P erceiving Spatial Audio Exp eriences
Op ening Up Sound Space .......................................... 347
A lain Bonar di and Paul Goutmann
Comp ositional Strategies to Create Exp eriences of Imaginary Spaces
through Headphone Listening ....................................... 359
F atma Neval T arım
Sonic Maze: Pressure-Sensitiv e Dra wing as an Affectiv e Sonic Practice ... 372
F e deric a Ciotti and Patrick Hartono
Influence of Late Rev erb eration and Ec ho es on Rh ythmic Drum Pla ying . 380
Joris A gator, Sølvi Ystad, Mitsuko A r amaki, A drien Vidal, and
R ichar d Kr onland-Martinet
P ap er Oral Session #9 (Sp ecial Session) - Exploring
A tmospheres: In terdisciplinary Phenomenologies
of Sound, Image, and Space
F rom Cyb ernetics to Immersivit y: Reframing Sound Space in
Electroacoustic and Audio visual Comp osition ......................... 391
Paulo C. Chagas and Ivana Petković L ozo
Spatial Sound Do es Not Exist: T o w ard a Con textual Mo del of Spatial
Meaning-Making .................................................. 409
Johannes Scherzer
P oster Session #1 - Music, Sound, and T ec hnology:
P erception, Computation, and Creativ e Practice
A Bit of Ev erything: A Cherry-Pic king Approac h to Massiv ely
Multic hannel Music Comp osition ................................... 423
Pierr e A lexandr e T r emblay
F eature-Based Mo delling of P erceiv ed Emotion in Film Music .......... 434
R uby Olivia Nagano Cr o cker and Gyor gy F azekas
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
4
Supp orting Narrativ e Comprehension in Programmatic Music through
Music and Ligh t .................................................. 447
A nna-Maria Christo doulou, Hugh A lexander von A rnim, and
A lexander R efsum Jensenius
Where to Put That Microphone? A Study of Sound Lo calization in
Am bisonics Recordings ............................................ 455
Maham R iaz, Jinyue Guo, Çağrı Er dem, and A lexander R efsum
Jensenius
Ethical Challenges in AI-Driv en Soundscap e Monitoring: Balancing
Ecological Insigh ts and Priv acy ..................................... 467
Sridhar an Sankar an
Listening is Better Understo o d with: Minds, Milieux, Mac hines & Musics 480
Etienne Thor et and Vinc ent L ostanlen
Implemen tation of a Brass Instrumen t Pla y er Classification Mo del
using F acial Recognition with T w o-Stage Classification Com bining
SVM and CNN ................................................... 494
Hir oya Miur a, Kazuto Inagaki, Hir oki W atanab e, and Y oshinari
T ake gawa
Scoring Instructional Utterances in Guitar Lessons using Seman tic
Lab els and LLM .................................. ................ 506
Nami Iino, Masaki Matsub ar a, Masatoshi Hamanaka, and Hide aki
T ake da
Euphrasis: An In teractiv e System for Music Preference V erbalization . . . . 519
A risa Okamur a, Kento W atanab e, T akayuki Nakatsuka, Kosetsu
Tsukuda, Tian Cheng, T omoyasu Nakano, and Masataka Goto
Audio visual P erception of Musical In teractions in Mixed Pieces with
the Karlax ............. .......................................... 532
Linglan Zhu, Mael Oudin, Benjamin L avastr e, and Stephen
McA dams
MuseLeader: T o w ard Music Editing through Time-series Seman tic
P arameters Con trol using Large Language Mo del ..................... 545
Ryosei Kawaguchi and Haruhir o Katayose
Computational Phrase Segmen tation of Ib erian F olk T raditions: An
Optimized LBDM Mo del ........................................... 557
A mir A bb as Or ouji, Nadia Carvalho, A ntónio Sá Pinto, and
Gilb erto Bernar des
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
5
Fine-Grained MIDI Expression T ranscription from Wind and String
Instrumen t Audio via Sim2Real T ransfer Learning .................... 570
Yifan Xie, Zixun Guo, and Mathieu Barthet
QuartSet: A String Quartet Dataset for T ranscription and Source
Separation of Real Instrumen t Recordings ............................ 582
Erika R umb old and Ge or ge Tzanetakis
Image Sonification as Unsup ervised Domain T ransfer .................. 596
Bálint L aczkó, Marie E. R o gnes, and A lexander R efsum Jensenius
T w o Journeys: Insigh ts on the Annotation of Large-Scale Optical Music
Recognition Datasets .............................................. 608
Pau T orr as, Martina Dvořáková, Carles Badal, Markéta Herzánová
Vlková, Ger ar d Asb ert, Jiří Mayer, A licia F ornés, and Jan Hajič Jr.
P oster Session #2 - Music, Sound, and T ec hnology:
P erception, Computation, and Creativ e Practice
The Sound of Arc het yp es: A Case Study in Immersiv e Sonic T ranslation
of the Ma jor Arcana .............................................. 623
T amas Czup er
Music, T ec hnology , Education and En vironmen tal Sustainabilit y: an
Ov erview of Findings from Recen t Researc h in the UK ................. 635
R oss M. Purves, Evangelos Himonides, and Emma A rnott
Con text-Dep enden t Lab eling of Impro visational Piano P erformance in
Early Childho o d Education and Therap y ............................ 647
R isa Kob ayashi, Kanae Suzuki, Nami Iino, Ikki Ohmukai, Hir oko
T er asawa, and Masaki Matsub ar a
P airing Disparate Datasets for Generativ e Outputs A cross Music and
Arc hitecture ...................................................... 661
R olandas Markevičius
Mo v emen t through Musical Space as a Generativ e Mo del .............. 678
Gr e g Corness
Bridging P ain tings and Music - Exploring Emotion-based Music
Generation Through P ain tings ...................................... 686
T anisha Hisariya, Huan Zhang, and Jinhua Liang
Ethical F ramew orks for Resp onsible Music AI: Balancing Creativit y ,
Ownership, and Cultural Impact .................................... 698
Sridhar an Sankar an
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
6
Pixel Lab eler: A Pixel-wise Annotation T o ol for A ctiv e Learning Music
Recognition ...................................................... 711
Carlos Pér ez-Sancho, A lejandr o Galan-Cuenc a, Juan P.
Martinez-Esteso, F r ancisc o J. Castel lanos, and A ntonio Javier
Gal le go
AI-Assisted Music Pro duction: A User Study on T ext-to-Music Mo dels . . 723
F r anc esc a R onchini, Luc a Comanduc ci, Simone Mar cuc ci, and
F abio A ntonac ci
Rob o-Sax Ensem ble System ........................................ 736
Gou Koutaki and Masatoshi Hamanaka
T raining a P erceptual Mo del for Ev aluating Auditory Similarit y in
Music A dv ersarial A ttac k .......................................... 748
Y uxuan Liu, R ui Sang, Peihong Zhang, Zhixin Li, and Shengchen Li
P arallelising the P olyphon y: High-P erformance Algorithms for Music
P attern Matc hing ............................ ..................... 760
A ntti L aaksonen and Kjel l L emstr öm
MUSE-Explainer: Coun terfactual Explanations for Sym b olic Music
Graph Classification Mo dels ........................................ 772
Baptiste Hilair e, Emmanouil Karystinaios, and Gerhar d Widmer
SP AR-Tim bre: Unlo c king Non-linear Dynamics in Music Audio with
Symmetric Pro jection A ttractor Reconstruction ....................... 784
Mateusz Solinski, W entao Hao, and Elaine Chew
Explicit T onal T ension Conditioning via Dual-Lev el Beam Searc h for
Sym b olic Music Generation ................................. ....... 795
Mar al Ebr ahimzadeh, Gilb erto Bernar des, and Seb astian Stob er
Bey ond the Beats: Can Do wn b eat T rac k ers Predict Hyp ermetre? ........ 807
Jose Esquivel and Jor dan B. Smith
P erceiv ed F emininit y in Singing V oice: Analysis and Prediction ........ . 817
Y uexuan Kong, Viet-A nh T r an, and R omain Henne quin
P oster Session #3 - Music, Sound, and T ec hnology:
P erception, Computation, and Creativ e Practice
Spatial Represen tation in 24 F r ames as a Condition for the Emergence
of A tmosphere .................................................... 826
Natacha Cyrulnik, Isab el le Singer, Je an-Pierr e Mor e au, Isotta
T r astever e, Isab el Castr o, Javier Elip e Gimeno, R osine Bénar d
O’Kel ly, Betina Hueto, and R émi A djiman
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
7
In tegrating Virtual Realit y in to Music P erformance for Concert Audiences 838
Elois Zelada and A nıl Çamcı
F rom Sound to Image: T o w ard an A esthetic of Saturation in P eter
T sc herk assky’s Outer Space ........................................ 848
Javier Elip e Gimeno, Isab el Castr o, and Charles Paiva Santana
Designing a Geolo cated Soundw alk Pla y er Using DIY Electronics: An
Underestimated Approac h to Site-sp ecific Sound Art and Researc h ...... 859
Thomas Denez, A ntoine Be aufort, and Charles V err on
Cross-Mo dal Analysis of Spatial-T emp oral Auditory Stim uli and
Human Micromotion when Standing Still in Indo or En vironmen ts ....... 871
Jinyue Guo, Jim T orr esen, and Alexander R efsum Jensenius
F rom P age to Place: Reviving F orgotten Hymns through AI in a
Site-Sp ecific Heritage Music Pro ject ................................. 883
R oss M. Purves
Can Y ou See the Sound? What Role do Visual Cues Pla y in Aiding
Comm unication for the Individual and b et w een Group Mem b ers in a
Brass Ensem ble Setting? ........................................... 893
No a Kurumi Nishizawa
F ostering Em b o died Sense of “Maai” in Musical Ensem bles: An
In teractiv e Media Approac h for Cultiv ating Collab orativ e Timing and
Spacing .......................................................... 905
Kar en Kur otaki, Kaoru Mikami, Y oshinari T ake gawa, and Masaki
Matsub ar a
A T rial Implemen tation of Augmen ted Realit y Pitc h Visualization in
Ensem ble Rehearsals: Insigh ts from Wind Instrumen t Pla y ers ........... 918
Y asumasa Y amaguchi
Challenges in Basso Con tin uo P erformance-to-Score Alignmen t ......... 924
Suhit Chiruthapudi, A dam Štefunko, Jan Hajič Jr., and Carlos
Eduar do Cancino-Chac ón
Real-Time Gesture Classification via Multi-Mo dal Sensor Data for
In tuitiv e P erformance Mapping ..................................... 937
Evan O’Donnel l and A tau T anaka
Filling MIDI V elo cit y using U-Net Image Colorizer .................... 949
Zhanhong He, David Co op er, Defeng Huang, and R ob erto T o gneri
Melo dy-Rendering Metho d Based on Generativ e Theory of T onal Music . . 961
Masatoshi Hamanaka, Keiji Hir ata, and Satoshi T ojo
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
8
Cluster Expressiv e Timing in P erformed Music with V Q-V AE .......... 973
Zihan Chai and Shengchen Li
Hierarc hical Sym b olic P op Music Generation with Graph Neural Net w orks 985
W en Qing Lim, Jinhua Liang, and Huan Zhang
A Case Study T o w ards Segmen t-A w are V ersion Identification ........... 997
Simon Hachmeier and R ob ert Jäschke
Non-Negativ e Matrix F actorization Decon volution for Multi-Instrumen t
T ranscription of Solo Drum P erformances ............................ 1005
Mikko Hakila and Kjel l L emstr öm
The Mo ving Soundscap e of Xic h uiting: Ritual Sound and Place Iden tit y
in Linhai, China .................................................. 1016
Y uling Chen
Demo Sessions
Singing Without a Larynx: A V o wel-Based Real-Time In terface for
Em b o died V o cal P erformance ....................................... 1027
Gala Lucia Gonzalez Barrios and Ic o Bukvic
Real-time Visualizer for Beats and Scratc hes of Breaking DJ
P erformances .................................................... 1031
Masatoshi Hamanaka
Ec ho F ossils: A F ossil-Inspired AI-Generated Am bient Soundscape
In teractiv e System for Emotional Resonance and Emotional So othing
(W ork in Progress) ................................................ 1036
Jingjing Sun, Ziqing Xu, and Ziming W angr
Undine Glo v es: A Hybrid CV-MIDI Wireless Gestural Con troller for
Mo dular Syn th Ecosystems ......................................... 1041
Bruno Newman
Author Index ................................................ 1046
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
9
ORAL P APER SESSIONS
Pro c. of the 17th In ternational Symposium on CMMR, London, UK, Nov. 3-7, 2025
10
Educational Profiling of Guitar T ablature:
T o ols to F oster Self-Guided Learning
Marcelo Ro dríguez and Anssi Klapuri
Y ousician, Helsinki, Finland
[email protected]
http://www.yousician.com
Abstract. Digital education platforms attract learners with a wide range
of ages and in terests. Platforms are hence required to accommo date guid-
ance accordingly . F or m usic education, sho wing learners what it is ab out
that phrase or transition that is c hallenging, b eing able to tell them whic h
skills they ha v e impro v ed or are lac king, suggest w a ys of practising, are
sough t-after features. That lev el of n uance puts significan t pressure on
man ual curation, calling for to ols to assist or automate coursew are pro-
duction. T o meet that need, this pap er in tro duces a system to enric h
annotation of guitar coursew are. More concretely , a means to automat-
ically profile guitar tablature, annotating it with educationally relev an t
tags and difficult y estimates, do wn to the note lev el. W e v alidate the
resulting descriptors on the task of predicting exp ert-assigned song syl-
labus lev els. W e rep ort a 0.94 Mean Squared Error, on a dataset of 5000
songs, eac h categorised in to one of ten lev els. W e also sho w ho w the de-
scriptors can b e used to visualise skill demands, and are in terpretable as
psyc homotor constrain ts, op ening the p ossibilit y to offer n uanced, p er-
sonalised instruction and feedbac k.
1 In tro duction
Characteristic to the man y services op erating in the EdT ec h domain, is the v ast
range of learner ages and goals. More imp ortan tly , p eople that gra vitate to w ards
these services t ypically express a desire for autonom y [24, 30, 42]. The learner is
hence exp ected to define, at least partially , ho w deep is deep enough, ho w hard
is hard enough, and ho w go o d is go o d enough. In that context, progress and
goal comm unication need to b e not just transparen t, but in tuitiv e. Relying only
on man ual curation to supp ort systems able to offer guidance, with that lev el
of n uance, has sho wn monetarily unsustainable [10], making to ols to assist or
automate coursew are pro duction indisp ensable.
In this pap er w e address that need. W e fo cus on guitar learning, giv en its
p opularit y and widespread demand [43, 56], and prop ose a system to automate
description of guitar coursew are. More concretely , a means to automatically pro-
file guitar tablature, annotating it with educationally-relev an t tags and difficult y
estimates, do wn to the note lev el. A profiled tablature can then express (1) a
v o cabulary of learning outcomes, flexible enough to serv e the comm unication
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
11
2 M. Ro drí g u e z, A. Klapuri
needs of learners of differen t abilities and in terests, and (2) c hallenges in terms
of instrumen tal tec hnique and psyc homotor difficult y .
2 Related W ork
The field of m usic EdT ec h researc h is v ast, ranging from mac hine-driv en p erfor-
mance ev aluation [19, 20, 25] and educational on tologies [3, 16, 46], to difficult y
estimation [2, 36–40, 52, 56]. F rom the latter, the w orks of Vélez et al. [56], and
Ramoneda et al. [36], are closest to ours. W e ha v e t w o w a ys in whic h this w ork
complemen ts theirs. First, the t yp e of m usic notation enco ding: [56] fo cus on
c hord-based reductions for guitar, and [36] fo cus on public-domain collections
for piano. W e add to that the space of tablature arrangemen ts for guitar, whic h
dominates the coursew are of platforms lik e Simply Guitar 1 and Y ousician 2 .
The second is the pursuit of in terpretabilit y . Represen tation learning is b y
no w a w ell-p erforming, mature tec hnology . Neural Net w ork (NN) approac hes
are the top p erformers in that domain. Unfortunately , it is not straigh tforw ard,
or in some cases ev en feasible, to in terpret what eac h dimension in a learn t
represen tation is describing [15, 51, 58, 61]. Much less summarize it in w ay that
is in tuitiv e to studen ts. F or educational applications that is a ma jor obstacle. It
hinders the abilit y to comm unicate progress, pro vide error diagnosis, et cetera.
Vélez et al. address this issue b y prop osing and v alidating a difficult y annotation
rubric. They sho w RNNs can successfully predict the aggregate difficult y , as w ell
as the difficult y of eac h rubric comp onen t. An elegan t solution, pairing an opaque
y et high-p erforming predictor, with a metho d of man ual annotation that is b oth
ric h (m ulti-dimensional) and easy to grasp. Extending their approac h b ey ond
c hord-based notation is, ho w ev er, not scalable. The m yriad of m usic textures
and time spans require rubrics of higher dimensionalit y , whic h in turn mak e
training data pro duction exp ensiv e. W e need additional automation that either
op erates b elo w the rubric lev el, or can suggest rubrics.
In [36] Ramoneda et al. pursue the latter for piano. The rubric comp onen ts
are annotated automatically , and an ordinal NN regressor is used to first com bine
the comp onen ts in to an aggregate of difficult y , and then predict suitable lev el
classes. The mo del is not only transparen t, but it also outp erforms the end-to-end
system they pursue in [39, 40]. This pap er extends that line of exp erimen tation
to guitar, expanding b oth on the n um b er of descriptors, and their affordance.
3 T ablature Profiling System
The con text of this pap er is a guitar learning platform. The platform pro vides
studen ts with access to songs, arranged at differen t difficult y lev els and orien ted
to w ards differen t st yles (e.g. melo dic pla ying, accompanimen t). The arrange-
men t’s notation can b e assumed to sp ecify , at least, the notes to b e pla y ed, their
m usical timing, and the string, fret, and finger with whic h to articulate them.
1 https://www.hellosimply.com/simply- guitar
2 https://yousician.com/
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
12
Educational Profiling of Guitar T ablature 3
Figure. 1: T ablature p rofiling example, f o r ' Under The Bridge ' b y the Red Hot Chili P epp ers.
Figure 1 depicts our profiling system in action. The repres en tation is dual.
On the one hand, w e ha v e m usic v o cabulary and instrumen tal tec hnique (tags).
On the other, psyc homotor demands (difficulties). The former is express ed cate-
gorically , as either nominal data (e.g. "barre c hords") or ordinal data (e.g. "small
jump", "big jump"). The latter is a m ultidimensional v ector. All descriptors are
time-stamp ed, making a tablature profile a tabular, time-series h ybrid.
V o cabulary T agging : Eac h tag is defined in collab oration with our team
of educ ato rs . T agging automation needs to supp ort tag terms exclusiv e to the
platform, and s cale to b oth new tag definitions, as w ell as tag d efinition c hanges.
W e sided for a rule-based system, as it pro vided a clean w a y to address clos ed-
w orld termin o l o gy and its up dates. The rules and tags are defined so as to
matc h the concepts and w ording use d in our instructional videos. The complete
tag se t has ∼ 650 unique tags, whic h includes, among others, instrumen tal tec h-
nique (e.g. ' slide ' , ' thick strings ' , ' partial barre ' ), r h ythm (e.g. ' 16ths ' ,
' off-beat ' , ' shuffles ' ), and ASCI I renditions of notation sym b ols (e.g. ' xx0232 ' ,
whic h is an op en D c hord v o cing, in standard-tuning).
Psyc homotor Difficult y : T he platform fo cuses on score-guided m usic p er-
formance. F or analytical purp oses w e can reduce that to a sequen tial task, in-
v olving three fundamen tal steps: read sym b ols, translate sym b ols to mo v emen ts,
execute mo v e men ts. W e discuss ho w to describ e c hallenges in that task c hain,
from a men tal and ph y s ical p ersp ec tiv e, in Section 4.
4 Describing Psyc homotor Difficult y
4.1 Motor Con trol
Motor c o n trol is resp onsible for planning and sync hronizing mo v emen t. W e tac kle
t w o con trol-related factors: predictabilit y (e.g. ho w often do y ou need to refer
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
13
4 M. Ro drí g u e z, A. Klapuri
to the notation), and coun ting (e.g. ho w to b est en train to the meter and/or
bac king-trac k). W e scop e out readabilit y , whic h in v olv es mapping notation to
pla ying action, as it requires not only kno wledge of the tablature, but also of
ho w the tablature is presen ted to the studen t (engra ving, user in terface).
W e implemen t mac hines for next-tok en prediction, do wn-b eat prediction, and
syncopation prediction. Our strategy is to lev erage the in ternal state of these ma-
c hines for difficult y description. Imp ortan tly , w e aim to dev elop them without the
the need for studen t data (e.g. pla ying histories), in order to free the description
from cold-start noise.
Motor Con trol Difficult y
Based On What’s Estimated
[1, 32] c hord shap e surprise
nec k p osition surprise
fingering + articulation surprise
rh ythmic surprise
[12, 27] offb eatness
[35] coun ting
T able 1: List difficult y descripto rs related to moto r control.
T able 1 summarises the descriptors. F or next-tok en prediction w e use an un-
sup ervised, v ariable-length N -gram mo del, implemen ted through PPM-C, with
in terp olated smo othing [32]. Our mo delling c hoice has b een empirically v alidated
in a wide range of m usic cognition tasks [33]. W e use it for surprisal profiling:
− log( p ( x | z )) with x current tok en, and z preceding tokens. Information-theoretic
surprisal has sho wn a strong causal link to m usic complexit y ratings [44]. It also
ranks amongst the most relev an t descriptors when predicting reading lags in lan-
guage [14, 47]. P arameters are tuned to minimise cross-en trop y in our catalogue.
W e use [12, 27] and [35] for syncopation prediction, and do wn b eat prediction,
resp ectiv ely . [35] finds the b est pulse stream that subsumes an input rh ythm. W e
use its in ternal represen tation to profile the degree of decorrelation b et w een no-
tated meter and rh ythm. [12, 27] pro vide (flat and hierarchical) w ays to estimate
saliency in b eat-to-note distances. W e use these mac hines to profile offb eatness.
As b efore, our mo delling c hoices ha v e extensiv e v alidation in p erception ex-
p erimen ts [11, 12]. P arameters prop osed in [49, 50] w ork b est in our v alidation
datasets.
4.2 Biomec hanics
W e aim to describ e fiv e factors that ma y complicate guitar p erformance: fretting
discomfort (e.g. when pressing strings), hand p osture discomfort (e.g. when hold-
ing a c hord), mo v emen t inefficiency (e.g. when transitioning from one p oin t on
the fretb oard to another), p ositioning inefficiency (e.g. when rearranging fingers
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
14
Educational Profiling of Guitar T ablature 5
so that they are placed on top of the strings that need to b e pressed), and co-
ordination inefficiency (e.g. when a pic king pattern do es not align with a metric
sub division pattern).
Biomec hanical Difficult y
Based On What’s Estimated
[13, 18, 55] [PD] elastic strain
[PD] pressing force
[MI] hand relo cation sp eed
[MI] finger relo cation sp eed
[PI] wrist rep ositioning sp eed
[PI] finger rep ositioning sp eed
[18, 26] [MI] hand relo cation sp eed
[PI] pic k rep ositioning sp eed
[CI] meter-pic king alignmen t
T able 2: List of biomechanical difficult y descripto rs. Stacks (top to b ottom): fretting-hand,
plucking-hand. Abb reviations: PD - P osture discomfo rt, MI - Movement inefficiency , PI -
P ositioning inefficiency , CI - Co o rdination inefficiency .
T o that end, w e implemen t mac hines whic h, giv en an input tablature, can
predict articulation sym b ols. W e mak e one to predict fretting-hand fingering,
based on [18], and another to predict pluc king direction, based on [26]. Both
frame articulation prediction as a searc h problem, where go o d fingers/pluc king
c hoices are those that minimize cum ulativ e biomec hanical cost. T o solv e the
searc h problem efficien tly , dynamic programming is used.
Our references are instances of pro of-of-principle w ork. In b oth cases w e
curated v alidation databases, so as to extend their pro cessing capacit y (e.g.
ha v e [18] supp ort p olyphonic input), and fine-tune their biomec hanical inter-
pretabilit y . T able 2 giv es a summary of the descriptors, and lists references to
the mo dels used to extend the represen tation capacit y of our original sources.
The reader can refer to App endix A and B for a more detailed description.
5 Exp erimen tal Setup
Rating systems and rubrics are widely used in p edagogy , as a means to mak e
learning goals and ev aluation criteria explicit [4, 59]. When paired, they offer
an efficien t w a y to compress and c hain p edagogical information, hence fostering
disco v ery and c hoice (whic h in turn ha v e strong links to p ositiv e engagemen t and
motiv ation [9, 60]). Our syllabus lev elling system is an example of one of suc h
to ols, making lev el prediction b oth a suitable v alidation task, and a v aluable
addition to our soft w are stac k. Our ob jectiv e is then exploring the relev ance of
differen t descriptor sets to lev el prediction. That includes the ones prop osed in
this pap er, as w ell as descriptors prop osed in previous researc h. In [31] we extend
v alidation exp erimen ts to p ersonalisation scenarios.
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
15
6 M. Ro drí g u e z, A. Klapuri
5.1 Lev el Annotated Dataset
F or our v alidation exp erimen ts, w e use a dataset of guitar song arrangemen ts,
notated as tablature, and stored as MusicXML. The dataset consists of 5000 ar-
rangemen ts, randomly sampled from our platform’s catalogue. The ground-truth
annotations consist of syllabus lev els, annotated b y exp ert educators. Levels are
annotated on a scale from 1 to 10.
Lev el Annotation Guidelines
Lev el Chords Riffs Melo dies
1 E, Em, Am, strum
do wn only
Single notes, 1st p os,
EAD strings
1st p os, GBE strings,
no string skips
2 C, G, basic do wn-up
strumming
Single notes, up to
fret 5 (no pinky)
Single notes, up to
fret 5 (no pinky)
5 Syncopation, arp eg-
gios
F aster, hammer-ons
and pull-offs
Hammer-ons, pull-
offs
8 F ull barre (ma j, min,
ma j7, m7, dom7)
F aster rh ythms, more
syncopation
Bending in tro duced
T able 3: Sample of guidelines used by experts to assign levels to song arrangements.
T able 3 pro vides a sample of the annotation guidelines. The guidelines are
clear cut for lo w lev els, and get progressiv ely more abstract for higher levels
(e.g. "faster", "more syncopation"). While the guidelines represen t a consensus,
the annotations themselv es do not: annotations are made b y single exp erts. Due
to the compactness and abstraction of the guidelines, w e can exp ect a non-
negligible lev el of disagreemen t. W e elab orate on the effects that that has on
prediction p erformance in Section 6.
5.2 Descriptors, Predictors, and Figures Of Merit
Prepro cessing : Difficulties are con tin uous v ariables. Eac h difficult y dimension
has a differen t range and a differen t distribution. W e quan tise their range to inte-
gers in 1-10, based on sample quan tiles. W e quan tise each dimension separately ,
using the en tiret y of our catalogue. Quan tisation b oundaries are chosen so that
the data is distributed uniformly among the ten buc k ets. W e also compute the
ranking of eac h tag o v er the entiret y of our catalogue. (The higher the rank, the
more instances.) The reader can refer to App endix A.3 for depictions.
Descriptors : In a profiled song, eac h note/c hord has 15 difficulty n umbers,
and b et w een 5-20 tags. T o summarise a song’s difficulty con tent, w e tak e the
cum ulativ e sum o ver eac h dimension, and also compute the a verage o ver eac h
dimension. A song difficult y v ector then has 15 + 15 dimensions. T o summarise a
song’s tag con ten t, w e make Bag of W ords (BoW) represen tations [54]. A BoW
v ector has o v er 600 dimensions. W e use truncation, where we impose a rank
limit (T opN tags), and dimensionalit y reduction (SVD, with TF-IDF w eigh ting).
Pro c. of the 17th International Symposium on CMMR, London, UK, No v. 3-7, 2025
16
Educational Profiling of Guitar T ablature 7
As b enc hmarks w e use the descriptors prop osed in [36, 56]. T able 4 pro vides
a summary . Benc hmark descriptors are prepro cessed in the same w a y as our
difficult y descriptors.
Baseline Descriptors
Vélez et al. [56] uncommonness of c hord, c hord finger p ositioning,
c hord fingering difficult y , song length in secs, right-
hand complexit y .
Ramoneda et al. [36] pitc h en tropy , pitc h-set LZ, pitc h range, a v erage
pitc h, a verage IOI.
T able 4: Benchma rk Descripto rs. The "rep etitiveness" of [56], and "displacement rate"
of [36] cannot b e computed with our input data. The first requires phrase-level boundary
annotations, which w e lack. The second is piano sp ecific. W e also omit "b eat difficult y",
as suggested b y [56].
Predictors : W e use a Multiv ariate Linear Regressor (MLR) and a Gradient
Bo osted Regressor (X GB). The former giv es us an idea of ho w well a simple
mo del do es on the task at hand, and the latter is an established industry stan-
dard. These c hoices also allo w a parallel with [31, 39].
Figures of Merit : W e use Mean Absolute Error (MAE) and Mean Squared
Error (MSE). Both measures are standard for regression. P erformance aggregates
are obtained b y a v eraging ov er all predictions (micro-a v erage). W e also measure
adjacen t accuracy (A CC), where w e count the n um b er of predictions whose
(absolute) class distance to the true class is no more than one. This score is
[0 , 1] b ounded, and unlik e the other t w o, higher is b etter.
6 Results And Discussion
T able 5 sho ws our exp erimen t results. V arian t 1 presen ts the p erformance of
regressors using the b enc hmark descriptors. V arian ts 2–4, stratify p erformances
b y tag and difficult y . V arian ts 5–6 presen t plateau p erformances for tags. Lastly ,
v arian ts 7–11 presen t com bination sets. F or all exp erimen ts the regressors use
factory defaults. W e pair dimensionalities to mitigate im balance effects. As a
lo w er-b ound, w e presen t the results of tw o naive baselines: Constant5 , whic h
alw a ys predicts lev el 5, and R andom which predicts lev els by sampling from a
N ( µ, σ 2 ) distribution tuned to our dataset.
W e rep ort a 0.94 MSE, obtained in v arian t 10, as the b est result. The pre-
dictor consists of an X GB regressor, with an input of 15 SVD-enco ded tag fea-
tures, and 15 difficult y features (a v erages, computed using quadratic mean). W e
prefer v arian t 10 o v er similar p erforming v ariants on the grounds of parsimon y
(lo w er input dimensionalit y), and generalizability (both regressors p erform w ell).
This predictor is in terpretable only through difficulties. If full in terpretabilit y is
needed, v arian t 11 can b e used.
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
17
8 M. Ro drí g u e z, A. Klapuri
Naiv e Baselines MSE MAE A CC
Random 10 . 7 ± 0 . 38 2 . 67 ± 0 . 05 0 . 22 ± 0 . 02
Constan t5 6 . 30 ± 0 . 19 2 . 08 ± 0 . 04 0 . 27 ± 0 . 01
Id Descriptors MLR X GB
MSE MAE A CC MSE MAE A CC
1 Benc hmark 1 . 97 ± 0 . 09 1 . 02 ± 0 . 03 0 . 68 ± 0 . 01 1 . 56 ± 0 . 07 0 . 96 ± 0 . 02 0 . 73 ± 0 . 02
2 T ags [ top 15] 2 . 99 ± 0 . 14 1 . 36 ± 0 . 04 0 . 60 ± 0 . 02 1 . 28 ± 0 . 07 0 . 79 ± 0 . 02 0 . 84 ± 0 . 02
3 T ags [ svd 15] 1 . 79 ± 0 . 13 1 . 02 ± 0 . 03 0 . 75 ± 0 . 02 1 . 28 ± 0 . 08 0 . 80 ± 0 . 03 0 . 84 ± 0 . 01
4 Diffs [ a v 15] 1 . 76 ± 0 . 06 0 . 99 ± 0 . 02 0 . 75 ± 0 . 01 1 . 36 ± 0 . 08 0 . 82 ± 0 . 03 0 . 83 ± 0 . 02
5 T ags [ top 100] 2 . 43 ± 0 . 12 1 . 19 ± 0 . 04 0 . 69 ± 0 . 02 0 . 92 ± 0 . 05 0 . 64 ± 0 . 01 0 . 89 ± 0 . 01
6 T ags [ svd 100] 1 . 55 ± 0 . 10 0 . 93 ± 0 . 03 0 . 80 ± 0 . 02 1 . 22 ± 0 . 07 0 . 85 ± 0 . 01 0 . 90 ± 0 . 01
7 top 30 + all 30 1 . 30 ± 0 . 05 0 . 83 ± 0 . 02 0 . 82 ± 0 . 02 1 . 01 ± 0 . 06 0 . 67 ± 0 . 03 0 . 88 ± 0 . 01
8 svd 30 + all 30 0 . 99 ± 0 . 05 0 . 71 ± 0 . 02 0 . 89 ± 0 . 01 0 . 90 ± 0 . 06 0 . 63 ± 0 . 02 0 . 90 ± 0 . 01
9 top 15 + a v 15 1 . 44 ± 0 . 07 0 . 87 ± 0 . 02 0 . 81 ± 0 . 02 1 . 09 ± 0 . 05 0 . 72 ± 0 . 02 0 . 87 ± 0 . 01
10 svd 15 + a v 15 1 . 07 ± 0 . 07 0 . 74 ± 0 . 03 0 . 87 ± 0 . 01 0 . 94 ± 0 . 04 0 . 66 ± 0 . 02 0 . 90 ± 0 . 01
11 top 75 + a v 15 1 . 31 ± 0 . 07 0 . 83 ± 0 . 03 0 . 83 ± 0 . 01 0 . 88 ± 0 . 05 0 . 62 ± 0 . 02 0 . 90 ± 0 . 01
T able 5: P erfo rmances as µ ± σ b etw een folds of 10-fold cross-validation.
Descriptor Imp ortance : In our exp erimen ts, tag data has a primary role.
A sophisticated predictor can use it effectiv ely . W e can see that in v arian ts 2–
4, where tags sho w more explanatory p o w er than difficulties, and in v arian t 5,
where the X GB regressor reac hes 0.92 MSE. That said, using a com bination
of tags and difficult y descriptors, w e get comp etitiv e p erformance with lo w er
input dimensionalit y (v arian ts 7–10). Imp ortan tly , when using descriptor com bi-
nations, ev en the h um ble linear regressor can reac h comp etitive lev el (v arian t 8).
The latter also rev eals that the difficult y space forms an orthogonal basis. (MLR
mac hines are kno wn to ha v e p erformance issues when the input is correlated,
visible in v arian ts 2-3, 5-6.)
A few parametrisation asp ects w orth noting. First, the tag descriptors pre-
ferred b y the X GB mac hine (T opX tags), are dominated b y abstractions (e.g.
' thick strings ' ). The dominance of abstractions in T opX sets consisten tly out-
p erformed v arian ts in which w e injected high ranking tags that more closely
represen t the notation (e.g. ' xx0232 ' , ' Amaj7 ' ). Second, for the difficult y descrip-
tors, kinematics driv e p erformance. W e attribute that to the fact that kinematics
ha v e a notational, quan tifiable coun terpart: temp o. Annotation guidelines that
refer to kinematic asp ects, e.g. "faster", ha v e less ro om for interpretation than
e.g. "more syncopation", leading to more consisten t lev el annotations. Third, w e
found that, when summarising difficult y o v er a whole song, length descriptors
(cum ulativ e sum) are less relev an t than means, which is in line with [56]. In our
exp erimen ts the quadratic mean outp erformed all classic Pythagorean means.
P erformance Upp er-Bound : A go o d n um b er of p erformances in T able 5
sho w an MSE ∼ 1. W e argue that that is lik ely b ordering human error. As men-
tioned in section 5.1, w e can exp ect disagreemen t among annotators. Disagree-
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
18
Educational Profiling of Guitar T ablature 9
men t translates in to i nconsisten t lab elling: one annotator ma y think that X
amoun t of, for instance, syncopation, m a k es an arrangemen t a lev el 5, while an-
other ma y think it mak es it a lev el 6. The compactness of the guidelines sho wn
in T able 3 mak e it reasonable to exp ect a minim um of a ± 1 lev el of disagree-
men t, more lik ely dominan t in higher lev els. Sustaining consisten t agreemen t,
o v er y ears, with teams that v ary o v er time, is understandably complex. Vélez
et al. disc uss the negativ e effect it has in comm unit y lead annotation [56]. W e
hence consider the exp ected lev e l of n o i se in our dataset is not exceptional, but
common. Whil e w e can not sa y an ything definitiv e, and m uc h less sp eculate ho w
our p erformances translate to other rubrics, guideline sets, or instrumen ts, w e
hop e that this discussion, ho w ev er brief, insp ir e s more researc h on this topic.
Skill Demands : Lastly , w e pro vide an example of ho w our des criptors enable
educational vis ualisations, and g ran ular skill diagnosis. Figure 2 sho ws a diffi-
cult y profile in con trast to a learner’s error plot. All o v erlaid on top of a song
timeline. The error plot corresp onds to the a v erage n um b er of errors learners
made, o v er m ultiple trials, for eac h note/c hord. W e use a 1K sample of learner
historicals to represen t the p opulation. In this example w e can see the temp o-
ral structure of errors correlates w ith that of difficulties. One p ossible use of the
profile is hence as a preview mec hanism, highligh ting whic h parts of the song are
more complex. W e aim to pursue exp erimen tal v alidation of these p ossibilities in
future w ork. Another use is as a w a y to m o dulate exp erience accum ulation. In
other w ords, correctly pla ying something difficult gran ts more exp erience p oin ts
than correctly pla ying something easy . In [31] w e pursue prediction of prima vista
p erformances based on that idea. W e tested v arious w a ys of quan tifying and ac-
cum ulating exp erience. Results sho w that accum ulating exp erience mo dulated
b y diffic ult y estimates p erforms b est.
Figure. 2: F retting-Hand difficult y p rofile (ba rs) vs. lea rner erro r data (line).
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
19
10 M. Ro dríguez, A. Klapuri
7 Conclusions
W e in tro duced an approac h to automatically profile guitar tablature, annotat-
ing it with in terpretable, educationally-relev an t tags and psyc homotor difficult y
estimates, do wn to single notes.
W e v alidated our descriptors on the task of predicting exp ert-assigned syl-
labus lev els. The b est predictor ac hiev es 0.94 Mean Squared Error on a dataset of
5000 songs, with lev els ranging from 1 to 10. More imp ortan tly , the same descrip-
tor set enables a linear regressor to ac hiev e a similar p erformance. The p erfor-
mance of b oth regressors compares fa v ourably to that ac hiev ed using descriptors
prop osed in previous w ork. W e also pro vided argumen ts for the h yp othesis that
our b est p erformances migh t b e at h uman lev el.
In future w ork w e plan to test the usefulness of our descriptors across differ-
en t p ersonalization scenarios, all with the o v erarc hing goal of aiding self-guided
learning. W e aim to explore automating summary descriptions and arrange-
men t naming, for example. As for the descriptors themselv es, w e plan to explore
ph ysically-informed neural net w orks [21], whic h mak e for an attractiv e c hoice
to tak e automation further, while retaining mo delling that is in trinsically in ter-
pretable. Also, w e plan to explore surprisal mo delling through transformers [14].
While higher (next-tok en prediction) p erformance can b e exp ected, exp erimen-
tation is needed to iden tify the training and parametrisation that b est matc hes
h uman p erception, as that is not a giv en [47, 48].
A c kno wledgmen ts. W e thank the anon ymous review ers for the insigh tful comments
and suggestions.
Disclosure of In terests. The authors ha v e no comp eting in terests to declare that
are relev an t to the con ten t of this article.
References
1. Ab dallah, S., Plum bley , M.: Information dynamics: patterns of exp ectation and
surprise in the p erception of m usic. Connection Science 21 (2-3), 89–117 (2009)
2. A cquilino, A., Sca v one, G.: Curren t state and future directions of tec hnologies for
m usic instrumen t p edagogy . F ron tiers in Psychology 13 (2022)
3. Alev en, V., K o edinger, K.: Design recommendations for in telligen t tutoring sys-
tems, c hap. Kno wledge Comp onen t (K C) Approaches to Learner Modeling, pp.
165–182. US Arm y Researc h Lab oratory (2013)
4. Allen, D., T anner, K.: Rubrics: T o ols for making learning goals and ev aluation
criteria explicit for b oth teac hers and learners. Life Sciences Education 5 (3), 197–
203 (2006)
5. A oki, T., F rancis, P .R., Kinoshita, H.: Differences in the abili ties of individual
fingers during the p erformance of fast, rep etitiv e tapping mo v ements. Experimental
Brain Researc h 152 (2), 270–280 (2003)
6. Bounou, O., P once, J., Carp en tier, J.: Learning system dynamics from sensory in-
put under optimal con trol principles. In: IEEE Conference on Decision and Con trol.
pp. 1885–1892 (2024)
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
20
Educational Profiling of Guitar T ablature 11
7. Chen, C.T., Gu, G.X.: Ph ysics-informed deep-learning for elasticit y: forw ard, in-
v erse, and mixed problems. A dv anced Science 10 (18), 2300439 (2023)
8. Den y er, R., Guillory , I., Cra wford, A.M., F ripp, R.: The guitar handb o ok. P an
b o oks (1992)
9. Duffy , M.C., Azev edo, R.: Motiv ation matters: In teractions b et w een ac hiev emen t
goals and agen t scaffolding for self-regulated learning within an in telligen t tutoring
system. Computers in Human Beha vior 52 , 338–348 (2015)
10. F erster, B.: In telligen t tutoring systems: What happ ened? https:
//elearningindustry.com/intelligent- tutoring- systems- what- happened
(2021), accessed: 2025-03-25
11. F ram, N.R., Berger, J.: Syncopation as probabilistic exp ectation: Conceptual, com-
putational, and exp erimen tal evidence. Cognitiv e Science 47 (12) (2023)
12. Gómez, F., Th ul, E., T oussain t, G.: An exp erimen tal comparison of formal mea-
sures of rh ythmic syncopation. In: ICMC (2007)
13. Grimes, D.R.: String theory-the ph ysics of string-b ending and other electric guitar
tec hniques. PloS one 9 (7) (2014)
14. Grutek e Klein, K., Meiri, Y., Sh ubi, O., Berzak, Y.: The effect of surprisal on read-
ing times in information seeking and rep eated reading. In: Conference on Compu-
tational Natural Language Learning. pp. 219–230 (2024)
15. Guillot, S., Prouteau, T., Dugué, N.: Sparser is b etter: one step closer to w ord
em b edding in terpretabilit y . In: Conference on Computational Seman tics. pp. 106–
115 (2023)
16. Hao, S.: Information tec hnology in tegration and p edagogical inno v ation in the dig-
ital transformation of m usic education. Applied Mathematics and Nonlinear Sci-
ences 9 (1), 1–19 (2024)
17. Heijink, H., Meulen bro ek, R.G.: On the complexit y of classical guitar pla ying:
functional adaptations to task constrain ts. Journal of Motor Beha vior 34 (4), 339–
351 (2002)
18. Hori, G., Saga y ama, S.: Minimax viterbi algorithm for hmm-based guitar fingering
decision. In: ISMIR. pp. 448–453 (2016)
19. Jiang, Y.: Exp ert and no vice ev aluations of piano p erformances: Criteria for
computer-aided feedbac k. In: ISMIR. pp. 367–374 (2023)
20. Ju, Y., W u, C.Y., Lorenzo, B.C., Y ang, J., Deng, J., F an, F., Lui, S.: End-to-End
Automatic Singing Skill Ev aluation Using Cross-A tten tion and Data Augmen tation
for Solo Singing and Singing With A ccompanimen t. In: ISMIR. pp. 493–500 (2024)
21. Karniadakis, G.E., Kevrekidis, I.G., Lu, L., P erdik aris, P ., W ang, S., Y ang, L.:
Ph ysics-informed mac hine learning. Nature Reviews Ph ysics 3 (6), 422–440 (2021)
22. Kimoto, Y., Oku, T., F uruy a, S.: Neurom uscular and biomec hanical functions sub-
serving finger dexterit y in m usicians. Scien tific Rep orts 9 (1) (2019)
23. K ot, M., Nagahashi, H., Szymczak, P .: Elastic mo duli of simple mass spring mo dels.
The Visual Computer 31 (10), 1339–1350 (2015)
24. Leah y , K.S., Smith, T.D.: The self-directed learning of adult m usic studen ts: A
comparison of teac her approac hes and studen t needs. Journal of Music Education
39 (3), 289–300 (2021)
25. Lerc h, A., Arth ur, C., P ati, K.A., Gururani, S.: Music P erformance Analysis: A
Surv ey. In: ISMIR. pp. 33–43 (2019)
26. de Lima, E.T., Ramalho, G.L.: On rh ythmic pattern extraction in b ossa no v a
m usic. In: ISMIR. pp. 641–646 (2008)
27. Longuet-Higgins, H.C., Lee, C.S.: The rh ythmic in terpretation of monophonic m u-
sic. Music P erception 1 (4), 424–441 (1984)
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
21
12 M. Ro dríguez, A. Klapuri
28. Magill, R., Anderson, D.I.: Motor learning and con trol. McGra w-Hill Publishing
(2012)
29. Mirakhorlo, M., V an Beek, N., W esseling, M., Maas, H., V eeger, H., Jonk ers, I.:
A m usculosk eletal mo del of the hand and wrist: model definition and ev aluation.
CMBBE 21 (9), 548–557 (2018)
30. Morris, T.H., Bremner, N., Sak ata, N.: Self-directed learning and studen t-cen tred
learning: a conceptual comparison. P edagogy , Culture & So ciet y pp. 1–20 (2023)
31. Müllersc hön, M., Ro dríguez, M., Klapuri, A., Cardin, C.: Pla y abilit y prediction
in digital guitar learning using in terpretable studen t and song represen tations. In:
ISMIR (2025)
32. P earce, M., Wiggins, G.: Impro v ed methods for statistical mo delling of monophonic
m usic. Journal of New Music Researc h 33 (4), 367–385 (2004)
33. P earce, M.T.: Learning to Listen, Listening to Learn: Music P erception and the
Psyc hology of Enculturation. Oxford Univ ersit y Press (2025)
34. P eña-Pitarc h, E., F alguera, N.T., Y ang, J.: Virtual h uman hand: mo del and kine-
matics. CMBBE 17 (5), 568–579 (2014)
35. P o v el, D.J., Essens, P .: P erception of temp oral patterns. Music P erception 2 (4),
411–440 (1985)
36. Ramoneda, P ., Eremenk o, V., D’Ho oge, A., P arada-Cabaleiro, E., Serra, X.: T o-
w ards explainable and in terpretable m usical difficult y estimation: A parameter-
efficien t approac h. In: ISMIR. pp. 520–528 (2024)
37. Ramoneda, P ., Jeong, D., Eremenk o, V., T amer, N.C., Miron, M., Serra, X.: Com-
bining piano p erformance dimensions for score difficult y classification. Exp ert Sys-
tems with Applications 238 (2024)
38. Ramoneda, P ., Lee, M., Jeong, D., V alero-Mas, J.J., Serra, X.: Can audio rev eal
m usic p erformance difficult y? insigh ts from the piano syllabus dataset. IEEE T rans.
on Audio, Sp eec h and Language Pro cessing (2025)
39. Ramoneda, P ., T amer, N.C., Eremenk o, V., Serra, X., Miron, M.: Score difficult y
analysis for piano p erformance education based on fingering. In: ICASSP. pp. 201–
205 (2022)
40. Ramoneda, P ., V alero-Mas, J.J., Jeong, D., Serra, X.: Predicting p erformance dif-
ficult y from piano sheet m usic images. In: ISMIR. pp. 708–715 (2023)
41. Rau, M.: A Linear-Chain Conditional Random Field Approac h to the Guitar Fin-
gering Problem. Master’s thesis, T ec hnisc he Univ ersität Darmstadt (2016)
42. Robinson, J.D., P ersky , A.M.: Dev eloping self-directed learners. American Journal
of Pharmaceutical Education 84 (3) (2020)
43. Ro driguez, R.C., Marone, V.: Guitar learning, p edagogy , and tec hnology: A his-
torical outline. So cial Sciences and Education Researc h Review 8 (2), 9–27 (2021)
44. Sauvé, S.A., P earce, M.T.: Information-theoretic mo deling of p erceiv ed m usical
complexit y . Music P erception 37 (2), 165–178 (2019)
45. Sc hmidt, R.A., Lee, T.D., Winstein, C., W ulf, G., Zelaznik, H.N.: Motor con trol
and learning: A b eha vioral emphasis. Human Kinetics (2018)
46. Sébastien, V., Sébastien, D., Conruyt, N.: Annotating w orks for m usic education:
prop ositions for a m usical forms and structures on tology and a m usical performance
on tology . In: ISMIR (2013)
47. Shain, C., Meister, C., Pimen tel, T., Cotterell, R., Levy , R.: Large-scale evidence
for logarithmic effects of w ord predictabilit y on reading time. National A cadem y
of Sciences 121 (10) (2024)
48. Shlegeris, B., Roger, F., Chan, L., McLean, E.: Language mo dels are b etter than
h umans at next-tok en prediction. T rans. on Mac hine Learning Research (2024)
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
22
Educational Profiling of Guitar T ablature 13
49. Shm ulevic h, I., P o v el, D.J.: Measures of temp oral pattern complexit y . Journal of
New Music Researc h 29 (1), 61–69 (2000)
50. Smith, L.M., Honing, H., et al.: Ev aluating and extending computational mo dels
of rh ythmic syncopation in m usic. In: ICMC (2006)
51. Stec k, H., Ek anadham, C., Kallus, N.: Is cosine-similarit y of em b eddings really
ab out similarit y? In: A CM W eb Conference. pp. 887–890 (2024)
52. Sébastien, V., Ralam b ondrain y , H., Sébastien, O., Conruyt, N.: Score analyzer:
Automatically determining scores difficult y lev el for instrumen tal e-learning. In:
ISMIR. pp. 571–576 (2012)
53. Th urston, D.P .: Elastic graphs. In: F orum of Mathematics, Sigma. v ol. 7 (2019)
54. T urney , P .D., P an tel, P .: F rom frequency to meaning: V ector space mo dels of se-
man tics. Journal of Artificial In telligence Researc h 37 , 141–188 (2010)
55. Ueda, N., Mase, K.: T rac king mo ving con tours using energy-minimizing elastic
con tour mo dels. In: ECCV. pp. 453–457. Springer (1992)
56. Vásquez, M.A.V., Baelemans, M., Driedger, J., Zuidema, W.H., Burgo yne, J.A.,
et al.: Quan tifying the ease of pla ying song c hords on the guitar. In: ISMIR. pp.
725–732 (2023)
57. Viterbi, A.: Error b ounds for con v olutional co des and an asymptotically optim um
deco ding algorithm. IEEE T rans. on Information Theory 13 (2), 260–269 (1967)
58. W ang, Y., Blei, D., Cunningham, J.P .: P osterior collapse and laten t v ariable non-
iden tifiabilit y . Neural Information Pro cessing Systems 34 , 5443–5455 (2021)
59. W esolo wski, B.C.: Understanding and dev eloping rubrics for m usic p erformance
assessmen t. Music Educators Journal 98 (3), 36–42 (2012)
60. W est, C.: Motiv ating m usic students: A review of the literature. Applications of
Researc h in Music Education 31 (2), 11–19 (2013)
61. Zhang, Y.L., T ang, C., Y u, L., Zhou, J., Li, L., Cui, Q., F an, F., Jiang, L., Zhao, X.:
Domain lev el in terpretabilit y: In terpreting blac k-b o x mo del with domain-sp ecific
em b edding. In: Conference on W eb Searc h and Data Mining. pp. 1102–1105 (2024)
62. Zheng, M., Y uan, Z., Zh u, W., Zhang, G.: A fast mass spring mo del solv er for
high-resolution elastic ob jects. Sim ulation 93 (10), 797–807 (2017)
App endix A Guitar P erformance As Optimal Planning
W e are in terested in quantitativ e, m ultiv ariate, in terpretable descriptors of gui-
tar pla ying difficult y . W e approac h that problem via sim ulation: Ha ve a virtual
guitarist pla y through a giv en tablature, and use its in te rnal represen tation as a
means to profile, note b y note, pla ying difficult y .
T o develop our virtual guitarist, w e ha ve it predict fretting-hand fingers, and
pluc king direction. Since w e deal with notation-guided p erformance, w e frame
prediction as a problem of optimal planning, along a prescrib ed path. The aim
is to find the sequence of finger/pluc king c hoices that minimises biomechanical
load. W e use dynamic-programming to tackle both problems.
The b enefit of this framing is dual. On the one hand, it pro vides a testable
framew ork to optimize difficult y descriptors. On the other, it enables estima-
tion in data-p o or situations, as man ually pro ducing notation with fingering and
pluc king directions is exp ensiv e, and hence rare. In this app endix w e in tro duce
the framew ork that underlies b oth approac hes, and in app endix B describ e dif-
ficult y mo dels. F or brevity , w e limit detailed description to fretting difficult y .
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
23
14 M. Ro dríguez, A. Klapuri
A.1 Scop e
W e assume a 6-string guitar, in standard tuning, with up to 24 frets. W e limit
the anatom y of our sim ulated guitarist to (1) a wrist-fingers mec hanism, mean t
to press strings, hereb y the fretting-hand, and (2) a wrist-plectrum mec hanism,
mean t to pluc k strings, hereb y the pluc king-hand. In tablature notation, guitar
strings are en umerated 1 to 6, in thinnest to thic k est order, and frets are en u-
merated 0 to 24, with fret 0 used to indicate that a string is pluc k ed but not
pressed. Our aim is to predict sym b ols in { 1 , 2 , 3 , 4 } , whic h en umerate fretting
fingers index to pinky , and sym b ols in { U , D } , whic h shorthand pluc king "up"
and "do wn", resp ectiv ely . Figure 3 sho ws an example.
e|---------0---|-8---7---5---|-4-----5-7---|-0---7---4---|-5-----------|
B|-------------|-------------|-------------|-------------|-------------|
G|-------------|-------------|-------------|-------------|-------------|
D|-------------|-------------|-------------|-------------|-------------|
A|-------------|-------------|-------------|-------------|-------------|
E|-------------|-------------|-------------|-------------|-------------|
4 3 1 1 2 4 4 1 2
D U D U D D U D U D U
Figure. 3: T ablature fragment. Finger and plucking symb ols a re notated b elo w the stave.
W e address tablature con taining melo dy , c hords, or a com bination of the t w o.
The latter includes p olyphon y , whic h describ es situations w ere notes can start
while others are still sounding. F or simplicit y , w e don’t tac kle those cases directly ,
and instead assume p olyphonic input can b e viably reduced to homophon y , where
sim ultaneously sounding notes ha v e the same start and end.
A.2 A ction Chain
Here w e in tro duce the sc hematisation of mo v emen t used to name difficult y de-
scriptors in T able 2 of the paren t do cumen t. Based on Sc hmidt [45, p. 48],
w e frame instrumen tal p erformance as serial mo v emen t. Serial mo v emen ts are
comp osed of a sequence of discrete actions, and are amenable to reductionist
analysis. F or instrumen t pla ying, b oth the order and timing of the actions are
imp ortan t. The discretisation that feels most natural is that whic h demarcates
sound-pro ducing actions: pla y c hord/note transition in to the next c hord/note,
and so on. W e refer to discretisation, at this lev el, as that of action units. T o
estimate the difficult y of an action unit, w e need to go a la y er deep er in to the
description of the mo v emen t tasks in v olv ed. F or this app endix, w e consider the
follo wing c hain of tasks:
1. P osition the fingers on top of the string/frets that require pressing
(adjusting grip and wrist angle as needed)
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
24
Educational Profiling of Guitar T ablature 15
2. Depress the strings
3. Excite the strings in order to pro duce sound
4. Release the string or strings b een held do wn b y lifting the fingers
5. Mo v e the hand from wherev er it is to (roughly) where it needs to b e next
T ask 3 can b e executed in man y w a ys: pluc king, tapping, hammer-on, or
pull-off. As stated in section A.1, for simplicit y w e limit the action space to
string pluc king, using a plectrum. The action c hain l isted ab o v e omits the set of
parallel (y et co ordinate d) mo v emen ts needed b y the pluc king-hand:
1. P osition the wrist on the bride or ab o v e the strings, as necessary
2. Mo v e the plectrum from wherev er it is to where it needs to b e
3. Pic k / Strum the strings in order to pro duce s ound
A.3 Example Output
Figure. 4: Distributions fo r each difficult y dimension.
In Figure 4 w e se e distributions for eac h difficult y dimension, considering
our platforms’ complete (guitar) catalogue. On the y-axis w e ha v e the difficult y
dimension names. The surprisal and en trainmen t dimensions corresp ond to those
listed in T able 1 of the paren t do cumen t. The fretting and p luc king dimensions
corresp ond to those of T able 2 of the paren t do c umen t. On the x-axis w e ha v e
the log 2 of the difficult y v alue. In some cases the di stributions are appro ximately
log -normal, as it is with hand relo cation sp eed. In others, distributions are closer
to long-tailed exp onen tial families, as it is with surprisal difficult y dimensions .
W e can see the naming is informed b y the sc hematization of mo v emen t prop osed
in the previous section. As an example, w e can see re-p ositioning relates to step
1 on the task c hain of action units, and p ostu r e discomfort relates to step 2.
In man y cases the difficult y v alue units are also in terpretable. Figure 5 sh o ws
an example. W e see a his tog ram for (ra w) difficult y v alues, as w ell as its inferred
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
25
16 M. Ro dríguez, A. Klapuri
Figure. 5: Example difficult y dimension: ra w values (ba rs) and quantisation cuts (lines).
quan tisation b oundaries (discussed in section 5.2 of the paren t do cumen t). Ra w
v alues corresp ond to relo cation v elo cities, in cen timetres p er second, whic h in-
v olv e mo ving the pic k so that it is ready to pluc k the n ext string. In most acous-
tic and elec tr ic guitars the distance b et w een strings is roughly 1 cen timetre.
So, lo w er quan tised lev els describ e situations where either the pic k is (mostly)
pluc king a single string, or when time is not a constrain t.
App endix B Predicting F retting-Hand Fingering
B.1 Searc h Space Represen tation
W e use a trellis g raph, depicted in Figure 6, as an abstraction of the searc h
space. The trellis represen ts notation-guided p erformance. Eac h v ertical la y er
corresp onds to a note/c hord, and eac h no de to a candidate fingering, from a finite
set. W e use t to denote la y er indexes, and i, j to denote the indexes of transitions
from one candidate to another, o v e r adjacen t la y ers. F ormally , w e assume a set of
no des S , represen ting the space of p erformer hand states, so that eac h s ∈ S has
a la y er l ( s ) ∈ { 1 , . . . , T } , and a set of arcs A ⊆ { ( s i , s j ): l ( s j )= l ( s i ) + 1 } , for
i, j ∈ [1 , | S | ) . A comple t e path through the trellis is a no de sequence ( s 1 , . . . , s n ) ,
suc h that l ( s 1 )=1 , n = T , a n d ev ery pair { ( s k , s k +1 ) } n − 1
k =1 ∈ A .
If w e then assume the cost of a complete path can b e scored c ( s 1 , . . . , s n ) =
P n − 1
k c ( s k , s k +1 ) , searc hing for the optimal path S ∗ can b e approac he d as a cost
minimization problem, efficien tly solv able using the Viterbi algorithm [57]. Our
trellis form ulatio n presupp oses a 1-1 c o rres p ondence b et w een note/c hords and
action units. While that is not alw a y s the case, in this app endix w e tak e it as
true, for brevit y .
B.2 Starting P oin t: Melo dy-Only F re tter
W e base our mo d elling on the w ork of Hori and Saga y ama [18]. They prop ose
an approac h to predict fretting-hand fingers, giv en a monophonic tablature. F or
monophon y a f r etting = ⟨ str ing , f r et, f ing er ⟩ sp ecifies the triplet needed to
articulate a note. T riplet eleme n ts follo w the definitions giv en in A.1. Op en-
string frettings are undefined, and hence discarded. It is as sumed the tablature
has temp o markings, so that onset tim e, in seconds, is computable throughout.
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
26
Educational Profiling of Guitar T ablature 17
t
i
j
Figure. 6: Depiction of a trellis graph.
The fingering problem is then finding go o d f ing er elements for eac h pre-
scrib ed ( str ing , f r et ) pair. Hori and Sagay ama mak e a probabilistic form ulation,
mo delling trellis no des as states of a Mark o v pro cess, with transition probabil-
ities a ij ( dt )= P ( X t = f r etting j | X t − 1 = f r etting i , dt ) , where dt is the time
a v ailable to transition from i to j , computed as onset t − onset t − 1 . Computing
− l og ( a ij ( dt )) yields a cost. A small cost v alue means mo ving from one p osture
to the next is easy , and a large v alue means it is difficult.
T ransition Mo del. A t the time of writing, Hori and Saga y ama rep orted a
lac k of a v ailable data to learn mo del parameters, and sough t out a design that
w ouldn’t require it. T o that end, they formulate transition probabilit y as
a ij ( dt ) ∝ L ( dx ; 0 , dt ) × p φ ( f ing er j ) . (1)
In Eq. 1 ∝ means prop ortional, and the left hand side is normalized so that
the summation with resp ect to j equals 1 for all i . The righ t hand side consists
of t w o terms, mo delling kinematic and ph ysiological factors, resp ectiv ely . The
first term is the Laplace distribution, with dx denoting the length of hand dis-
placemen t along the fretb oard, and, as b efore, dt denoting the time a v ailable to
transition. Due to its hea vy tails, L ( · ) appro ximates the kind of zoned mo v emen t
c haracteristic of guitar pla ying, where small distances are t ypical, but large ones
(outliers) are not uncommon [17]. It also in tegrates dt as a scale parameter. As a
result, the mo del accoun ts for a (widely v alidated) notion: time constrain ts mak e
accurate mo v emen t less feasible [28, Ch. 7]. T o compute p osition estimates,
h ( f r etting )= f r et − f ing er + 1 , (2)
is used. Equation 2 estimates the index finger p osition, whic h in guitar pla ying
dictates the p osition of the hand. The form ula quan tifies a p osture guide, p opular
in guitar p edagogy , known b y the mnemonic "one-fret-p er-finger" [8]. With that,
the densit y function for L ( · ) can b e expressed as
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
27
18 M. Ro dríguez, A. Klapuri
L ( dx ; 0 , dt ) = 1
2 dt exp | h ( f r etting j ) − h ( f r etting i ) |
dt . (3)
The second term in Equation 1 is a ph ysiological prior, aiming to accoun t
for finger length and mobilit y constrain ts at arriv al point j . The prior is elicited
sub jectiv ely , with p ( φ = f ing er ) obtained through { f ( φ ) | φ ∈ { 1 , 2 , 3 , 4 } , f (1) =
0 . 4 , f (2) = 0 . 3 , f (3) = 0 . 2 , f (4) = 0 . 1 } . In that setting fretting candidates that
use the index finger are the easiest, and those using the pinky finger are the
hardest. The suggested setting agrees with findings in finger dexterit y [5, 22].
B.3 Generalizing: P olyphonic F retter
W e extend [18], aiming to w ards p olyphon y , where a f r etting = ⟨⟨ str ing , f r et,
f ing er ⟩ , ... ⟩ , with | f r etting |∈ [1 , 6] . As b efore, onset time is deriv ed from temp o
markings, and op en-string frettings are discarded. F or optimal path searc h w e
assume cost c ij is decomp osable in to mo v ement and posture comp onents, as
c ij = α c m ( s i , s j ) + (1 − α ) c p ( s j ) , (4)
where α is a mixing w eigh t. W e address c m and c p in turn, using the same frame-
w ork as b efore, with a mo del inferable from displacemen t, and one defined with
tunable w eigh ts. W e k eep the probabilistic frame, but only as an abstraction.
In practice that means w e don’t normalize la y ers so that they add up to 1, and
drop additiv e constan ts.
Mo v emen t Difficult y
W e need to accoun t for kinematic cases in v olving transitions from/to frettings
with m ultiple pressed strings. F ollo wing A.1, w e limit our scop e to the mo vemen t
of hand, fingers, and wrist. In order to accommo date the dimensionalit y increase
(in anatomical space), w e start b y restating our represen tation, and con tinue with
the approac h to cost computation.
F retboard Space. W e mo del the fretb oard as a N 2 motion plane, with ( str ing ,
f r et ) co ordinates. The space is discrete, non-negativ e, and equidistan t. W e can
then pro ject our guitarist’s hand on to that plane as an ordered p oin t set { q n } 4
n =1 ,
consisting of fingertip p ositions q ∈ N 2 . The set is ordered index to pinky , so
that n + 1 retriev es the f ing er corresp onding to p osition q n .
T o fully sp ecify a p oin t set, w e need the p ositions of b oth pressing and
idle fingers. As with 2, w e mak e use of the "one-fret-p er-finger" guide to ex-
trap olate p ositions. Let us assume a giv en fretting, con taining candidate fin-
gers for a prescrib ed set of ( str ing , f r et ) p ositions, and represen ted in our
searc h space b y a giv en no de s ∈ S . In that setting, idle fingers, if any , are
assigned ( str ing av , f r et idle ) p ositions. The str ing av co ordinate, constan t for
ev ery idle finger, is the a v erage o v er the strings in s , rounded to the nearest
string. If s sp ecifies one pressed string, eac h f r et idle is computed by recursing
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
28
Educational Profiling of Guitar T ablature 19
f r et next − f ing er idle + f ing er next , starting from the closest prescrib ed p osition.
If it sp ecifies man y , w e in terp olate b et w een prescrib ed fret p ositions, or cap to
the nearest one, as applicable.
Hand, wrist, and fingers are considered as separate mo ving comp onen ts. The
first is c haracterized using a single p osition ( str ing av , f r et index ) . The latter
subscript means the fret is that corresp onding to the index finger. As a reminder:
in guitar pla ying the index dictates the p osition of the hand. F or the wrist w e
use a scalar angular p osition. W e compute unit v ectors for the index and pinky ,
deriv e a rotation matrix, and then the p osition angle. F or the fingers we use the
full p oin t set.
Mo v emen t Difficult y . W e can generalise Equation 1 so that
p m ∝ Y
c
L ( d c ; 0 , dt ) Y
n
p φ ( f ing er n ) , (5)
calculates mo v emen t probabilit y , where d c denotes the length of displacemen t,
for the c th mo ving comp onen t, b ounded to 3 comp onen ts, and f ing er n denotes
the n th (pressing) finger at arriv al p oin t, b ounded to 6 (string) p ositions. T aking
− l og ( p m ) w e can express mo v emen t cost as
c m = dx
dt − X
n
log( p φ ( f ing er n ) )+ A log( dt )+ B , (6)
where A, B denote constan ts, and dx = d h + d f + d w linearises displacemen t
length, as the aggregate for hand, fingers, and wrist, resp ectiv ely . d h is mo delled
as taxicab distance, d f is mo delled as the a verage of individual finger displace-
men ts, using euclidean distance and quadratic mean, and d w is mo delled as
normalised angular displacemen t. There is no theoretical underpinning for these
c hoices, they w ere selected on the basis of exp erimen tal p erformance.
P arametrizations and Practicalities
1. If the transition in v olv es fretting pairs with only one pressed string, w e set p φ as
Hori et al. If the transition in v olv es mo ving from/to frettings with multiple pressed
strings, altering the setting so that f (1) = f (3) = 0 . 3 , and f (2) = f (4) = 0 . 2 ,
p erformed b est in our v alidation dataset.
2. W e found log-compressing d h impro v es generalisabilit y .
3. F or fingering prediction w e use the aggregate computed in Equation 6. F or the tasks
explored in the paren t pap er eac h term is returned separately . That includes all
displacemen t comp onen ts, as w ell as the bias term. All are in terpretable as sp eed
constrain ts, with the latter referred to as "rep ositioning difficult y" in Figure 4.
P osture Difficulty
W e mo del p osture difficulty taking in to accoun t t w o factors, that of holding a
p osture, and that of pressing strings. W e in tro duce our approac h in stages, first
defining a suitable represen tation space, then a mo del for finger-pairs, and lastly
a mo del for a full p osture.
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, Nov. 3-7, 2025
29
20 M. Ro dríguez, A. Klapuri
Elastic Space. W e use an elastic graph G = ( V , E ) to represen t the space of
p osture-related elastics. W e define G to b e undirected and simple. The graph
is elastic in that edges represen t ideal springs [53]. Let us then define a map
γ : V 7→ R 2 to em b ed the graph on to the guitar fretb oard, with ( str ing , f r et )
dimensions, measurable in cen timetres. An y fretting can hence b e represen ted
as a subgraph F with { v k } n =6
k =1 ∈ V and { ( v k , v k +1 ) } n − 1
k =1 ∈ E . Imp ortan tly , eac h
v ∈ F tak es a lab el equal to the candidate f ing er considered for that p osition,
and eac h e ∈ E adopts the equilibrium length of its corresp onding finger pair.
The equilibrium length is the spacing b et w een fingers, when the hand is in resting
p osition, whic h, for exp ository purp oses, w e assume to b e kno wn.
Finger-P air Discomfort. F or an y { v , u }⊆ F , w e can calculate the pair’s
p osture probabilit y as
p pair ∝ N ( d ε ; 0 , 1) × p σ ( v , u ) , (7)
where N is a Gaussian distribution, accoun ting for the probabilit y of holding
the p osture, and p σ is a prior accoun ting for the probabilit y of pressing the
strings. In b oth cases a lo w er probabilit y can b e in terpreted as higher discom-
fort, and consequen tly higher difficult y . The term d ε = ( ∥ v − u ∥ − r v u ) /r vu
denotes displacemen t-related strain, with v , u as the p osition v ectors asso ciated
with eac h no de in fretb oard space, and r v u the equilibrium distance for the fin-
gers corresp onding to the no de pair, expressed as a radius. The unit Gaussian
c hoice means the mo del assumes elastics under static equilibrium: displacemen ts
are not a function of time, and m uscle stresses arise from deformation. T aking
− l og ( p pair ) w e get
c pair = d ( v , u | r v u ) 2 + log( p σ ( v , u )) + C, (8)
where the first term on the righ t expresses d ε computation as a distance function,
and C denotes a constan t.
Finger-P air Strength. Let us assume a kno wn prior p φ defined for eac h finger
φ ∈ { 1 , 2 , 3 , 4 } . A lo w probabilit y means pressing a string, with the fingertip, is
difficult, and a high probabilit y means it is easy . On that grounding, w e define
the probabilit y for finger-pairs as
p σ ( v , u ) = ( p av g ( v , u ) f ing er v = f ing er u
| 1
str ing v − string u |× p av g ( v , u ) f ing er v = f ing er u
where p av g ( v , u ) = p p ( φ = f ing er v ) × p ( φ = f ing er u ) , and f ing er v ,u ∧ str ing v ,u
are, resp ectiv ely , the finger lab el and string co ordinate of the corresp onding sub-
script. F or f ing er v = f ing er u we ha v e a smo othing condition, to accoun t for
barre tec hnique. The smo other ev enly distributes fingertip probabilit y mass o v er
eac h pressed string.
Pro c. of the 17th International Symposium on CMMR, London, UK, No v. 3-7, 2025
30
Educational Profiling of Guitar T ablature 21
Hand Discomfort. F or eac h v ∈ F w e define a neigh b ourho od η ( v ) , containing
all pairings of v to ev ery v ertex u on a thinner string. F ormally , η ( v )= { u ∈ F |
str ing v > string u } . With that,
c p = X
n
max
η ( v n ) { U e ( v n , u ) + log( p σ ( v n , u )) } , (9)
calculates the o v erall discomfort of holding a p osture, while depressing m ul-
tiple strings, as the maxim um of individual finger-pair discomforts, where n is
b ounded to 6 pressing positions (strings). This form ulation uses p oten tial energy
U e , considering that
U e ( v , u ) ∝ d ( v , u | r v u ) 2 = w 1 ( ε f r et ) 2 + w 2 ( ε str ing ) 2 .
In the righ t-most expression w e fo cus on string and fret orien tations separately ,
with ε = ( l − L ) /L denoting normal strain, L length at rest, l deformation length,
and w 1 , w 2 tunable w eigh ts.
P arametrizations and Practicalities
1. Defining prior p φ as { f ( φ ) | φ ∈ { 1 , 2 , 3 , 4 } , f (1) = 0 . 4 , f (2) = 0 . 2 , f (3) = 0 . 2 ,
f (4) = 0 . 2 } , as w ell setting w 1 , w 2 = 2 , p erformed b est in out v alidation dataset.
2. W e mo del guitar fretb oard dimensions using manufacturer a v erages. W e settled for
a middle ground b et w een acoustic and electric guitar. The reader can refer to [41]
for metho ds to calculate fret p ositions in cen timetres, giv en fretb oard dimensions.
3. W e mak e a parametric mo del of resting p osition distances, for eac h finger pair.
The mo del consists of 12 fret-wise distances, and 12 string-wise distances. All
distances are sp ecified in centimetres. The model takes in to accoun t orienta tion.
That is, whic h finger presses the thick er string, and whic h the thinner, can make
the same distance b et w een strings comfortable in one case, and very uncomfortable
in another. Rest distances w ere first calculated using publicly av ailable data on
ph ysiological av erages for hands, as w ell as adduction / ab duction angle ranges,
and then fine-tuned exp erimen tally , on the basis of v alidation dataset performance.
4. F or fingering prediction w e use the aggregate computed in Equation 9, for the
task explored in the paren t pap er, w e return eac h comp onent in the summation
separately . The comp onen ts are in terpretable as p osture discomfort and pressing
discomfort, resp ectiv ely .
B.4 Outlo ok
The approac hes presen ted in this app endix answ er to situations of data scarcit y .
W e fo cus on domain-informed mo del design, and aim for low dimensionalit y .
One area in whic h w e (strongly) enforced that fo cus is on the representation of
the hand-guitar in teraction space, where we fa v oured planar appro ximations. In
this last section w e outline directions to extend the approac hes presen ted. W e
cen tre on w a ys to enric h the represen tation space, while still assuming data to
train/v alidate mo dels is limited.
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, Nov. 3-7, 2025
31
22 M. Ro dríguez, A. Klapuri
Assuming enough data for v alidation, but not for training
In our ev aluation exp eriments, w e sa w planar appro ximations p erformed w orse
for p osture difficult y mo delling. In short, finger-pair flexibilit y thresholds c hange
if a third (or fourth) finger is also pressing a string, and our mo del do es not
accoun t for that. One w a y to extend the planar approac h to tac kle that problem,
is to reform ulate neigh b ourho o d construction. That can b e done, for instance, by
framing it as a graph-partitioning problem, op erating on v ertex simplexes [23, 62].
A differen t alternative is to allo w for a three-dimensional represen tation. F or that
approac hes of inv erse kinematics [34], or in v erse dynamics [29], can b e tested.
These alternativ es come with a substan tial increase in computation time, as w ell
as a need for careful calibration (hand-guitar in teraction is a nic he application
area, so it is not unreasonable to exp ect unsatisfactory off-the-shelf p erformance).
Assuming enough data for v alidation, and limited data to train
T w o recent areas of researc h in mac hine learning lend themselv es w ell to the
problem at hand. One in tegrates kno wledge of Ph ysics, in the form of estab-
lished go v erning equations, using them as a w a y to regularize Neural Net w orks.
F or instance, [7] train a mo del to reconstruct the Y oung’s mo dulus field of a
heterogeneous ob ject, based only on a partial measuremen t of the ob ject’s ax-
ial displacemen t field. The other area fo cuses on dynamical systems, aiming to
learn state-space represen tations. F or instance, [6] prop osed an approac h to map
non-linear dynamics to a high-dimensional space, where the dynamics of state
and con trol are linearly separable. F or training, these approac hes w ould require
m usic scores aligned to a recording of p erformance mo v emen t, or force dynamics
(e.g. via motion capture data, or strain-gauge data, resp ectiv ely).
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, Nov. 3-7, 2025
32
Prop osal of a Singing Instruction Supp ort System
with an Inhalation Restriction F unction to
Demonstrate the Difference in Ph ysical Abilit y
b et w een a T eac her and Studen t in Online Lessons
Hiro y a Miura 1 , Akiy o Obara 2 , Hiroki W atanab e 2 ,
Y oshinari T ak ega w a 2 , and Keiji Hirata 2
1 RIKEN Cen ter for A dv anced In telligence Pro ject
2 F uture Univ ersity Hak o date
[email protected]
Abstract. In the teac hing of singing, breathing con trol is an imp or-
tan t comp onen t, the observ ation of whic h has con v en tionally required
a teac her’s long exp erience and refined p erception. Ho wev er, in online
lessons where visual information is limited, there is the problem that
ph ysical differences are hard to p erceiv e and it is difficult to obtain the
information necessary to pro vide instruction. In this researc h, whic h uses
a residual respiratory v olume visualization system based on previous re-
searc h, w e v erify the system’s usefulness for supp orting teac hing, as w ell
as implemen ting a new function to restrict inhalation v olume based on
a studen t’s lung capacit y , and ev aluating the effects of this function. Ex-
p erimen tal results implied that the prop osed system is useful as a to ol to
enable teac hers to recognize the ph ysical difference b et w een themselves
and their studen ts.
Keyw ords: T eac hing Supp ort · Singing Lecture · Visualization.
1 In tro duction
The prom ulgation of online lessons during the Co vid-19 pandemic led to the v eri-
fication of the effectiv eness of learning and teac hin g in suc h lessons in comparison
to face-to-face lessons. It has b een found that online lessons ha v e a higher degree
of studen t satisfaction than face-to-face lessons, and ha v e the effect of increasing
autonom y and concen tration [1–4]. In addition, it has b een p oin ted out that, b y
enabling remote lessons, online lessons can p oten tially help solv e the problems
of a lac k of teac hers and mismatc hes b et w een studen ts and teac hers [5, 6]. F rom
this, w e consider that online lessons are not simply a replacemen t for face-to-face
lessons, and will con tin ue to b e c hosen long after the end of the pandemic.
All righ ts remain with the authors under the Creativ e Commons A ttribution
4.0 In ternational License (CC BY 4.0).
Pro c. of the 17th In t. Symp osium on Computer Music Multidisciplinary Researc h,
London, United Kingdom, 2025
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, Nov. 3-7, 2025
33
Miura et al.
On the other hand, when compared to face-to-face lessons, in online lessons
it is difficult to obtain information needed for instruction [1]. Stabilized breath-
ing is essen tial for singing, but judging whether a studen t’s breathing pro cess,
whic h tak es place in ternally , is appropriate requires min ute observ ation and long
exp erience. Ho w ev er, in an online lesson, the teac her c hec ks the studen t’s p er-
formance on a monitor, making it difficult to observ e b o dy mo v emen ts when
compared to a face-to-face lesson. F urthermore, it is difficult to grasp a stu-
den t’s actual heigh t and c hest size. A ccordingly , it is also difficult to surmise
what v o cal v olume and breathing duration a studen t is capable of.
The aim of this researc h is to construct a system to pro vide teac hing supp ort
in online lessons b y visualizing residual respiratory v olume during singing. Here,
residual respiratory v olume refers to the amoun t of breath remaining when a
p erson sings. W e confer on the teac her a fresh viewp oin t for instruction, b y pro-
viding information ab out an individual studen t’s breathing, and v erify whether
this con tributes to teac hing supp ort. Also, w e construct a function (residual
respiratory v olume displa y function) that aids the comprehension of ph ysical
difference, b y comparing the resp ectiv e residual respiratory v olumes of the stu-
den t and the teac her. In particular, w e ev aluate the function that restricts the
v olume of the teac her’s inhalation (inhalation restriction function). This func-
tion enables the teac her to share the sensation of the studen t’s breathing and
exp erience the ph ysical difference b et w een them.
2 Related Researc h
Lã et al. [7] determined that real-time visual feedbac k of p erformance information
is effectiv e at enabling more efficien t mastery of singing. Sev eral mastery supp ort
systems that displa y acoustic information during singing ha v e actually b een
prop osed [8, 9]; more recen t in terfaces explicitly incorp orate breath guidance
and m ulti-sensor analyses for p edagogical feedbac k [10, 11]. Ho w ev er, there are
few cases that fo cus on breathing during singing.
In general, to measure breathing during singing the follo wing conditions m ust
b e met: (1) measuremen t can b e carried out while the singer is standing still, (2)
measuremen t m ust b e carried out safely , (3) measuremen t equipmen t is small and
easy to use, (4) measuremen t has a minimal effect on the mo v emen ts in v olv ed in
breathing. As a means that satisfies these requiremen ts, there is the metho d of
using a sensor to measure the b o dy mo v emen ts that accompan y singing. In this
metho d, the con traction of the c hest or ab domen, whic h accompanies breathing,
is measured using a device suc h as an inductor [12] or stretc h/p ositional sen-
sors with learned mappings from thoracoab dominal motion to v olume [13, 14].
The w a v eform of this b o dy mo v emen t and the w a veform of breathing ha ve been
sho wn to ha v e a strong relationship to the w a v eform measured b y a flo w meter
and to ob jectiv e respiratory patterns in trained singers [15]. F oundational w ork
further established the decomp osition of rib cage and ab dominal con tributions
to breathing v olume, whic h underpins mo dern respiratory inductance pleth ys-
mograph y and related approac hes [16]. In this researc h, w e decided to target the
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
34
Singing Instruction Supp or t System with an Inhalation Restriction F unction
Stretch Sensor
M5stack System Configuration Installation
Example
Fig. 1: W earing Breathing Measuremen t Equipmen t
teac her and displa y residual respiratory v olume to supp ort teac hing. W e v erify
whether this is effectiv e teac hing supp ort b y presen ting the residual respiratory
v olume to a teac her who is pro viding singing instruction. In addition, w e imple-
men t a new function to imp ose a restriction on the teac her in accordance with
the difference in the residual respiratory v olumes of the teac her and the studen t.
3 Prop osed System F unctions
In this researc h, our goal is to design and implemen t a teac hing supp ort system
that aims, in particular, to enable mastery of breathing in singing. Sp ecifically ,
this is realized b y displa ying residual respiratory v olume, obtained from infor-
mation gathered b y sensors attac hed to the b o dy .
3.1 Residual respiratory v olume estimation function
The structure of the breathing measuremen t system and an example of w earing
the device are sho wn in Figure 1. Stretc h sensors are used to measure c hest
mo v emen t. A stretc h sensor is a v ariable resistor that expands and con tracts
lik e rubb er. The sensors, whic h are fixed on to b elts that can b e attac hed and
remo v ed using buc kles, are attac hed to t w o places: the c hest and ab domen. The
sensor v alues are sen t to the master PC via an M5Stac k, a compact, all-in-one
dev elopmen t b oard designed for rapid protot yping and em b edded applications.
M5Stac k and the master PC are connected b y Blueto oth.
An outline of the residual respiratory v olume estimation part of the system
is presen ted in Figure 2. Reception of the data from the breathing measuremen t
system, and the implemen tation of the displa y screen sho wing graphs etc., w as
carried out in Pro cessing. Estimation of residual respiratory v olume w as imple-
men ted in Python. The sensor data receiv ed from the breathing measuremen t
system is sen t to the estimation part via consecutiv e text files. As prepro cessing
for estimation, the sensor data offset is adjusted.
The estimation part estimates residual respiratory v olume from the data
receiv ed from the sensors on the c hest and ab domen. The KNN (K Nearest
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
35
Miura et al.
Respiration
Measurement
Offset
Calculation
Writing
Reception
Reading
Predict
(KNN)
Writing
Reading
・
Processing
Display Section
Reception Section
Estimation Section
Offset.txt
rawData.txt
predictIO.txt
train.csv
Processing
Processing Python
UI screen
Fig. 2: Ov erview of Residual Respiratory V olume Estimation P art
Neigh b our) metho d is used for estimation. The KNN metho d selects k learning
data that are close to input data with unkno wn ob ject v ariables, and obtains
the a v erage v alue of ob ject v ariables of the learning data to estimate the ob ject
v ariable of the input data. In this researc h, the instan taneous sens o r v alues from
the c hest and ab domen are feature v alues and the residual res pira tory v olume is
the ob ject v ariable. F or the k parameter, w e cross v alidated the learning data for
eac h exp erimen t participan t and selected the v alue with the smalles t err or. Here ,
error refers to the RMS E (Ro ot Me a n Squared Error). The RMSE is obtained
b y the follo wing form ula:
R M S E = v
u
u
t 1
n
n
X
i =1
( y i − ˆ y i ) 2
– n : No. of samples
– y i : Estimated v alue
– ˆ y i : A ctual v alue
T o estimate r e sidual respiratory v olume w e correlated the sens o r v alues and
residual respiratory v olume and constructed datasets. A dataset is constructed
eac h time the system user puts on the sensor b elt. The user breathes in to a
spirometer, whic h measures lung c a p a c it y , and the sensor v alues during exha-
lation are recorded. In addition, the fluctuating scale on the spirometer is cap-
tured w ith a video camera. The measuremen t range of the spirometer is b et w een
1000mL and 7000mL, in degrees of 100mL. After measuremen t, w e correlated
the sensor v alues and residual respiratory v olume for ev ery 100mL.
3.2 Resi d ual respiratory v olume displa y fun cti on
This function displa ys the system user’s residual respiratory v olume in real time.
The left side of Figure 3 is an o v erview of the basic screen . The basic screen
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
36
Singing Instruction Supp or t System with an Inhalation Restriction F unction
waveform of residual respiratory volume
residual
respiratory (%)
expressing
residual respiratory
with color
musical score
Fig. 3: Screenshot of Residual Respiratory V olume Displa y F unction (Left) and
Screenshot of Inhalation Restriction F unction (Righ t)
comprises a part that plots the w a v eform of residual respiratory v olume, a part
that displa ys residual respiratory v olume b y p ercen tage and color c hange, and a
part that displa ys the m usical score. The w a v eform and the score are displa y ed
mo ving across the screen from righ t to left, while the p ercen tage and the color
c hange along with the latest residual respiratory v olume. This mak es it p ossible
to see the exten t to whic h the singer is using their breath at eac h p oin t on the
score.
3.3 Inhalation v olume restriction function
A visual-based con trol metho d that mak es the teac her adjust their inhalation
v olume while comparing it with the studen t’s residual respiratory v olume enables
restriction of breathing without c hanging the temp o or v olume of singing. By
lo oking at the target inhalation v olume and the state of their o wn inhalation, on
the screen, the teac her can understand the exten t to whic h they m ust restrict
inhalation in order to sing with a residual respiratory v olume equal to that of
the studen t.
T o restrict the teac her’s breathing, w e consider a metho d of limiting the max-
im um v alue of inhalation v olume to that of the studen t’s lung capacit y . Ho w ev er,
it is difficult to restrict inhalation just b y one’s o wn sensation. Therefore, w e im-
plemen ted a function to displa y a blac k-colored b elt on the basic screen. The
righ t side of Figure 3 presen ts an example of the displa y screen when this func-
tion is in use.
The ratio of the studen t’s lung capacit y to the teac her’s lung capacit y is tak en
as the standard for the maxim um inhalation rate, and the area on the graph
ab o v e the standard is colored blac k. F or example, if the teac her’s lung capacit y
is 5000mL and the studen t’s lung capacit y is 2500mL, the part of the graph ab o v e
the 50% mark will b e colored blac k. If the teac her’s inhalation v olume exceeds
the standard, the teac her is notified b y a w arning sound. By breathing in suc h
a w a y as not to exceed the standard, the teac her can artificially exp erience ho w
it feels to p erform with the residual respiratory v olume of the studen t.
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
37
Miura et al.
T able 1: Exp erimen t P articipan t Data
ID 1 2 3 4 5 6 7 8 9
Gender m f m m m f m f f
Lung capacit y (mL) 4400 2700 3550 5700 5100 2700 6500 1850 1500
Error (mL) 69.9 56.7 134.4 44.0 46.5 49.1 73.9 162.1 75.0
k 2 2 4 2 2 2 2 2 2
When singing, there are cases in whic h the timing means that one cannot
inhale a sufficien t amoun t in accordance with one’s lung capacit y . Ho w ev er, this
is due to a lac k of proficiency and not to a ph ysical limitation. A ccordingly , it
is though t that, rather than altering the standard with eac h breath, taking the
studen t’s lung capacit y as the maxim um inhalation v olume p ermitted to the
teac her is appropriate for the purp ose of this function.
4 Ev aluation of the Inhalation V olume Restriction
F unction
The inhalation v olume restriction function describ ed in section 3.3 w as ev aluated
in terms of whether breathing can b e restricted appropriately . The exp eriment
participan ts w ere presen ted with m ultiple standards of p ermitted maxim um in-
halation v olume, and ask ed to adjust their breathing accordingly . The function
w as ev aluated b y comparing the difference b et w een the standard and the actual
inhalation v olume in the cases of using and not using the function.
4.1 Exp erimen t conditions and pro cedure
Nine univ ersit y studen ts participated in this exp erimen t. Their information is
presen ted in T able 1. As exp eriment procedure, first, the lung capacity of the
participan ts w as measured once in order to construct learning data. Next, the
target v alue of maxim um inhalation v olume w as presen ted and breathing w as
measured. The measuremen t pro cess w as as follo ws. First, participan ts practiced
breathing sev eral times. Afterw ards, while listening to a metronome set to 60
BPM, participan ts inhaled once ev ery 4 b eats, a total of 5 times. There w ere
6 target v alues: 50%, 60%, 70%, 80%, 90%, and 100%. Eac h target v alue w as
tested with 5 inhalations p er participan t, and this 6-v alue blo c k w as p erformed
under t w o conditions (with and without the function), yielding 30 inhalations
p er condition (60 total p er participan t). Regarding the order in whic h the target
v alues w ere presen ted, 100% w as alw a ys presen ted first, after whic h the remaining
5 v alues w ere presen ted in random order.
After measuremen t for the six target v alues had b een carried out with the
participan ts not using the inhalation v olume restriction function, measuremen t
w as carried out in the same manner with the participan ts using the function. The
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
38
Singing Instruction Supp or t System with an Inhalation Restriction F unction
order of the target v alues w as the same as in the case of not using the function.
Analysis w as carried out b y the follo wing pro cess. First, the p eak v alue of the
fiv e inhalations for eac h standard w as obtained. The absolute p ercen tage error
w as calculated for eac h p eak. The absolute p ercen tage error is calculated with
the follo wing form ula.
δ = | T ar g et − y peak
V C × 100 |
– T ar g et : Standard v alue(%)
– y peak : Estimated v alue of p eak(mL)
– V C : Lung capacit y (mL)
When the n um b er of samples is n , the MAPE (Mean Absolute P ercen tage
Error) is calculated with the follo wing form ula.
M AP E = 1
n
n
X
i =1
δ i
4.2 Results and consideration
W e obtained the MAPE of inhalation v olume up to the presen ted standard v alue,
and compared the case of using the inhalation v olume restriction function and
the case of not using the function. The resulting v alues w ere 18.19% without
the function, and 14.24% with the function. The n um b er of samples w as 270.
Regarding the MAPE in the cases of using and not using the function, when
a t w o-sided t test w as carried out, the t-statistic w as 3.49, p< .01, indicating
a significan t difference. A ccordingly , it can b e said that breathing using the
inhalation restriction function decreases the error b et w een the actual inhalation
v olume and the target v alue.
The p ercen tage of samples in whic h the exp erimen t participan t’s breathing
w as b elo w the target v alue w as 56.7% in the case of without function and 74.1%
in the case of with function. When a c hi-squared test w as carried out, there
w as sho wn to b e a significan t difference ( χ 2 = 17.31, p< .01). In addition, w e
extracted only samples in whic h breathing surpassed the target v alue, and cal-
culated the MAPE of those samples. The resulting v alues w ere 18.31 (n=117) for
without function and 10.34 (n=70) for with function. When a t w o-sided W elc h’s
t test w as carried out, the t statistic w as 4.46, p< .01, sho wing a significan t differ-
ence. A ccordingly , it can b e said that the inhalation v olume restriction function
is effectiv e at inducing the user to limit their inhalation v olume so as not to
exceed the target v alue. F rom this, it can b e understo o d that the inhalation
v olume restriction function appropriately imp oses a restriction on the teac her in
corresp ondence with the studen t’s lung capacit y .
5 V erification of the T eac hing Supp ort Effect of the
Inhalation V olume Restriction F unction
A user study w as conducted to v erify the teac hing-supp ort effect of the inhala-
tion v olume restriction function. The exp erimen t recreated an online lesson and
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
39
Page 48 could not be converted
This page was skipped, but the rest of the document is available.
HTTPSConnectionPool(host='convert.identific.com', port=2053): Read timed out. (read timeout=60)
Singing Instruction Supp or t System with an Inhalation Restriction F unction
5.3 Results
When singing using the function that restricts inhalation v olume, the exp eri-
men t participan ts made observ ations, listed b elo w, from the viewp oin ts of (1)
in tro ducing a w areness that is differen t from usual, (2) ph ysical difference b e-
t w een oneself and the studen t, (3) usefulness regarding teac hing supp ort, and
(4) usefulness regarding online lessons. Note that the text in brac k ets within the
quotations is con text added b y the authors.
In tro ducing a w areness that is differen t from usual When a teac her sang
with restricted breathing, it w as found that they b ecame a w are of differen t as-
p ects from usual and made new realizations. F or example, commen ts suc h as
“ P aying atten tion (to how m uch I inhaled) made me realize that I prob-
ably need to use things lik e the surrounding m uscles more...I usually
dep end on breathing, but b ecause I couldn’t do that I felt lik e I w as
using other parts of m y b o dy .” (P articipan t 1)
sho w ed that the participan ts w ere con triving new w a ys to ac hiev e the b est p os-
sible v o calization under the breathing restrictions. This rev eals ho w, despite the
fact that the restriction actually imp osed b y the inhalation v olume restriction
function is only on the amoun t of inhalation, singing under this restriction in-
creases the burden on the m uscles used for v o calization as w ell as the accuracy
with whic h those m uscles are used. Th us it can b e said that this enables a teac her
to recognize anew exactly whic h parts of the b o dy are b eing used when singing.
In addition, the follo wing commen t w as obtained regarding the displa y of the
teac her’s o wn w a v eform.
“I realized that I don’t use m uc h breath, as I had a lot to spare. I w as
thinking I’d try to inhale once ev ery quarter rest (but I w as able to sing
without inhaling that often).” (P articipan t 2)
Lo oking at their o wn w a v eform enabled this teac her to recognize the gap b et w ee n
the breathing they had supp osed they w ere doing b efore and the w a y they w ere
actually breathing. This resulted in the teac her c hanging their b eha vior, limiting
their inhalation v olume in accordance with the fact that they had a surplus
residual respiratory v olume. This suggests that using the system pro vided new
realizations regarding a teac her’s o wn p erformance and the w a y they use their
o wn b o dy .
Ph ysical difference b et w een oneself and the student The exp erimen t
sho w ed that singing under restriction made the teac hers feel the ph ysical dif-
ference b et w een themselv es and the studen ts. The teac hers not only gained an
ob jectiv e understanding of the difference in n umerical v alues, but, b y singing in
a state of restriction, also actually exp erienced the strain felt b y the studen ts.
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
41
Miura et al.
“Because y ou sing while making sure not to inhale a lot, y ou really feel
that it is actually quite differen t from y our o wn normal sensation.” (P ar-
ticipan t 1)
“Because y ou’re made to sing under restriction, y ou can exp erience ho w
it feels for p eople who can’t sing prop erly or breath prop erly and it
reminds y ou of ho w y ou used to b e. P erhaps it enables y ou to rev ert to
ho w y ou w ere bac k when y ou couldn’t sing prop erly y et, lik e when y ou
w ere a c hild. P erhaps y ou can understand ho w it feels in the stage when
y ou can’t do it y et.” (P articipan t 2)
These commen ts indicate that limiting a teac her’s inhalation v olume enables
them to ph ysically exp erience the sensation of the studen ts, who ha v e a smaller
lung capacit y . F urthermore, the commen ts implied the p ossibilit y that the re-
striction pro vides a c hance for teac hers to remem b er the difficult y of singing
when they w ere b eginners themselv es.
Usefulness regarding teac hing supp ort W e found that using this system
mak es the teac her notice things that enable more substan tial instruction.
“I think that for p eople who ha v e b een singing for a while, rather than
p eople who ha v e just started learning to sing, it could b e in teresting to
use this (system) to c hec k, b y oneself, whether one’s o wn breathing is
correct.” (P articipan t 1)
This participan t suggests that b eing able to c hec k one’s o wn residual respira-
tory v olume is more appropriate for exp erienced singers than b eginners. This is
though t to b e b ecause exp erienced singers can judge from residual respiratory
v olume whether or not breathing is correct, and adjust their breathing accord-
ingly . This commen t implies that the approac h of this researc h, whic h, as stated
in Section 2, targets teac hers rather than the singing studen ts themselv es, is
appropriate.
“I m yself am capable of expression, so I w an t m y studen t to sing with
more expression...but I can’t sa y exactly ho w the studen t should pro duce
that m usical expression. When y ou exp erience suc h situations y ou gradu-
ally learn ho w to explain what to do. (The inhalation v olume restriction
function) is one means that could b e utilized in instruction regarding
breathing.” (P articipan t 2)
This commen t p oin ts out that usually the difference b et w een the m usical expres-
sion sough t after b y a teac her and the expression of whic h a studen t is capable
is reconciled as the teac her accum ulates teac hing exp erience. This implies that
using the system can mak e teac hers a w are of a difference that normally tak es a
long time to recognize.
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
42
Singing Instruction Supp or t System with an Inhalation Restriction F unction
“(My w a v eform) do esn’t mo v e while I’m holding a note, so I though t
that I w asn’t really using breathing. I think that’s the biggest difference
(b et w een me and the studen t). I think I just sing using force. That’s all
I use to pro duce a sound.” (P articipan t 2)
This commen t demonstrates a participan t using the visual breathing informa-
tion, something that cannot normally b e seen, to compare with the studen t and
consider the factors b ehind the difference b et w een them.
Usefulness regarding online lessons The exp erimen t participan ts considered
the system useful for online lessons, as it pro vides supplemen tary information
that is difficult to attain b y sigh t.
“(Ev en in a remote lo cation) y ou can feel (the ph ysical difference)... The
blac k band and so on (on the screen) sho w y ou that there really is a
difference, but without that, ev en if y ou knew that the studen t’s lung
capacit y (migh t b e smaller than y our o wn), I don’t think y ou w ould or
could b ear it in mind (as m uc h as when the standard is displa y ed).”
(P articipan t 1)
This commen t sho ws that, in online lessons, although it is not p ossible to com-
pare with the lung capacit y of a studen t in a remote lo cation, ha ving the standard
that represen ts residual respiratory v olume as a function to restrict inhalation
enables the teac her to feel the ph ysical difference, ev en from a remote lo cation.
“(Online lessons are on the increase but) b ecause y ou can’t hear the ra w
sound, I didn’t really trust them, but I think regarding breathing, lik e
in this researc h, and things relating to ho w y ou use y our b o dy , if y ou
can lo ok at those things on a sp ecialized in terface I think it’s go o d.”
(P articipan t 2)
This participan t w as sk eptical ab out the effects of online lessons when compared
to face-to-face lessons, regarding instruction and impro v emen t. Nev ertheless, the
participan t ac kno wledged that using this system, whic h is sp ecialized for observ-
ing breathing while singing, can partially solv e the problem of the difficult y of
pro viding instruction in online lessons.
5.4 Consideration
F rom these results, it is though t that the inhalation restriction function is effec-
tiv e in terms of pro viding the user with new insigh ts. Sp ecifically , singing under
restriction causes the user to re-assess ho w m uc h breath they are using and when,
whic h promotes c hanges in b eha vior to reduce w asted breath. In addition, the
restriction led to re-recognition of the w orking of the m uscles and the accuracy
with whic h they are used, asp ects that w ere previously sub conscious.
Regarding the effect of enabling teac hers to feel the ph ysical difference b e-
t w een themselv es and their studen ts, it w as found that b eing restricted made
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
43
Miura et al.
it easy to exp erience a feeling of not b eing able to use as m uc h breath as one
had imagined, giving the teac hers a more solid understanding of their studen ts’
strain and limitations. In addition the exp erimen t results implied the p ossibilit y
that teac hers can recall the difficult y of p erforming as a b eginner. F or an ad-
v anced teac her, recreating an underdev elop ed p erformance is not easy . Ho w ev er,
it is p ossible that using this function will enable man y more teac hers to remem-
b er ho w they used to b e as b eginners. This connects to teac hers b eing able to
iden tify more sp ecific p oin ts that studen ts need to impro v e on, and suggesting
effectiv e instruction metho ds.
The exp erimen t results implied that, in online lessons, whic h are limited b y
the difficult y of directly grasping the studen t’s ph ysical state, the inhalation
v olume restriction function mak es it p ossible, ev en online, to pro vide instruction
fo cused on the studen t’s breathing and the w a y they use their b o dy . F urthermore,
the usefulness of this function has the p oten tial to motiv ate teac hers who ha v e
only ev er giv en face-to-face lessons to try giving online lessons, th us con tributing
the further p opularization of online lessons.
F rom the ab o v e, it is though t that the inhalation v olume restriction function
is useful for teac hing supp ort. In particular, b y pro viding new v alues in online
lessons and the instruction of c hild studen ts who differ greatly from their teac hers
in terms of ph ysicalit y .
6 Conclusion
In this researc h w e constructed a singing supp ort system with an inhalation v ol-
ume restriction function, with the aim of supp orting teac hing in online singing
lessons. W e implemen ted a new function that restricts a teac her’s inhalation v ol-
ume b y displa ying a gauge, to enable teac hers to c hec k the ph ysical difference
b et w een themselv es and studen ts. When teac hers breathed while using this func-
tion, the difference b et w een the teac hers’ inhalation v olume and the target v alue
w as smaller than when not using the function, sho wing that it is effectiv e at
leading users to limit inhalation. When this function w as used b y teac hers with
exp erience of teac hing singing, it w as found that the teac hers w ere able to share
the studen ts’ sensations. In addition, it w as implied that the system con tributes
to the conception of new instruction metho ds suited to the studen t.
It is though t that, hereafter, the prop osed system could b e impro v ed through
deplo ymen t in actual classro om settings and through b eing used b y teac hers with
greater teac hing exp erience. F urthermore, conducting an exp erimen t in whic h
the system is used while con v ersing with a studen t in real time w ould enable
ev aluation of whether the teac hing supp ort leads to the impro v emen t of the
studen t. In future w ork, w e will also examine in tegration with common online
lesson platforms, ev aluation in group-lesson scenarios, and longitudinal studies
to assess learning outcomes and usabilit y at scale.
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
44
Singing Instruction Supp or t System with an Inhalation Restriction F unction
References
1. Marcos San tos, T eac hing v oice online: reflection and a practical advice, Bulletin of
the T ransilv ania Univ ersit y of Braso v. Supplemen t Series VI I I: P erforming Arts,
V ol.13, No.62, pp.267–276 (2021).
2. Qiangfu Y u, F actors Influencing Online Learning Satisfaction, F rontiers in Psyc hol-
ogy , V ol.13, pp.852360 (2022).
3. W enxia Meng, Lili Y u, Chao Liu, Na P an, Xiao y an P ang, and Y an Zhu, A system atic
review of the effectiv eness of online learning in higher education during the CO VID-
19 pandemic p erio d, F ron tiers in Education, V ol.8, Article 1334153 (2023).
4. Zahra Mohammadi Zenouzagh, Wilfried A dmiraal, and Nadira Saab, Learner au-
tonom y , learner engagemen t and learner satisfaction in text-based and m ultimo dal
computer mediated writing en vironmen ts, Education and Information T ec hnologies,
V ol.28, No.11, pp.14283–14323 (2023).
5. Luis M. Dos San tos, Online learning after the CO VID-19 pandemic: Learners’ mo-
tiv ations, F ron tiers in Education, V ol.7, 879091 (2022).
6. Xiao c hen Zhang and Junk ai Zhang, A new model of vocal music teac hing in the
con text of In ternet distance learning, In ternational Journal of W eb-Based Learning
and T eac hing T echnologies (IJWL TT), V ol.19, No.1, pp.1–12 (2024).
7. Filipa M. B. Lã and Mauro B. Fiuza, Real-Time Visual F eedbac k in Singing P ed-
agogy: Curren t T rends and F uture Directions, Applied Sciences, V ol.12, No.21,
p.10781 (2022).
8. T omo y asu Nak ano, Masatak a Goto, and Y uzuru Hiraga, MiruSinger: A Singing Skill
Visualization In terface Using Real-Time F eedbac k and Music CD Recordings as
Referen tial Data, Pro ceedings of the IEEE In ternational Symp osium on Multimedia
W orkshops (ISM W orkshops), pp.75–76 (2007).
9. Dorien Hopp e, Makik o Sadak ata, and Peter Desain, Dev elopmen t of real-time visual
feedbac k assistance in singing training: a review, Journal of Computer Assisted
Learning, V ol.22, No.4, pp.308–316 (2006).
10. Ziyue Piao and Gus Xia, Sensing the Breath: A Multimo dal Singing T utoring
In terface with Breath Guidance, Pro ceedings of the In ternational Conference on
New In terfaces for Musical Expression (NIME), (2022).
11. Eft yc hios Angelakis, Natalia K otsani, and Anastasia Georgaki, T o w ards a Singing
V oice Multi-Sensor Analysis T o ol: System Design, and Assessment Based on V ocal
Breathiness, Sensors, V ol.21, No.23, p.8006 (2021).
12. Mattias Heldner, Marcin Wło darczak, P eter Branderud, and Johan Stark, The Re-
spT rac k system, SEF OS 2019 (Sto c kholm), Pro ceedings/Abstracts, pp.1–3 (2019).
13. Akio Y amamoto, Hiro yuki Nak amoto, et. al., Developmen t of a Measurement De-
vice Using a Sheet Stretc h Sensor for Chest W all Motion, The Japanese Journal of
Rehabilitation Medicine, V ol.55, No.4, pp.348–357 (2018).
14. Rém y Dumond, Stev en Gastinger, Hala Ab dul Rahman, Alexis Le F auc heur,
P atrice Quin ton, Haitao Kang, and Jacques Prioux, Estimation of respiratory v ol-
ume from thoracoab dominal breathing distances: comparison of t wo models of ma-
c hine learning, Europ ean Journal of Applied Ph ysiology , V ol.117, pp.1533–1555
(2017).
15. Sauro Salomoni, W olb ert v an den Ho orn, and Paul Hodges, Breathing and Singing:
Ob jectiv e Characterization of Breathing P atterns in Classical Singers, PLOS ONE,
V ol.11, No.5, e0155084 (2016).
16. Kazutoshi K onno and James Mead, Measuremen t of the separate v olume c hanges
of rib cage and ab domen during breathing, Journal of Applied Ph ysiology , V ol.22,
No.3, pp.407–422 (1967).
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
45
GuitarFlo w: Realistic Electric Guitar Syn thesis
F rom T ablatures via Flo w Matc hing and St yle
T ransfer
Jac kson Loth 1 [0009 − 0002 − 8709 − 1218] , P e dro Sarmen to 1 , 2 [0000 − 0002 − 4518 − 0194] ,
Mark Sandler 1 [0000 − 0002 − 5691 − 8107] , and Mathieu
Barthet 1 , 3 [0000 − 0002 − 9869 − 1668]
1 Cen tre for Digital Music, Queen Mary Univ ersity of London
{j.j.loth,p.p.sarmento,mark.sandler,m.barthet}@qmul.ac.uk
2 Music.AI
3 Aix-Marseille Univ CNRS PRISM
Abstract. Music generation in the audio domain using artificial in tel-
ligence (AI) has witnessed steady progress in recen t y ears. Ho w ev er for
some instrumen ts, particularly the guitar, con trollable instrumen t syn-
thesis remains limited in expressivit y . W e in tro duce GuitarFlo w, a mo del
designed sp ecifically for electric guitar syn thesis. The generativ e pro cess
is guided using tablatures, an ubiquitous and in tuitiv e guitar-sp ecific
sym b olic format. The tablature format easily represen ts guitar-sp ecific
pla ying tec hniques (e.g. b ends, muted strings and legatos), whic h are
more difficult to represen t in other common m usic notation formats suc h
as MIDI. Our mo del relies on an in termediary step of first rendering
the tablature to audio using a simple sample-based virtual instrumen t,
then p erforming st yle transfer using Flow Matc hing in order to transform
the virtual instrumen t audio in to more realistic sounding examples. This
results in a mo del that is quick to train and to perform inference, requir-
ing less than 6 hours of training data. W e presen t the results of ob jectiv e
ev aluation metrics, together with a listening test, in whic h we sho w sig-
nifican t impro v emen t in the realism of the generated guitar audio from
tablatures.
Keyw ords: Flow matc hing · St yle transfer · Guitar · Synthesis · Audio
effects.
1 In tro duction
Recen t adv ances in generativ e audio systems ha v e giv en rise to impressiv e text-
to-m usic systems that allow users to generate full songs from simple text prompts
[14] [15]. While this is v ery in teresting from a science and tec hnology p ersp ectiv e,
it is arguably lac king as a creativ e to ol due to a lack of fine-grained con trol
o ver the m usic missing. Music generation systems with full con trol o v er the
notes do exist, but t ypically in the sym b olic domain [39] [30], requiring a w ay
to transform a sym b olic m usic score in to audio. While this can b e ac hiev ed
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
46
2 J. Loth et al.
through virtual instrumen ts to v arying degrees of success, guitars are particularly
expressiv e instrumen ts whic h are difficult to represen t through the standard
Musical Instrumen t Digital In terface (MIDI) [33]. Guitar tim bre itself can ev en
b e p erceiv ed differen tly when p erformed in differen t pla ying st yles [32]. Previous
attempts at syn thesising expressiv e m usic instrumen ts ha v e largely fo cused on
MIDI or represen tations deriv ed from MIDI [45] [24].
This pap er explores the task of syn thesising electric guitar b y pro cessing
syn thetic audio rendered from an expressiv e sym b olic m usical represen tation.
Instead of MIDI, w e use guitar tablatures (see Figure 1) to represen t the m usi-
cal con ten t of a desired audio syn thesis. This allo ws us to incorp orate expressiv e
pla ying tec hniques suc h as slides, b ends, hammer-ons, etc, something that other
m usic instrumen t syn thesis mo dels struggle to do without additional condition-
ing mec hanisms. W e also in tro duce an in termediary step of first rendering a
guitar tablature to audio using a quic k and simple sample-based virtual instru-
men t. Our mo del p erforms st yle transfer on this simple audio rendering and
transforms it in to more realistic sounding audio. W e use ‘st yle‘ here to refer to
v arious p erformance-related qualities whic h distinguish natural and more syn-
thetic p erformances. The st yle transfer is ac hiev ed using Flo w Matc hing [28], a
recen t generativ e mo deling paradigm. This metho d allo ws us to greatly simplify
the complexit y of the training and inference pip eline, requiring less data and less
time to train.
The con tributions of this pap er are summarized as: (1) GuitarFlo w, a no v el
mo del and metho dology for realistic electric guitar syn thesis from guitar tab-
latures using Flo w Matc hing and st yle transfer; (2) an ev aluation of the mo del
using b oth ob jectiv e metrics and a sub jectiv e listening test whic h sho ws the suc-
cess of GuitarFlo w in transforming audio rendered with a virtual instrumen t to
sound more realistic; (3) a public rep ository 4 of co de to allo w other researc hers
to replicate and extend the researc h. This pap er demonstrates the p oten tial of
the tec hnique in greatly lo w ering data and computational requiremen ts when
training generativ e audio mo dels.
2 Bac kground
2.1 Guitar T ablatures
Guitar tablatures (refer to Figure 1), also kno wn as tabs, are sym b olic repre-
sen tations of guitar m usic and ha v e seen increased atten tion in recen t y ears due
to their abilit y to easily represen t guitar-sp ecific expressions [38]. In con trast to
MIDI, whic h simply represen ts a note’s pitc h and v elo cit y o v er time, tabs repre-
sen t b oth the fret and string n um b er of a guitar. They can also supp ort expressiv e
pla ying tec hniques suc h as b ends, hammer-ons, pull-offs, strum directions, and
more. T abs ha v e seen an increase in atten tion from the MIR researc h comm u-
nit y in the past few y ears in areas suc h as guitar tablature generation [39] [40],
automatic guitar transcription [46] [8] and tablature prediction from MIDI [11].
4 h ttps://gith ub.com/Jac kJamesLoth/GuitarFlo w
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
47
GuitarFlo w 3
Fig. 1. Example o f a guitar tablature, from the Guitar Pro editing soft w are.
2.2 St yle T ransfer
Broadly sp eaking, st yle transfer in v olv es transforming the “st yle” of a signal
while preserving the “con ten t” of signal. The task is p opular in the image do-
main [16] [49] . A common assumption is that the audio of an instrume n t can
b e br ok en up in to “con ten t”, referring to the pitc h, le ngth and arguably loudness
or in tensit y of the notes, and “st yle”, referring to the actual sound of the in-
strumen t and expressiv eness of the p erformer pla ying those notes. Some m usical
st yle transfer w orks us e a GAN-based approac h, training a generator to c o n v ert
an audio input to the desired tim bre [22] [48]. Others use an auto enco der to
pro duce disen tangled represen tations of c o n ten t and st yle [35] [1]. As with gen-
erativ e tasks recen tly , Diffusion has also b ecome a p opular c hoice [6] [21 ], while
differen tial digital signal pro cessing (DDSP) [13] pro vides an alternativ e to all
of these mo dels b y allo wing trainable DSP functions.
2.3 Music Instrumen t Syn thesis
Instrumen ts suc h as guitar ha v e traditionally b een sim ulated b y mo deling the
ph ysics of the instrumen t and strings [27]. More recen tly , mo dels su c h as W a v eNet
[43] op ened the do or to neural generation conditione d on sym b olic m usical rep-
resen tations [19] [25]. DDSP has seen a lot of use creating audio syn thesis mo d-
els [2] [41] due to its flexibilit y . Despite its large data requiremen ts, diffusion
has also b ecome a p opular metho d for neural audio generation. While m uc h
w ork is fo cused on generating full song mixes [15] [14 ] , syn thesising sym b olic
m usical represen tations suc h as MIDI allo ws m uc h finer m usical con trol o v er the
syn thesised audio. Ha wthorne et al. [18] trained a T ransformer-based Diffusion
mo del for m ulti instrumen t syn thesis, whic h w as in turn expande d o n b y Kim
et al. [24] for sp ecific ally acoustic guitar. While the mo dels sho w ed promising
results, the y w ere held bac k b y lac k of annotated data. Maman et al. [34] used
automatic MIDI transcription systems to help address this problem while adding
conditioning on ins tr u men t tim bre. All of these metho ds ha v e the dra w bac k of
requiring large amoun ts of data and necessitating significan t compu ta tional re-
sources to train. F or example, the training in [3 4] to ok 350 hours o v er three
p o w erful Nv id ia A100 GPUs.
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
48
4 J. Loth et al.
3 Metho dology
3.1 Flo w Matc hing
Flo w Matc hing (FM) [28] is a generativ e mo deling paradigm whic h pro vides
a w a y to efficien tly train contin uous normalizing flo ws (CNF). It has recen tly
pic k ed up in terest among audio researchers [17] [44]. In tuitiv ely , the idea is to
learn the direction to mo v e a p oin t sampled from some data distribution o v er
time in order to transform it in to a sample from a differen t data distribution.
More formally , w e would lik e to transform a data distribution p 0 ( x 0 ) to another
data distribution p 1 ( x 1 ) b y learning a v elo cit y field v t whic h p oin ts in the di-
rection in whic h w e w ould like to transform a point at a giv en time t ∈ [0 , 1] .
Giv en a v ector field u t whic h generates a target probability flo w p t , w e learn
a time-dep enden t flo w whic h matc hes p t . T o do this, we consider the ordinary
differen tial equation (ODE):
dx t = v t ( x t ) dt (1)
Flo w Matc hing attempts to learn this time-dep enden t v ector field v t b y min-
imising a conditional Flo w Matc hing (CFM) loss [42]. This is generalized to
arbitrary source and target distributions, making CFM an in tuitiv e and p o w er-
ful metho d for transforming data distributions. W e adopt rectified flo w [29] b y
simply sampling x 0 and x 1 from our data and calculating the in terp olated p oin t
x = (1 − t ) x 0 + tx 1 and target v elo cit y u t ( x | z ) = x 1 − x 0 . W e then train a neural
net w ork θ b y minimising the following:
L = || v θ ( t, x ) − ( x 1 − x 0 ) || 2
2 (2)
Once w e ha v e learned our v θ ( x, t ) , we can use an ODE solv er [3] to appro xi-
mate the solution to Equation 1 o v er some set of discrete time steps to transform
source data to the target distribution.
3.2 Metho d
Our approac h fo cuses on mo deling the direct input (DI) signal, the ra w output
of an electric guitar, rather than the amplified and distorted signal whic h is
t ypically heard in recordings. This allo ws us to simplify the task and offload
distortion pro cessing to digital amplifier mo dels [47] [4].
Our training pip eline in volv es first rendering syn thetic audio from a guitar
tablature using a virtual instrumen t. This and corresp onding real guitar DI
audio are then enco ded in to a latent space using Music2Laten t [36], a pretrained
auto enco der. The audio is broken in to four second c h unks in order to k eep the
audio length consisten t and k eep the data size reasonably small. Our mo del
then learns a mapping from a laten t distribution p 0 ( x 0 ) of syn thetic guitar DI
rendered b y our virtual instrumen t to a laten t distribution p 1 ( x 1 ) of real guitar
DI. W e can then use x 0 ∼ p 0 and x 1 ∼ p 1 , corresp onding to the syn thetic
and real guitar DI resp ectiv ely , for training. This mak es the mo del effectively
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
49
GuitarFlo w 5
a one-to-one st y le transfer mo del, as the real a n d syn thetic guitar recordings
are required to con tain exactly the same m usical con ten t (i.e. notes, timings ,
expressiv e tec hniques, etc.).
3.3 Mo d e l
Our mo del, titled GuitarFlo w, is primarily based on the UNet [37] arc hitecture,
consisting of f our do wns ampling and upsampling la y ers 5 . T o condition the mo del
on t , x 0 and t are c oncatenated together in the feature lev el. The mo del then
outputs predicted flo w v elo cities v t , whic h are used alongside real flo w v elo cities
u t to calculate mean squared error (MSE) loss. Figure 2 sho ws the training and
inference pip eline.
Fig. 2. T raining and in f erence using G u itarF l o w.
F or inferenc e, an ODE solv er [3] is us ed to in tegrate across the learned v elo c-
it y field across 100 discrete tim e steps. W e found the Dormand-Prince metho d [9]
to yield significan tly higher qualit y results compared to a more standard Euler
ODE solv er.
4 Exp erimen ts
4.1 Data and T raining
The GO A T data s et [31] w as used to train and ev aluate our mo del. This dataset
con tained ro u g h ly 5.75 hours of real guitar DI audio recorded at 44.1kHz and
5 h ttps://g ith ub.com/clemk oa/u-net
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
50
6 J. Loth et al.
the corresp onding tablature annotations. The p erformances w ere p erformed b y
three guitarists on four differen t guitars, largely co v er ro c k and indie ro c k genres
and include a wide v ariet y of expressiv e tec hniques. The guitar tablatures w ere
rendered in to audio using the Realistic Sound Engine (RSE) virtual instrumen t
in Guitar Pro 7 6 .
The data w as first prepro cessed using Music2Laten t in to individual four-
second audio c h unks prior to training. The mo del w as then trained on a NVIDIA
R TX A5000 GPU for 50 ep o c hs (3,900 total steps) with a batc h size of 64 and
learning rate of 0 . 0001 , whic h to ok roughly 12 min utes in total. While this is
a surprisingly small n um b er of training steps, w e found it to b e sufficien t for
con v ergence.
4.2 Ev aluation
Both the Guitar Pro and GuitarFlo w audio are ev aluated against the real DI
audio in the test split of GO A T. Because electric guitar is almost alw a ys heard
through the distortion of a guitar amplifier (whether subtle or hea vy distortion)
rather than the pure DI, all of the ev aluation audio w as separately rendered
using a digital guitar amplifier 7 plugin. W e then ev aluate b oth the DI and am-
plifier conditions of the ev aluation audio. Through this, w e test t w o h yp otheses:
H 1 : audio from GuitarFlo w sounds more realistic than the original Guitar Pro
rendered audio as DI; and H 2 : audio from GuitarFlo w sounds more realistic
than the original Guitar Pro rendered audio when rendered using a distorted
guitar amplifier.
F réc het Audio Distance (F AD) [23] is a commonly used metric to measure
audio similarit y , while Kernel Audio Distance (KAD) [5] w as recen tly prop osed as
a more flexible alternativ e whic h do es not rely on a normalit y assumption of the
audio em b eddings. W e calculate b oth F AD and KAD to ev aluate the closeness of
the syn thetic and transformed audio to the real guitar DI recordings, using the
real audio as the paren t distribution. Sev eral differen t em b eddings [20] [12] [26]
[10] [7] are used in these calculations. F ollo wing Ha wthorne et al. [18], w e also
calculate the reconstruction em b edding distance b et w een the real DI audio and
b oth the Guitar Pro audio and GuitarFlo w output audio. This giv es a b etter
metric of ho w closely the actual audio is to the in tended target, unlik e F AD
and KAD whic h w ork o v er a distribution of audio em b eddings. This distance is
obtained b y calculating the F rob enius norm for b oth em b eddings, and a v eraged
o v er all time frames.
A simple listening test w as also conducted with 16 participan ts (13 male, 3
female, with an a v erage age of 27.8 y ears) to ev aluate the mo del outputs. W e
selected 15 four-second outputs whic h co v ered single notes, c hords, and pla ying
tec hniques suc h as b ends and m uted strings. P articipan ts w ere ask ed to rate eac h
6 h ttps://www.guitar-pro.com/blog/p/14545-signature-sounds-explained-guitar-pro-
7
7 The “crunc h” amplifier and default cabinet IR w ere used from
h ttps://neuraldsp.com/plugins/arc het yp e-nolly
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
51
GuitarFlo w 7
T able 1. Comparison of the original syn thetic Guitar Pro virtual instrumen t (GP)
and GuitarFlo w. W e calculate F réc het Audio Distance (F AD), Kernel Audio Distance
(KAD) and reconstruction distance (Recon. Dist.) on b oth the DI and the DI rendered
through a guitar amplfier. Best v alues for each metric in eac h ro w mark ed in b old.
Condition Em b edding Mo del F AD ↓ KAD ↓ Recon. Dist. ↓
GP GuitarFlo w GP GuitarFlo w GP GuitarFlow
DI
V GGish 2 . 71 2 . 35 8 . 16 5 . 72 0 . 79 0 . 90
CLAP 207 . 26 120 . 74 17 . 65 4 . 65 0 . 09 0 . 08
P ANNs 16 . 83 8 . 48 17 . 30 4 . 84 0 . 96 0 . 66
EnCo dec 39 . 12 14 . 74 22 . 89 9 . 70 1 . 82 1 . 40
Op enL3 64 . 71 37 . 98 8 . 20 3 . 33 1 . 56 1 . 72
Amplifier
V GGish 1 . 96 0 . 77 4 . 17 5 . 28 1 . 09 1 . 10
CLAP 80 . 25 33 . 41 8 . 73 8 . 61 0 . 05 0 . 04
P ANNs 13 . 90 5 . 18 6 . 51 9 . 99 1 . 00 0 . 74
EnCo dec 11 . 76 4 . 39 8 . 20 13 . 77 1 . 57 0 . 93
Op enL3 48 . 93 20 . 62 5 . 29 5 . 06 1 . 77 1 . 42
in terms of realism, whic h w e define as ho w close the audio resem bles a h uman
pla ying guitar. All audio examples w ere normalised to -9dB RMS prior to the
amplifier in order to ensure that the gain staging w as consisten t. P articipan ts
w ere paid with a £10 Amazon v ouc her.
Finding a baseline mo del to compare against is difficult as most syn thesis
mo dels use MIDI as a input sequence whic h cannot replicate an y of the n umer-
ous expressiv e tec hniques, creating an unfair comparison with GuitarFlo w. Man y
common st yle transfer approac hes are also difficult to compare due to mono-
phonic constrain ts [1] [13] or large computational and data requiremen ts [21] [6].
Since our primary goal in this w ork is to establish the feasibilit y of flo w-matc hing-
based transfer in the con text of sym b olic-to-real audio syn thesis, w e fo cus on
ev aluating the impro v emen t that the mo del mak es compared to the in termedi-
ary syn thetic audio within that pip eline. W e leav e a comprehensiv e comparison
with other st yle transfer metho ds as an imp ortan t direction for future w ork.
5 Results and Discussion
5.1 Ob jectiv e Metrics
The F AD, KAD and reconstruction distances are presen ted in T able 1. Gui-
tarFlo w generally p erforms m uch better than Guitar Pro, particularly in the
DI condition. The amplifier condition is a bit more mixed, with GuitarFlo w
struggling to impro ve on Guitar Pro in the KAD metric. Ho w ev er, the F AD and
reconstruction distance results are still strong in the amplifier condition. It is
p ossible that the clipping and distortion from the amplifier remo v ed some infor-
mation from the DI whic h affected the calculated em b eddings in a w a y that the
KAD metric is sensitiv e to.
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, Nov. 3-7, 2025
52
8 J. Loth et al.
5.2 Listening T est
The listening test mean opinion score (MOS) results are sho wn in Figure 3.
F riedman tests rev ealed signifi can t differences b et w een groups in b oth the DI
( χ 2 (3) = 267 . 114 , p<. 001 ) and amplifier ( χ 2 (3) = 107 . 153 , p<. 001 ) condi-
tions. A pairwise Wilco xon signed-rank test w as then p erformed for b oth con-
ditions using a Bonferro n i- c o rrec ted α = 0 . 0167 . F or the DI stim uli, significan t
differences w e r e fo u nd b et w een the real and GP ( p < . 001) groups and the real
and GuitarFlo w ( p<. 001) groups. F or the a m plifier stim u li, significan t dif-
ferences w ere found b et w een the real and GP ( p<. 001) groups, the real and
GuitarFlo w ( p<. 001) groups and the the GP and GuitarFlo w ( p<. 001) groups.
Fig. 3. Bo xp ot with mean indicators for the MOS results of b oth the DI and amplifier
conditions of the l iste n ing test. Mean scores i n eac h group m ark ed b y whi te dots.
While the GuitarFlo w mo d el is only barely p erceiv ed as more realistic than
the Guitar Pro audio in the DI scenario ( H 1 ), it is p erceiv ed as significan tly
more realistic when run through the distortion of a guitar amplifier ( H 2 ). It is
in teresting that the realism MOS w ould increase after the distortion is applied
to the signal, as this distortion clips the signal and lose s information. This also
seems to con trast the KAD results of the amplifier condition, though th is discrep-
ancy could b e due to a limitation of the KAD metric, the em b eddings used, or
simply a difference b et w een what eac h result is measuring. W e als o see the MOS
scores of the Guitar Pro and GuitarFlo w stim uli increas e and the real stim uli
decrease in the ampli fier condition. In their p ost-surv ey remarks, one participan t
noted that some of the examples had “c haracteristic sounds of neural syn thesis”.
W e theorise that this amplifier distortion p r o cess helps the participan ts to fo cus
more on the realism of the example, in the con text of its con ten t and tim bre.
This is b ecause the amplifier distortion pro cess hide s the aforemen tioned ne ural
artifacts that w ould otherwise cause the MOS to b e lo w ered, as in the DI case,
and distract listeners from the ta s k at hand.
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
53
GuitarFlo w 9
5.3 Sub jectiv e Analysis
A careful sub jectiv e listening analysis of the audio examples rev ealed that Gui-
tarFlo w seems to particularly excel at recreating strumming c hords. This is
particularly relev an t giv en that most commercially a v ailable guitar virtual in-
strumen t soft w are notoriously struggle with this tec hnique. Ho w ev er, the mo del
app ears to struggle m uc h more when generating single notes, creating ob vious
neural artifacts whic h are not presen t during strummed c hords. Due to this dis-
crepancy , w e h yp othesize that this could b e due to the notes of the real guitar DI
not b eing p erfectly aligned to the notes in the Guitar Pro rendering. As c hords
are strummed, they inheren tly ha v e a lo oser timing windo w compared to single
notes. This should b e addressed in a future listening study whic h fo cuses on
strumming vs. single notes.
5.4 Limitations & F uture W ork
As in man y deep learning related w orks, data is the main limiting factor in
our approac h. While the exp erimen t sho w ed great promise despite its data con-
strain ts, additional data p oin ts from a more v aried p o ol of guitars and guitarists
could p oten tially allo w for b etter sound quality and generalisabilit y , as well as
explicit st yle con trollabilit y . Unfortunately , obtaining paired examples of guitar
audio and tablatures is an exp ensive and time consuming process. Pretraining
on syn thetic data or using unpaired training metho ds ma y help address this.
The annotation alignmen t issue ma y ha v e also adv ersely affected the final audio
qualit y . Ho w ev er, the exp erimen t undertak en alone is not sufficien t to clarify this
h yp othesis. A dditionally , the listening test only measures a broad “realism" of
the syn thesis qualit y , and th us w e are unable to gain an y insigh t to an y more fine-
grained asp ects of the results suc h as the tim bre, note accuracy or h uman-lik e
v ariabilit y .
6 Conclusion
In this pap er w e presen ted GuitarFlo w, a no v el metho dology and mo del for syn-
thesising realistic electric guitar from guitar tablatures. This mo del mak es use
of laten t Flo w Matc hing to p erform st yle transfer on a basic audio render of
the tablature, allo wing the mo del to b e trained quic kly on an extremely small
amoun t of data while still generalising to unseen data. This approac h w as jus-
tified through sev eral ob jectiv e metrics and a listening test. W e hop e that our
w ork will inspire more researc hers to in v estigate Flo w Matc hing for generativ e
audio, as w ell as w ork on generativ e systems whic h allo w for increased m usical
con trol.
A c kno wledgmen ts. This work is supported by the EPSR C UKRI Cen tre for Doctoral
T raining in Artificial In telligence and Music (Gran t no. EP/S022694/1) and UKRI -
Inno v ate UK (Pro ject no. 10102804).
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
54
10 J. Loth et al.
Disclosure of In terests. The authors ha v e no comp eting in terests.
References
1. Caillon, A., Esling, P .: RA VE: A v ariational auto enco der for fast and high-qualit y
neural audio syn thesis, http://arxiv.org/abs/2111.05011
2. Casp e, F., McPherson, A., Sandler, M.: DD X7: Differen tiable FM Syn thesis of
Musical Instrumen t Sounds. Pro ceedings of the In ternational So ciet y for Music
Information Retriev al (ISMIR) (2022)
3. Chen, R.T.Q.: torc hdiffeq (2018), https://github.com/rtqichen/torchdiffeq
4. Chen, Y.H., Y eh, Y.T., Cheng, Y.C., W u, J.T., Ho, Y.H., Jang, J.S.R., Y ang,
Y.H.: T o w ards Zero-Shot Amplifier Mo deling: One-to-man y Amplifier Mo deling
via T one Em b edding Con trol. Pro ceedings of the In ternational So ciet y for Music
Information Retriev al Conference (ISMIR) (2024)
5. Ch ung, Y., Eu, P ., Lee, J., Choi, K., Nam, J., Chon, B.S.: KAD: No More F AD! An
Effectiv e and Efficien t Ev aluation Metric for Audio Generation. arXiv:2502.15602
(2025), https://arxiv.org/abs/2502.15602
6. Comanducci, L., An tonacci, F., Sarti, A.: Tim bre T ransfer Using Image-to-Image
Denoising Diffusion Mo dels. In: Proceedings of the International Society for Music
Information Retriev al Conference (ISMIR) (2023)
7. Cramer, A.L., W u, H.H., Salamon, J., Bello, J.P .: Lo ok, Listen, and Learn More:
Design Choices for Deep Audio Em b eddings. In: IEEE In ternational Conference
on A coustics, Sp eec h and Signal Pro cessing (ICASSP). pp. 3852–3856 (2019)
8. Cwitk o witz, F., Hirv onen, T., Klapuri, A.: F retNet: Contin uous-v alued pitch con-
tour streaming for p olyphonic guitar tablature transcription. In: IEEE In terna-
tional Conference on A coustics, Sp eec h and Signal Pro cessing (ICASSP). pp. 1–5.
IEEE (2023)
9. Dormand, J.R., Prince, P .J.: A family of Em b edded Runge-Kutta F ormulae. Jour-
nal of Computational and Applied Mathematics 6 (1), 19–26 (1980)
10. Défossez, A., Cop et, J., Synnaev e, G., A di, Y.: High Fidelit y Neural Audio Com-
pression. T ransactions on Mac hine Learning Researc h (2022)
11. Edw ards, D., Riley , X., Sarmen to, P ., Dixon, S.: MIDI-to-T ab: Guitar T ablature
Inference via Mask ed Language Mo deling. In: Pro ceedings of the In ternational
So ciet y for Music Information Retriev al Conference (ISMIR) (2024)
12. Elizalde, B., Deshm ukh, S., Al Ismail, M., W ang, H.: CLAP: Learning Audio Con-
cepts F rom Natural Language Sup ervision. In: IEEE In ternational Conference on
A coustics, Sp eec h and Signal Pro cessing (ICASSP). pp. 1–5 (2023)
13. Engel, J., Han trakul, L.H., Gu, C., Rob erts, A.: DDSP: Differen tiable Digital Sig-
nal Pro cessing. In: In ternational Conference on Learning Represen tations (ICLR)
(2020)
14. Ev ans, Z., P ark er, J.D., Carr, C., Zuk o wski, Z., T a ylor, J., Pons, J.: Long-F orm Mu-
sic Generation with Laten t Diffusion. In: Pro ceedings of the In ternational So ciet y
for Music Information Retriev al Conference (ISMIR) (2024)
15. F orsgren, S., Martiros, H.: Riffusion - Stable Diffusion for Real-time Music Gener-
ation (2022), https://riffusion.com/about
16. Gat ys, L.A., Ec k er, A.S., Bethge, M.: Image St yle T ransfer Using Con volutional
Neural Net w orks. In: IEEE Conference on Computer Vision and P attern Recogni-
tion (CVPR). pp. 2414–2423. IEEE (2016)
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
55
GuitarFlo w 11
17. Guan, W., W ang, K., Zhou, W., W ang, Y., Deng, F., W ang, H., Li, L., Hong,
Q., Qin, Y.: LAFMA: A Laten t Flo w Matc hing Mo del for T ext-to-Audio Gener-
ation. In: In tersp eec h 2024. pp. 4813–4817 (2024). https://doi.org/10.21437/
Interspeech.2024- 1848
18. Ha wthorne, C., Simon, I., Rob erts, A., Zeghidour, N., Gardner, J., Manilo w, E.,
Engel, J.: Multi-instrumen t Music Syn thesis with Sp ectrogram Diffusion. Pro ceed-
ings of the In ternational So ciet y for Music Information Retriev al (ISMIR) (2022)
19. Ha wthorne, C., Stasyuk, A., Rob erts, A., Simon, I., Huang, C.Z.A., Dieleman,
S., Elsen, E., Engel, J., Ec k, D.: Enabling F actorized Piano Music Mo deling and
Generation with the MAESTR O Dataset. In ternational Conference on Learning
Represen tations (ICLR) (2018)
20. Hershey , S., Chaudh uri, S., Ellis, D.P .W., Gemmek e, J.F., Jansen, A., Mo ore, C.,
Plak al, M., Platt, D., Saurous, R.A., Seyb old, B., Slaney , M., W eiss, R., Wilson,
K.: CNN Arc hitectures for Large-Scale Audio Classification. In: In ternational Con-
ference on A coustics, Sp eec h and Signal Pro cessing (ICASSP) (2017)
21. Huang, H., Man, J., Li, L., Zeng, R.: Musical Tim bre St yle T ransfer with Diffusion
Mo del. P eerJ Computer Science 10 (2024)
22. Huang, S., Li, Q., Anil, C., Bao, X., Oore, S., Grosse, R.B.: Tim breT ron: A
W a v eNet(CycleGAN(CQT(Audio))) Pip eline for Musical Tim bre T ransfer. In: In-
ternational Conference on Learning Represen tations (ICLR) (2019)
23. Kilgour, K., Zuluaga, M., Roblek, D., Sharifi, M.: F réc het audio distance: A
reference-free metric for ev aluating m usic enhancement algorithms. In: In tersp eech
2019. pp. 2350–2354 (2019). https://doi.org/10.21437/Interspeech.2019- 2219
24. Kim, H., Choi, S., Nam, J.: Expressiv e A coustic Guitar Sound Syn thesis with
an Instrumen t-Sp ecific Input Represen tation and Diffusion Outpain ting. In: IEEE
In ternational Conference on A coustics, Sp eec h and Signal Pro cessing (ICASSP).
pp. 7620–7624 (2024)
25. Kim, J.W., Bittner, R., Kumar, A., Bello, J.P .: Neural Music Syn thesis for Flexi-
ble Tim bre Con trol. In: IEEE In ternational Conference on A coustics, Sp eec h and
Signal Pro cessing (ICASSP). pp. 176–180 (2019)
26. K ong, Q., Cao, Y., Iqbal, T., W ang, Y., W ang, W., Plum bley , M.D.: P ANNs:
Large-Scale Pretrained Audio Neural Net w orks for Audio P attern Recognition .
IEEE/A CM T ransactions on Audio, Sp eec h, and Language Pro cessing 28 , 2880–
2894 (2020)
27. Laurson, M., Erkut, C., Välimäki, V., Kuusk ank are, M.: Methods for Mo deling
Realistic Pla ying in A coustic Guitar Syn thesis. Computer Music Journal 25 (3),
38–49 (2001)
28. Lipman, Y., Chen, R.T., Ben-Ham u, H., Nic k el, M., Le, M.: Flo w matc hing for gen-
erativ e mo deling. In: 11th In ternational Conference on Learning Represen tations
(ICLR) (2023)
29. Liu, X., Gong, C., Liu, Q.: Flo w straigh t and fast: Learning to generate and transfer
data with rectified flo w. In: The Elev en th In ternational Conference on Learning
Represen tations (ICLR) (2023)
30. Loth, J., Sarmen to, P ., Carr, C., Zuk o wski, Z., Barthet, M.: ProgGP: F rom Gui-
tarPro T ablature Neural Generation T o Progressiv e Metal Pro duction. In: The 16th
In ternational Symp osium on Computer Music Multidisciplinary Researc h. T oky o,
Japan (2023)
31. Loth, J., Sarmen to, P ., Sark ar, S., Guo, Z., Barthet, M., Sandler, M.: GO A T:
A Large Dataset of P aired Guitar Audio Recordings and T ablatures. In: 26th
In ternational So ciet y for Music Information Retriev al (ISMIR) Conference (2025)
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
56
12 J. Loth et al.
32. Loth, J., Mamou-Mani, A., Barthet, M.: Pla ying St yle Affects Steel-String A coustic
Guitar Tim bre. 3rd In ternational Conference on Tim bre (2023)
33. Loth, J., Sarmen to, P ., Sark ar, S., Barthet, M.: Analysis of MIDI as Input Repre-
sen tations for Guitar Syn thesis. In: DMRN+19: Digital Music Researc h Net w ork
One-da y W orkshop (2024)
34. Maman, B., Zeitler, J., Müller, M., Bermano, A.H.: Multi-Asp ect Conditioning
for Diffusion-Based Music Syn thesis: Enhancing Realism and A coustic Con trol.
IEEE/A CM T ransactions on Audio, Sp eec h, and Language Pro cessing (2024)
35. Noam Mor, Lior W old, A.P ., T aigman, Y.: A Univ ersal Music T ranslation Net w ork.
In: In ternational Conference on Learning Represen tations (ICLR) (2019)
36. P asini, M., Lattner, S., F azek as, G.: Music2Laten t: Consistency Auto enco ders for
Laten t Audio Compression. Pro ceedings of the In ternational So ciet y for Music
Information Retriev al Conference (ISMIR) (2024)
37. Ronneb erger, O., Fisc her, P ., Bro x, T.: U-net: Con v olutional net w orks for biomedi-
cal image segmen tation. In: In ternational Conference on Medical image computing
and computer-assisted in terv en tion. pp. 234–241. Springer (2015)
38. Sarmen to, P .: Guitar T ablature Generation with Deep Learning. Ph.D. thesis,
Queen Mary Univ ersit y of London (August 2024)
39. Sarmen to, P ., Kumar, A., Carr, C., Zuk o wski, Z., Barthet, M., Y ang, Y.H.:
DadaGP: a Dataset of T ok enized GuitarPro Songs for Sequence Mo dels. In: Pro-
ceedings of the In ternational So ciet y for Music Information Retriev al Conference
(ISMIR). pp. 610–618 (2021)
40. Sarmen to, P ., Kumar, A., Chen, Y.H., Carr, C., Zuk o wski, Z., Barthet, M.: GTR-
CTRL: Instrumen t and Genre Conditioning for Guitar-F o cused Music Generation
with T ransformers. In: Pro c. of the Ev oMUSAR T Conf. pp. 260–275. Springer
(2023)
41. Shier, J., Casp e, F., Rob ertson, A., Sandler, M., Saitis, C., McPherson, A.: Dif-
feren tiable Mo delling of P ercussiv e Audio with T ransien t and Sp ectral Syn thesis.
Con v en tion of the Europ ean A coustics Association (2023)
42. T ong, A., F atras, K., Malkin, N., Huguet, G., Zhang, Y., Rector-Bro oks, J., W olf,
G., Bengio, Y.: Impro ving and generalizing flo w-based generativ e mo dels with mini-
batc h optimal transp ort. T ransactions on Mac hine Learning Researc h pp. 1–34
(2024)
43. v an den Oord, A., Dieleman, S., Zen, H., Simon y an, K., Vin y als, O., Gra v es, A.,
Kalc h brenner, N., Senior, A., Ka vuk cuoglu, K.: W a v enet: A generativ e model for
ra w audio. In: 9th ISCA W orkshop on Sp eec h Syn thesis W orkshop (SSW 9). p. 125
(2016)
44. W ang, Y., Guo, W., Huang, R., Huang, J., W ang, Z., Y ou, F., Li, R., Zhao, Z.:
F rieren: Efficien t video-to-audio generation net w ork with rectified flo w matc hing.
A dv ances in Neural Information Pro cessing Systems 37 , 128118–128138 (2024)
45. Wiggins, A., Kim, Y.: A differen tiable acoustic guitar mo del for string-sp ecific
p olyphonic syn thesis. In: 2023 IEEE w orkshop on applications of signal pro cessing
to audio and acoustics (W ASP AA). pp. 1–5. IEEE (2023)
46. Wiggins, A., Kim, Y.E.: Guitar T ablature Estimation with a Con v olutional Neu-
ral Net w ork. In: Pro ceedings of the In ternational So ciet y for Music Information
Retriev al Conference (ISMIR). pp. 284–291 (2019)
47. W righ t, A., Damsk ägg, E.P ., Juv ela, L., Välimäki, V.: Real-Time Guitar Amplifier
Em ulation with Deep Learning. Applied Sciences 10 (3), 766 (2020)
48. Y ang, J., Cinquin, T., Soros, G.: Unsup ervised Musical Tim bre T ransfer for Notifi-
cation Sounds. In: IEEE In ternational Conference on A coustics, Sp eec h and Signal
Pro cessing (ICASSP). pp. 3735–3739 (2021)
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
57
GuitarFlo w 13
49. Zh u, J.Y., P ark, T., Isola, P ., Efros, A.A.: Unpaired Image-to-Image T ranslation
Using Cycle-Consisten t A dv ersarial Net w orks. In: IEEE In ternational Conference
on Computer Vision (ICCV). pp. 2242–2251 (Oct 2017)
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
58
All right s remain with the autho rs under the Creati ve Commons Attrib ution 4.0
International Licens e (CC BY 4.0 ).
Proc. of the 17th I nt. Symposium on Computer Music Multi disciplinar y Research,
London, United Kin gdom, 2025
The Impact of Nonlinear Distortion on the Perceived
Brightness and Roughness of Electric Guitar Chords
Richard Be rrebi 1* , Agathe Mangialomi ni 1,2* , Etienne Thoret 3+ [0000 - 0002 - 8214 - 6278] and
Richard Kr onland Marti net 1+ [ 0000 - 0002 - 7325 - 4920 ]
1 Aix - Mars eil le Unive rsi ty, CNRS, Pe rcep tion Rep res enta tio n Ima ge Sound Musi c ( PRISM)
UMR7061
2 Univers ity of Tour s, CNRS , Cent re d’ études supér ieures de l a Renai ssance (CESR)
UMR7323
3 Aix - Mars eil le Unive rsi ty, CNRS, In sti tut de Neur osci ence s d e l a Ti mone (INT ) UMR7 289
[email protected]
[email protected]
* The two aut hors con tribute d equall y to th is work
+ Senior co - authorship
Abstrac t. When it come s to quali fy the nat ure of audio effects, such as nonlinear
distortion, it is not uncommon to observe artists or audio engineers using seman-
tic descriptors such as “warmth”, “smooth”, “r ough” or “ bright ” , to give a spe-
cific indication on the timb re itself , a ltered by those effects. One of t he many
ways of modifyi ng sound is the nonl inear dis torti on synthes is techni que, also
called waveshaping. W hi le a lot of different hardware i mplementations of distor-
tion effects have their own specificities, each is designed to gene rate new partial s
thanks to a non linear function. On the other side, it is unclear to which distortion
parameter(s) "roughness" and "brightness" refer to , and we don't know ye t if
those characteri stics have a cl ose relati on t o nonline ar distorti on. In this study ,
we investigate the effect of the nonlinear distortion induced by a guitar pedal, the
Boss DS - 1, on t he perceived bright ness and roughness. Firs t, an acoust ic analysis
reveals how the various controls of the guitar pedal affect the brightness and the
roughness of a guitar chord. Then, a listening test confirmed that perceived evo-
lutions of brightness and roughness are coherent with their related p resets related
to them.
Keyword s: Brigh tness , Rou ghness, Waveshapi ng, Nonl inear Distor tion, Tim-
bre, Guitar.
1 Introduction
Nonlinear d istorti on is used in many modern genres such as Jazz, or Rock music,
and is mostly coupled with an electric guitar . M any ro ck subg enres, su ch as metal [10]
or punk [20] , oft en use this effect with a specific chord selection , such as major cho rds
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
59
R. Be r r e b i a n d a l .
( like the largely used “power chord ”) . These arti stic choices in the m usic production of
this “extreme music” cou ld induce strong feelings in the listener, like, for instance if
somebody wants to enhance his mood or to match his anger [29] . How ever, we may
ask whether these emotions are driven by consi stent pat terns that connect listener affect
to sound prod uction. We consider the specific case of the use of dis torted guitar.
For audio effect s such as nonlinear distortion, it is common to describe the resulting
timbre by using adjectives or qualifiers. A large lexicon of semantic descriptors has
been est ablished, including adjectives such as “crunchy”, “bright” or “full” [25] . Each
of them has a relations hip with s pecific setti ngs. For i nstance, “fizzy” and “bright” are
related to “increased treble” settings. As a first glance, it is noticeable that nonlinear
distortion strengthens the higher overtones of a guitar : in consequence, the beating p ar-
tials increase the auditory dissonance , which lead to increase the sensation of roughness
[10] , and, on the other hand , the high frequencies of the guitar ’s timbre are more
charged, which could be assimilated as “bright” [4] .
Regarding nonlinear distorti on, as Hartman [8] says, one can talk about it when new
partials are added to the output si gnal spectrum (while, in the temporal domain, for a
periodic signal, its original period is preserved, and only the amplitude is clipped). This
technique is particularly useful for altering the timbre of musical instruments [14] . As
Roads [27] sum marizes, to create nonlinear distortion, we have to pass the input signal
into a “shaping function”, whose x - axis is the level of the i nput signal, and the y - axis
is the level of the output signal. A ccording to the shape of the nonlinear function, if t he
input amplitude value is superior to the output amplitude value, the waveform is
clipped, distortion appears and adds new partials (called “modulation products”, ac-
cording to Roads) into the spectrum. Most common shaping functions [14, 27] behave
linearly for input level values around 0, while par tials are generated as the input level
increases . Some examples of these functions are hyperbolic tangent, sinusoid or expo-
nential. Each has it s own features and will generate partials in a different way.
One of the consequen ces of the introdu ction of those p artials , is the appearan ce of
strong energy in the high frequencies of the spectrum [10] . It is often difficult to char-
acterize the signal of timbre i nstruments processed by nonlinear distortion, due to its
non - reversible and unpredictable behavior [24] . This is why it is interesting to analyze
this phenomenon from a perceptual perspective. Timbre descriptors, also called “fea-
tures”, have been developed [19, 22, 23] to quantify the spectro - tempora l characteris-
tics of musical timbres. Some of them could give us indications on the brightness of a
timbre.
Brightness is ofte n defined a s “ the amount of high - frequency content perceived in
the spectrum of a sound ”. It i s strongly correlated with the acoustic feature of spectral
centroid [7, 18] , which indicates the barycenter of the spectrum (with equivalent ener-
gies at the right and at the lef t of thi s bar ycenter). Spectral centr oid i s expr essed as [22] :
∑ 𝑓 ( 𝑛 ) 𝑥 (𝑛)
! "#
$% &
∑ 𝑥 (𝑛 )
! "#
$% &
(1)
With x( n) the wei ght ed freq uenc y value (bin nu mber) , and f( n) the ce nter frequ ency
of that bi n. It’s not uncommon to use nonlinear effects t o control t he brightness of the
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
60
Th e i m p a c t o f n o n l i n e a r d i s t o r t i o n o n th e p e r c e p tio n o f e le c tr ic gui t a r c hor ds tim b r e
timbre [4, 38] , sometimes coupled with spectral equalization [31] (it could b e an expla -
nation for the semantic association between “bright” and “increased treble”, as men-
tioned earlier). Brightness is not the only acoustic feature that we can highlight as a
result from nonlinear synthesis techniques: inharmonicity [17] and increased auditory
roughness [5] could also be noticed.
Roughness was first mentioned by Helmholtz [9] , by considering the effect produced
by closely spaced high - frequency partials in the spectrum. Indeed, rou ghness sensation
depends on t he interval between two beating tones. According to Lichte [16] , in musical
timbres and complex tones, roughness is present when consecutive partials are above
the sixth and is a function of the location of such partials . From a perceptual point of
view, roughness is often associated to “sensory dissonance” [28] , and is emphasized
when two sinusoids lie wi thin t he same critical band [6, 37] (the way that our ear per-
ceives the spectr al content, and could be modelled by a filter - bank, like the well - known
Bark scale [36] or Mel scale [32]). It is sometimes associated to a rubbing sensation
and physical friction [34] . On the spectrum, the closer two frequencies are, the slower
the beating, and the timbre will be perceived as “smooth”. On the other hand, when two
frequencies diverge into separate critical bands, roug hness sensation will attenuate and
two separate pitch es will begin to be distinguishable [34] . Zwicker and Fastl proposed
a model to compute perceived roughness [3 7] with, fo r un it reference , o ne 1 kHz sinus-
oid of 60 dB 100% modulated by a f requency of 70 Hz in amplitude, to define one
Asper . This model i s based on th e observanc e of r oughness i n a f requency ba nd be-
tween 70 Hz and 300 Hz, with the modulation frequency ( 𝑓 '() ) multiplied by the mod-
ulation depth ( ΔL ):
𝑅 = 𝑓 '() ∫ Δ𝐿 * ( 𝑧 ) 𝑑𝑧
+, - ./01
& (2)
Considerin g the liter ature, t here is a lack of studies linking nonlinear dist ortion and
auditory roughness , even if Lichte briefly mentions it [16] . For b rightness, some authors
studied nonlinear distortion to manipulate the perceived brightness in the musical tim-
bre [1] . Some of them studied precis ely the effect of distorted guita r on brightness and
sharpness [33] , but did not linger on roughness . In its “Metal Music Studies”, Jan - Peter
Herbst [11] studied some acoustic features (including brightness and roughness) of the
distorted guit ar [10] , by taking into account the “pleasan tness” of the chord. How ever,
those studies don’t delimit how the controls of a distortion guitar pedal affect those two
acoustic feat ures. This is what this study is all about: how the controls of an existing
model of nonlinear distort ion affect the perceived brightness and roughness of the tim-
bre of a guitar chord, and how to determine a rel ationship between them?
To set a realis tic musical cont ext, we took one of t he most used dist ortion pedals by
rock guitarists: the Boss DS - 1. This pedal is analogic, and as we will see lat er, we must
have a full ability to control the parameter s. An accurate way to reproduce this effect
into a digital environment is the “white - box approach” (for analog reproductions i n
digital), to get a reali stic emulation of this distort ion pedal [35] . This modelling ap-
proach is based on the knowledge of the circuits and the components of t he Boss DS -
1, which are opened t o the public [39] , allowing the developmen t of a controllable im-
plementation , very close to the reality of an an alogue Boss DS - 1 [2] . This effec t has
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, Nov. 3-7, 2025
61
R. Be r r e b i a n d a l .
three important controls: the “distortio n” control, which is a pre - gain (corresponding to
the Gain+filt er block in fi gure 1) before the passage of the signal through t he nonlinear
function (a n approximation of the hyperbolic tangent [2] , corresponding to the Sat uring
nonlin block in figure 1), the “tone” control, that splits nonlinearly the signal in two
bands (bass and t he treble, cor responding to the Tone filter block in figure 1) and carries
out the balance between them, and then the “level” control, that acts like a volume
control, to readjust t he possible volume loss .
Fig. 1. Block diagra m of distorti on pedal [35]
In the following, two acoustic spaces for both brightness and roughness will be de-
fined, in which their evolution will be analyzed acco rding to the Boss DS - 1 controls.
Next, a perc eptual ex periment t hat confir ms the previ ous resu lts found will be deta iled
and will conclude on how these cont rols affect brightness and roughness. Throughout
this study, brightness and roughness will be considered separately and a relation be-
tween those two characteristics of the timbre will not b e established.
2 Acoustic Analysis
2.1 Design and proce dure
For the acoustic analysis, we decided to use through the well - known analysis by syn-
thesis methodolog y established by Risset and Wessel [26] . We considered a guitar
chord and applied various quantified Boss DS - 1 pr esets to it. Those sounds are call ed
“measurements” in this study. This set of sounds were used to build up an acoustic
space for both brightness and roughness, using spectral centro id and the roughness
model mentioned previously. Measu rements con siste d of 2420 rec ordings of only one
one - shot guitar chord cut from the first few seconds of the file “directin-
put_Set1_maj.wav” [21] , which is an A# major (A# -D - F). The gui tar used for this
chord was the Ibanez PF300, recorded directly from the input (not from an amplifier,
to avoid any additive nonlinear distortion) using the Audient iD14 audio interface. To
make those 2420 measu remen ts from this guitar chord, we took a MATLAB imple-
mentati on of the Boss DS - 1 [2] , and we quantified each parameter (“distortion”, “tone”
and “level”) with a step of 0.1. The parameters “distortion” and “tone” have a range of
values between 0 and 1, and t he “level” parameter has a range of values between 0 and
2. After quantization, we obtained a total amount of 2420 com binat ions, or “presets” of
the Boss DS - 1 implementation, by excluding all the cases where the “level” parameter
is equal to 0 (the level parameter is acting as the output volume control, so if its value
is set to 0, no signal will be emitted). We applied those presets on this chord to get 2420
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
62
Th e i m p a c t o f n o n l i n e a r d i s t o r t i o n o n th e p e r c e p tio n o f e le c tr ic gui t a r c hor ds tim b r e
different sounds. The total durati on of each measurement is 3 882 ms. The loudness of
these mea surements w as calculated using the Loudn ess K - weighted Full Scal e (LKFS,
ITU - R BS.1770) and e qualized in LKFS valu e of - 18.
2.2 Timbre des criptors
Method . For brightness modeling, the spectral centroid of the Timbre Toolbox [23]
was u sed . The latter is computed for each 25 ms time frame, using the ERB - spaced
gammatone filter bank decompositi on of the signal to obtain a representati on close to
perceptual reality. As a time - varying descriptor, spectral centroid (in Hz) was summa-
rized through the statistics of the median , as a measure of central tendency.
For roughness modeling, t he model of Zwicker & Fastl [37] was used through the
acousticRoughness MATLAB funct ion, based on ISO 532 - 1 [12] . Jus t like
brightness, the roughness model is a time - varying model (in A spers) , so it w as summa -
rized through the statistics of the median .
All t hes e values wer e represented in two acoustic measurement spaces, to get an
overview of the variation in brightness and roughness as a functi on of the parameters
of the Boss DS - 1 in the F igures 2 and 3.
Brightness analysis. For brightness, the “tone” par ameter of the Boss DS - 1 is playing
a crucial role and i s responsible of the shifting of the spectrum barycenter. At each
“tone” value (figure 2), the spectral centr oid is moving one step f urther. The “distor-
tion” parameter pla ys a supporting role in the increasing of t he brightness, especiall y
when it reaches ext reme values from 0. 6 - 0.7 to 1 (figure 2) . This effect is mostly visible
at high “tone” values (t he movement become s steeper and steeper as the “ton e” in-
creases). This is probably because, as a cros sover filter (balancing bass and trebles), the
“tone” keeps more and more high frequencies for values above 0.5. The “level” control
(not visible in figure 2, for the sake of clarity and readability) does n’t play a specific
role in the behavior of the brightness, which shows th at its sole purpose is to adjust the
output volume. To su mmarize, brightness increases with the “ distortion ” amount, in
line with the “tone ” setting, as it was prev iously mentione d in introduction [4] .
Pro c. of the 17th In ternational Symposium on CMMR, London, UK, Nov. 3-7, 2025
63
R. Be r r e b i a n d a l .
Fig. 2. Evoluti on of brightness (median spec tral centroid, i n Hz) as a f unction of the Boss DS - 1
parameters: dis tortion (x - axis) and tone (in different colors for each curve), and level fixed at 1.
Roughness analysis . For roughness, the analyses ar e less str aightforward. For mid -
range “tone” settings (figure 3), e.g., 0.3, 0.4, 0.5 and 0.6, the shape of the plot is the
same (the maxim um of roughness is reached for low “distortion” values), while the
“level” doesn’t play a crucial role. Also, for very small and very high “tone” values,
e.g., as 0 to 0.2, and 0.7 to 0.9 ( 0.2 and 0.9 in figure 3), the shapes of the plots have
some similarities between them: there is a con sequent gap between small and very large
(after 0.5) “distortion” values. For these values, the “level” control seems to contribute
to the generation of roug hness: for instance, for “tone ” values like 0 and 0.1, the gap is
quickly reached for high values of “level”. This gap tends t o increase with the “tone”
setting (as we see a “peak” of roughness at the end of each plot gradually shrinking).
As th e audi tory roughness was descri bed i n the int roducti on, for ext remely low or ex-
tremely high “tone” values, and at high levels of distortion, it is logical that rou ghness
increases: the number of new beating tones (dis tortion products) increases, while a very
small portion of the spectrum (in the bass or the treble range) is retained by the extreme
filtering, which reduces the total bandwidth of the output signal. Al l these factors in-
crease the “chances” of observing and perceiving roughness.
Pro c. of the 17th In ternational Symposium on CMMR, London, UK, Nov. 3-7, 2025
64
Th e i m p a c t o f n o n l i n e a r d i s t o r t i o n o n th e p e r c e p tio n o f e le c tr ic gui t a r c hor ds tim b r e
Fig. 3. Evolution of roughness (median roughness, in Aspers) in functi on of the pa rameters of
the Boss DS - 1 implementat ion: distorti on (x - axis), level (y - axis) and tone fixed at 0.2, 0.5, 0.6
and 0.9.
3 Experiment
3.1 Metho d
Participants . 12 listeners were recruited wit h an average age of 27.92 years (M=26;
SD=8.68 y.o,) and a range of 29 years. They were nat ive French speakers or spoke
French fluently and did not pr esent any heari ng impair ment. 3 of them are guitarist s
with an average age of 24 years (M=26; SD=3.46 y.o,) and a range of 6 years. 6 of them
listen to rock music with an average age of 28.67 years (M=26.5; SD=6.56 y.o,) and a
range of 19 years.
Design. Li steners were tested individual ly in a soundp roof booth. Stimuli were pre-
sented on S ennheiser HD 800 S h eadphones using an Apple MacBook Pro 2024 and an
audio interface Focusrite Scarlett 2i2 (with a fixed volume). The graphical user inter-
face was develop ed using MATLA B 2024.
Procedure. We divid ed t he ex peri ment in tw o par ts ( of 1 0 st imuli for eac h), to ev alua te
separately the brightness and the roughness characteristics of the timbre. In each part,
the participan ts evaluated the perce ived brightn ess and perceived ro ughness diffe rence
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
65
R. Be r r e b i a n d a l .
for each stimuli pair. A nume ric scale from 0 to 10 was used: 0 for “com pletely identi-
cal”, and 10 for “completely different”. Each part of the experiment had a total of 45
pairs to evaluate. Each block proceeded in three phases: the first one was a famili ariza-
tion phase, where the participants listened to each stimulus (in a random order for every
participant). The second phase was the tr aining phas e: the participants evaluated two
test pairs to fam iliarize themselves with the interface. The third phase wa s the pairwise
evaluation, where the participants evaluated the 45 pair s of stimuli on the scale from 0
to 10. Between each block, the participants were entitled to a break. At the end of the
experiment, the participants completed a questionnaire to provide their age, their sex,
their strong hand, indicating if their profession is related to s ound and acou stics, if they
play a musical instr ument and what genres of music t hey listen to.
Stimuli. We se lect ed 1 0 st imuli for the per cept ual evalua tion of b rig htne ss a nd 10 oth er
different stimuli for t he perceptual evaluation of roughness from the 2420 s ounds pre-
viously made in sect ion 2.1. For the select ion of each sti muli group, we carefully chosen
t he preset values associat ed with the sound, for maximum variance. Also, to avoid in-
troducing bias, we ensured that for theses pre - settings, the growth of brightness a nd the
growth of roughness were not correlated. Indeed, for the brightness pr eset s, brightness
grows at each stimulus index, but not the roughness associated wit h t hese same presets,
and vice versa. A cli ck at the end of the st imuli was audible, so we made a volume fade
out with a decreasing exponent ial of 120 milli seconds to remove it .
3.2 Data ana lysis
Inter - listeners agreement. Before runni ng the analysis, we ch ecked the reli ability of
participants answers by computing the inter - rater reliability and by computin g intra-
class correlat ion coefficients . For brightness, an inter - rater correlation coefficient of
0.685 was obtained, using Pearson correlation with 𝑑𝑓 1 = 1 43 , which is a fairly high
value for behavioral ratings. We computed intraclass correlation usi ng ICC(3, 1),
ICC(2,1) ( 𝐹 = 31 . 14 , 𝑝 = 7. 11 ∗ 10 " ##2 ) and ICC(k,1) coefficients: we got 0.6916,
0.6689 and 0.9569, which was enough to consider t he sample as coherent, generaliza-
ble, and reliable. For roughness, we got an inter - rater correlation coefficient of 0.53,
using Pearson correlation wit h 𝑑𝑓 1 = 1 43 , which is good too for behav ioral ratings. We
computed intraclass correlation usi ng ICC(3,1), ICC(2,1) ( 𝐹 = 14 . 51 , 𝑝 = 1. 55 ∗
10 " 34 ) and ICC(k,1) coefficients: we got 0.5205, 0.4695 and 0.9068, which give us
less variability betw een the stimuli, and some disagreem ents.
Multi dimen sion al sc ali ng. The t wo blocks (bri ghtness and ro ughness) were anal yzed
using nonmetric MDS [13, 30] through the evaluation phase of each participant (a sym-
metric 10x10 dissimil arity matrix for the 10 stimuli, with 45 evaluatio ns). For its capa-
bility to evaluate an entire sample of participant s , INDSCA L [3, 15] was used to model
the nonmetric MDS. INDSCAL has no specificities b ut has weights on each dimension
for each listener, and allow s for differences among individual listeners or latent classes
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
66
Th e i m p a c t o f n o n l i n e a r d i s t o r t i o n o n th e p e r c e p tio n o f e le c tr ic gui t a r c hor ds tim b r e
of listeners, respectivel y [19] . For both brightness and roughness block, I NDSCAL was
computed on 2 dimensions.
Fig. 4. INDSCAL group stimulus space for both perceptive evaluation (with in bl ue the stimulus
number and in red the val ues of the control s: “d” for “dist ortion”, “t” for “tone” and “l” for
“level”).
To validate th e coherence between th e acoustic analy sis in secti on 2.1 and the per-
ceptual experiment, a correlation was computed between the distances of the measure-
ments of our stimuli (10 sti muli of bright ness and 10 stimuli of roughnes s) and with the
Eu clidean distances computed in the INDSCAL of each block. For bri ghtness, a great
correlation ( 𝜌1 = 0.739, p < 5 ∗ 10 "5 , 𝑑𝑓 = 43 , CI95% [0.534 , 0.871]) was found for
the entire sample. If the rock music listeners are isolated from the sample, the distances
correlate better ( 𝜌 = 0.815, p = 0, 𝑑𝑓 = 8 , CI95% [0.641, 0.914]). If the same is done
by isolating the guitarists from t he sample, distances correl ate still ( 𝜌 = 0.672, p <
1. 015 ∗ 10 "3 , 𝑑𝑓 = 8 , CI95% [0.416 , 0.837 ]). Fo r roug hness, perceptual evaluation
correlated with measurements too ( 𝜌 = 0.5378, p < 0.0005, 𝑑𝑓 = 43 , CI9 5% [0.303 ,
0.718]). If the rock music listeners are isolated f rom the sample, t he distances corr elate
a bit more ( 𝜌 = 0.467, p < 1.4 ∗ 10 "6 , 𝑑𝑓 = 8 , CI95% [0.199, 0.666]). If the same is
done by isolating the guitarists from the sample, distances cor relate again ( 𝜌 = 0.442, p
< 2.6 ∗ 10 "6 , 𝑑𝑓 = 8 , CI95 % [0.158, 0.661]). All correlatio ns were computed using
Spearman correlations, and 95% confidence intervals were computed by boots trapping
the correlations, using 2000 bootstrap samples. The stress obtained for the INDSCAL
of brightness has a value o f 0.176, and for the INDSCAL of roughness, 0.209. (the
entire sample).
At the end of ou r analysis, for each block (brightness and roughness), a space in two
dimensions was established, matching one or two controls of the Boss DS - 1:
For brightness, dimension 1 fits wit h “t one” control ( 𝜌 = 0.76 2, p < 0.011, 𝑑𝑓 = 8 ,
CI95% [ - 0.213, 0.968]), while dimension 2 fits with “distortion” control ( 𝜌 = - 0.796, p
< 0.006, 𝑑𝑓 = 8 , CI95% [ - 0.973, 0.138]).
For roughness, dimension 1 fits with “dis tortion” cont rol ( 𝜌 = 0.975, p < 1. 55 ∗
10 "3 , 𝑑𝑓 = 8 , CI95% [0.90 1, 1]).
For bri ghtness, those cor relations confirmed our hypothesis where we stat ed that
“tone” and “distortion” controls were responsible for shifting the spectr al centroid to
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
67
R. Be r r e b i a n d a l .
the right (on the spectrum), as dimension 1 correlated with spectral centroid values ( 𝜌
= 0.697, p < 0.035 , CI95% [0 .15, 0.9 87]), and same for dimension 2 ( 𝜌 = - 0.709, p <
0.03, 𝑑𝑓 = 8 , CI95% [ - 1, 0.09]). If the rock music listeners are is olated fr om the initial
sample, and an INDSCAL is computed, dimension 1 fits more with “tone” control ( 𝜌
= 0.976, p < 1.5 ∗ 10 "3 , 𝑑𝑓 = 8 , CI95% [0.6 73, 1]) and is correlated with spectral
centroid values ( 𝜌 = 0.98 8, p = 0, CI95% [0.677, 1]). If the same is do ne for the gui tar-
ists, dimension 2 fits a bit more with “distortion” control ( 𝜌 = - 0.889, p < 5.8 ∗ 10 ", ,
𝑑𝑓 = 8 , CI95% [ - 0.991, 0.123]) and is correlated with spectral cent roid values ( 𝜌 = -
0.964, p = 0, CI95% [ - 1, - 0.585]) .
For roughness, it’s a bit more dif ficult t o match the dissimilarity ratings wit h the
acoustic m easurements, but in most cases, roughness increases dr astically with the “dis -
tortion” contro l, as dimens ion 1 correlate d with roug hness values ( 𝜌 = 0.721, p < 0.0 25,
𝑑𝑓 = 8 , CI95% [0.121, 0.963]). Surprisingly, if the rock music listeners are isolated
from the initial sample, and an INDSCAL is computed, dimension 2 fits with “tone”
control ( 𝜌 = 0.96, p < 1 . 07 ∗ 10 "4 , 𝑑𝑓 = 8 , CI95% [0.824, 0.99 7]), but not with roug h-
ness val ues ( 𝜌 = 0.406, p < 0.25, 𝑑𝑓 = 8 , C I95% [ - 0.42, 0. 899]). All corr elations w ere
computed using Spearman correlat ions, and 95% confidence intervals were computed
by bootstrapping the cor relations, using 2000 bootstrap samples.
4 Discussion and conclusion
In this study, the evolution of the b rightness and roughness of a guitar chord processed
through a Boss DS - 1 implementation was invest igated, consideri ng its var ious settings.
Two aco ustic spaces using measurements (median spectral cent roid and median roug h-
ness) were built, and then a pairwise comparison experiment using two subsets of sti m-
uli (one for br ightness and one for r oughness) was conducted. As a conclusion of this
experiment analysis , two dimensions (linked to the “distor tion” and “tone” controls)
were ident ified as pr edominant for brigh tness, while for ro ughness, a single di mension
(linked to the “distortion” control) was ide ntified as predominant.
Results of t he ac oustic analysis show n that the evolution of brightness is highly re-
lated to the “ tone ” control, and somehow, backed by the “ distortion ” control . It makes
sense that a crossover filter is responsible of the high - frequency distrib ution in the spec-
trum. However, in the results of the experim ent, we notice d that the participants did not
show any predominance between the “ tone ” and “ distortion ” controls, as the 𝜌 coeffi-
cients and the p - values for both dimensions are very close. We can wonder whether the
increase of partials affects th e perception of low - band filtering. On the o ther hand , and
by distinguishing the sample population (the guitar pl ayers and the rock music listen-
ers), is it remarkable to see that guit ar players perceive the impact of the “tone” control
on brightness better than the whole sample, and rock music listeners percei ve the i m-
pact of the “distor tion” control on bri ghtness better.
For the roughness results, the evolut ion of roughness is not linear as the evolut ion of
brightness is, simply because the phenomenon of roughness increases when the band-
width is red uced with a large n umber of clos e partial s. Roughness can be easily created
if we set up the pedal w ith very s mall or v ery high values of “tone” (the bandwidth will
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, Nov. 3-7, 2025
68
Th e i m p a c t o f n o n l i n e a r d i s t o r t i o n o n th e p e r c e p tio n o f e le c tr ic gui t a r c hor ds tim b r e
be reduced, as the bass or the treble becomes predominant at the expense of the other)
and hi gh values of “distorti on” (to induce more clipping). Compared to the results ob-
tained with the experiment, it is noticeable that participants weren’t affected by the
“tone” and only per ceived the action of the “distortion”, which means that the introduc-
tion of new distortion products is highly predominant in the assessment. Just like bright-
ness, by distinguishing the sample population (the guitar players and the rock music
listeners), is it noticeable to see that rock music listeners perceive the impact of the
“tone” control on roughness bett er than the whole sample.
In both cases, a “bright” or a “rough” timbre could be created by incr easing the
amount of new distortion products, and then, to distinguish a “rough” timbre from a
“bright” timbre, the filtering (bass/treble balance) should be adjusted by taste . K eeping
only a very small part of the bass or the treble will make a rougher timbre, while keep-
ing more treble w ill make a brighter tim bre.
Here, we didn’t discuss much about the “level” contr ol, which is not prominent in
the ca se of the brightness analysis as it is for roughness (which seem s to have the same
effect than the “distortion” control in section 2.2), because of the loudness equalization
detailed in section 2.1. Participants didn’t find any r elation between this control and the
acoustic features of bright ness and roughness.
At the begi nning of t his stud y, we asked ou rselves how distor tion peda l parameter s
affect timbre characteristics, such as brightness and roughness. Our ac oustic and per-
ceptual analysis allowed us to determine a close relationship between increasing bright -
ness and increasing “distortion” and “tone” controls (w ith a high predominance for the
“tone” in the acoustic analysi s), and same for roughness with the “distortion” control,
and very small or very high values of “tone” ( with a high predominance for the “di s-
tortion” in perceptu al analysis).
As t he c ontent of this study is a work in p rogress, to make our resul ts more r obust,
we must increas e our sample of part icipants ( and specifically rock listeners and guitar-
ists ) . It might be interesting to use other timbre descr iptors or roughness models for our
measure ments and est ablish a more complete cor relat ion with more INDSCAL dimen-
sions. In the future, we will analyze the relationship between nonlinear distortion and
the “musical gesture” of the guitarist, to pursue our global study on t he relation between
“extreme rock music” and the affect of the listener, using our results on brightness and
roughness, as we think they are closely rela ted.
Ackn owledgm ent s. We woul d like to ex press our s inc ere gratitude to Sol ve ig Serre an d Luc
Robè ne for their valuabl e insight s, co nstru ctive feedb ack , an d gener ous su ppor t thro ugho ut the
de velopm en t of t his wor k. The i r ex pertise an d en c our ag ement have been instru me nt al in sh a pi ng
the dire ction and qua lit y of this rese arch. Resea rch suppor te d by grants ANR - 16- CO NV -000 2
(I LC B), ANR -24-CE 38 - 417 5 (PS I ND ) , an d the Exc elle nce I niti ative of Ai x -Uni ver sity
(A *M IDE X) . This project has receive d fina ncial su ppor t from the CNRS thro ugh the MI T I in-
terd isci plina ry pr ogr am s . AM docto ra l fellows hip is su ppor te d by a CNRS 80 -prim e fellows hip .
Discl osure of Int erest s. The a uthors declare no conf lict o f inte rest.
Pro c. of the 17th In ternational Symposium on CMMR, London, UK, Nov. 3-7, 2025
69
R. Be r r e b i a n d a l .
References
1. Beaucha mp, J.: Synt hesis by Spectral Ampli tude and “Bri ghtnes s” Matching of Ana-
lyzed Mu sical Instrum ent Tones. J. A udio Eng. S oc. (1981).
2. Brojbo iu, D.: Si mplifi ed Mode lling of t he Bos s DS - 1 Distort ion Pedal. ( 2019).
3. Carrol l, J.D., Chang, J. - J.: Analysis of individual differences in multidimensional sca-
ling via an n - way gen erali zati on of “Eckar t - Young” deco mposit ion. Psy chomet rika. 35 , 3, 283 –
319 (1970). htt ps://doi.or g/10.1007/BF02310791.
4. Enderby, S., Stabl es, R.: A Nonli near Method fo r Manipulat ing Warmth and
Brightn ess. (2017) .
5. Fastl, H., Zwicker, E.: Psychoacoustic s. Spr inger Berlin Heidel berg, Berlin, Heidel-
berg (2007). ht tps://doi. org/10.1007/978 -3- 540 - 68888 - 4.
6. Fletcher, H.: Auditory Patt erns. Rev. Mod. Phys . 12, 1, 47 – 65 (1940).
https://doi .org/10.1103/RevModPhys.12.47.
7. Grey, J.M., Gordon, J.W.: Percep tual effect s of sp ectr al modi ficat ions on music al
timbres. J. Ac oust. Soc. Am . 63, 5, 1493 – 1500 (1978). https:/ /doi.org/10.1121/ 1.381843.
8. Hartma nn, W.M.: Pri ncipl es of Musical Acousti cs. Spri nger New York, New York,
NY (201 3). http s:// doi. org/ 10.10 07/978 -1- 4614 - 6786 - 1.
9. Helmhol tz, H.L. F.: On th e Sens atio ns of Tone . (1 954).
10. Herbs t, J . - P.: Dist ortion a nd Rock Guitar Harmony. Music Percept. 3 6, 4, 33 5 – 352
(2019). https://doi.org/10.152 5/mp.2019.36 .4.335.
11. Herbs t, J . - P.: The Cambridge Companion to Metal Music . (2023).
12. ISO 532 - 1: 2017: Acoustics — Meth ods f or calc ula tin g loudne ss — Part 1: Zwicker
metho d. ( 2017 ).
13. Krusk al, J.B.: Multi dimensi onal Scalin g by Optimiz ing Goodn ess of Fit to a Nonmetr ic
Hypothe sis. Psyc hometr ika. 29, 1, 1 – 27 (1964). https: //doi.org/10. 1007/BF02289565.
14. Le Brun, M.: Digital Waveshapi ng Synt hesis. Audio Eng. Soc . (19 79).
15. Leeuw, J. de, Mair, P.: Multidimens ional Scalin g Using Majorization : SMACOF in R.
J. Stat. Softw. 31, 1 – 30 (2009). https: //doi.org/10. 18637/jss.v031.i 03.
16. Lichte, W.H.: Attribut es of complex tones. J. Exp. Psycho l. 28, 6, 455 – 480 (1941).
https://doi .org/10.1037/h0053526.
17. MacC all um, J. , Ein bond , A.: Real - Time Analysis o f Sensory Dissonance . In: Kron-
land - Martin et, R. et al. (ed s.) Comput er Musi c Model ing a nd Retr ie val . Sense of Sounds . pp.
203 – 211 Springer Berlin Heidelberg, Berli n, Heidelberg (2008). htt ps://doi. org/10.1007/978 -3-
540 - 85035 - 9_13.
18. McAda ms, S. e t a l. : Discri mina tio n of musi cal i nst rum ent so unds r esy nth esi zed wi th
simplified spectrotemporal parameters. J. Acoust. Soc. Am. 105, 2, 882 – 897 (1999).
https://doi .org/10.1121/1.426277.
19. McAda ms, S.: Th e Per cept ual Repr ese nta ti on of Timbr e. In : Si eden bur g, K. et al . (ed s.)
Timbre: Acoustics , Per ceptio n, and Cognit ion. pp. 2 3 – 57 Springer Inte rnational Publishing,
Cham (201 9). h ttps: //doi .org/ 10.100 7/978 -3- 030 - 14832 - 4_2.
20. Pearson, D.M.: Extre me Hardcore Punk and t he Analyti cal Challe nges of Rhythm,
Riffs, and Timbre in Punk Music. Musi c Theory Onli ne. 25, 1, (2019).
https://doi .org/10.30535/mto.25. 1.5.
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
70
Th e i m p a c t o f n o n l i n e a r d i s t o r t i o n o n th e p e r c e p tio n o f e le c tr ic gui t a r c hor ds tim b r e
21. Pedroza, H. et al.: Guitar - TECHS: An Electr ic Gui tar Dataset Covering Techniques ,
Musi cal Exce rpt s, Ch ord s and Sca les Usin g a D iver se A rra y of Hard ware , ht tp: //a r-
xiv.org/abs/2501. 03720, (2025). htt ps://doi.or g/10.48550/arXiv.2501.03720.
22. Peeters, G.: A large set of audio feature s for sound descr iption (si milarity and cl assifi-
cation) in the CUIDADO project. (2004).
23. Peeters, G. et al. : The Timbre Toolbox: Extracting audio descript ors fr om musical si-
gnals. J. Acoust. Soc. Am. 130, 5, 2902 – 2916 (2011). htt ps://doi.org/ 10.1121/1.3642604.
24. Puckette, M.: The Theor y and Technique of Electr onic Music. (2007).
25. Rice, T. : An Inves tigat ion into Non - linear Distortion Pedals with a Focus on Establis-
hing a Lexicon t o Accurately Descri be Them.
26. Risse t, J. - C., Wes sel, D.L. : Explora tion of Timbr e b y Ana lysis an d S ynthesi s. In: The
Psychology of Musi c. p p. 113 – 169 Elsevier (1999) . htt ps://doi.or g/10.1016/B978 - 012213564 -
4/50006 - 8.
27. Roads, C.: A Tutoria l on Non - Linea r Dis tortio n or Waveshapi ng Syn thesis. Comput.
Musi c J . 3 , 2 , 2 9 ( 1979) . htt ps: //d oi. org/ 10. 2307 /368 0281 .
28. Sethares, W.A.: Tuning, timbre, spectrum, sc ale. Sprin ger, London (2005).
29. Sharman, L., Dingle, G.A. : Ext reme Metal Music and Anger Proces sing. Front. Hum.
Neurosc i. 9 , (2 015). htt ps:// doi. org/1 0.3389 /fnhum. 2015. 00272.
30. Shepard, R.N.: The analysis of proxi mities: Mult idimensional scaling with an unknown
distance function. II. Psychometrika. 27, 3, 219 – 246 (1962).
https://doi .org/10.1007/BF02289621.
31. Stables, R. et al.: Semant ic Des cription of Timbral Transf ormations in Music Produc-
tion. In: Procee dings of the 24th ACM international confere nce on Multimedia. pp. 337 – 341
ACM, Amst erdam T he Net herlan ds (2 016). htt ps:// doi. org/1 0.1145 /29642 84.296 7238.
32. Stevens, S.S. et al.: A Scale fo r the Measurement of the Psycholo gical Magnitu de Pitch.
J. Acoust. Soc. Am. 8, 3, 185 – 190 (1937). https://doi. org/10.1121/1.1915893.
33. Tsumoto, K. et al .: The effe ct of harmoni c over tones in r elation to “sharpn ess” f or
perception of brightness of dis torted guitar timbre. Pres ented at the 172nd Meeting of the Acous-
tical Society of America , H onolulu, Ha waii (2016 ). https://doi.org/1 0.1121/2.00 00380.
34. Wal lmar k, Z. T.: Ap pra is ing Ti mbr e: Embo dim ent an d Aff ect a t the Th res hol d of Mu-
sic and Noise. (2014).
35. Yeh, D.T. et al.: SIMPLIFIED, PHYSICALLY - INFOR MED MODE LS OF
DISTORTION AND OVERDRIVE GUITAR EFFE CTS PEDALS. (200 7).
36. Zwicker, E.: Subd ivision of the Audible Frequency Range into Cr itical Bands (Fre -
quenzgruppen). J. Acoust. Soc. Am. 33, 2, 248 – 248 (1961). htt ps://doi.or g/10.1121/1.1908630.
37. Zwicker, E., Fastl , H.: Psych oacousti cs. Sp ringer Berlin Heidelber g, Berl in, Heidel-
berg (1999). ht tps://doi.or g/10.1007/978 -3- 662 - 09562 - 1.
38. AES J ournal For um » Synth esis by Spectr al Amplitu de a nd “Br ight ness” Matchi ng of
Analyze d Musi cal Instr ument Tones, htt ps:/ /secur e.aes .org/ forum/ pubs/ journ al/?e lib=38 37, last
accessed 2025/04/29.
39. Electr oSmash - Bos s DS1 Dist ortion Analy sis, htt ps:// www.elect rosmash. com/boss -
ds1 - analysis, last accessed 2025/05/14.
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, Nov. 3-7, 2025
71
Exploring individualization of sound stim uli
Salomé Sudre 1 , 2 [0009 − 0005 − 5514 − 8912] , Sølvi Y stad 2 [0000 − 0001 − 9022 − 9690] , Ric hard
Kronland-Martinet 2 [0000 − 0002 − 7325 − 4920] , Laetitia P etit 1 [0000 − 0003 − 3501 − 2952] ,
and Mitsuk o Aramaki 2 [0000 − 0001 − 6518 − 374 X ]
1 Aix Marseille Univ, LPCPP (Lab oratoire de Psyc hologie Clinique, de
Psyc hopathologie et de Psyc hanalyse) (EA 3278), 13621 Aix-en-Pro vence, F rance
2 Aix Marseille Univ, CNRS, PRISM (P erception, Representations, Image, Sound,
Music), 31 Chemin J. Aiguier, 13402 Marseille Cedex 20, F rance
[email protected]
Abstract. There is extensiv e literature on the effects of binaural b eats
on cognitiv e states, how ev er, the findings are often con tradictory . This
inconsistency ma y b e due to the use of fixed, high-pitched frequencies in
exp erimen tal proto cols, which can mak e the sounds unpleasan t to listen
to. The aim of this article is to explore the individualization of spatial-
ized sounds. T o do so, participan ts w ere asked to create a sound of their
c hoice by selecting b oth the fundamen tal and b eat frequencies. W e ex-
amined differen t types of sounds, including monaural and binaural b eats,
panning b eats, and alternate b eeps. The results rev eal individual v ari-
abilit y in the p erception and preference of sounds, although some general
trends emerged, suc h as a preference for low frequencies. Eac h t yp e of
sound ev oked specific imagery , ranging from mec hanical ob jects to at-
mospheric scenes, highligh ting the imp ortance of emotional and sym b olic
dimensions in studies of h uman auditory p erception. In conclusion, these
preliminary findings supp ort the idea that, in the absence of univ ersal
rules regarding sound preference and p erception, an individualized ap-
proac h could offer deep er insigh t into the effects of binaural beats on
cognitiv e states.
Keyw ords: Spatialized sound · Perception · Individualized sound.
1 In tro duction
This pap er presen ts an ongoing exploratory stud y on the individualization of
sound stim uli in the context of their influence on the men tal state of the listener.
F or that purp ose, w e inv estigated in a previous study the phenomenon of binaural
b eats (BB) that arises when t wo pure tones with sligh tly differen t frequencies
are presen ted separately into eac h ear [14]. The listener p erceiv es a third sound,
All righ ts remain with the authors under the Creativ e Commons A ttribution
4.0 In ternational Li cense (CC BY 4.0).
Pro c. of the 17th In t. Symp osium on Computer Music Multidisciplinary Research,
London, United Kingdom, 2025
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
72
S. Sudre et al.
c haracterized b y a carrier frequency that is the a v erage of the t w o tones, and a
mo dulation frequency corresp onding to the difference b et w een them. BB ha v e
b een prop osed as a means of influencing the men tal and cognitiv e state of the
listener, b y mo dulating electrical brain activit y (for a review, see [6,1,2,8]).
F rom an acoustic p ersp ectiv e, it can b e sho wn that these b eats are equiv alen t
to the mo v emen t of a virtual sound source around the listener. In our previous
study , w e highligh ted similar effects b et w een these t w o t yp es of auditory stim u-
lation (BB and the sounds ev oking mo v emen t of a virtual sound source). These
stim uli, presen ted discretely or con tin uously , with a carrier frequency of 440 Hz
and b eat frequencies of 6 or 40 Hz, promoted an impro v ed state of relaxation,
asso ciated with c hanges in brain electrical activit y measured b y EEG, compared
to non-spatialized auditory stim ulation (monaural b eats, [14]).
Numerous studies ha v e in v estigated the effects of binaural b eats on atten-
tion, but the findings remain inconclusiv e. Although some rep ort impro v emen ts
in alertness and atten tion p erformance [3,9,13], others ha v e not b een able to
replicate these effects [14,10,4], and no consensus has emerged. W e b eliev e that
this disparit y of findings could b e due to the use of stim uli constructed with fixed
frequencies and iden tical for all participan ts. In line with these considerations,
studies sho w ed that when participan ts are ask ed to pro duce a regular rh ythm at
a pace that suits them, they tend to pro duce a temp o a v eraging around 120 bpm
(i.e. 500 ms, [12]). Ho w ev er, this phenomenon, kno wn as pr eferr e d temp o , v aries
b et w een individuals [7,11,12]. It also w as sho wn that p eople prefer to listen to
rh ythms that matc h their preferred temp o [11].
The curren t pap er therefore presen ts a study allo wing participan ts to con-
struct their o wn sounds b y adjusting the auditory parameters (frequency and
b eat v alues) themselv es. Our main h yp othesis is that within the studied p op-
ulation, there is a tendency to w ard sp ecific auditory attributes, though with
in teresting v ariabilit y reflecting individual sp ecificit y . In the longer term, our
h yp othesis is that b y adapting sound stim uli to eac h participan t, the observ ed
effects of auditory stim uli on emotional and cognitiv e states w ould b e enhanced.
2 Materials and Metho ds
2.1 P articipan ts
Sixteen participan ts (5 females, 11 males), with ages ranging from 20 to 36 y ears
(M = 25.1, SD = 4) to ok part in the exp erimen t. The participation criteria
w ere the follo wing: normal or corrected vision and hearing, no neurological his-
tory . W ritten informed consen t w as obtained from all participan ts in v olv ed in
the study , whic h w as review ed and appro v ed b y the ethics committee of Aix-
Marseille Univ ersit y . P articipan ts w ere informed that they w ere free to lea v e the
exp erimen t at an y time and that their data w ould b e treated anon ymously .
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
73
Individualization of sound stim uli
2.2 Stim uli
The four t yp es of stim uli used in our previous study [14] w ere in v estigated. In
particular, the participan ts w ere ask ed to construct the follo wing four stim uli b y
acting on sp ecific auditory parameters:
– Monaural b eats (no spatial attributes): the generated sound is a pure tone
of fundamen tal frequency f 0 and mo dulated in amplitude b y a cosine func-
tion of frequency f b defined as ‘b eating frequency’. The resulting sound is
presen ted sim ultaneously to b oth ears.
– Binaural b eats (BB): t w o pure tones of differen t frequencies presen ted sep-
arately in eac h ear. The frequency f 0 corresp onds to the a v erage of the
frequencies of the t w o pure tones, while the b eating frequency f b is equal to
the difference b et w een them.
– P anning b eats: a pure tone of fundamen tal frequency f 0 is mo dulated in
amplitude b y a sp ecific panning la w whic h is a sine function of frequency f b
in the left ear and a cosine function of frequency f b in the righ t ear.
– Alternate b eeps: a sequence of b eeps presen ted alternately in eac h ear at
frequency rate f b . A single b eep is 80 Hz amplitude-mo dulated pure tone of
fundamen tal frequency f 0 .
F or eac h t yp e of stim uli, the participan ts could mo dify 2 parameters, i.e. the
fundamen tal frequency f 0 and the b eating frequency f b , to construct their o wn
sounds, as detailed in the next section.
2.3 Pro cedure
F or eac h t yp e of stim uli (exp erimen tal conditions), participan ts w ere ask ed to
construct a sound b y c ho osing, on a graphical in terface, the v alue of the funda-
men tal frequency (b et w een 20-1000 Hz) and the b eat frequency (b et w een 0-40
Hz). A ‘Gain’ slider w as also prop osed to the participan ts to adjust the in ten-
sit y of the sound only if they cannot hear the sound in a comfortable w a y . The
only instruction giv en to the participan ts w as: ‘Cr e ate the sound of your choic e’ .
The c hoice of instructions w as v alidated b y a pilot study conducted b efore this
exp erimen t. If participan ts ask ed for more information, they w ere told that they
could mak e the sound they w an ted, based on the criteria of their c hoice. Once
the participan ts w ere satisfied with the sound they created, they mo v ed on to the
next trial b y clic king the ‘V alidate’ button. There w as no time limit.
The sounds w ere presen ted through headphones (AK G K52) and w ere gen-
erated in real time using the Max/MSP soft w are (h ttps://cycling74.com/), ac-
cording to the parameters set b y the participan ts. Eac h participan t designed the
four stim uli. The order of the condition presen tation w as randomized.
The participan ts w ere then in vited to tak e part in a semi-structured in terview
with the exp erimen ter, aimed at understanding ho w they approac hed the sound
construction pro cess and the reasons b ehind their c hoices. The in terview had the
same structure for eac h participan t. The in terview b egan with a general question,
in whic h the exp erimen ter ask ed participan ts to describ e their exp erience during
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
74
S. Sudre et al.
the sound creation pro cess. Then, eac h sound they created w as presen ted to
them. F or eac h sound, the follo wing questions w ere ask ed:
– Can you describ e this sound?
– If you had to find an image or a wor d to describ e this sound, what would it
b e?
– Why did you cho ose this sound r ather than another?
After that, questions ab out an y difficulties encoun tered during the exp erimen t
w ere ask ed. The exp erimen ter also ask ed whic h of the 4 sounds they constructed
they preferred, whic h one they lik ed the least, and wh y . The audio in terview
w as recorded with the prior consen t of the participan ts for the recording. The
recordings w ere then transcrib ed using the noScrib e soft w are (v0.6.2, [5]) and
man ually corrected afterw ards.
3 Preliminary results
3.1 In tra-class correlation
W e computed in traclass correlation co efficien t (ICC) on the carrier and b eating
frequencies data for eac h t yp e of stim uli. ICC is generally used to assess the
consistency of quan titativ e measuremen ts made b y differen t ev aluators. It com-
pares the v ariabilit y of differen t ratings of the same ob ject to the total v ariation
across all ratings and all ob jects. The ICC v alues v ary b et w een 0 and 1. ICC
estimates w ere calculated using irr pac k age for R based on single rater, absolute
agreemen t, 2-w a y mixed-effects mo del, with 16 raters (participan ts) and 2 rated
ob jects (fundamen tal and b eating frequency v alues). Results sho w ed a mo derate
agreemen t b et w een participan ts for the monaural [ICC = 0.651; F (1, 29.8) =
29.3] and BB ([ICC = 0.641; F (1,30) = 30.9] and alternate b eeps [ICC = 0.613;
F (1,29.9) = 26.1]. F or the panning b eats condition, w e found a go o d agreemen t
with ICC = 0.815 [ F (1,29.8) = 67].
3.2 Sound preference
Concerning preferred sound (T able 1), no clear preference emerged but there
is a tendency for BB with 31,25% of participan ts who c hose it as preferred
sound. In terestingly , w e found the same tendency for BB with 37, 5% for the
less preferred sound. T o examine the participan ts preferences, w e calculated the
median fundamen tal and b eating frequencies for eac h t yp e of stim uli: 1) Monau-
ral b eats: median f 0 = 157 . 68 , median f b = 4 . 88 ; 2) BB: median f 0 = 202 . 8 ,
median f b = 3 . 89 ; ; 3) P anning b eats: median f 0 = 134 . 98 , median f b = 2 . 43 ; 4)
Alternate b eeps: median f 0 = 233 . 59 , median f b = 6 . 73 ). These median v alues
allo w ed us to divide our sample in to t w o groups: those who selected frequencies
b elo w the median and those who selected frequencies ab o v e it. T able 2 sho ws, for
eac h preferred sound, the p ercen tage of participan ts who c hose fundamen tal and
b eat frequencies either ab o v e or b elo w the median. P articipan ts who preferred
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
75
Individualization of sound stim uli
monaural sounds or BB mostly selected fundamen tal and b eat frequencies b e-
lo w the median (i.e., lo w er frequencies relativ e to our sample). In con trast, when
panning b eats w ere the preferred sound, there did not app ear to b e a clear dif-
ference in preference based on the fundamen tal or b eat frequencies. F or those
who preferred alternate b eeps, the selected fundamen tal frequency w as predom-
inan tly b elo w the median. Ho w ev er, the b eat frequency c hosen did not app ear
to influence preference in this condition.
T able 1. P ercen tage of preferred and less preferred sound with resp ect to the t yp e of
stim uli.
Sound Preferred sound (in %) Less preferred sound (in %)
Monaural b eats 18,75 18,75
Binaural b eats 31,25 37,5
P anning b eats 25 25
Alternate b eeps 25 18,75
T able 2. P ercen tage of preferred sound dep ending on fundamen tal and b eating fre-
quencies.
Monaural b eats Binaural b eats P anning beats Alternate beeps
f 0 < median 66,67 80 50 75
f 0 > median 33,33 20 50 25
f b < median 66,67 80 50 50
f b > median 33,33 20 50 50
W e did the same pro cess with the less preferred sound. The results sho w ed
that when monaural b eats are the less preferred sound, the c hosen fundamen tal
frequency is ab o v e the median (T able 3, 66.67%) and the b eating frequency is
b elo w the median (66.67%). A similar tendency w as observ ed when alternate
b eeps are less preferred. No effect of fundamen tal or b eating frequency app ears
for the panning b eats. Ho w ev er, when BB are less preferred, 83.33% and 66.67%
of the participan ts c hose fundamen tal and b eating frequencies ab ov e the median,
resp ectiv ely .
T able 3. P ercen tage of less preferred sound dep ending on fundamen tal and b eating
frequencies.
Monaural b eats Binaural b eats P anning beats Alternate beeps
f 0 < median 33,33 16,67 50 33,33
f 0 > median 66,67 83,33 50 66,67
f b < median 66,67 33,33 50 66,67
f b > median 33,33 66,67 50 33,33
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
76
S. Sudre et al.
3.3 Ev o cation of sounds
W e also analyzed the resp onses of participan ts to the question: “If you had to
describ e this sound with a wor d or an image, what would it b e?”. Based on their
answ ers, w e iden tified five categories of ev o cations:
– Sounds that ev ok e an ob ject (e.g., “metronome”);
– Sounds that ev ok e an atmosphere or mo o d (e.g., “it feels lik e w e’re in the
future, shifting dimensions”);
– Biological or living-related sounds (e.g., “flat electro cardiogram”, “cat purring”);
– Sounds that ev ok e a sensation (e.g., “it feels lik e I’m b et ween t w o vibrating
w alls”);
– Sounds that ev ok e an action (e.g., “it’s the sound made when you rub a w et
finger on a glass”).
T able 4 presen ts the p ercen tage of resp onses that fall in to eac h category , b y
condition. The monaural b eats most frequen tly evok ed an ob ject for 47.62% of
the participan ts. In particular, they often brough t to mind alarm-like ob jects or
ev en ts suc h as "b eeping sounds", "flashing lights", "sirens", or "w a ves", suggest-
ing a sense of alternation or something that disapp ears and returns. In con trast,
BB ga v e rise to more v aried asso ciations: 29.41% of the participan ts describ ed
them as biological sounds (e.g. "a bum bleb ee", "cardiac, blo o d"), 23.53% as an
atmosphere (e.g. "underground", "p ost-ap o calytpic films"), and 17.65% as ev ok-
ing an ob ject (e.g. "a clock", "a pure sin usoid"). P anning b eats w ere most often
asso ciated with an atmosphere (42. 86% of the participan ts, e.g. "industrial",
"lab oratory"), and nearly half of them sp ecifically describ ed a science-fiction-lik e
atmosphere(e.g. "alien films from the 80s", "w e’re in the future, w e’re c hanging
dimension"). Finally , alternate b eeps ev ok ed an ob ject for 36.84% of participan ts
and an atmosphere for 31.58%. They w ere often asso ciated with mac hinery (e.g.,
"engine", "helicopter blades", "w ashing mac hine spin cycle" or "rob ot computer
noises").
T able 4. P ercen tage of resp onse dep ending on the category of the ev o cation.
Categories Ob ject A tmosphere Biological Sensation Action
Monaural b eats 47,62 23,81 14,28 0 0
Binaural b eats 17,65 23,53 29,41 11,76 5,88
P anning b eats 28,57 42,86 0 9,52 4,76
Alternate b eeps 36,84 31,58 10,53 10,53 5,26
4 Discussion
Our main h yp othesis w as that, similar to the concept of a pr eferr e d temp o , eac h
individual has a preferred fundamen tal frequency and b eat frequency . F urther-
more, w e h yp othesized that, despite b eing sub jectiv e, there would be some de-
gree of agreemen t or consensus within the p opulation around a sp ecific frequency
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
77
Individualization of sound stim uli
range. The computation of ICC v alues (section 3.1) rev ealed mo derate to go o d
lev els of agreemen t b et w een participan ts. These results suggest a certain consis-
tency in the frequency v alues c hosen b y participan ts. The ICC w as particularly
high in the panning condition, whic h is also the one where the p erception of
mo v emen t w as strongest and spatial cues w ere most salien t. This salience lik ely
facilitated agreemen t among participan ts b y making frequency adjustmen t eas-
ier. In the other conditions, ICCs w ere mo derate, indicating a greater influence
of individual factors in sound construction compared to the panning condition.
Regarding preference, resp onses w ere relativ ely balanced across the differen t
conditions. This result supp orts the idea that the effect of an auditory stim ulus
on a listener’s psyc hological state requires an individualized approac h. Indeed, it
is difficult to imagine that a sound w e dislik e could help us relaxing or impro ving
our atten tion. The fact that BB w ere b oth the most and the least lik ed b y partic-
ipan ts clearly illustrates this p oin t. Ho w ev er, when examining preferences based
on fundamen tal frequency , it app eared that participan ts c hose lo w-frequency
sounds. Con v ersely , the least lik ed sounds w ere generally higher pitc hed. As
for b eating frequency , preferred monaural and BB w ere mostly lo w frequency . In
con trast, b eating frequency did not seem to influence the preference for alternate
b eeps. In terestingly , when monaural b eats and alternating b eeps w ere the least
lik ed sounds, their b eating frequencies w ere also lo w frequency , whereas dislik ed
BB tended to ha v e b eating of high frequencies. These patterns did not apply to
the panning condition, where neither fundamen tal nor b eat frequency app eared
to correlate with participan ts preference. T ak en together, the results suggested
that fundamen tal frequency pla ys a significan t role in auditory preference: for a
sound to b e appreciated, it m ust generally ha v e a lo w fundamen tal frequency . It
is also notew orth y that the median fundamen tal frequency for preferred BB w as
202 Hz. This v alue is half the v alue t ypically used in most BB studies, whic h
often cen ter around 400 Hz. Th us, lo w ering the fundamen tal frequency of BB in
future exp erimen ts ma y enhance their effectiv eness.
The rep orted ev o cations induced b y sound w ere ric h and div erse, reflecting
the complexit y of the p erceptual exp erience induced b y the auditory stim uli.
Ho w ev er, eac h t yp e of sound seemed to elicit sp ecific kinds of men tal imagery .
Monaural b eats and alternating b eeps mainly ev ok ed ob jects, although the na-
ture of these ob jects differed b et w een the t w o conditions. Binaural b eats, on
the other hand, generated more v aried asso ciations, including biological sounds.
P anning sounds primarily ev ok ed atmospheric impressions, particularly related
to science fiction. The ric hness of these ev o cations highligh ts the imp ortance of
considering not only the acoustic prop erties of sounds but also their emotional
and sym b olic dimensions in studies on h uman p erception.
Altogether, these results suggest that there are no univ ersal rules regarding
preference or the imagery ev ok ed b y auditory stim uli that can b e applied to all
listeners, reinforcing the need for an individualized approac h. This could help,
in particular, to impro v e the effects of binaural b eats, and more generally of
spatialized sounds, on cognitiv e states and to partly explain the con tradictory
results rep orted in the literature on this topic. It is imp ortan t to note that this
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
78
S. Sudre et al.
study is still ongoing. The obtained tendencies will ha v e to b e consolidated with
more participan ts. The next step of this study is to ev aluate the individualization
of sound stim uli in the con text of their impact on the men tal states of the
listeners. F uture researc h could also explore ho w these ev o cations and preferences
ev olv e o v er time ev en for a giv en listener. Indeed, the sound pro ductions collected
here reflect participan ts at a sp ecific momen t in time. It w ould therefore b e
in teresting to examine ho w these pro ductions shift or dev elop, as a reflection of
individuals in constan t ev olution.
A c kno wledgmen ts. W e thank Jo celyn Rozé for his help in dev eloping the sound
sitm uli and the in terface. This w ork, carried out within the Institute of Con v ergence
ILCB (ANR-16-CONV-0002), has b enefited from supp ort from the F renc h go v ernment
(F rance 2030), managed b y the F rench National Agency for Researc h (ANR) and the
Excellence Initiativ e of Aix-Marseille Univ ersit y (A*MIDEX).
Disclosure of In terests. The authors ha v e no comp eting in terests to declare that
are relev an t to the con ten t of this article.
References
1. Basu, S., Banerjee, B.: P oten tial of binaural b eats in terv en tion for impro ving mem-
ory and atten tion: insigh ts from meta-analysis and systematic review. Psyc hological
Researc h 87 (4), 951–963 (Jun 2023). h ttps://doi.org/10.1007/s00426- 022- 01706- 7,
h ttps://link.springer.com/10.1007/s00426- 022- 01706- 7
2. Chaieb, L., Wilp ert, E.C., Reb er, T.P ., F ell, J.: Auditory Beat Stimulation and
its Effects on Cognition and Mo o d States. F ron tiers in Psyc hiatry 6 , 70 (2015).
h ttps://doi.org/10.3389/fpsyt.2015.00070, h ttps://www.fron tiersin.org/article/10.
3389/fpsyt.2015.00070
3. Colzato, L.S., Barone, H., Sellaro, R., Hommel, B.: More atten tional fo cusing
through binaural b eats: evidence from the global–lo cal task. Psyc hological Re-
searc h 81 (1), 271–277 (Jan 2017). h ttps://doi.org/10.1007/s00426- 015- 0727- 0,
h ttp://link.springer.com/10.1007/s00426- 015- 0727- 0
4. Cresp o, A., Recuero, M., Galv ez, G., Begoña, A.: Effect of Binaural
Stim ulation on A tten tion and EEG. Arc hiv es of A coustics 38 (4), 517–
528 (Dec 2013). h ttps://doi.org/10.2478/aoa- 2013- 0061, h ttp://journals.pan.pl/
dlibra/publication/116825/edition/101556/con ten t
5. Dröge, K.: noScrib e. AI-p o w ered Audio T ranscription (2024), h ttps://gith ub.com/
k aixxx/noScrib e
6. Garcia-Argiba y , M., San ted, M.A., Reales, J.M.: Efficacy of binaural auditory b eats
in cognition, anxiet y , and pain p erception: a meta-analysis. Psyc hological Research
83 (2), 357–372 (Mar 2019). h ttps://doi.org/10.1007/s00426- 018- 1066- 8, h ttp://
link.springer.com/10.1007/s00426- 018- 1066- 8
7. Hammersc hmidt, D., F rieler, K., Wöllner, C.: Sp on taneous Motor T emp o: Inv es-
tigating Psyc hological, Chronobiological, and Demographic F actors in a Large-
Scale Online T apping Exp erimen t. F rontiers in Psyc hology 12 , 677201 (Jun 2021).
h ttps://doi.org/10.3389/fpsyg.2021.677201, h ttps://www.fron tiersin.org/articles/
10.3389/fpsyg.2021.677201/full
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
79
Individualization of sound stim uli
8. Ingendoh, R.M., P osn y , E.S., Heine, A.: Binaural b eats to en train the brain? A
systematic review of the effects of binaural b eat stim ulation on brain oscillatory
activit y , and the implications for psyc hological researc h and in terv en tion. PLOS
ONE 18 (5), e0286023 (Ma y 2023). h ttps://doi.org/10.1371/journal.p one.0286023,
h ttps://dx.plos.org/10.1371/journal.p one.0286023
9. Lane, J.D., Kasian, S.J., Ow ens, J.E., Marsh, G.R.: Binaural Auditory Beats Af-
fect Vigilance P erformance and Mo o d. Ph ysiology & Beha vior 63 (2), 249–252
(Jan 1998). h ttps://doi.org/10.1016/S0031- 9384(97)00436- 8, h ttps://linkingh ub.
elsevier.com/retriev e/pii/S0031938497004368
10. Leistik o, N.M., Madanat, L., Y eung, W.K.A., Stone, J.M.: Effects of gamma
frequency binaural b eats on atten tion and anxiet y . Curren t Psyc hology 43 (6),
5032–5039 (F eb 2024). h ttps://doi.org/10.1007/s12144- 023- 04681- 3, h ttps://link.
springer.com/10.1007/s12144- 023- 04681- 3
11. McAuley , J.D., Jones, M.R., Holub, S., Johnston, H.M., Miller, N.S.: The time
of our liv es: Life span dev elopmen t of timing and ev en t tracking. Journal of Ex-
p erimen tal Psyc hology: General 135 (3), 348–367 (2006). h ttps://doi.org/10.1037/
0096- 3445.135.3.348, h ttps://doi.apa.org/doi/10.1037/0096- 3445.135.3.348
12. Mo elan ts, D.: Preferred T emp o Reconsidered (2002)
13. Reedijk, S.A., Bolders, A., Colzato, L.S., Hommel, B.: Eliminating the A tten-
tional Blink through Binaural Beats: A Case for T ailored Cognitiv e Enhancemen t.
F ron tiers in Psyc hiatry 6 , 82 (2015). h ttps://doi.org/10.3389/fpsyt.2015.00082,
h ttps://www.fron tiersin.org/article/10.3389/fpsyt.2015.00082
14. Sudre, S., Kronland-Martinet, R., P etit, L., Rozé, J., Y stad, S., Aramaki, M.: A new
p ersp ectiv e on binaural b eats: In v estigating the effects of spatially mo ving sounds
on h uman men tal states. PLOS ONE 19 (7), e0306427 (Jul 2024). h ttps://doi.org/
10.1371/journal.p one.0306427, h ttps://dx.plos.org/10.1371/journal.p one.0306427
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
80
B ouba, K i k i , and t he Movem en t B et w een: T ow ard A
Sens ori m o t or H ypot hes i s o f Sound - Shape C ros s - m odal
C orres pondences
Lou C ou st on 1 , 2 , M i t suko A r a m a ki 2 , S ol vi Y st a d 2 , R i c ha r d K r onl a nd- Ma r t i ne t 2 a nd
Et i e nne Thor e t 1
1 Aix -Marseille Univ ersity, CNRS, Institut de Neurosciences d e la Timone (INT) , UMR7289
2 Aix -Marseille Univ ersity, CNRS, Per ception Representation Image Sound Music (PRISM)
UMR7061 , Marseille , Fran ce
[email protected]
[email protected]
Abstract . This study investigates th e role of implicit motor dynamics in
shaping cross-modal c orrespondences that underlie the Bouba/Kiki effe ct.
Building on the hypothesis of a sh ared amodal code linking movement , sound ,
and shape , we conducted two behaviour al ex periments in which participants
associat ed nonwords with either visual geometric shapes or friction sounds
derived from hand-dr awn movements . In both modalities , we observed
systematic m appings consistent with the Bouba/Kiki pattern . Crucially, we
manipulated the relative phase (RP) betw een the horizontal and vertical
components of movem ent trajectories, a k ey parameter describing spatial
coordination . Results reveal that RP significantly predi cts participants’
judgments across mod alities, forming a continuous perc eptual gradient from
" Kiki " (angular, phas e ≈ 0/π ) to " Bouba" (rounded , phase ≈ π /2). These findings
suggest that per ceptual judgments are modul ated not solely by sensory featur es,
but by an internal motor representation of the movement that could have
generat ed the stimulus. These preliminary findings support the ex istence of a
motor-based , amod al code that structures cross-modal correspondences .
Keywords: Cross-modal correspond ences, sound synthesis , friction sounds,
bouba/kiki effect .
1 I n tr od u c ti on
S ounds f i l l our da i l y l i ve s l i ke a soundt r a c k a t t a c he d t o our m ove m e nt s. Eve r y
m ove m e nt , a st e p, a br ush, a n i m pa c t , i s a nsw e r e d by a vi br a t i on, a r e sona nc e .
Move m e nt be c om e s a sour c e of sound, a nd sound, a w i t ne ss t o m ove m e nt .
Thi s or ga ni c r e l a t i onshi p be t w e e n m ove m e nt a nd sound spa ns a ge s a nd d i sc i pl i ne s.
S i nc e A nt i qui t y, G r e e k t hi nke r s a ssoc i a t e d c ol or s w i t h sounds, pe r c e i vi ng i n pa i nt i ng
a n e c ho of m usi c . I n t he f i e l d of da nc e , t hi s c onne c t i on i s e ve n m or e obv i ous. The
hum a n body r e sponds s pont a ne ousl y t o m usi c , a r e a l i t y obse r va bl e f r om e a r l y
c hi l dhood ; ba bi e s t he m se l ve s na t ur a l l y sw a y t o t he r hyt hm of sounds. Thi s
Pro c. of the 17th In ternational Symposium on CMMR, London, UK, Nov. 3-7, 2025
81
2 L . Couston and al .
phe nom e non w a s f or m a l i z e d by F r a i sse [ 1 ] , w ho show e d t ha t m us i c c a n ha ve a
dyna m oge ni c e f f e c t , m e a ni ng i t pr om ot e s spont a ne ous m ot or a c t i va t i on.
A l t hough t he se pa r a t i on be t w e e n da nc e a nd m usi c ha s be e n a n a r e a of
e x pe r i m e nt a t i on i n c ont e m por a r y da nc e , not a bl y w i t h t he e x pl or a t i on of s i l e nc e by
Me r c e C unni ngha m or Tr i sha B r ow n [ 2 , 3 ] , i t i s a l so e sse nt i a l t o r e c a l l t ha t t he
r e l a t i onshi p be t w e e n body a nd s ound c a n ope r a t e i n t he o t he r di r e c t i on. I nde e d, i f
m u si c c a n gui de da nc e , t he body c a n a l so be t he dr i ve r of sound. I ni t i a l l y, i n a n
i n t ui t i ve w a y, a m usi c i a n m a y a da pt t he i r p l a yi ng t o t he i m pul se s of a m ovi ng body.
The n, m or e f or m a l l y, t ha nks t o t e c hnol ogi c a l t ool s. The a dve nt of i nt e r a c t i ve de vi c e s
ope ns ne w pe r spe c t i ve s, w he r e m ove m e nt pa r a m e t e r s, suc h a s spe e d, di r e c t i on, or
c ur va t ur e , c a n be t r a nsl a t e d i nt o sound m a t e r i a l , i nf l ue nc i ng t i m br e , i n t e nsi t y, or
t e x t ur e of t he sound.
The se c or r e sponde n c e s be c om e t he c or e of num e r ous a r t i st i c a nd sc i e nt i f i c
r e se a r c h. S om e c hor e ogr a phi c w or ks i nc or por a t e t hi s r e a l - t i m e i nt e r a c t i on be t w e e n
m ove m e nt a nd sound, suc h a s R - A - U - X - A by A i na A l e gr e , w he r e t he m us i c i a n
sha pe s a sound m a t e r i a l i n d i r e c t di a l ogue w i t h t he da nc i ng body. Thi s a ppr oa c h
sugge st s t ha t a be t t e r unde r st a ndi ng of t he l i nks be t w e e n m ove m e nt a nd sound
pe r c e pt i on c oul d ope n ne w pa t hs f or c r e a t i on, but a l so, m or e br oa dl y, f or our
unde r st a ndi ng of t he de e p c onne c t i ons be t w e e n se nsor y a nd m ot or m oda l i t i e s.
Thi s e x pl or a t i on of t he l i nk be t w e e n m ove m e nt a nd sound l e d us t o f or m ul a t e a
br oa de r hypot he si s : w ha t i f t he pe r c e pt i on of a vi sua l sha pe or a sound w a s not sol e l y
a m a t t e r of vi si on or he a r i ng, bu t a l so i nvol ve d si m ul a t i ng t he m ove m e nt t ha t m i gh t
ha ve pr oduc e d t he m [ 4 , 5 , 6 ] ?
Thi s i de a i s suppor t e d by se ve r a l w or ks i n c ogni t i ve psyc hol ogy a nd ne ur osc i e nc e .
F r e yd [ 7 ] show e d t ha t m e m or i z i ng a n i m a ge i nvol vi ng m ove m e nt , suc h a s a j u m p, i s
a c c om pa ni e d by a m e nt a l a nt i c i pa t i on of t he m ove m e nt , sugge st i ng t ha t t he br a i n
s i m ul a t e s t he a c t i on be yond t he obse r ve d sc e ne . O t he r s t udi e s, suc h a s Longc ha m p e t
a l . [ 8 ] , r e ve a l e d t ha t pe r c e i vi ng l e t t e r s or know n sha pe s a c t i va t e s m ot or a r e a s
i nvo l ve d i n t he i r pr oduc t i on. The se r e sul t s c onve r ge t ow a r d t he i de a t ha t vi sua l
pe r c e pt i on of a sha pe e nga ge s a m o t or s i m ul a t i on of t he m ove m e nt t ha t ha s dr a w n i t .
W e t he n dr a w a pa r a l l e l w i t h a udi t or y pe r c e pt i on: i f w e he a r a sound r e sul t i ng f r om
f r i c t i on, suc h a s a pe nc i l on pa pe r or a n obj e c t dr a gge d on t he gr ound, c oul d w e a l so
s i m ul a t e t he m ove m e nt t ha t pr oduc e d i t ? A nd c oul d t hi s m e nt a l l y r e c onst r u c t e d
m ove m e nt e voke a sha pe ?
The w or ks of Thor e t e t a l .[ 9 , 10 , 11 ] de ve l op t hi s i de a w i t hi n t he f r a m e w o r k of
m u l t i m oda l pe r c e pt i on of m ove m e nt s a nd sha pe s. F i r st , Thor e t e t a l . [ 9 ] de m ons t r a t e d
t ha t he a r i ng a dr a w i ng sound a l l ow s pe r c e i vi ng t he dyna m i c c ha r a c t e r i st i c s of t he
m ove m e nt t ha t pr oduc e d i t , not a bl y t he s pe e d pr of i l e , t ha t i s, how t he spe e d of t he
m ove m e nt va r i e s ove r t i m e . Thi s pr of i l e i s e sse nt i a l t o r e c ogni z e a nd di f f e r e nt i a t e
sha pe s by t he i r dyna m i c s, l i nki ng sound t e m por a l i t y a nd gr a phi c t r a j e c t or y. The n,
Thor e t e t a l . [ 10 ] show e d t ha t sound i nde e d i nf l ue nc e s t he se nsor i m ot or
sync hr oni z a t i on w i t h a vi sua l m ove m e nt , f or e x a m pl e , by bi a si ng t he r e pr oduc t i on of
a n e l l i pt i c a l t r a j e c t or y t ow a r d a m or e c i r c ul a r sha pe de pe ndi ng on t he sound
pr ope r t i e s a nd a ssoc i a t e d ve l oc i t y pr of i l e . F i na l l y, Thor e t e t a l . [ 11 ] hi ghl i ght e d t ha t
t he se phe nom e na a r e not m e r e l y se nsor y, but r e f l e c t t he i m pr i nt of a n a c t i on
Pro c. of the 17th In ternational Symposium on CMMR, London, UK, No v. 3-7, 2025
82
Bouba , Kiki , and the Mov ement Between 3
r e pr e se nt a t i on i n pr ot ot ype ge om e t r i c sha pe s. The y t hus sugge st t ha t he a r i ng a n
e l l i pt i c a l m ove m e nt a c t i va t e s a c om m on pr ot ot ypi c a l r e pr e se nt a t i on i n t e gr a t i ng
sound, v i si on, spe e d pr of i l e , a nd m o t or s i m ul a t i on, w hi c h st r uc t ur e s sha pe pe r c e pt i on.
B a se d on t he se w or ks, w e f or m ul a t e d our hypot he si s: t he r e e x i st s a sha r e d a m oda l
c odi ng be t w e e n sha pe , sound, a nd m ove m e nt , w hi c h unde r l i e s c r oss- m oda l
c or r e sponde n c e s suc h a s i n t he w e l l - know n B ouba - K i ki e f f e c t [ 12 ] . Thi s e f f e c t
de not e s a sys t e m a t i c a ssoc i a t i on be t w e e n nonse nse w or ds ( suc h a s " B ouba " a nd
" K i k i " ) a nd ge om e t r i c sha pe s ( r ounde d a nd a ngul a r , r e spe c t i ve l y) , r obust l y obse r ve d
a c r oss c ul t ur e s, l a ngua ge s, a nd a ge s [ 13 ] . A r e c e nt a r t i c l e by W i n t e r a nd P e r l m a n [ 14 ]
r e vi e w s t he m a i n hypot he se s pr opose d t o e x pl a i n t hi s phe nom e non, di st i ngui shi ng
t w o m a i n d i r e c t i ons: t he a r t i c ul a t or y e x pl a na t i on a nd t he a c oust i c e x pl a na t i on. I n t h i s
pe r spe c t i ve , F or t e a nd S c hw a r t z [ 15 ] pr opose a n e c ol ogi c a l a ppr oa c h t ha t goe s
be yond t hi s oppos i t i on by r oot i ng t he B ouba - K i ki e f f e c t i n t he se nsor y e x pe r i e nc e of
t he physi c a l w or l d: t he sounds of " B ouba " a nd " K i k i " w oul d r e spe c t i ve l y e voke t he
sounds pr oduc e d by r ound obj e c t s ( r a t he r l ow - pi t c he d a nd c ont i nuous) or a ngul a r
obj e c t s ( r a t he r h i gh- p i t c he d a nd di sc ont i nuous) , e st a bl i shi ng a n i c oni c a ssoc i a t i on
be t w e e n a udi t or y pe r c e pt i on a nd vi sua l sha pe . Mor e ove r , t he a r t i c ul a t or y hypot he si s,
suppor t e d by V a i ni o e t a l .[ 16 ] , posi t s t ha t m ot or c onf i gur a t i ons i nvo l ve d i n pr oduc i ng
t he sounds ( suc h a s l i p r oundi ng f or " B ouba " or t ongue t e nsi on f or " K i k i " ) a l so
c ont r i but e t o vi sua l l y e voki ng c e r t a i n sha pe s.
C om pl e m e nt i ng t he se a ppr oa c he s, w e f or m ul a t e d a m ot or hypot he si s: t he se c r oss-
m oda l c or r e sponde nc e s c oul d r e l y on a m e nt a l r e pr e s e nt a t i on of t he m ove m e nt
ne e de d t o pr oduc e t he sha pe or sound, i n ot he r w or ds, on a m ot or s i m ul a t i on. F or
e x a m pl e , " K i ki " c oul d e voke a n a ngul a r , f a st , j e r ky m ove m e nt , c l ose t o t ha t ne e de d t o
dr a w a poi nt e d f i gur e , w hi l e " B ouba " w oul d c or r e spond t o a c ont i nuous, c i r c ul a r ,
f l ui d m ove m e nt . I n ot he r w or ds, be yond a c oust i c or vi sua l f e a t ur e s, i t i s t he sha r e d
m o t or dyna m i c s t ha t c oul d st r uc t ur e t he se a ssoc i a t i ons.
To t e st t hi s hypo t he si s, w e c onduc t e d a f i r st e x pe r i m e nt i n w hi c h pa r t i c i pa nt s
a ssoc i a t e d s t i m ul i , e i t he r vi sua l sha pe s or t he i r c or r e spondi ng f r i c t i on sounds, w i t h t he
nonse nse w or ds B ouba or K i ki . The sound s w e r e synt he si z e d f r om ha nd- dr a w n
t r a j e c t or i e s, c onve r t i ng m ove m e nt spe e d i nt o t i m br e va r i a t i ons, f ol l ow i ng a m ode l
ba se d on C ona n e t a l . [ 17 ] . Th i s a l l ow e d us t o a sse ss w he t he r si m i l a r m ot or dyna m i c s,
pe r c e i ve d t hr ough d i f f e r e nt m oda l i t i e s, w oul d yi e l d c om pa r a bl e a ssoc i a t i ons.
R e sul t s r e ve a l e d a l i nk be t w e e n pa r t i c i pa nt s’ j udg m e nt s a nd t he r e l a t i ve pha se ( R P )
of t he unde r l yi ng m ove m e nt , a pa r a m e t e r de sc r i bi ng t he c oor di na t i on be t w e e n
hor i z ont a l a nd ve r t i c a l m ove m e nt c om pone nt s. Low e r pha se s pr oduc e m or e a ngul a r
t r a j e c t or i e s; hi ghe r pha se s r e sul t i n r ounde r one s.
To i nve st i ga t e t he e f f e c t of t hi s pa r a m e t e r m or e syst e m a t i c a l l y, w e de si gne d a
se c ond e x pe r i m e nt i n w hi c h R P w a s pr ogr e ssi ve l y va r i e d t o ge ne r a t e a c ont i nuum
f r om a ngul a r t o r ounde d sha pe s, w i t h c or r e spondi ng f r i c t i on s ounds. Thi s a ppr oa c h,
i n spi r e d by t he m ove m e nt a na l yse s m a de by W a m a i n e t a l . [ 18 ] , a l l ow e d us t o t e st t he
i m pa c t of R P m odu l a t i on on B ouba / K i ki j udgm e nt s, a nd w he t he r i t ge ne r a l i z e s a c r oss
m oda l i t i e s.
Pro c. of the 17th In ternational Symposium on CMMR, London, UK, Nov. 3-7, 2025
83
4 L . Couston and al .
Toge t he r , t he se t w o e x pe r i m e nt s a i m e d t o e va l ua t e w he t he r R P c a n s t r uc t ur e c r oss-
m oda l c or r e spond e nc e s be t w e e n sha pe , sound, a nd non- w or ds, a nd t o e x pl or e a
poss i bl e s ha r e d, a m oda l , m o t or c ode unde r l yi ng t he B ouba - K i ki e f f e c t .
Fig . 1 . Historical shap es of the Bouba-Kiki effect [ 19 ]
2 Exp e r i me n t 1: Exp l or i n g th e Bou b a/ K i k i Effe c t fr om th e
V i s u al S h ap e s a n d th e S ou n d of D r aw i n g M ove me n ts
Thi s f i r st e x pe r i m e nt e x pl or e d t he e x i st e nc e of a B ouba - K i ki e f f e c t i n t he a udi t or y
m oda l i t y, r e l yi ng not on l i ngu i st i c sounds but on f r i c t i on sounds de r i ve d f r om t he
t r a c i ng of ge om e t r i c sha pe s. The goa l w a s t o de t e r m i ne w he t he r t he se sounds,
c a r r yi ng t he dyna m i c s of a gr a phi c m ove m e nt , c oul d be spont a ne ousl y a ssoc i a t e d
w i t h t he nonse nse w or ds “ B ouba ” a nd “ K i ki ,” s i m i l a r l y t o w ha t i s c l a ssi c a l l y
obse r ve d w i t h v i sua l sha pe s.
To do so, w e de si gne d t w o a ssoc i a t i on t a sks: a vi sua l t a sk, i n w hi c h ge om e t r i c
sha pe s w e r e pr e se nt e d on sc r e e n, a nd a n a udi t or y t a sk, i n w hi c h pa r t i c i pa nt s ha d t o
j udge sounds ge ne r a t e d f r om t he se sa m e t r a c i ngs. Th i s dua l a ppr oa c h a l l ow e d us t o
c om pa r e t he e f f e c t s of m oda l i t y on t he obse r ve d a ssoc i a t i ons w hi l e w or ki ng f r om a
c om m on ba si s — t ha t of t he dr a w n m ove m e nt a nd, m or e i m por t a nt l y, w he t he r sounds
a l one , c a r r yi ng on l y t he dyna m i c of m ove m e nt , w e r e suf f i c i e nt t o c onve y suc h a n
a ssoc i a t i on.
2.1 M e t h od s
R e c or d i n g t h e S h a p e s. W e be ga n by r e c or di ng t he dr a w i ng of ge om e t r i c sha pe s on a
t ouc hpa d. F or t hi s, w e use d t he l i bi nput t oo l unde r U bunt u, w hi c h a l l ow s c a pt ur i ng
i npu t de vi c e e ve nt s ( suc h a s poi n t e r m ove m e nt s) . W e de ve l ope d a sc r i pt t o c a pt ur e
P O I N TER _MO TI O N e ve nt s a nd sa ve t he m i nt o a .c sv f i l e . Thi s f i l e c ont a i ne d, f or
e a c h t i m e s t e p, t he poi n t e r ’ s di spl a c e m e nt i n x a nd y r e l a t i ve t o t he pr e vi ous i nst a nt .
F r o m t he se da t a , t he t r a j e c t or i e s of t he m ove m e nt s pe r f or m e d on t he t ouc hpa d w e r e
r e c onst r uc t e d. The r e c or di ngs st a r t e d w he n t he f i nge r t ouc he d t he pa d a nd e nde d
w he n i t l e f t . N o i nst r uc t i ons w e r e gi ve n r e ga r di ng spe e d or dur a t i on, t o pr e se r ve na t u -
r a l va r i a bi l i t y i n m ove m e nt e x e c ut i on. Tw e l ve pse udo- r a ndom sha pe s w e r e dr a w n by
t he f i r st a ut hor , f or m i ng t he ba si s of t he st i m u l i . The se sha pe s w e r e c ha r a c t e r i z e d by
t he x a nd y posi t i on s ove r t i m e . W e a r e a w a r e t ha t t he f a c t t ha t a l l t he sha pe s w e r e
dr a w n by a s i ngl e pe r son c oul d i n t r oduc e a bi a s i n t he da t a . H ow e ve r , w e c onsi de r
t h i s a s a f i r st st e p t o e st a bl i sh a ba se l i ne , a nd f ut ur e w or k w i l l a i m t o di ve r si f y t he
sour c e of dr a w i ngs t o a c c ount f or i ndi vi dua l d i f f e r e nc e s i n ge st ur e e x e c ut i on.
Pro c. of the 17th In ternational Symposium on CMMR, London, UK, No v. 3-7, 2025
84
Bouba , Kiki , and the Mov ement Between 5
V i s u al S t i m u l i . V i sua l st a t i c st i m ul i w e r e ge ne r a t e d f r om t he se m a nua l t r a c i ngs. The
12 sha pe s w e r e di vi de d i n t o t w o br oa d c a t e gor i e s: r ounde d a nd a ngul a r . Ea c h sha pe
r e pr e se nt e d a n e x t r e m e of t he vi sua l spe c t r um i n t e r m s of m ove m e nt dyna m i c s, a l l ow -
i ng u s t o c r e a t e a se t of 24 vi sua l s t i m ul i .
Fig . 1 . Ex ample of two visual stimuli (Boub a - Kiki)
A u d i t or y S t i m u l i . Ea c h v i sua l s ha pe w a s a ssoc i a t e d w i t h a synt he si z e d f r i c t i on sound
de si gne d t o r e f l e c t t he dyna m i c s of t he m ove m e nt t ha t pr oduc e d t ha t sha pe .
W e ba se d our a ppr oa c h on t he w or k of C ona n e t a l . [ 17 ] , w hi c h s how e d t ha t t he
t i m br e of f r i c t i on sounds, suc h a s a pe nc i l on pa pe r , i s r e l a t e d t o t he ve l oc i t y of t he
m ove m e nt , a nd t he sur f a c e r oughne ss. The i r m ode l c onsi de r s sound a s t he r e sul t of
i n t e r a c t i on be t w e e n a sour c e ( t he pe nc i l ) a nd a r e sona t or ( t he pa pe r a nd t a bl e ) , w i t hi n
a n e c ol ogi c a l pe r c e pt i on pe r spe c t i ve . F ol l ow i ng t hi s a ppr oa c h, w e ge ne r a t e d a udi t or y
s t i m ul i f r om t he m a nua l l y dr a w n t r a j e c t or i e s. F or e a c h s ha pe , w e e x t r a c t e d t he
ve l oc i t y pr of i l e , obt a i ne d a s t he de r i va t i ve of t he posi t i ons, a nd use d i t t o dyna m i c a l l y
f i l t e r a p i nk noi se w i t h a l ow pa ss f i l t e r w hose c ut of f f r e que nc y w a s l i ne a r l y m a ppe d
t o t he f r e que nc y of t he m ove m e nt .
The r e sul t i ng si gna l w a s t he n c onvol ve d w i t h a w oode n i m pa c t sound t o i n t r oduc e
a r e sona nt c om pone nt , f ol l ow i ng t he sour c e –f i l t e r m ode l pr opose d by Thor e t
[ 17 , 20 , 21 ] . Th i s pr oc e ss a l l ow e d us t o pr oduc e sounds i n w hi c h t he ki ne m a t i c s of t he
m ove m e nt a r e e nc ode d i n t he spe c t r a l c ont e nt , si m ul a t i ng r e a l i st i c f r i c t i on. I n t ot a l , 24
dr a w i ng s ounds w e r e ge ne r a t e d, e a c h c or r e spondi ng t o one of t he 24 r e c or de d v i sua l
sha pe s.
2.2 Par t i c i p an t s
The s t udy i nvo l ve d ni ne pa r t i c i pa nt s ( m e a n a ge : 32.2 ± 18.6 ye a r s o l d) , i nc l udi ng f i ve
m e n a nd f our w om e n. A l l w e r e r i gh t - ha nde d a nd w or e he a dphone s dur i ng t he
e x pe r i m e nt . S i x w e r e non- m usi c i a ns or non- spe c i a l i st s i n s ound, t w o w e r e m usi c i a ns,
a nd one w a s a sound pr of e ssi ona l .
2.3 T ask an d Pr oc e d u r e .
The e x pe r i m e nt w a s c onduc t e d on l i ne vi a t he Te st a bl e p l a t f or m [ 22 ] . P a r t i c i pa nt s
w e r e a ske d t o be i n a qui e t , d i st r a c t i on- f r e e e nvi r onm e nt a nd t o w e a r he a dphone s.
Ea c h pe r f or m e d t w o t a sks, one vi sua l a nd one a udi t or y, pr e se nt e d i n r a ndom i z e d
Pro c. of the 17th In ternational Symposium on CMMR, London, UK, Nov. 3-7, 2025
85
6 L . Couston and al .
or de r . Ea c h t a sk w a s pr e c e de d by t w o f a m i l i a r i z a t i on e x a m pl e s t o i nt r oduc e t he
B ouba / K i ki a ssoc i a t i on pr i nc i pl e i n bot h v i sua l a nd a udi t or y m oda l i t y.
V i s u al T ask . A ge om e t r i c sha pe a ppe a r e d a t t he c e nt e r of t he sc r e e n. P a r t i c i pa nt s
ha d t o i nd i c a t e w he t he r t he sha pe e voke d t he non se nse w or d “ B ouba ” or “ K i ki ” by
c l i c ki ng on one of t w o but t ons on t he sc r e e n ( “ B ouba ” on t he l e f t , “ K i ki ” on t he r i ght )
or usi ng ke yboa r d a r r ow s ( l e f t = B ouba , r i ght = K i ki ) .
A u d i t or y T ask . A f r i c t i on sound c or r e spondi ng t o a sha pe ’ s t r a c i ng w a s pl a ye d.
W i t h no i m a ge show n, pa r t i c i pa nt s w e r e a ske d t o i m a gi ne t he unde r l yi ng sha pe a nd
t he n c hoose “ B ouba ” or “ K i ki ” ba se d on w ha t t he sound e voke d, usi ng t he sa m e
r e sponse m e t hod a s i n t he vi sua l t a sk.
S t i m u l i w e r e pr e se nt e d one by one , w i t h a f i x e d 800 m s i n t e r - s t i m u l us i nt e r va l
a f t e r e a c h r e sponse . The w hol e e x pe r i m e nt l a st e d a bout 10 m i nut e s.
2.4 D at a A n al ysi s
The m a i n goa l of t hi s e x pe r i m e nt w a s t o a sse ss w he t he r B ouba / K i ki a ssoc i a t i ons a r e
c onsi st e nt a c r oss v i sua l a nd a udi t or y m oda l i t i e s, a nd w he t he r t he y c a n be e x pl a i ne d
by ge st ur a l pa r a m e t e r s. The R P be t w e e n t he x ( t ) a nd y ( t ) ( 1) c om pone nt s of t he m ove -
m e nt w a s m or e spe c i f i c a l l y a na l yz e d. Thi s pa r a m e t e r c a pt ur e s t he t e m por a l r e l a t i on -
sh i p be t w e e n hor i z ont a l a nd ve r t i c a l m ove m e nt c om pone nt s. R P w a s c om pu t e d
t hr oughout t he t r a j e c t or y, a nd va l ue s a t t he t i m e of r e sponse w e r e use d f or a na l ysi s.
R P ( t ) = ϕ y ( t ) − ϕ x ( t ) ( 1 ) ( w he r e ϕ x ( t ) e t ϕ y ( t ) a r e t he i nst a nt a ne ous pha se s of t he x a nd y
s i gna l s a t t i m e t , r e spe c t i ve l y.)
W e f i r st e x a m i ne d w he t he r r e sponse pa t t e r ns w e r e c onsi s t e nt a c r oss m oda l i t i e s f or
t he sa m e st i m u l us. To t hi s e nd, r e sponse s w e r e b i na r i z e d ( " B ouba " = 0, " K i ki " = 1) , a nd
t he m e a n r e sponse s pe r st i m ul us w e r e c om pa r e d be t w e e n vi sua l a nd a udi t or y c ondi -
t i ons.
To e x pl or e how R P r e l a t e s t o pe r c e pt ua l j udgm e nt s, w e vi sua l i z e d t he di s t r i but i on
of R P usi ng v i ol i n pl ot s a c r oss t he c r osse d c ondi t i on s of Moda l i t y ( vi sua l v s. a udi -
t or y) a nd pe r c e i ve d S ha pe ( B ouba vs. K i ki ) . F i na l l y, a t w o- w a y A N O V A w a s c on -
duc t e d t o t e st t he e f f e c t s of R P a nd m oda l i t y, a s w e l l a s t he i r i nt e r a c t i on, on pa r t i c i -
pa nt s’ r e sponse s.
2.5 R e su l t s
The a ve r a g e r e sponse s a c r oss m oda l i t i e s r e ve a l e d c onsi st e nt B ouba / K i ki a ssoc i a t i ons
( F i g. 1 ( A ) ) . I n t he a udi t or y c ondi t i on, sounds de r i ve d f r om r ound m ove m e nt s ( c a t e -
gor y “ b” ) w e r e m or e f r e que nt l y j udge d a s B ouba ( M= .31, S D = .47) , w hi l e t hose f r om
a ngul a r m ove m e nt s ( c a t e gor y “ k” ) w e r e m or e of t e n l a be l e d a s K i ki ( M= .66, S D = .48) .
The a na l ysi s, w hi c h use d R a t i ona l i z e d A r c si n U ni t sc or e s ( F i g. 1 ( A ) ) , i ndi c a t e s t ha t
t he st i m u l i w e r e ge ne r a l l y pe r c e i ve d a s e x pe c t e d, a l t hough s om e st i m ul i w e r e m or e
a m bi guous t ha n ot he r s. Thi s va l i ds t he possi b i l i t y of r e c a l l i ng t he B ouba / K i k i e f f e c t
f r om t he sounds e voki ng t he m ot i on of som e one dr a w i ng ge om e t r i c sha pe s.
Pro c. of the 17th In ternational Symposium on CMMR, London, UK, Nov. 3-7, 2025
86
Bouba , Kiki , and the Mov ement Between 7
Fig . 1 . (A) Mean Rationalized Arcsine Unit (RAU) scor es for each sensory modality a cross dif -
ferent stimuli - (B) Violin Plot Showing Average D ensity Distributions of RP by Modality
(Visual , Audio) and Shape Label (Bouba, Kiki) — Ex periment 1
W e t he n e x a m i ne d how R P c ont r i but e s t o t he se j udgm e nt s. A vi ol i n pl ot of t he pha se
di st r i but i on ( F i g. 1 ( B ) ) show s d i st i nc t pha se pa t t e r ns a ssoc i a t e d w i t h B ouba a nd K i ki
r e sponse s. N ot a bl y, t he di s t r i but i on r e ve a l s a pe a k a r ound 0 r a di a ns f or K i ki r e -
sponse s a nd a r ound π / 2 r a di a ns f or B ouba r e sponse s, i nd i c a t i ng d i st i nc t pha se pr e f e r -
e nc e s. The se pa t t e r ns w e r e c onsi st e nt a c r oss m oda l i t i e s, sugge st i ng t ha t R P se r ve s a s
a c r oss- m oda l pe r c e pt ua l c ue .
A t w o- f a c t or A N O V A c onf i r m e d a s t r ong m a i n e f f e c t of R P on c a t e gor i z a t i on
( F ( 7) = 49.82, p < .001) , w i t h a s m a l l e r but s i gni f i c a nt e f f e c t of m oda l i t y ( F ( 1) = 12.49,
p= .0004) . The si gn i f i c a nt i nt e r a c t i on ( F ( 7) = 3.33, p= .0027) i ndi c a t e s t ha t t he i nf l ue nc e
of R P va r i e s a c r oss se nsor y m oda l i t i e s.
2.6 D i sc u ss i on
Ex pe r i m e nt 1 c onf i r m s t ha t t he B ouba –K i ki e f f e c t e x t e nds t o t he a udi t or y m oda l i t y,
w i t h f r i c t i on sounds f r om m ove m e nt s i nf l ue nc i ng s ha pe j udgm e nt s si m i l a r l y t o vi sua l
s t i m ul i . Thi s suppor t s t he i de a of a sha r e d ge st ur a l ba si s a c c e ssi bl e a c r oss se nse s, pos -
s i bl y r e f l e c t i ng se nsor i m ot or pr oc e sse s unde r l yi ng pe r c e pt i on.
The c l e a r l i nk be t w e e n R P a nd j udgm e nt s i n bot h m oda l i t i e s r e i nf or c e s t hi s vi e w .
H ow e ve r , a s t h i s e x pe r i m e nt w a s e x pl or a t or y, f ur t he r w or k i s ne e de d t o i de nt i f y t he
spe c i f i c m ove m e nt pa r a m e t e r s i nvol ve d, m ot i va t i ng t he m or e t a r ge t e d a ppr oa c h of
Ex pe r i m e nt 2.
3 Exp e r i me n t 2: Effe c t of Th e R e l ati ve P h a s e o n Bou b a/ K i k i
A s s oc i ati on s
The f i r st e x pe r i m e nt use d f r e e l y dr a w n sha pe s, l e t t i ng na t ur a l va r i a t i ons i n t r a c i ng dy -
na m i c s spont a ne ousl y i nf l ue nc e t he r e sul t s. I n c ont r a st , t he se c ond e x pe r i m e nt a i m e d
t o pr e c i se l y c ont r ol a nd m a ni pul a t e t he R P , a ke y dyna m i c pa r a m e t e r t o t e s t w he t he r i t
a l one gove r ns t he c r oss- m oda l c or r e sponde n c e .
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
87
8 L . Couston and al .
W e ge ne r a t e d a c ont i nuum of st i m ul i f r om a si ngl e r e f e r e n c e m ove m e nt by
gr a dua l l y m od i f yi ng t he R P be t w e e n t he x ( t ) a nd y( t ) . Th i s pr oduc e d t r a j e c t or i e s
r a ngi ng f r om ve r y c i r c ul a r t o ve r y a ngul a r sha pe s. W e he r e by sought t o c onf i r m t ha t i t
i s i nde e d t he spa t i a l st r uc t ur e of t he m ove m e nt , a nd not a n i sol a t e d vi sua l or a c oust i c
pr ope r t y, t ha t dr i ve s B ouba / K i ki j udgm e nt s.
3.1 M e t h od s
G e n e r at i on of t h e S t i m u l u s C o n t i n u u m . W e se l e c t e d si x ge om e t r i c sha pe s f r om t he
s t i m ul i of Ex pe r i m e nt 1, t ha t w e r e no t a m bi guous i n t he r e sul t s.
F or e a c h dr a w i ng, w e a ppl i e d a H i l be r t t r a nsf or m se pa r a t e l y t o t he x a nd y si gna l s
t o e x t r a c t t he i r e nve l ope s a nd i nst a nt a ne ous pha se s. R P w a s c om pu t e d a s t he
di f f e r e nc e , a t e a c h t i m e poi nt , be t w e e n t he i nst a nt a ne ous pha se s of y a nd x . Thi s
m e a sur e r e f l e c t s t he c yc l i c c oor di na t i on be t w e e n t he t w o spa t i a l di m e nsi ons.
To c r e a t e a c ont i nuum of 15 va r i a t i ons f or e a c h or i gi na l sha pe ( F i g. 1 ) , w e sys t e m -
a t i c a l l y m odi f i e d t he R P be t w e e n t he x a nd y c om pone nt s of t he m ove m e nt . S t a r t i ng
f r om a n or i g i na l t r a j e c t or y, w e a dde d a c onst a nt pha se of f se t ( Δ φ ) t o t he i nst a nt a ne ous
pha se of t he y- c om pone nt , a nd t he n r e c om put e d a m od i f i e d y- si gna l ( y′ ( t ) ) . Thi s
pr oc e ss w a s r e pe a t e d f or 15 va l ue s of Δ φ , l i ne a r l y spa c e d be t w e e n –π a nd π , r e sul t i ng
i n 15 d i st i nc t va r i a t i ons of t he i ni t i a l sha pe . Th i s m a ni pul a t i on a l t e r e d spa t i a l c oor di -
na t i on, t he r e by c ha ngi ng t he posi t i on t r a j e c t or i e s a nd t he i r ve l oc i t y pr of i l e s, w hi c h i n
t ur n a f f e c t e d bot h t he pe r c e i ve d sha pe ( r ounde r or m or e a ngul a r ) a nd t he dyna m i c s
e nc ode d i n t he a ssoc i a t e d sounds.
Fig . 1 . Ex ample of a continuum made of 15 variations of a single shape
S yn t h e si s of A u d i t or y S t i m u l i . The m odi f i e d t r a j e c t or i e s ( x ( t ) , y’ ( t ) ) w e r e t he n
use d t o synt he si z e sounds, f ol l ow i ng t he sa m e m e t hod a s i n Ex pe r i m e nt 1.
3.2 Par t i c i p an t s
The e x pe r i m e nt i nvol ve d 9 pa r t i c i pa nt s ( m e a n a ge : 30.6 ± 12.6 ye a r s) , i nc l udi ng 5
m e n a nd 4 w om e n, a l l r i ght - ha nde d. A m ong t he m , 4 w e r e ne i t he r m usi c i a ns nor
sound spe c i a l i st s, 3 w e r e m usi c i a ns, a nd 2 w e r e sound pr of e ssi ona l s. A l l pa r t i c i pa nt s
r e por t e d usi ng he a dphone s dur i ng t he e x pe r i m e nt .
3.3 T ask an d Pr oc e d u r e
The e x pe r i m e nt w a s c onduc t e d on l i ne usi ng t he Te st a bl e pl a t f or m [ 23 ] , unde r
c ondi t i ons si m i l a r t o t ho se of Ex pe r i m e nt 1. P a r t i c i pa nt s w e r e a ske d t o be i n a qui e t
e nvi r onm e nt a nd t o use he a dphone s. Tw o t a sks, one v i sua l a nd one a udi t or y, w e r e
pr e se nt e d i n r a ndom i z e d or de r , e a c h pr e c e d e d by t hr e e f a m i l i a r i z a t i on t r i a l s. Ea c h
Pro c. of the 17th In ternational Symposium on CMMR, London, UK, No v. 3-7, 2025
88
Bouba , Kiki , and the Mov ement Between 9
t a sk c onsi st e d of a c ont i nuu m of 15 va r i a t i ons ge ne r a t e d f r om 6 di f f e r e nt sha pe s,
t o t a l i ng 90 st i m ul i pe r m oda l i t y. I n t he vi sua l t a sk , sha pe s w e r e pr e se nt e d w i t h no
t i m e l i m i t , a nd pa r t i c i pa nt s c oul d r e spond f r e e l y. I n t he a udi t or y t a sk , e a c h f r i c t i on
sound w a s pl a ye d i n f ul l ( a bout 4 se c onds) , a nd a ga i n, r e sponse s w e r e m a de w i t hout
t i m e c onst r a i nt s. A f i x e d 800 m s i nt e r va l se pa r a t e d e a c h st i m ul us. D ue t o t he l a r ge r
num be r of s t i m ul i , t he t ot a l e x pe r i m e nt dur a t i on w a s e st i m a t e d a t a r ound 20 m i nut e s.
3.4 D at a A n al ysi s
The m a i n obj e c t i ve of t hi s e x pe r i m e nt w a s t o e va l ua t e t he i nf l ue nc e of t he R P
be t w e e n t he x ( t ) a nd y( t ) c om pone nt s of t he gr a phi c m ove m e nt on B ouba / K i ki
j udg m e nt s, pa r t i c ul a r l y i n t he a udi t or y m oda l i t y.
A s i n Ex pe r i m e nt 1, w e ge ne r a t e d v i ol i n pl o t s t o vi sua l i z e t he d i st r i but i on of R P
a c r oss t he c r osse d Moda l i t y × S ha pe c ondi t i ons. I n a ddi t i on, t he m e a n r e sponse
( B ouba vs. K i ki r a t e ) w a s p l ot t e d a s a f unc t i on of t he dom i na nt pha se , de f i ne d a s t he
pe a k of t he pha se di s t r i but i on, i .e ., t he m ost f r e que nt bi n , t o be t t e r c a pt ur e how
spe c i f i c R P va l ue s dr i ve s pe r c e pt ua l j udgm e nt s.
R e sponse s w e r e st a t i st i c a l l y a na l yse d us i ng a t w o- w a y A N O V A , w i t h Moda l i t y
( vi sua l vs. a udi t or y) a nd R P ( a s a c ont i nuous va r i a bl e , bi nne d f or a na l ysi s) a s f a c t or s,
t o a sse ss m a i n e f f e c t s a nd t he i r i nt e r a c t i on.
3.5 R e su l t s
R P s i gni f i c a nt l y i nf l ue nc e d B ouba / K i ki j udgm e nt s, a l t hough t he e f f e c t a ppe a r e d m or e
m ode r a t e t ha n i n Ex pe r i m e nt 1. Thi s t r e nd w a s f i r st obse r va bl e i n F i g. 1 ( B ) , w hi c h
sugge st s a c onsi st e nt pa t t e r n a c r oss bo t h m oda l i t i e s: c e r t a i n R P t e nde d t o be m or e f r e -
que nt l y a ssoc i a t e d w i t h B ouba , a nd ot he r s w i t h K i ki . D ue t o t he sha pe c ont i nuum
use d i n t h i s e x pe r i m e nt , t he ke y R P va l ue s obse r ve d i n Ex pe r i m e nt 1 a r e l e ss c l e a r l y
de f i ne d he r e . N e ve r t he l e ss, F i g. 1 ( B ) sugge st s di s t i nc t t e nde nc i e s: K i ki r e sponse s
w e r e ge ne r a l l y m or e f r e que nt a r ound R P ne a r 0 a nd π , w he r e a s B ouba r e sponse s
t e nde d t o c l ust e r a r ound π / 2.
Fig . 1 . (A) Mean respons e by dominant RP for each modality - (B) Violin Plot Showing Aver -
age Density Distributions of RP by Modality (Visu al, Audio) and Shape Label (Bouba , Kiki)
— Ex periment 2
Pro c. of the 17th In ternational Symposium on CMMR, London, UK, Nov. 3-7, 2025
89
10 L . Couston and al .
P l ot t i ng t he m e a n r e sponse ( c ode d a s 0 f or B ouba a nd 1 f or K i ki ) a s a f unc t i on of t he
dom i na nt R P ( F i g. 1 ( A ) ) r e ve a l e d a c om m on t r e nd a c r oss m oda l i t i e s. Mor e i m por -
t a nt l y, i t show s t ha t R P i nf l ue nc e s pa r t i c i pa nt s' j udgm e nt s i n a c ont i nuous m a nne r ,
r a t he r t ha n c a t e gor i c a l l y or di sc r e t e l y.
A t w o- w a y A N O V A c onf i r m e d a s t r ong m a i n e f f e c t of R P on sha pe j udgm e nt s
( F ( 3) = 1.17, p < .001) , w hi l e m oda l i t y show e d no s i gni f i c a nt m a i n e f f e c t ( F ( 3) = .65,
p= .057) . H ow e ve r , a s i gni f i c a nt i nt e r a c t i on ( F ( 4) = 3.34, p= .00038) sugge st s t ha t t he
i nf l ue nc e of R P i s st r onge r a nd m or e c onsi st e nt i n t he v i sua l m oda l i t y, i n l i ne w i t h t he
obse r ve d be ha vi our a l pa t t e r ns.
3.6 D i sc u ss i on
Thi s e x pe r i m e nt c onf i r m s t ha t R P c ont i nuousl y m odul a t e s B ouba –K i ki j udgm e nt s
a c r oss m oda l i t i e s r a t he r t ha n a c t i ng a s a c a t e gor i c a l c ue . The s t r onge r a nd m or e
c onsi st e nt e f f e c t i n t he vi sua l m oda l i t y m a y r e f l e c t di f f e r e n c e s i n st i m ul us qua l i t y, a s
our st i m ul i ’ s a udi t or y r e nde r i ng c oul d i nt r oduc e m or e va r i a bi l i t y. The se f i nd i ngs
hi gh l i ght t he i m por t a nc e of i m pr ovi ng t he ge ne r a t i on m ode l f or t he sha pe c ont i nuum
a nd f ur t he r i nve st i ga t i ng how st i m ul u s pr ope r t i e s a f f e c t pe r c e pt i on.
A m bi gu i t i e s obse r ve d a t i nt e r m e di a t e R P va l ue s e m pha si z e t he c om pl e x i t y of t he
unde r l yi ng dyna m i c s a nd t he ne e d f or c ont i nue d r e se a r c h t o r e f i ne st i m ul us de si gn
a nd be t t e r unde r st a nd c r oss- m oda l i nf l ue nc e s.
4 G e n e r al D i s c u s s i on
The t w o e x pe r i m e nt s r e por t e d he r e e x pl or e d t he r ol e of i m pl i c i t m ove m e nt dyna m i c s
i n t he B ouba - K i ki e f f e c t , u si ng t he sounds pr oduc e d by dr a w i ng m ove m e nt s t o
i nve st i ga t e non- w or d- sha pe a ssoc i a t i ons. The f i ndi ng s sugge st t ha t be yond se nsor y
f e a t ur e s, pa r t i c i pa nt s m a y r e l y on i nt e r na l r e pr e se nt a t i ons of m ove m e nt unde r l yi ng t he
s t i m ul us.
Ex pe r i m e nt 1 show e d t ha t f r i c t i on sound s de r i ve d f r om m ove m e nt - ba se d dr a w i ngs
t e nd t o e l i c i t B ouba / K i ki a ssoc i a t i ons si m i l a r t o t hose obse r ve d i n t he vi sua l do m a i n.
Thi s c onve r ge nc e suppor t s t he i de a of a sha r e d ge st ur a l subst r a t e a c c e ssi bl e a c r oss
m oda l i t i e s, c onsi st e nt w i t h m o t or s i m ul a t i on t he or i e s. Ex pe r i m e nt 2 e x a m i ne d t hi s
f ur t he r by m a ni pul a t i ng t he R P be t w e e n m ove m e nt c om pone nt s. R e su l t s show e d t ha t
t h i s pa r a m e t e r c ont i nuousl y sha pe d j udgm e nt s: R P s ne a r 90° ( π / 2) w e r e a ssoc i a t e d
w i t h B ouba , w hi l e va l ue s ne a r 0 or 180° ( π ) a l i gne d w i t h K i k i . Thi s sugge st s
se nsi t i vi t y t o t he t e m por a l c oor di na t i on of m o t i on a c r oss bot h se nsor y m oda l i t i e s.
The se f i nd i ngs, a l t hough ve r y pr e l i m i na r y r e ga r di ng t he sm a l l pa r t i c i pa nt sa m pl e
s i z e , e x pa nd t r a di t i ona l a r t i c ul a t or y or a c oust i c a c c ount s by i n t r oduc i ng a
se nsor i m ot or di m e nsi on. R a t he r t ha n be i ng ba se d sol e l y on spe e c h- r e l a t e d f e a t ur e s,
t he obse r ve d a ssoc i a t i ons m a y r e f l e c t a n i m pl i c i t r e c onst r uc t i on of t he m ove m e nt
dyna m i c s be hi nd t he st i m ul us. Thi s m ot or pe r spe c t i ve d i f f e r s f r om l i ngu i st i c
i n t e r pr e t a t i ons, a s i t i nvol ve s non- se m a nt i c , physi c a l m ove m e nt c ue s. I t a l i gns w i t h
Pro c. of the 17th In ternational Symposium on CMMR, London, UK, Nov. 3-7, 2025
90
Bouba , Kiki , and the Mov ement Between 11
e m bodi e d c ogni t i on t he or i e s, i n w hi c h pe r c e pt i on dr a w s on i nt e r na l si m ul a t i ons of
a c t i on.
I n sum , i m pl i c i t m ove m e nt dyna m i c s, suc h a s R P , a ppe a r t o sha pe c r oss- m oda l
c or r e sponde n c e s. The B ouba - K i ki e f f e c t m a y t hus r e f l e c t a de e pe r m ot or pr i nc i pl e ,
gr ounde d i n our e m bodi e d e x pe r i e nc e a nd s ha r e d a c r oss se nsor y syst e m s.
5 C on c l u s i on
Thi s s t udy of f e r s ne w i nsi gh t s i nt o t he m e c ha ni sm s unde r l yi ng t he B ouba - K i ki e f f e c t ,
de m onst r a t i ng t ha t i m pl i c i t m ot or dyna m i c s pl a y a c e nt r a l r ol e i n a ssoc i a t i ons
be t w e e n vi sua l sha pe s, f r i c t i on sounds, a nd non- w or ds. O ur r e sul t s c onf i r m t ha t t he
R P of m ove m e nt s m odu l a t e s pa r t i c i pa nt s' j udgm e nt s c ont i nuousl y a nd c onsi st e nt l y,
a c r oss bot h v i sua l a nd a udi t or y m oda l i t i e s. Th i s c onve r ge n c e suppor t s t he i de a of a
sha r e d a m oda l c odi ng syst e m ba se d on i m pl i c i t m o t or s i m ul a t i on t ha t l i nks
pe r c e pt i on, sound, a nd m ove m e nt .
Mor e br oa dl y, t h i s w or k ope ns pr om i s i ng a ve nue s f or unde r st a ndi ng how t he br a i n
i n t e gr a t e s a nd l i nks se nsor y m oda l i t i e s t hr ough dyna m i c m ove m e nt r e pr e se nt a t i ons. I t
sugge st s t ha t pe r c e pt i on i s not m e r e l y t he pa ssi ve r e c e pt i on of i nf or m a t i on, but
i nvo l ve s t he a c t i ve r e c onst r uc t i on of t he m ove m e nt s be hi nd pe r c e i v e d st i m ul i .
6 A c k n ow l e d gme n ts
R e se a r c h suppor t e d by gr a nt s A N R - 16- C O N V - 0002 ( I LC B ) , A N R - 24- C E38- 4175
( P S I N D ) , a nd t he Ex c e l l e nc e I ni t i a t i ve of A i x - U ni ve r si t y ( A *MI D EX ) . Thi s w or k
w a s c a r r i e d out a s pa r t of t he H A R MO N I E pr oj e c t of t he P EP R I C C A R E. I t w a s
suppor t e d by t he F r e nc h N a t i ona l R e se a r c h A ge nc y ( A N R - 23- P EI C - 0002) unde r t he
F r a nc e 2030 pr ogr a m . Thi s pr oj e c t ha s r e c e i ve d f i na nc i a l suppor t f r om t he C N R S
t hr ough t he MI TI i nt e r di sc i pl i na r y pr ogr a m s.
R e fe r e n c e s
1 . Fraisse , P. , Oléron, G . , Paillard , J. : I . Les Effets Dynamogéniques de La Musiqu e. Étude
Ex périmental e. L’année Psychologique , 53(1) , 1–34 (1953). https://doi. org/10 . 3406/
psy . 1953. 8685
2 . Merce Cunningham , John Cage, 1948, Silent Prayer
3 . Trisha Brown, 1979, Glacial Decoy
4 . Decety , J. , Jeannerod , M. , Prablanc, C . (1989). The timing of m entally represented actions .
Behavioural Brain Research, 34(1), 35–42.
5 . Jeannerod , M . (1995) . M ental imagery in the motor cont ex t. Neuropsychologia , 33(11) ,
1419–1432 .
6 . Jeannerod , M . (2011) . L a fabrique des idées . Odil e Jacob .
Pro c. of the 17th In ternational Symposium on CMMR, London, UK, Nov. 3-7, 2025
91
12 L . Couston and al .
7 . Freyd , J. J . : The Mental Representation of Movement W hen Static Stimuli Are Viewed. Per -
ception & Psychophysics , 33(6) , 575–588 (1983). https://doi. org/10 . 3758/BF03202940
8 . Longcamp , M . , Tanskanen , T . , Hari, R. : The Imprint of Action: Motor Cortex Involv ement
in Visual Perception of H andwritten Letters. NeuroImage , 33(2), 681–688 (2006) . https://
doi . org/10. 1016/j . neuroimage. 2006 . 06. 042
9 . Thoret , E . , Aramaki, M. , Kronl and-Martinet, R. , Velay , J . -L. , Ystad , S . : From Sound to
Shape: Auditory Perception of Dr awing Movements . Journal of Ex periment al Psychology:
Human Perception a nd Performance , 40(3) , 983–994 (2014).
https://doi . org/10. 1037/a0035441
10 . Thoret , E. , Aramaki, M . , Bringoux , L. , Ystad , S. , Kronl and-Martinet, R . : Seeing Circles
and Drawing Ellipses: When Sound Biases Reproduction of Visual Motion . PLoS ONE,
11(4) , e0154475 (2016). https://doi. org/10 . 1371/journal. pone . 0154475
11 . Thoret , E . , Aramaki, M. , Bringoux , L . , Ystad , S. , Kronl and-Martinet, R. : Hearing Elliptic
Movements Reveals the Imprint of Action on Prototypical Geometries . Cognition (2023).
https://doi . org/10. 1016/j . cognition. 2023 . 105478. Preprint: https://doi . org/
10 . 1101/2022. 10 . 30. 514456
12 . Köhler , W . (1929) . Gest alt Psychology . New York, NY: Liveright .
13 . Ćwiek , A. et al . (2022). Th e bouba/kiki effect is robust across cultures and writing systems .
Phil . Trans . R . Soc . B , 377:20200390 .
14 . Winter , B. (2025) . The Size and Shape of Sound: The Role of Articulation a nd Acoustics in
Iconicity and Crossmod al Correspondences. Journal of th e Acoustical Society of America .
https://doi . org/10. 1121/10 . 0036362
15 . Forte , M. , Schwartz, J . -L. (2022) . Resolving the Bouba-Kiki Eff ect Enigma by Rooting
Iconic Sound Symbolism in Physical Properties of Round and Spiky Obje cts.
https://doi . org/10. 1038/s41598-022-23623-w
16 . Vainio , L . , Vainio, M. (2021). Sound-Action Symbolism . https://doi . org/10. 3389/
fpsyg . 2021. 718700
17 . Conan , S . , Thoret, E. , Aramaki , M . , Derrien , O . , Gondre, C. , Ystad , S . , Kronland-Martinet,
R . : An Intuitive Synthesizer of Continuous-Interaction Sounds: Rubbing , Scratching, and
Rolling . Computer Music Journal , 38(4), 24–37 (2014) . https://doi. org/10 . 1162/
COMJ_a_00266
18 . Wamain , Y. , Tallet , J. , Zanone, P . -G. , Longcamp, M . : Biological Geometry Perception:
Visual Discrimination of E ccentricity Is Related to Individual Motor Preferences . PLoS
ONE , 6(6) , e20568 (2011). https://doi . org/10. 1371/journal . pone. 0015995
19 . Source: Monochrome v ersion 1 June 2007 by Bendž. Vectori zed with Inkscape -- Qef (talk)
21:21 , 23 June 2008 (UTC). Drawn by Andrew Dunn, 1 Octob er 2004 .
20 . Gaver , W. W (1993). How do we hear in the world? Ex plor ations in ecological acoustics .
Ecological Psychology, 5(4), 285–313.
21 . Van Den Doel , K . , Kry, P. G . and Pai, D. K . , “FoleyAutomatic: physically-based sound
effects for int eractive simulation and animation , ” in Proceedings of the 28th annual
conferen ce on Computer graphics and int eractive techniqu es. ACM, 2001, pp. 537–544.
22 . First ex periment: https://tstbl . co/231-518
23. Second ex periment: https://tstbl. co/530-569
Pro c. of the 17th In ternational Symposium on CMMR, London, UK, Nov. 3-7, 2025
92
Bet w een Bouba and Kiki lies Grougi: Mo delling
the Rough Middle Ground of Sound–Shap e
Corresp ondences
Asterios Zac harakis 1 [0000 − 0003 − 2024 − 6924] ,
K onstantinos V elenis 1 [0009 − 0006 − 7299 − 4763] , and
Emilios Cam b ourop oulos 1 [0000 − 0002 − 6887 − 4287]
Aristotle Univ ersit y , Thessaloniki, Greece
[email protected]
Abstract. Cross-mo dal corresp ondences b et ween audition and vision
ha v e b een extensiv ely studied in relation to pitc h and loudness, but tim-
bre remains more c hallenging to capture empirically and mo del acous-
tically due to its m ultidimensional nature. Nev ertheless, recen t researc h
has sho wn that tim bral qualities suc h as brigh tness, roughness, and clar-
it y can b e reflected in the visual prop erties of images. Building on the
w ell-kno wn Bouba/Kiki effect, the presen t study extends this framew ork
to m usical tim bres and in v estigates visual forms that span a con tin uum
ranging from full roundness to high angularit y . W e h yp othesised not
only that visually induced roughness w ould o ccup y the p erceptual mid-
dle ground b et w een these t w o extremes, but also that this categorisation
w ould b e meaningfully reflected in auditory p erception. T o test this, 27
participan ts w ere ask ed to asso ciate 30 short sound stim uli with visual
shap es using a morphing in terface that enabled con tin uous transforma-
tion from a circle to a jagged, starlik e form. The results rev ealed strong,
non-random asso ciations b et ween sounds and the three shape categories
(smo oth, rough, and sharp). A coustic mo delling further supp orted these
distinctions, sho wing that established mo dels of auditory roughness and
sharpness effectiv ely predicted participan ts’ assignmen ts of the sound
stim uli within these categories. T ogether, these findings offer empirical
bac king for p erceptually grounded mappings b et w een sound and image,
with p oten tial applications in artistic creation, education, and in teractiv e
media.
Keyw ords: timbre · roughness · sharpness · cross-modal corresp on-
dences · shap es.
All righ ts remain with the authors under the Creativ e Commons A ttribution
4.0 In ternational License (CC BY 4.0).
Pro c. of the 17th In t. Symp osium on Computer Music Multidisciplinary Researc h,
London, United Kingdom, 2025
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
93
A. Zac harakis et al.
1 In tro duction
Musical sounds —ev en when exp erienced in isolation from their original con-
text— can con v ey extra-m usical meaning [15,8], often b y ev oking qualities that
extend b ey ond audition in the strict sense [15]. Gro wing in terest in tim bre se-
man tics o v er the past fifteen y ears has adv anced our understanding of ho w suc h
qualities are cognitiv ely organised and what acoustic features underlie them (for
an o v erview of the domain un til 2019 see [33,25]). Some of the most notable
con tributions ha v e organised tim bre descriptions in to seman tic categories, man y
of whic h in v olv e cross-mo dal loans. As early as 1941, Lic h te iden tified brigh tness
and roughness as imp ortan t dimensions in the description of tim bre —qualities
that dra w from the visual and tactile domains, resp ectiv ely [12]. Since then,
the prominence of brigh tness and roughness dimensions (and their v ariations)
within tim bral seman tics has b een reinforced b y n umerous studies across div erse
linguistic bac kgrounds [5,4,27,2,39,23,22]. A t the same time, it should not b e
o v erlo ok ed that researc h has also rev ealed n uanced v ariations within these pri-
mary dimensions, as w ell additional seman tic categories altogether suc h as, for
example, ful l , wo o dy , br assy/metal lic , hol low , op en , p er cussive , etc. [37,20,19].
Giv en that a considerable p ortion of tim bral seman tics originates from cross-
mo dal corresp ondences [30], it has b een argued that systematically exploring
suc h asso ciations offers a fruitful path for future researc h on extra-m usical mean-
ing con v ey ed b y sound [25]. Indeed, a gro wing b o dy of recen t w ork has fo cused on
p oten tial links b et w een tim bral and visual or tactile qualities [24,32,31,34,11,21]
and ev en b et w een tim bral and olfactory qualities [35,36].
While brigh tness is arguably the most w ell-established seman tic dimension
of tim bre, the tactile-related qualit y of roughness has dra wn sc holarly atten tion
since the time of Helmholtz [7], and consisten tly emerges as the most widely
agreed-up on attribute across linguistic groups and lev els of m usical exp ertise
[39,37,32,38,22]. Roughness constitutes one of the principal qualities of surfaces
and is primarily assessed through touc h. That said, certain asp ects of roughness
can also b e ev aluated visually [6]. With this in mind, w e recen tly explored p o-
ten tial relationships b et w een v arious graphical manifestations of roughness (i.e.,
roughness, noisiness, gran ularit y , and sharpness) and sound stim uli sp ecifically
designed to con v ey these qualities in the auditory domain [40] using a fiv e-
alternativ e forced-c hoice (5AF C) exp erimen tal design. Our data sho w ed that
the audio-image corresp ondences w ere ab o v e c hance o v erall. A t the same time,
the rough and gran ular visual categories could b e acoustically differen tiated us-
ing an audio feature originating from the Mo dulation P o w er Sp ectrum (MPS),
kno wn to predict auditory roughness [3,22] and Zwic k er’s sharpness mo del [41].
Em b oldened b y these preliminary results, w e hereb y pro ceed to a simplified
v ariation of this exp erimen t that is partly inspired b y the Bouba/Kiki effect
[9,17,16], whic h has demonstrated the systematic mapping b et w een w ords and
shap es. A bit more than a decade ago, A deli and colleagues [1] also expanded
the Bouba/Kiki paradigm from w ord-shap e to sound-shap e asso ciations. Their
w ork demonstrated that softer tim bres tended to b e asso ciated with rounder,
Bouba-st yle forms, while harsher ones w ere link ed to spikier, Kiki-st yle shap es.
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
94
Bouba-Grougi-Kiki
Ho w ev er, this w ork iden tified asso ciations b et w een sound and shap es primarily
at the source-cause lev el (e.g., piano, saxophone, gong) rather than with sp ecific
acoustic correlates. In the curren t exp erimen t, participan ts w ere giv en the option
to adjust the prop erties of a shap e rather than c ho osing from a predefined set
of images. Th us, using a single slider, it w as p ossible to morph a full circle
in to a spikier shap e comp osed of 20 trap ezoids, whic h could gradually ev olv e
in to triangles b y reducing the length of their shorter side (see Figure 1). The
underlying assumption w as that this con tin uous visual transformation w ould
progress from smo oth to sharp, passing through an in termediate stage p erceiv ed
as rough —m uc h lik e ho w the lexical blend Gr ougi , p ositioned b et ween Kiki
and Boub a , ev ok es a rough qualit y . This design pro vided a neat and efficien t
w a y to collect data on all three p erceptual qualities (smo othness, roughness and
sharpness) using a single, in tuitiv e in terface, describ ed in greater detail in the
Metho d section.
These in v estigations in to audio-visual corresp ondences are part of a broader
researc h pro ject en titled Soundsk etc her [29], whic h aims to automatically trans-
late audio in to p erceptually meaningful visual represen tations for the creation
of aural scores in b oth creativ e and educational con texts.
2 Metho d
2.1 Sound stim uli
The exp erimen tal dataset consisted of 30 short audio stim uli (1.5–3 seconds),
dra wn from v arious sound libraries (e.g., F reesound, MUMS [14]), adapted from
past studies [13], or designed from scratc h to equally represen t the three sonic
qualities of in terest (smo oth, rough, and sharp; 10 stim uli p er category). T o con-
struct this set, the researc h team, together with undergraduate studen ts from
a m usic cognition mo dule, compiled a broader p o ol of candidate sounds, whic h
w ere then discussed and ev aluated through informal listening across the three
seman tic categories, ultimately yielding the final 30 stim uli. The selected stim-
uli w ere subsequen tly loudness equalised through an iterativ e listening pro cess
conducted within the researc h team.
2.2 Visual v ariable
The visual resp onse elemen t w as a shap e-morphing in terface presented in the
form of a con tin uous slider (Figure 1). The base shap e w as a p erfect circle. As
participan ts mo v ed the slider to the righ t, the shap e gradually morphed in to
increasingly jagged, star-lik e forms with higher angular mo dulation. This trans-
formation w as con tin uous and allo w ed for fine-grained adjustmen ts across the
full range of p ossible shap es. The slider corresp onded to a hidden scale ranging
from 0 to 20 that w as not rev ealed to the participan ts. This morphing con tin uum
w as designed to pro vide participan ts with a straigh tforw ard and in tuitiv e w a y
to explore the transition from round to angular shap es. The basic assumption
here w as that the midp oin t b et ween complete roundness (i.e., smoothness) and
full angularit y (i.e., sharpness) w ould signify roughness.
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
95
A. Zac harakis et al.
Fig. 1. Visual shap e con tin uum reflecting textural morphing from smo oth to rough and
finally sharp. P articipan ts adjusted a slider to select the shap e that b est represen ted
the sound they h ea rd . The shap e gra d ually transformed from a smo oth circle (left), to
rough undulating con tours (cen ter), and finally to sharp, p oin ted spik es (righ t). The
n u m b e rs sho wn b eneath eac h image represen t the corresp onding sli der v alue, noting
that thi s w as not visi ble to participan ts.
2.3 P articipan ts
A total of 27 participan ts to ok part in the study . The participan t p o ol w as div erse
in terms of gender and m usical bac kground. Mos t participan ts iden t i fied as male
(19), follo w ed b y female (8), with one participan t c ho osing not to disclose their
gender.
Musical e x p erience lev els v aried considerably: the sample included 9 pro-
fessional m usicians, 5 semi-professionals, 7 amateurs and 6 non-m usicians. The
ma jo rit y of participan ts rep orted pla ying a m usical instrumen t, with instrumen t
t yp es including piano, guitar, and electric bass, to more sp ecialised ones suc h as
the double bass.
Regarding liste ning equipmen t, most participan ts used closed-bac k (o v er-
ear) headphones, while others used on-ear headphones or studio monitors. All
participan ts rep orted normal hearing and normal or corrected-to-normal v is ion.
P articipation w as v olun tary and no comp ensation w as pro vided.
2.4 Exp e rimen tal pro cedure
The exp e r im en t w as conducted online via a bro wser-based in teractiv e in terface.
In eac h trial, participan ts w ere presen ted with an initial circle shap e represen ting
the smo othne ss and a slid er b eneath it at v alue 0. T heir task w as to adjust the
shap e of the image b y mo ving the slider un til they felt it b est represen ted the
qualit y of eac h sound stim ulus. They could trigger audio pla ybac k b y clic king
a button lab eled “Press "A" or clic k to pla y the sound”. A t the b e ginning of
the session, participan ts w ere presen ted with three practice sounds to familiarize
themselv es with the p r o cedure and the in terface. These p r ac tice trials w ere ex-
cluded from the data analysis. In the main exp erimen t, the presen tation order of
the 30 audio stim uli w as randomised for eac h participan t to coun terbalance p o-
ten tial order effects. There w ere no restrictions on ho w man y times participan ts
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
96
Bouba-Grougi-Kiki
could listen to a sound b efore making their judgmen t. Once they w ere satisfied
with their c hoice, they clic k ed the “Next Sound” button to pro ceed. This pro cess
w as rep eated for all 30 randomised trials follo wing the initial practice phase.
Throughout the exp erimen t, the system recorded the slider p osition (whic h
corresp onded to a sp ecific shap e configuration), the asso ciated sound ID, and the
timestamp of the selection. The full set of audio stim uli, along with the shap e
generation co de used in the exp erimen t, is publicly a v ailable on GitHub.
3 Results
3.1 Analysis of resp onses
The primary goal of the presen t analysis w as to determine whether the b e-
ha vioural resp onses exhibited systematic trends. T o this end, the original con tin-
uous v ariable (ranging from 0 to 20) w as transformed in to a categorical v ariable
b y dividing the scale in to three regions: [0–5] for smo oth, [6–15] for rough, and
[16–20] for sharp. This categorisation w as informed b y b oth visual insp ection of
the stim uli’s shap e transformations (see Figure 1) and b y the goal of optimising
Cramér’s V and the fit of the acoustic mo dels (see b elo w).
Subsequen tly , a Chi-square on the con tingency table of resp onses (T able 1)
test rev ealed a significan t deviation from uniformit y in category c hoices across
sound stim uli, χ 2 (58) = 478 . 89 , p < 0 . 00001 , indicating a systematic, non-
random pattern in ho w sounds w ere attributed to the three categories. The
effect size, measured b y Cramér’s V ( V = 0 . 544 ), reflects a relativ ely strong
structure in these category assignmen ts [18].
T able 1. Con tingency table sho wing participan t selections (coun ts) by seman tic cate-
gory for eac h stim ulus.
Stim ulus Smo oth Rough Sharp Stim ulus Smo oth Rough Sharp
mo og-noise 4 15 8 sharp 1 5 21
qua v er-p ok es 1 6 20 tuba-rough 3 15 9
gran ulator 7 18 2 ebass 23 3 1
dissonan t-harm 2 7 18 fly 4 17 6
rain-noise 6 10 10 glass 21 3 3
storm-noise 4 17 6 guitar-rough 2 16 9
anc hor-c hain 2 17 8 guitar-smo oth 23 2 2
scary-rasp 3 2 22 harp 1 9 17
gran ular 6 15 6 saxophone 2 19 6
b o w ed-cym bal 2 7 18 sharp-sa w 6 5 16
cello 21 4 2 sharp-kirk 8 8 11
flute 17 6 4 smo oth-bass 24 2 1
hammond 19 5 3 smo oth-sine 23 3 1
pan-flute 17 7 3 smo oth2 20 3 4
SP1Désin tégrations 5 5 17 steelsea-sharp 7 13 7
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
97
A. Zac harakis et al.
A t the individual lev el, Figure 2 illustrates the strength of asso ciations b e-
t w een eac h sound and the three resp onse categories. Subfigure 2a presen ts the
significan t uncorrected standardised residuals from the Chi-square test of inde-
p endence, while Subfigure 2b sho ws the significan t standardised residuals after
F alse Disco v ery Rate (Benjamini-Ho c h b erg) correction, whic h enhances sensi-
tivit y to meaningful relationships while con trolling false disco v eries. The uncor-
rected co efficien ts rev eal 27 out of 30 significan t asso ciations, while ev en the
stricter corrected co efficien ts still indicate significan t asso ciations for 24 out of
30 sounds with one of the visual categories.
Smooth Rough Sharp
Categories
moog-noise
quaver-pokes
granulator
dissonant-harm
rain-noise
storm-noise
anchor-chain
scary-rasp
granular
bowed-cymbal
cello
flute
hammond
pan-flute
sharp
tuba-rough
ebass
fly
glass
guitar-rough
guitar-smooth
harp
saxophone
sharp-saw
sharp-kirk
smooth-bass
smooth-sine
smooth2
steelsea-sharp
Sounds
Uncorrected Standardised Residuals
3.742
2.443
3.093
2.443
4.392
3.742
4.392
4.717
4.392
3.418
2.085
3.096
2.759
2.759
2.085
2.085
2.759
2.422
3.433
3.825
3.147
4.503
3.147
2.809
4.164
2.809
2.47
-2.754
-2.429
-2.429
-2.104
-2.429
-2.754
-2.104
-2.429
-2.754
-2.429
-2.295
-2.295
-2.295
-2.274
-2.274
-2.613
-2.274
-2.613
-2.613 -3
-2
-1
0
1
2
3
[]
(a)
Smooth Rough Sharp
Categories
moog-noise
quaver-pokes
granulator
dissonant-harm
rain-noise
storm-noise
anchor-chain
scary-rasp
granular
bowed-cymbal
cello
flute
hammond
pan-flute
sharp
tuba-rough
ebass
fly
glass
guitar-rough
guitar-smooth
harp
saxophone
sharp-saw
sharp-kirk
smooth-bass
smooth-sine
smooth2
steelsea-sharp
Sounds
FDR-Corrected Standardised Residuals
3.742
2.443
3.093
2.443
4.392
3.742
4.392
4.717
4.392
3.418
3.096
2.759
2.759
2.759
2.422
3.433
3.825
3.147
4.503
3.147
2.809
4.164
2.809
2.47
-2.754
-2.429
-2.429
-2.429
-2.754
-2.429
-2.754
-2.429
-2.613
-2.613
-2.613
-3
-2
-1
0
1
2
3
Not Significant
(b)
Fig. 2. Heatmap of the standardised residuals from the Chi-square test of indep en-
dence, with blue indicating negativ e v alues and y ello w represen ting p ositiv e v alues,
where non-significan t v alues are sho wn in white. (a) Significan t matchings for the un-
corrected standardised residuals. (b) Significan t matc hings after F alse Disco very Rate
(Benjamini-Ho c h b erg) correction.
3.2 Auditory mo delling
After establishing a strong systematic relationship b et w een sounds and visual
categories, w e sough t to in terpret these relationships through acoustic predictors.
T o this end, w e computed conditional probabilities reflecting the lik eliho o d of
eac h visual category b eing c hosen giv en a sp ecific sound stim ulus, P(Visual | Au-
ditory). Resp onses w ere aggregated in to a 30 × 3 con tingency table, capturing the
distribution of selections across conditions. A m ultinomial regression w as then
p erformed on this con tingency structure to assess the exten t to whic h acoustic
predictors systematically influenced visual c hoices.
The mo delling explored some audio features that ha v e b een link ed with au-
ditory roughness and sharpness. The roughness mo dels b y V assilakis [28] and
Sethares [26] w ere implemen ted using the MIR T o olb o x [10] (windo w size: 40ms,
50% o v erlap). A dditionally , the roughness and sharpness mo dels based on the
w ork b y F astl & Zwic k er [41] w ere computed using the MA TLAB functions
acousticRoughness and acousticSharpness. F urthermore, a metric based on the
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
98
Bouba-Grougi-Kiki
Mo dulation P o w er Sp ectrum (MPS), kno wn to predict auditory roughness [3,22],
w as also calculated. The MPS is the 2D-F ourier transform of a sp ectrogram that
quan tifies the distribution of mo dulation energy across the temp oral and sp ectral
dimensions. In this study , the 2D F ourier transform w as applied to a loudness
sp ectrogram estimated using MA TLAB’s acousticLoudness function (windo w
size: 50 ms, hop size: 2ms), whic h implemen ts Zwic k er’s metho d for loudness
estimation (ISO 532-1) [41], rather than a sp ectrogram deriv ed from a standard
F ourier transform. MPS-roughness w as calculated as the ratio of energy within
the 30–150 Hz range of the temp oral mo dulation axis to the total p ositiv e mo d-
ulation energy , using a sligh t v ariation of the metric prop osed b y Rosi et al.
[22].
A m ultinomial regression mo del w as then emplo y ed to predict visual category
probabilities through the audio features. A log transformation w as applied to the
predictor table X , X log = log 10 ( X + ε ) , with ε = 10 − 4 to suppress p ositiv e
outliers in the audio features and impro v e the stabilit y and in terpretabilit y of re-
gression co efficien ts. Standardisation (z-scoring) w as subsequen tly p erformed to
ensure comparabilit y across features. Mo del ev aluation relied on log-lik eliho o d-
based metrics, including McF adden’s R 2 and Nagelk erk e’s R 2 , to assess explana-
tory p o w er. Rather than relying on discrete classifications, predictions w ere de-
riv ed from estimated probabilit y distributions o v er the visual categories, pro vid-
ing a more con tin uous in terpretation of the relationship b et w een predictors and
resp onse tendencies.
T able 2. Multinomial regression results predicting the selection of the shap e categories
smo oth and rough (vs. sharp) based on acoustic features. The sign and magnitude of
the b eta co efficien ts indicate ho w the probabilit y of selecting a giv en category is in-
fluenced b y eac h feature, accompanied b y 95% confidence interv als and significance
lev els. Mo del fit: Log-Lik eliho o d ( L ) = -12.77; McF adden’s R 2 = 0.61, Nagelk erke’s
R 2 = 0.83; Prediction A ccuracy = 100%; Log Loss = 11.18. These metrics reflect
the mo del’s go o dness-of-fit, classification p erformance, and uncertain ty in predictions.
Lik eliho o d Ratio T est (LR T) p-v alues: V assilakis’ roughness = 0.0001, Zwic k er’s sharp-
ness = 0.0000, indicating that b oth predictors are statistically significan t.
Predictors Betas p-V alue 95% CI
Cat. 1: [Smo oth]
In tercept -337.7 <.0001 [-343.05 , -332.38]
V assilakis’ roughness -845.3 <.0001 [ -849.5, -841.09]
Zwic k er’s sharpness -2608 <.0001 [-2616.3, -2599.6 ]
Cat. 2: [Rough]
In tercept 775.7 <.0001 [770.67 , 780.76]
V assilakis’ roughness 189.6 <.0001 [ 185.66 , 193.6]
Zwic k er’s sharpness -1120.5 <.0001 [-1127 , -1114]
Sev eral pairwise com binations of the fiv e examined descriptors resulted in
significan t and comparable mo dels, all of whic h featured one of the roughness
descriptors as the first predictor and Zwicker’s sharpness as the second. The
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
99
A. Zac harakis et al.
mo del rep orted here (T able 2) yielded the highest McF adden’s R 2 (.612) and
features V assilakis roughness as the first predictor and Zwicker’s sharpness as
the second. This mo del also demonstrated meaningful in terpretabilit y , as all of
the b etas w ere statistically significan t. The b eta co efficien ts presen ted in T able 2
reflect ho w the t w o auditory predictors influence the lik eliho o d of a shap e falling
in to the first t w o categories (i.e., smo oth and rough), relativ e to the reference
category (sharp).
A ccording to the mo del, the lik eliho o d of a stim ulus falling in to the smo oth
category o v er the sharp category increases substan tially as V assilakis’ roughness
decreases, and ev en more so as Zwic k er’s sharpness decreases. In con trast, while
higher V assilakis’ roughness increases the lik eliho o d of the rough category rela-
tiv e to sharp, lo w er Zwic k er’s sharpness con tin ues to increase the lik eliho o d of
rough o v er sharp.
-2.5 -2 -1.5 -1 -0.5 0 0.5 1 1.5
Vassilakis' roughness
-2.5
-2
-1.5
-1
-0.5
0
0.5
1
1.5
Zwicker's Sharpnes
Class Separation in Log-Transformed Feature Space
moog-noise
quaver-pokes
granulator
dissonant-harm
rain-noise
storm-noise
anchor-chain
scary-rasp
granular
bowed-cymbal
cello
flute
hammond
pan-flute
sharp
tuba-rough
ebass
fly
glass
guitar-rough
guitar-smooth
harp
saxophone
sharp-saw
sharp-kirk
smooth-bass
smooth-sine
smooth2
steelsea-sharp
Smooth
Rough
Sharp
Fig. 3. Separation of visual categories in the feature space defined b y V assilakis’ rough-
ness and Zwic k er’s sharpness (log-transformed). Eac h p oin t represents a stim ulus, col-
ored and shap ed according to the most probable visual category (smo oth, rough, or
sharp) based on participan t resp onses. The clear separation b et w een regions illustrates
the strong con tribution of these auditory features to the categorisation of shap e re-
sp onses.
While the magnitude of the mo del’s b eta co efficien ts and the p erfect clas-
sification accuracy migh t raise concerns ab out p oten tial o v erfitting, the t w o-
dimensional feature space defined b y the log-transformed V assilakis’ roughness
and Zwic k er’s sharpness (Figure 3) sho ws a clear separation b et w een the visual
categories. The fact that stim uli asso ciated with eac h category —based on the
most frequen t visual assignmen t b y participan ts— o ccup y distinct regions in the
feature space supp orts the conclusion that these auditory features pla y a mean-
Pro c. of the 17th In ternational Symposium on CMMR, London, UK, No v. 3-7, 2025
100
Bouba-Grougi-Kiki
ingful role in shaping p erceptual categorisation. Nonetheless, the robustness of
this mo del should ideally b e confirmed through indep enden t v alidation.
A cross m ultiple 5-fold cross-v alidation runs, the m ultinomial regression mo del
trained on auditory features ac hiev ed mean classification accuracies ranging from
76.7% to 86.7%, with an o v erall a v erage of approximately 80.0% ( ± 11.4%). The
corresp onding mean log loss v alues a v eraged around 9.6 ( ± 1.9), indicating rela-
tiv ely lo w predictiv e confidence despite high classification accuracy . This suggests
that the auditory features offer meaningful discriminatory p o w er for modelling
sound–shap e corresp ondences, although the elev ated log loss p oin ts to limited
certain t y in the probabilit y estimates, lik ely a consequence of the small dataset
size.
4 Discussion
These preliminary findings offer promising insigh ts in to cross-mo dal corresp on-
dences b et w een vision and audition. In particular, our h ypothesis of a visual
con tin uum linking smo othness to sharpness via roughness app ears to b e sup-
p orted b y the categorisation and analysis of our visual v ariable. As men tioned
in the Results section, the discretisation of the con tin uous visual v ariable in to
three distinct regions ( smo oth , r ough , and sharp ) w as guided b y a com bina-
tion of visual insp ection of the generated shap es and iterativ e refinemen t based
on optimisation of the Cramér’s V v alues (here quan tifying the strength of as-
so ciation b et w een visual categories and the sound stim uli) and go o dness-of-fit
measures from the subsequen t acoustic mo delling. A Chi-square test of inde-
p endence iden tified audio-visual asso ciations, whic h exceeded c hance lev els b oth
o v erall and across individual stim uli. Indeed, the distribution of our 30 audio
stim uli across the three visual categories w as appro ximately ev en (see Figure
2) and broadly aligned with our initial assumptions on p erceiv ed smo othness,
roughness and sharpness. Th us, our results extend the w ell-kno wn Bouba/Kiki
effect from sp eec h-based sounds to m usical tim bres, in agreement with previous
researc h [1], while also rev ealing an in termediate perceptual category linked to
roughness.
Based on this finding, w e subsequen tly sough t to ac hieve an acoustic mod-
elling of the visual categories emplo ying a n um b er of established mo dels of audi-
tory roughness and sharpness. The mo delling further strengthened our assump-
tions on the distinctions ac hiev ed betw een smo othnes s, roughness and sharpness.
Sounds with higher roughness v alues, as measured b y V assilakis’ mo del [28], w ere
more lik ely to b e assigned to the r ough category and less lik ely to b e assigned
to smo oth , relativ e to sharp . A dditionally , higher v alues of Zwic k er’s sharpness
[41] reduced the lik eliho o d of assignmen t to b oth smo oth and r ough , again rel-
ativ e to sharp . This pattern is clearly illustrated in the scatterplot of Figure 3,
where stim uli with higher probabilities of b eing rated as sharp (green diamonds)
tend to cluster in the upp er-righ t region of the plot, c haracterised b y both high
Zwic k er’s sharpness and high V assilakis’ roughness. Stim uli with a higher lik eli-
ho o d of b eing categorised as rough (red squares) app ear just b elo w and further
to the righ t, defined b y lo wer acoustic sharpness but ev en higher acoustic rough-
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
101
A. Zac harakis et al.
ness v alues on a v erage. Finally , stim uli most lik ely to b e rated as smo oth (blue
circles) o ccup y the b ottom-left region, mark ed b y b oth lo w acoustic roughness
and lo w sharpness. Note that eac h of the three categories con tains the 10 stim uli
w e initially assumed to represen t them, with the sole exception of ste else a-sharp .
Although acoustically it b elonged to the sharpness group (see Fig. 3), whic h
w as also our assumption, participan ts rated it with a higher probabilit y of b eing
rough.
The m ultinomial regression used for acoustic mo delling pro duced a mark edly
stronger outcome compared to a similar mo del in our recen t w ork [40]. This im-
pro v emen t is exp ected, giv en that here w e limited the visual categories to three
(rather than fiv e previously) and increased the n um b er of audio stim uli to 30
(compared to 18 b efore). These c hanges already create more fa v ourable con-
ditions for mo delling. Moreo v er, the impro v emen t ma y also stem from a more
careful selection of audio stim uli that b etter represen t the qualities under exam-
ination, as w ell as from the fact that the con tin uous visual morphing paradigm
constitutes a more efficien t metho d for acquiring visual information.
Although the m ultinomial regression mo del incorp orating V assilakis’ rough-
ness and Zwic k er’s sharpness is excellen t at predicting the most lik ely resp onse
category , its confidence remains limited, suggesting lo w generalisabilit y of the
results. F uture researc h should aim to include larger audio stim ulus sets, ac-
quire larger scale empirical data and p oten tially explore additional acoustic
factors to deep en our understanding of sound-shap e relationships within the
smo othness–roughness-sharpness con tin uum. Approac hing the problem from the
opp osite p ersp ectiv e, participan ts could b e pro vided with a sound syn thesis en-
gine and ask ed to create tim bres that matc h a v ariet y of distinct shap es. This
w ould allo w us to directly map sound syn thesis parameters on to visual categories.
Mean while, within the framew ork of the Soundsk etc her pro ject [29], these
findings already pro vide v aluable insigh ts that con tribute to the dev elopmen t
of p erceptually meaningful mappings b et w een sounds and images. By grounding
these mappings in empirically supp orted cross-mo dal corresp ondences, this w ork
helps pa v e the w a y for more in tuitiv e audio-visual in teractions, enhancing b oth
creativ e exploration and user exp erience.
A c kno wledgmen ts. W e w ould lik e to thank the undergraduate studen t assistan ts
who con tributed to the creation and selection of sound stim uli, as w ell as to participant
recruitmen t. W e w ould also like to thank Anastasia V ak atari for her assistance with
data acquisition and prepro cessing. This researc h w as carried out within the framew ork
of the National Reco v ery and Resilience Plan Greece 2.0, funded b y the Europ ean Union
– NextGenerationEU (Implemen tation b o dy: HFRI).
References
1. A deli, M., Rouat, J., Molotc hnik off, S.: Audio visual corresp ondence b et w een m u-
sical tim bre and visual shap es. F ron tiers in h uman neuroscience 8 , 352 (2014)
2. Alluri, V., T oiviainen, P .: Effect of enculturation on the semantic and acoustic
correlates of p olyphonic tim bre. Music P erception 29 (3), 297–310 (2011)
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
102
Bouba-Grougi-Kiki
3. Arnal, L.H., Flink er, A., Kleinsc hmidt, A., Giraud, A.L., P o epp el, D.: Human
screams o ccup y a privileged nic he in the comm unication soundscap e. Curren t bi-
ology 25 (15), 2051–2056 (2015)
4. v on Bismarc k, G.: Sharpness as an attribute of the tim bre of steady sounds. A cus-
tica 30 , 17–24 (1974)
5. v on Bismarc k, G.: Tim bre of steady tones: A factorial in v estigation of its v erbal
attributes. A custica 30 , 146–159 (1974)
6. Di Stefano, N., Sp ence, C.: Roughness p erception: A m ultisensory/crossmo dal p er-
sp ectiv e. A tten tion, P erception, & Psyc hoph ysics 84 (7), 2087–2114 (2022)
7. v on Helmholtz, H.L.F.: On the Sensations of T one as a Ph ysiological Basis for the
Theory of Music. New Y ork: Do v er (1954), 4 edn. (1877), english translation b y A.
J. Ellis
8. K o elsc h, S.: T o w ards a neural basis of pro cessing m usical seman tics. Ph ysics of Life
Reviews 8 (2), 89–105 (2011)
9. Köhler, W.: Gestalt psyc hology . New Y ork, NY: Liv erigh t (1929)
10. Lartillot, O., T oiviainen, P .: A Matlab to olb o x for m usical feature extraction from
audio. In: In ternational conference on digital audio effects. v ol. 237, p. 244. Bor-
deaux (2007)
11. Lembk e, S.A.: Hearing triangles: P erceptual clarit y , opacit y , and symmetry of
sp ectrotemp oral sound shap es. The Journal of the A coustical So ciet y of Amer-
ica 144 (2), 608–619 (2018)
12. Lic h te, W.: A ttributes of complex tones. J. Exp. Psyc h 28 , 455—-480 (1941)
13. Noble, J., Thoret, E., Henry , M., McA dams, S.: Seman tic dimensions of sound
mass m usic: mappings b et w een p erceptual and acoustic domains. Music P erception
38 (2), 214–242 (2020)
14. Op olk o, F., W apnic k, J.: McGill Univ ersit y master samples collection on D VD
(2006)
15. P atel, A.D.: Music, language, and the brain., c hap. Meaning, pp. 299–254. Oxford
Univ ersit y Press (2008)
16. Ramac handran, V.S., Hubbard, E.M.: Hearing colors, tasting shap es. Scien tific
American 288 (5), 52–59 (2003)
17. Ramac handran, V.S., Marcus, Z., Ch unharas, C.: Bouba-kiki: Cross-domain reso-
nance and the origins of synesthesia, metaphor, and w ords in the h uman mind. In:
Multisensory p erception, pp. 3–40. Elsevier (2020)
18. Rea, L.M., P ark er, R.A.: Designing and conducting surv ey researc h: A comprehen-
siv e guide. John Wiley & Sons (2014)
19. Reymore, L.: Characterizing protot ypical m usical instrumen t tim bres with tim bre
trait profiles. Musicae Scien tiae 26 (3), 648–674 (2022)
20. Reymore, L., Huron, D.: Using auditory imagery tasks to map the cognitiv e lin-
guistic dimensions of m usical instrumen t tim bre qualia. Psyc hom usicology: Music,
Mind, and Brain 30 (3), 124 (2020)
21. Reymore, L., Lindsey , D.T.: Color and tone color: audio visual crossmo dal corre-
sp ondences with m usical instrumen t tim bre. F ron tiers in Psyc hology 15 (2025).
https://doi.org/10.3389/fpsyg.2024.1520131
22. Rosi, V., Arias Sarah, P ., Houix, O., Misdariis, N., Susini, P .: Shared men tal rep-
resen tations underlie metaphorical sound concepts. Scien tific Rep orts 13 (1), 5180
(2023)
23. Rosi, V., Houix, O., Misdariis, N., Susini, P .: In v estigating the shared meaning of
metaphorical sound attributes: brigh t, w arm, round, and rough. Music P erception
39 (5), 468–483 (2022)
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
103
A. Zac harakis et al.
24. Saitis, C., W allmark, Z.: Tim bral brigh tness p erception in v estigated through m ul-
timo dal in terference. A tten tion, P erception, & Psyc hophysics 86 (6), 1835–1845
(2024)
25. Saitis, C., W einzierl, S.: The seman tics of tim bre. In: Sieden burg, K., Saitis, C.,
McA dams, S., P opp er, A.N., F a y , R.R. (eds.) Tim bre: A coustics, Perception, and
Cognition, pp. 119–149. Springer (2019)
26. Sethares, W.A.: T uning, tim bre, sp ectrum, scale. Springer Science & Business Me-
dia (2005)
27. Stepánek, J.: Musical sound tim bre: V erbal description and dimensions. In: Pro-
ceedings of the 9th In ternational Conference on Digital Audio Effects (D AFx-06).
pp. 121–126. Citeseer (2006)
28. V assilakis, P .N.: P erceptual and ph ysical prop erties of amplitude fluctuation and
their m usical significance. Ph.D. thesis, Univ ersit y of California, Los Angeles (2001)
29. V elenis, K., Zac harakis, A., Cam b ourop oulos, E.: Soundsk etc her: Graphical repre-
sen tation of sound based on crossmo dal audio visual corresp ondences. In: Pro ceed-
ings of the 22nd Sound and Music Computing Conference (SMC2025). Univ ersit y
of Music and P erforming Arts, Graz, Austria (2025)
30. W allmark, Z.: A corpus analysis of tim bre seman tics in orc hestration treatises.
Psyc hology of Music 47 (4), 585–605 (2019)
31. W allmark, Z.: Seman tic crosstalk in tim bre p erception. Music & Science 2 ,
2059204319846617 (2019)
32. W allmark, Z., Allen, S.E.: Presc ho olers’ crossmo dal mappings of tim bre. A tten tion,
P erception, & Psyc hoph ysics 82 , 2230–2236 (2020)
33. W allmark, Z., Kendall, R.A.: Describing sound: The cognitiv e linguistics of tim bre.
In: Dolan, E.I., Rehding, A. (eds.) The Oxford Handb o ok of Tim bre, pp. 578–
608. New Y ork, NY: Oxford Univ ersit y Press (2018). https://doi.org/10.1093/
oxfordhb/9780190637224.013.14
34. W allmark, Z., Nghiem, L., Marks, L.E.: Do es tim bre mo dulate visual p erception?
exploring crossmo dal in teractions. Music P erception 39 (1), 1–20 (2021)
35. Zac harakis, A.: Sonic b ouquet: Deco ding cross-mo dal corresp ondences b et w een
tim bre and scen t. Music & Science 7 (2024). https://doi.org/10.1177/
20592043241274258
36. Zac harakis, A.: The p o etry of senses: exploring seman tic mediation in tim bre-aroma
corresp ondences. F ron tiers in Psyc hology 16 , 1520046 (2025)
37. Zac harakis, A., P astiadis, K.: Revisiting the luminance-texture-mass mo del for
m usical tim bre seman tics: A confirmatory approac h and p ersp ectiv es of extension.
Journal of the Audio Engineering So ciet y 64 (9), 636–645 (2016)
38. Zac harakis, A., P astiadis, K.: In teraction b et w een time-v arying tone inharmonicit y ,
fundamen tal frequency and sp ectral shap e affects felt tension and tim bral seman-
tics. A cta A custica 5 , 21 (2021)
39. Zac harakis, A., P astiadis, K., Reiss, J.D.: An in terlanguage study of musical tim bre
seman tic dimensions and their acoustic correlates. Music P erception 31 (4), 339–
358 (2014)
40. Zac harakis, A., Vlaikidou, A., V elenis, K., Cam b ourop oulos, E.: F acets of audio-
visual roughness: from gran ularit y to sharpness. In: Pro ceedings of the 22nd Sound
and Music Computing Conference (SMC2025). Univ ersit y of Music and P erforming
Arts, Graz, Austria (2025)
41. Zwic k er, E., F astl, H.: Psyc hoacoustics: F acts and mo dels, vol. 22. Springer Science
& Business Media (2013)
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
104
Cross - modal coupling between pi tch and the perception
of space , and its modulation by contextual pri ors.
Zofia Anna Hoł ubowska 1,2 [0009 - 0002 - 1113 - 4509] , Marle n Herk enberg 1 and Marc
Schönwiesner 1,3 [ 0000 - 0002 - 2023 - 1207 ]
1 Institute for Biology, Fa culty of L ife Sciences , Unive rsity of Le ipzig, Ta lstr aße 33, 04103
Leipzig, Germany
2 Max - Planck - Institute for Human and Co gnitive Brain Science s, Stephanstr aße 1a, 04103
Leipzig, Germany
3 Depart ment o f Psy cholo gy, Uni versi ty of Montreal, Pavillon Marie - Victo rin , 90 avenue Vin-
cent d'Indy , Mont ré al , Quebec, Ca nada
holubowska @ cbs.mpg.de
Abstrac t. Frequenc y - elevation mapping (FEM) describes how listeners system-
atically mislocalize sounds in v ertical space: high - pitched sounds appear elevated
while low - pitched sounds seem lower than their actual positions. This phenome-
non shapes how we experience sound in space, yet its underlyin g mechanism s
remain debated , with some suggesting the main influenc e of statistical learning
based on sounds naturally occurring in the environment, whereas others support-
ing the main linguistic or cultural contribution .
We i nve sti ga ted FEM us ing vi oli n, fl ute , and spectrally - match ed ar tif ici al s ounds
across nine frequencies (200 – 1600 Hz) and five el evation angles. Twenty - nine
participants local ized these sounds while we manipulated both spectra l cent roid
and musical context to disentangle local and global context contributions to the
effect .
Our resul ts show that FEM is primarily driven by fundamental frequency rather
than spectral content. While spectral centroid slightly influenced elevation per-
ception at low f requencies, this ef fect vanished at higher frequencies. Crucial ly,
music al ins tru ments sh owed si gn ificantly weaker FEM effects compared to arti-
ficial sounds with identical spectral properties. Local context also mattered: the
interval betw een succe ssive notes influ enced spa tial perception .
These findings sho w that musical k nowled ge can out weigh fundamen tal p sycho-
acoustic expectations . Recognit ion of a music al instrume nt as a coherent sound
source capable of producing multiple pitches reduces FEM effect. This demon-
strates how musical and contextual understanding actively shape our perception
of audi tory s pace, t ransforming abstract frequency relat ionships into a mean ing-
ful spatial experience w ithin musical env ironment.
Keyword s: spatial perception, c ross - modal per cept ion, freq uency - elevation
mappi ng .
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
105
2 Ho ł u b o ws k a , He r k e n b e r g a n d S c h ö n wi e s n e r
1 Introduction
When we talk about sounds, cross - modal connect ion s are omnipresent , as we borrow
vocabulary from other sensor y modalit ies to d escribe our exp erience [1] . For ins tance,
w hen we say, that one sound was “ higher ” than another , we typi cally refer t o an i ncrease
in fu ndamental frequency, not to a change in spatial location . Crucially, this mapping
is not just a linguistic convention – it reflects a perceptual phenom enon. The f reque ncy -
elevation mapping (FEM) , also know n as the Pratt effect , was first desc ribed in 193 0
[2] , when participants hearing pure tones ranging from 256 to 4096 Hz showed linear
relation between the pitch of presented sound and its perc eived spatial location, namely
sound of higher frequency were perceived higher in space. Rof fler and Butle r [3] late r
confirmed Pratt’s effect using complex tones of nine fundamental frequencies (250 –
7200 Hz). They found t hat high - frequency tones w ere perceived several degrees higher
than low - frequency ones, and this bias persisted despite changes in spatial position of
parti ci pants and the presence or absence of visual cues. Notably, it has been fo und that
children , as young as 4 - mo n th - old infants showed t he same pitch – height bias [4] , argu-
ing against a purely linguistic or learned origin (however see [5] for different conclu-
sions) .
Recent work suggests that FEM arise s from natural scene statistics , rather than
linguistical associations . Par ise, Knorre, and Ernst [6 ] analysed real - world sounds and
found a robust statistical correlation between pitch and source elevation : high - fre-
quency sounds in natur e are more likely to come from above than low - frequency
sounds . More over, they show that the anatomy of t he ear reinforces the mapping, as the
filtering properties of the pinnae reflec t the higher frequ ency - higher location mapping.
Thus, the explanation of the FEM cannot be deri ved solely from linguistic or cultural
associations t hrough music notation, for instance , but is also reflected in statistics of
the natural environ ment and the anatom y of the audito ry system.
In our daily experience, w e are surr ounded by statistical learning on m any lev-
els. Although we mig ht be exposed to the genera l natural scene statistics of higher
sounds co ming from elevated locations, we also learn, that single sound sources have
stable spatial locations . For example, a single violin produces sounds ranging from 200
to 2 600 Hz [7] . Thus, we have an additional global prior of a stable sound source of a
musical instr ument, with c hanging sound frequency .
On a mo re loca l level , pit ch perce ption i n musica l seque nces i nvolves not only
absolute frequenc ies but also rel ative relations hips between success ive tones. The same
frequency could be perceived di fferently depending on the precedi ng musical context.
This conte xtual depende ncy might influence frequency - elevation mapping by these rel-
ative pitch r elationships rather than operating solely on abs olute frequenc ies . Musical
intervals – the pitch distance between conse cutive notes – could therefore modulate
spat ial perception by establishing a directional context that guides elevation judgments.
Evidence from audi o - visual studies supports this, demonstrat ing t hat sequent ial pitch
changes can influence the perceived motion of visual objects [8] , suggesting tha t per-
ception of direction in pi tch and space is dynamically linked.
While F EM is a we ll - established effect , its sensitivity to contextual knowledge
remains unclear. We hypothesized that both global priors (e.g., knowledge of sound
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
106
Cr o s s - mo d a l c o u p l i n g b e t w e e n p i t c h a n d th e p e r c e p tio n o f sp a c e 3
source stability ) and local priors (e.g., pitch relations between tones) modulate FEM
strength. To test this, w e compared spatial localization of real musical instruments –
recognized as coherent, pitch - varying sources – with spectra lly - matche d ar tific ial
sounds . This design disentangles the effects of acoustics, sound source identity, and
context on frequency - space mapping in music.
2 Methods and Materials
2.1 Participant s
Twenty - nine healthy participants without sel f - reported hea ring deficits to ok part in the
study (19 females, 9 males, 1 diverse; M age = 24.55 ± 6.04) . They received a compensa-
tion for participation in the study (10€ per hour). Participants gave a written consent
before the study. Musicality of participants was assessed using t he Goldsmith Music
Sophistication I ndex (Gold - MSI; [9]). This study was carrie d out with the approva l of
the local ethics com mittee.
2.2 Stimuli
Participants were presented with tones of twelve differ ent fundamental fr equencies,
each lasting one second. These frequ encies were organ ized into fou r bins, each c entered
around one of t he fol lowing values : 208 Hz, 554 Hz, 1480 Hz, and 3951 Hz. For each
central frequency, two additional tones were included – one a semit one below and one
a semitone above – resulting in three freque ncies per bin. The frequency bins were later
used for data aggregation and for testing the m ain ef fects of pitch on sound locali zation
performance.
Experimental stimuli were created in four conditions: violin , flute , complex tone :
harmonic complex tones with fixed ha rmonic st ructure (7 harmoni cs of a mplitude - 10,
- 20, - 30, - 40, - 50, - 60, and - 70 dB), and synthetic viol in : generated as harmonic com-
plexes whose component frequencies and relative amplitudes were deri ved from spec-
tral analyses of recorded violin tones, thereby replicating the instrument’s harmonic
profile . The selection of the stimuli allowed us to group stimuli based on two criteria .
They can be grouped based on the ma tched spectral centroid, where vi olin and synthetic
violin have similar spectral cent roid, and s ystematically higher than flute and complex
tone ( Fig.1 A and B ). Alternatively, the sounds can be grouped based on musica l vs.
artificial sound contrast ( violin - flute and complex tone - synthetic violin ) . The spectral
centroid for each tone, a weighted average frequency of a spect rum, was computed
using slab package in python [10] .
Artific ial sounds ( complex tone and synthetic violin ) were created using slab
package [10] . Musical soun ds were genera ted in MuseSco re (MuseSco re Studio 4.4.2)
as MIDI files and recorded with the instruments’ library of Garage Band (vers ion
10.4.12) .
Pro c. of the 17th In ternational Symp osium on CMMR, London, UK, No v. 3-7, 2025
107
[Document text truncated for crawler view.]