Full text
Scoring Instructional Utterances in Guitar Lessons using Semantic Labels and LLM Nami Iino1,2[0000→0001→7806→7131],MasakiMatsubara 1,3[0000→0003→1950→683X], Masatoshi Hamanaka2[0000→0002→0249→4900],and Hideaki Takeda4[0000→0002→2909→7163] 1University of Tsukuba, Japan [email protected] 2RIKEN Center for Advanced Intelligence Project, Japan 3Keio University, Japan 4National Institute of Informatics, Japan Abstract. This study examines the utility of semantically grounded labels in private music instruction, focusing on how they capture instructional intent and identify important teaching utterances. Our prior work introduced a framework for semantic analysis of classical guitar lessons, annotating teacher utterances with six Instructional Content Labels (ICL): Giving Subjective Information, Giving Objective Information, Asking Question, Giving Feedback, Giving Practice, and Giving Advice. In this study, we extend this framework by developing an ICLweighted scoring method that combines utterance length with semantic weights to highlight instructionally significant discourse. We also reinterpret ICL categories for real-time spoken instruction, assigning higher weights to actionable guidance. To validate this approach, we compared our scoring outputs against rankings generated by multiple large language models—GPT-4.5, GPT-4o, and Claude Opus 4—across 24 classical guitar lessons. All models showed significantly stronger alignment with ICL-weighted scores than length-only baselines. Claude Opus 4 achieved near-perfect correlation (ω= 0.993), while GPT-4.5 also demonstrated strong alignment (ω= 0.902). These findings suggest that ICLweighted scoring can capture instructional priorities and that generalpurpose LLMs may approximate domain-specific judgments. The framework may provide a foundation for automated instructional analysis in music education. Keywords: Instructional content ·Guitar lessons ·Semantic scoring · Large language models. All rights remain with the authors under the Creative Commons Attribution 4.0 International License (CC BY 4.0). Proc. of the 17th Int. Symposium on Computer Music Multidisciplinary Research, London, United Kingdom, 2025 Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 506
N. Iino et al. 1Introduction Lessons in instrumental music involve a unique form of interactive communication between teachers and students. This communication occurs through playing sounds, verbal feedback and dialogue, and singing (see Fig. 1). Such lessons often involve repeated, step-by-step instruction on the same piece of music over multiple sessions. While students can record each lesson using audio or video, reviewing this material is often too time-consuming to be practical. This makes it di!cult to reflect on or fully understand previous instruction. Therefore, developing an approach that enables the structured review of past lessons remains akeychallenge. To address this issue, we previously developed a framework for the semantic analysis and annotation of teacher utterances in one-to-one classical guitar lessons. The aim was to clarify the instructional intentions embedded in verbal feedback across multiple sessions [8,9]. In this framework, teacher utterances were categorized from conceptual and instructional perspectives using six semantic labels, revealing structural patterns within lesson content. While these studies provided valuable insights into the semantic characteristics of music instruction, they remained theoretical and did not contribute directly to practical lessons. E"ective methods for supporting reflection, exploration, and reuse of lesson content by teachers and students are still lacking. This study aims to reinterpret Instructional Content Labels (ICL), originally designed for annotation, as functional indicators of instructional intent. We introduced an ICL-weighted scoring method to prioritize instructional utterances in private guitar lessons and evaluate its consistency using both human-annotated and machine-generated rankings. As a result, we found that the ICL-weighted scores were consistent with the pedagogical intuitions reflected in the language models’ rankings, with GPT-4.5 and Claude Opus 4 showing strong correlations. Claude, in particular, achieved near-perfect agreement across 24 lessons, suggesting that general-purpose LLMs can reliably approximate domain-specific instructional judgments. The contribution of this study is as follows: –We reinterpret manually annotated semantic labels (ICL) in music lessons as functional indicators of instructional intent. –We evaluate the generalizability of ICL using large language models, revealing their potential for automated instructional analysis. We position this study as a first step toward automated, pedagogically aware analysis of instructional discourse in music education. While our current focus is on verbal instruction, future work could integrate multimodal elements—such as gestures, posture, and acoustic environment. This paper is organized as follows. Section 2 gives an overview of earlier research on instructional aspects. Section 3 presents our dataset of multiple classical guitar lessons and investigates the usefulness of ICL. Section 4 presents the proposed ICL-weighted scoring method and evaluates its validity by comparing Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 507
Scoring Instructional Utterances in Guitar Lessons Fig. 1. Audio information in classical guitar lessons. it to large language model outputs. It also discusses the results. Finally, Section 5concludesthepaperandoutlinesdirectionsforfutureresearch. 2RelatedWork To understand how teachers convey instructional intent, it is essential to consider teacher–student interaction, the structure of pedagogical discourse, and emerging applications of language models in education. This section reviews these foundational areas. 2.1 Teacher–Student Interaction in Music Lessons Research on one-to-one music instruction has revealed complex dynamics of pedagogical communication. Daniel et al. [3] demonstrated that individual lessons tend to be teacher-dominated, limiting student-initiated interaction. Conversely, Yamamoto et al. [19] observed that in professional-level instruction, musical expression emerges through collaborative teacher-student exchanges. Hasumi et al. [5] further illustrated how teachers guide collaborative phrase development in jazz improvisation contexts. However, most studies focus on single-session interactions. In practice, instrumental instruction unfolds across multiple sessions, with teachers providing iterative guidance on the same repertoire. This cumulative dimension of teaching and learning remains underexplored. The shift to online instruction during the COVID-19 pandemic has further underscored the importance of verbal communication. Hrabluk [6] found that teachers view online lessons as pedagogically e"ective despite technological constraints, highlighting the central role of spoken instruction. Furthermore, recent meta-synthesis research has identified four crucial aspects for e"ective musical learning: framing of teaching, taking learners’ perspectives, sca"olding strategies, and representation of sounding music [4]. These findings provide a broader context for understanding how verbal instruction contributes to overall pedagogical e"ectiveness. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 508
N. Iino et al. 2.2 Semantic Annotation in Instructional Contexts Educational discourse analysis has employed various frameworks to categorize instructional communication. Traditional approaches include Zhukov’s [21] threecategory system (modeling, instruction, feedback) and Simones et al.’s [16] analysis of gesture use across student proficiency levels. Classroom discourse research has extensively studied Initiation-ResponseEvaluation (IRE) and Initiation-Response-Feedback (IRF) patterns [17], which often dominate traditional educational settings. Understanding these established patterns provides crucial context for interpreting the pedagogical significance of di"erent utterance types in specialized instructional contexts like music lessons. CROCUS (CRitique dOCUmentS) [15,14,11] is a collection of instructional feedback on instrumental performance in the form of written critiques. Currently, it covers three instruments—oboe, piano, and classical guitar—with a total of 635 critiques provided by 49 teachers for 213 performance recordings5.Tovalidate its utility, the researchers developed a set of domain-specific semantic tags to annotate these critiques, demonstrating high relevance and inter-rater reliability. In the oboe dataset, inter-rater agreement between two annotators reached Cohen’s ω=0.96. Our ICL framework draws directly on CROCUS, while adapting its methodology to the context of real-time spoken instruction to address the challenges of annotating spontaneous pedagogical discourse. 2.3 Large Language Models in Educational Discourse The emergence of large language models has created new opportunities for automated educational analysis. There are many types of applications such as enhancing the quality of teaching plans [7], writing feedback generation [20] and students’ assignments [2]. In instructional discourse analysis, Jensen et al. [10] developed systems for extracting pedagogically salient teacher utterances, while Ma et al. [13] applied generative models to real-time learner monitoring. Recent studies have demonstrated that LLMs can achieve human-level performance in analyzing classroom dialogue, particularly in identifying discourse patterns and providing automated feedback [18,12]. Complementary research has applied NLP-based semantic similarity analysis to classroom discourse [1], showing that automated methods can capture subtle aspects of teacher–student interaction quality. Our study extends this trajectory by examining whether LLMs can consistently identify pedagogical importance in structured lesson data, using humanderived semantic annotations as validation criteria. 3LessonDataandInterpretationofInstructionalLabels This section presents the dataset of classical guitar lessons developed in previous work and outlines the semantic labeling scheme applied to teacher utterances. 5https://masaki-cb.github.io/crocus/ Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 509
Scoring Instructional Utterances in Guitar Lessons We then examine the semantic usefulness of these labels and reinterpret them for the real-time context of spoken instruction. 3.1 Overview of the Lesson Dataset Our dataset consists of 24 private classical guitar lessons targeting seven pieces, with at least two sessions dedicated to each piece. The lessons were performed by three teacher-student pairs (two instructors and three students in total), resulting in approximately 8.5 hours of recorded instruction in total. Each lesson is indexed by a song ID (e.g., LS1) and a session number (e.g., LS1-1 for the first session of song LS1). The dataset is available on GitHub6,excludingrawaudio for privacy reasons. This dataset was previously annotated with semantic labels for teacher utterances in our previous study [9]. In this study, we build upon that framework to explore the practical utility and interpretation of these labels. 3.2 Instructional Content Label (ICL) The Instructional Content Label (ICL) refers to the semantic labels mentioned above. It categorizes semantic elements, such as the teacher’s intended message, and is based on textual critique research (CROCUS). We adjusted our definitions for the real-time spoken context of music lessons because they often feature overlapping concepts and incomplete phrasing. The ICL consist of the following labels: –Giving Subjective Information (GSI): teacher providing general and/or specific conceptual information based on teacher’s subjectivity. –Giving Objective Information (GOI): teacher providing general and/or specific conceptual information based on objectively referable events or concepts. –Asking Question (AQ): enquiring. –Giving Feedback (GF): teacher evaluation of a student’s applied and/or conceptual knowledge. –Giving Practice (GP): providing suggestions of ways to practice a particular passage or discussing a practicing schedule. –Giving Advice (GA): giving a specific opinion or recommendation without demonstration or modelling to guide the student’s action towards the achievement of certain specific musical aims. To ensure annotation reliability, two researchers independently labeled a subset of the lesson data, achieving 94.5% inter-annotator agreement for ICL assignments. This high agreement rate demonstrates the consistency and applicability of the ICL framework to real-time spoken instruction, despite the inherent challenges of annotating spontaneous speech compared to written critiques. 6https://github.com/guitar-san/guitar-lesson-data Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 510
N. Iino et al. 0 10 20 30 40 50 60 70 GSI GOI AQ GF GP GA Number of Labels ICL Labels Best10% Worst10% Crowdsourcing 0 10 20 30 40 50 60 70 GSI GOI AQ GF GP GA Number of Labels ICL Labels Best10% Worst10% Performers Fig. 2. Histogram of ICL in best 10% and worst 10% of critique documents. 3.3 ICL Patterns in Useful Instructional Feedback Matsubara et al. [14] showed that certain ICL labels appear more frequently in critiques of oboe performances within the CROCUS dataset when rated as highly useful. In particular, tags such as GA and GP were strongly associated with higher perceived usefulness. We analyzed the distribution of ICL in the top and bottom 10% of usefulness scores in classical guitar critiques (N=252), based on both third-party and performer evaluations. The results are shown in Figure 2(a). As was the case with the oboe, the top 10% had more labels (GOI,GF,GP,andGA)thanthebottom 10% (more than twice as many). This suggests that usefulness scores tend to increase with the amount of labeled information in critique texts. Figure 2(b) shows the results of the performers’ analysis of usefulness scores. Compared to the crowdsourcing results, the number of labels tended to decrease overall. This suggests that third parties, such as crowdsourcers, are more strongly influenced by the amount of information in the critique text. Conversely, this implies that, for the performers, having more information alone is insu!cient. 3.4 Reinterpreting ICL Based on Annotator Di!erences While our analysis using CROCUS (Section 3.3) revealed a correlation between certain ICL labels and higher usefulness ratings, similar patterns were not observed in our lesson dataset. This discrepancy can be attributed to di"erences in the nature of the data (spoken, real-time instruction vs. written critique) and in the profiles of the annotators. The ICL in CROCUS were annotated by two musicians (Group A) who had general musical experience, but no background in classical guitar. In contrast, the lesson data in this study were annotated by two researchers (Group B). One of these researchers is a professional classical guitarist and teacher. When comparing the annotated labels between Group A and Group B, we observed inconsistencies in the application of specific labels. For example, utterances labelled as GOI (objective information) by Group A were often labelled Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 511
Scoring Instructional Utterances in Guitar Lessons as GF (feedback) by Group B, particularly when the utterance was evaluative in tone. Similarly, the distinction between GA (advice) and GP (practice) was interpreted di"erently: Group A tended to classify general instructional suggestions as GP, whereas Group B reserved GP for concrete, repeatable actions and labelled broader guidance as GA. These findings suggest that real-time instruction’s usefulness depends more on directive content than on information quantity. As such, we reinterpreted the label distribution by prioritising the functional categories GF and GP, which reflect direct pedagogical interventions. In addition, we remapped ambiguous GOI utterances to GF and GA utterances to GP, in order to better align with the communicative intent observed in the lesson context. 4SemanticScoringandEvaluation Building on the label interpretation adjustments presented in Section 3.4, we have developed a method for assigning importance score to each utterance. 4.1 Instructional Importance Scoring Method To identify pedagogically salient utterances across the 24 classical guitar lessons6, we developed a scoring method that incorporates both utterance length and semantic label relevance. Since the number of utterances varied widely between lessons (total: 1669 teacher utterances), we designed a weighted importance score that reflects both the quantity and instructional value of each utterance. Let Ldenote the character count of a given utterance segment and tL the total character count of all utterances in the same lesson (e.g. LS1-1). The relative length of the utterance is given by L tL →100,representingitsproportion of the lesson’s total verbal content. Each utterance is also annotated with one Instructional Content Label (ICL), for which we assign a semantic weight w: GF=5, GP=4, GA=3, GOI=2, GSI=1, AQ=0. These weights were determined based on the reinterpretation of ICL categories discussed in Section 3.4. In that analysis, many utterances initially labeled as GOI were functionally closer to GF utterances, and utterances labeled as GA often aligned with GP.Asaresult,GF emerges as the most frequent and pedagogically central category, with GP occupying a nearly equivalent level of importance. We therefore assign the highest weights to GF (=5) and GP (=4), followed by GA (=3) and GOI (=2), which remain relevant but less directive. Finally, GSI (=1) and AQ (=0) are weighted lowest, as they primarily provide impressions or elicit responses rather than direct instruction. The final importance score is calculated as: importance =!L tL →100"→w This score assigns greater importance to utterances that are both relatively long and semantically more instructive. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 512
N. Iino et al. 4.2 Comparison with Length-only Baseline We first investigated the e"ectiveness of our instructional importance scoring by comparing it with a length-only baseline. Specifically, we analyzed: – ICL-weighted Score (importance):Usesbothlengthandlabelweight – Length-only Score: Ignores semantic labels We applied each method to a lesson (LS1-3), which was selected at random. For each method, we extracted the top 5 utterances by importance score and analyzed their distribution, content, and perceived teaching usefulness. To investigate the e"ect of semantic weighting on identifying pedagogically significant utterances, we compared the top five utterances from lesson LS13selectedusingtwoscoringmethods.Table1presentstheutterances,their instructional labels, text length, and computed importance score. All utterances were translated from Japanese to English for clarity. While both methods identified the same utterance as most important—a long advisory statement on performance pacing—the remaining selections diverged significantly. The ICL-weighted method surfaced utterances that included GP and GF,suchas: –“Try using the whole arm. That excessive motion isn’t necessary.” (GF) –“The ring finger, it’s tricky unless you align it correctly.” (GP) These utterances, though shorter, received higher importance score due to their pedagogical function. In contrast, some of the selected utterances under the Length-based setting lacked clear directive or corrective intent. such as: –“That D note just now was really nice” (GSI) To further validate our semantic scoring scheme, we calculated Spearman’s εbetween the ICL-weighted and length-based rankings for each lesson, using the rank() method with ’average’ tie-breaking. The average correlation across all lessons was ε=0.868,indicatingpartialagreementbutalsohighlightingthe distinct contribution of semantic weighting. Especially, the presence of GP and GF labels among the top-ranked results suggests that combining utterance length with semantic role may lead to more pedagogically meaningful prioritization. 4.3 LLM-Based Usefulness Analysis To test whether LLMs are capable of recognizing utterances that are teachingfocused in guitar instruction, we compared their output with two benchmarks: our proposed instructional importance score and a simple baseline based solely on utterance length. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 513
Scoring Instructional Utterances in Guitar Lessons Table 1. Top-5 Utterances for Lesson LS1-3 under Each Scoring Method (English Translation) Rank Method Utterance (translated excerpt) ICL Length Importance 1 ICL-weighted You played at a fairly fast tempo like before, so from now on, try to listen to each note and shape the musical story — no need to rush; let’s aim for a more careful performance. GA 147 19.36 2ICL-weightedAlso,somenotesaregettingtoo soft, so feel free to bring each one out more clearly. I get it, but I want to fix that. GF 76 15.01 3 ICL-weighted It feels like you’re focusing only on the fingers — try using the whole arm. That excessive motion isn’t necessary. GF 65 12.84 4ICL-weightedYes,exactly—especiallythering finger, it’s tricky unless you align it correctly. GP 64 10.54 5 ICL-weighted If you could adjust it a little more... it’s not just about playing well, it’s about conveying the music to the audience. GA 80 10.54 1 Length-based You played at a fairly fast tempo like before, so from now on, try to listen to each note and shape the musical story — no need to rush; let’s aim for a more careful performance. GA 147 19.36 2Length-basedIthinkit’sokaytolettherhythm breathe a bit — even if the overall tempo is the same, you can add expression internally. GSI 125 8.23 3 Length-based If you could adjust it a little more... it’s not just about playing well, it’s about conveying the music to the audience. GA 80 10.54 4Length-basedAlso,somenotesaregettingtoo soft, so feel free to bring each one out more clearly. I get it, but I want to fix that. GF 76 15.01 5 Length-based That D note just now was really nice — much better than before. GSI 76 5.00 Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 514