scieee AI-readable full text Open interactive document viewer

Implementation of a Brass Instrument Player Classification Model Using Facial Recognition with Two-Stage Classification Combining SVM and CNN

Miura, Hiroya; Inagaki, Kazuto; Watanabe, Hiroki; Takegawa, Yoshinari

Abstract

In wind music, beginners often have difficulty selecting an instrument that suits them. This is particularly true for brass instruments, as physical characteristics such as lip shape and bone structure influence playing aptitude, making appropriate instrument selection crucial. In this study, we developed a model for identifying suitable brass instrument players using facial recognition with a two-stage classification approach. This approach combines global classification using a Support Vector Machine (SVM) and local classification using a Convolutional Neural Network (CNN). Specifically, we implemented a deep learning model that estimates the appropriate instrument on the basis of lip shape and bone structure by analyzing profile photos of professional orchestra brass musicians. The experimental results demonstrated the effectiveness of the proposed classification model, achieving an accuracy of 72% in classification. Through this study, we aim to create an environment in which beginners are assigned suitable instruments, thereby enhancing their motivation.

Full text

Implementation of a Brass Instrument Player Classification Model Using Facial Recognition with Two-Stage Classification Combining SVM and CNN Hiroya Miura1, Kazuto Inagaki2,HirokiWatanabe 2,andYoshinariTakegawa 2 1RIKEN Center for Advanced Intelligence Project 2Future University Hakodate [email protected] Abstract. In wind music, beginners often have difficulty selecting an instrument that suits them. This is particularly true for brass instruments, as physical characteristics such as lip shape and bone structure influence playing aptitude, making appropriate instrument selection crucial. In this study, we developed a model that identifies a suitable brass instrument for a given player using facial recognition with a two-stage classification approach. This approach combines global classification using a Support Vector Machine (SVM) and local classification using a Convolutional Neural Network (CNN). Specifically, we implemented a deep learning model that estimates instrument suitability for a player on the basis of lip shape and bone structure by analyzing profile photos of professional orchestra brass musicians. The experimental results demonstrated the effectiveness of the proposed classification model, achieving an accuracy of 72% in classification. Through this study, we aim to create an environment in which beginners are assigned suitable instruments, thereby enhancing their motivation. Keywords: Musical Instrument Selection ·Musical Instrument Beginner ·Deep Learning. 1Introduction One of the most common reasons for taking up a wind instrument is to participate in extracurricular school activities such as wind bands or orchestras, which play a significant role in introducing students to musical performance. In school wind bands, ensembles typically consist of brass and woodwind instruments, complemented by percussion, providing students with their first experience of instrumental performance [1]. When joining a school band, students All rights remain with the authors under the Creative Commons Attribution 4.0 International License (CC BY 4.0). Proc. of the 17th Int. Symposium on Computer Music Multidisciplinary Research, London, United Kingdom, 2025 Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 494 Miura et al. usually choose their preferred instrument on the basis of introductory sessions. However, this selection process is generally conducted shortly after joining, and changing instruments later is often difficult. To clarify this issue, we preliminarily surveyed 41 individuals aged 10 to 50 who had played wind instruments for at least three years. The results revealed that 95.12% of respondents had their instrument assigned immediately after joining their school band. Additionally, 26.82% had switched instruments, while 7.14% had wished to switch instruments during their school years, and 51.21% had wished to switch instruments upon advancing to a higher level of education. These findings suggest that beginners are often not assigned the most suitable instrument for them, which may contribute to slower progress and decreased motivation for playing. In general, wind instrument playing techniques can be classified into four types (Figure 1, left), each requiring distinct movements and usage of the orbicularis oris muscle [2]. For every instrument, there is an optimal embouchure - a combination of the lips, tongue, teeth, jaw, and facial muscles - that facilitates proper sound production. For brass instruments in particular, an appropriate embouchure is crucial to establish. Since mouthpiece sizes vary between instruments, embouchure formation significantly affects tone quality and pitch range during performance (Figure 1, right) [2]. Additionally, dental occlusion or misalignment can make a comfortable embouchure difficult to maintain, potentially hindering musical development [3]. Therefore, finding the most suitable instrument for each player is of the utmost importance. The objective of this study is to develop a brass instrument player classification model using facial recognition, employing a two-stage classification approach that combines global classification using a Support Vector Machine (SVM) and local classification using a Convolutional Neural Network (CNN). Specifically, we implement a deep learning model that estimates a musician’s assigned instrument on the basis of lip shape and bone structure, using profile photos of professional orchestra wind instrument players. As the first phase of our proposed system, this study focuses on classifying brass instruments, specifically trumpet, horn, trombone, and tuba. This distinction is made because brass and woodwind instruments involve significantly different muscle usage and skeletal structures during performance, requiring separate sets of physical characteristics for classification. Furthermore, brass instruments vary notably in pitch range and mouthpiece size, making these factors critical indicators for evaluating player suitability. By incorporating physical attributes into the classification process, our proposed system is expected to enhance accuracy in instrument suitability assessments and enable instruments to be more appropriately assigned on the basis of each player’s physical characteristics. 2RelatedResearch 2.1 Instrument Selection Process Proper instrument selection is influenced by physical characteristics (such as lip size and shape, jaw structure, and overall physique), musical abilities (including Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 495 Implementation of a Brass Instrument Player Classification Model PitchRange FFrenchHorn Tenor-Bass Trombone BTrumpet ♭ BTuba ♭ MouthpieceRimDiameter(mm) HighLow E1 F4 E2 B4 ♭ F1 C5 E3 B5 ♭ 33 31 29 27 25 23 21 19 17 17~19 16~1725~28 30~33 Large Small BrassInstruments 1. Mouthpiece:CupShape Tuba Trombone Horn Trumpet WoodwindInstruments Clarinet Saxophone 2.SingleWoodenReed Oboe Bassoon 3.WoodenDoubleReed Flute 4.Mouthpiece:OpeningShape Fig. 1. Classification of Wind Instrument Playing Techniques (Left) and Mouthpiece Sizes with their Corresponding Pitch Ranges for Brass Instruments (Right) sense of rhythm and pitch perception), and external factors (such as advice from teachers and family) [4] [5]. Previous studies have indicated that when educators accurately assess students’ physical traits and musical abilities and assist them in selecting an appropriate instrument, it can enhance their performance success rate, maintain motivation, and increase retention in instrumental practice. Therefore, providing appropriate feedback and guidelines from educators serves as a foundation for students to select the most suitable instrument, contributing to the overall improvement of music education quality. Furthermore, Fortney et al. investigated the key factors influencing middle school students’ instrument choices and demonstrated that an instrument selection process considering physical characteristics can enhance students’ motivation and long-term engagement in playing [6]. Specifically, they emphasized that the size and shape of an instrument should match a student’s physical attributes, particularly physique and lip shape. This consideration is particularly crucial for brass instruments, where lip structure and strength play a significant role in sound production, and compatibility with the mouthpiece directly impacts the success of instrument selection. Additionally, the influence of family and teachers’ opinions on students’ instrument choices has been highlighted. However, a major challenge remains: many educators find physical characteristics difficult to quantitatively evaluate, often relying on experience and intuition rather than objective assessment methods. 2.2 Methodologies for Instrument Selection Support Various methodologies leverage a player’s physical characteristics to support instrument performance. Samuel et al. studied the vibration properties of the lips and their control mechanisms in brass instrument playing in detail [7]. This research clarified the interaction between lip vibrations and instrument shape, demonstrating that the required lip movements differ among brass instruments, affecting both playing difficulty and player suitability. Furthermore, by using physical modeling and simulations, the study quantitatively evaluated how lip movements influence acoustic properties. These findings highlight the significance of lip shape and motion in instrument selection and suggest that appropriate instrument assignment can contribute to improved playing technique. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 496 Miura et al. Boyette et al. proposed an instrument support method specifically designed for musicians with physical constraints [8]. Their approach involved customizing instrument fittings on the basis of physical characteristics such as lip and facial structure and upper limb morphology. The goal was to reduce chronic pain and physical strain while maintaining sound quality and playing flexibility. The study was based on the concept of “instrument adaptation,” where customized fittings were applied to actual musicians, and their impacts on acoustic properties and playing comfort were assessed. This approach provides a practical framework for instrument support that considers individual physical attributes. These studies collectively indicate that instrument selection and performance support based on physical characteristics are effective in developing methodologies tailored to each musician’s unique attributes. 3 Data Collection and Feature Extraction In this study, we aim to develop a model that accurately classifies brass instrument players on the basis of lip shape and skeletal structure. To achieve this, we collected images of wind instrument players belonging to professional orchestras and wind ensembles in Japan and categorized them in accordance with their instrument type. Subsequently, we detected landmarks in and extracted features from the image data. 3.1 Facial Image Data Collection In this study, we collected image data from the official websites of professional orchestras and wind ensembles in Japan, as well as from concert performance photos and musicians’ profile pictures. The dataset includes images from 42 ensembles, comprised of 420 musicians, including members of the NHK Symphony Orchestra3,theTokyoSymphonyOrchestra 4,andtheTokyoKoseiWindOrchestra5. The instrument distribution was as follows: trumpet (n=120), horn (n=110), trombone (n=100), and tuba (n=90). The decision to restrict image collection to musicians in Japan was made to minimize environmental factors related to instrument selection, given the shared cultural and educational background among Japanese musicians. Additionally, non-Japanese musicians may exhibit differences in skeletal structure and facial features, which could introduce variability in the model. By limiting the dataset to Japanese musicians, we aimed to improve the model’s accuracy by reducing potential confounding factors. The collected images were categorized on the basis of instrument type: trumpet (Tp), horn (Hr), trombone (Tb), and tuba (Tuba). For preprocessing, all image files were standardized to the JPEG format. Since the collected images contained upper-body portraits and background elements, we adjusted them to focus exclusively on the musicians’ faces. This was 3https://www.nhkso.or.jp/about/member 4https://tokyosymphony.jp/aboutTSO/member 5https://www.tkwo.jp/about/players Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 497 Implementation of a Brass Instrument Player Classification Model achieved using Dlib to detect facial landmarks and extract the face region. The detected face regions were cropped and saved as new images. For cases where automatic cropping was unsuccessful or incomplete, images were manually adjusted. Finally, to ensure compatibility with the model’s input requirements, all cropped images were resized to 200⇥200 pixels. 3.2 Landmark Data Detection To extract detailed features from facial image data, we focused on lip shape and overall facial structure, which are particularly important for wind instrument performance. For this purpose, we performed facial landmark detection using the Dlib library. Dlib provides a high-precision method for detecting 68 facial landmark points, allowing for detailed capture of facial position and shape characteristics. We detected all 68 landmark points on each face and specifically extracted the coordinates of landmarks corresponding to the lips (points 48⇠67). Additionally, we calculated various features (including distances between landmarks, lip angles, bilateral symmetry, and area ratios) and saved these values in a CSV file. The calculation methods for each feature are described below. Acquisition of Landmark Coordinates:Wedetected68landmarkpointsfor the entire face and extracted the coordinates of the lip landmarks (points 48⇠67). The lip landmarks are defined as follows. Let (xi,y i)represent the coordinates of landmark i. lip_landmarks ={(xi,y i)|i= 48,49,...,67} Distances Between Landmarks:Torepresenttheshapeofthelipsindetail, we calculated the Euclidean distances between the lip landmarks. The distance dij between each pair of landmarks is defined by the following equation. This distance reflects the relative characteristics of lip size and shape. dij =q(xixj)2+(yiyj)2,i,j2{48,...,67} Angle Features: The angle ✓formed by three key lip landmarks - the left corner of the lips (landmark 48), the upper lip center (landmark 62), and the right corner of the lips (landmark 54) - was calculated using the following steps: (1) Vector from the left corner to the upper lip center (v1): v1=(x48 x62,y 48 y62) (2) Vector from the right corner to the upper lip center (v2): v2=(x54 x62,y 54 y62) (3) Angle between the vectors (✓): cos ✓=v1·v2 kv1kkv2k,✓= arccos(cos ✓) Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 498 Miura et al. Here, v1·v2represents the dot product of the two vectors, while kv1kand kv2kdenote their respective norms (magnitudes). This angle serves as a feature representing the degree of lip opening and overall lip shape. Symmetry:Tomeasuretheleft-rightsymmetryofthelips,wecomputedthe absolute differences in the x-coordinates between corresponding landmark pairs on the left (landmarks 48⇠57) and right (landmarks 54⇠67) sides of the lips. The symmetry index is defined as: symmetry = 9 X i=0 |x48+ix54i| Area Ratio: The contour areas of both the entire face (landmarks 0⇠67) and the lips (landmarks 48⇠67) were computed using OpenCV. ·Face area (Aface): Aface =cv2.contourArea(face_contour) ·Lip area (Alip): Alip =cv2.contourArea(lip_contour) ·Lip-to-face area ratio (lip_ratio) : lip_ratio =Alip Aface ,A face >0 4ProposedModel In this study, we develop a classification model for brass instrument players on the basis of the image data and feature sets described in the previous section. To achieve this, we adopt a two-stage classification approach that combines global classification using SVM with landmark data and local classification using CNN with facial image data. 4.1 Classification Using SVM with Landmark Data For classification based on facial landmark data, we employed SVM. Since the landmark data consists of 235-dimensional feature vectors (coordinates: 136, distances: 20, angle: 30, symmetry: 30, area ratio: 9), which are lower in dimensionality than image data, and the dataset size is relatively small, we considered SVM to be an effective at preventing overfitting while maintaining classification efficiency. The SVM model was trained to classify musicians into four categories: trumpet, horn, trombone, and tuba. To address class imbalance, we expanded the feature set by scaling each class’s data, ensuring that each class contained 300 samples. The hyperparameters utilized in the SVM model for this study consist of a linear kernel, a regularization parameter C set to 1.0, and a random seed of 42. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 499 Implementation of a Brass Instrument Player Classification Model GlobalClassificationUsingSVM ClassificationUsingLandmarkDatawithSVM Tp Tb Hr Tuba Tp-Tb Group Hr-Tuba Group LandmarkExtractionandFeatureCalculation -LandmarkCoordinates -DistancesBetween Landmarks -AngleFeatures -Symmetry -Lip-to-FaceAreaRatio  Tuba Tb Hr Tp 0 1 0 1 LocalClassificationUsingCNN ClassificationUsingFacialImageDatawithCNN FacialImageData  -Tp-Hr:Class0 -Tb-Tuba:Class1 Tp-Tb Group Hr-TubaGroup 01 01 Tp-TbGroup Hr-TubaGroup x y -NHKSymphonyOrchestra -TokyoSymphonyOrchestra Others DataCollection Facial Image InputLayer (150,150,3) ConvolutionalLayer (148,148,32) PoolingLayer (74,74,32) Conversionto1D (7,7,128)→(6272) dropout FullyConnected Layer OutputLayer ×4 ×2 0 1 … … Flatten Fig. 2. Overview of the Proposed Model 4.2 Classification Using CNN with Facial Image Data For classification based on facial image data, we employed CNNs. CNNs are highly effective at automatically extracting image features through convolutional layers and are widely used in image classification tasks, with numerous studies reporting high performance. Given these advantages, we expected CNNs to yield reliable results in our study as well. For classification, we grouped trumpet and horn (which have similar pitch ranges) into Class 0, and trombone and tuba into Class 1, performing a binary classification task. To augment the dataset, random rotation, translation, and horizontal flipping were applied to the images, generating 1,000 samples per class. The hyperparameters used for the CNN model in this study include a number of filters set to (32, 64, 128), a kernel size of (3, 3), a pooling size of (2, 2), a dropout rate of 0.5, a learning rate of 0.0001, 100 epochs, and a batch size of 20. 4.3 Design of the Combined SVM and CNN Model In this study, we designed a two-stage classification model that integrates SVM and CNN for facial recognition. This approach leverages the complementary strengths of both methods to suitably balance classification accuracy and computational efficiency. SVM efficiently processes numerical feature data such as facial landmark coordinates with high speed and stability; however, it struggles to extract complex structural and contextual features directly from images. Conversely, CNN excels at automatically extracting detailed features from image data but has high computational costs and is highly dependent on dataset size for optimal performance. By combining these methods, our approach utilizes the strengths of both while mitigating their individual limitations. An overview of the proposed model is shown in Figure 2. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 500 Miura et al. Table 1. Comparison of Accuracy Across Methods Classification / Questionnaire SVM CNN Proposed Model Methods Value Avg. Value Avg. Value Avg. Value Avg. Two-Class Tp, Hr 0.60 0.57 0.65 0.61 0.82 0.79 - Classification Tb, Tuba 0.54 0.58 0.75 Tp 0.31 0.28 0.57 0.56 0.22 0.21 0.77 0.72 Four-Class Hr 0.34 0.72 0.23 0.72 Classification Tb 0.23 0.63 0.19 0.71 Tuba 0.24 0.32 0.17 0.66 Specifically, we adopted a hierarchical approach in which SVM performs global classification, grouping musicians into either trumpet/trombone or horn/ tuba, while CNN handles fine-grained classification within each group. This approach enables computational resources to be efficiently allocated by conducting an initial coarse classification at a low cost and then utilizing CNN’s high discriminative power for detailed classification. Moreover, SVM effectively captures static features of the lips and jaw, whereas CNN complements this by analyzing contextual facial features and individual variations. As a result, the proposed method is expected to enhance both the generalizability and classification accuracy of the model. Additionally, this two-stage classification method addresses challenges that are difficult to solve with a single algorithm. For instance, using SVM alone may lead to misclassification due to its dependence on landmark data, while CNN alone imposes a high computational burden, limiting its practicality. The proposed model overcomes these issues, achieving fast and accurate classification performance. Therefore, this model demonstrates high performance and practical applicability in the classification of brass instrument players, making it a promising solution for real-world applications. 5ExperimentalResults In this experiment, we evaluated the effectiveness of the proposed method by comparing classification accuracies across four approaches: subjective classification based on a questionnaire, SVM alone, CNN alone, and the proposed two-stage classification method. The results are summarized in Table 1. 5.1 Classification Results from the Questionnaire To assess the accuracy of human classification, we surveyed 41 individuals aged 10 to 50 who had played wind instruments for at least three years. Participants were shown facial images only (no instruments visible) of professional brass musicians (trumpet, horn, trombone, and tuba) and asked to identify each musician’s instrument from four multiple-choice options. A total of 20 images (five per instrument) were used. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 501 Implementation of a Brass Instrument Player Classification Model The overall accuracy rate was 28%, with frequent misclassification occurring among instruments with similar pitch ranges. In particular, horn players were often misclassified as trumpet or trombone players, likely due to the horn’s wide pitch range. Additionally, when the classification task was simplified into two categories - high-pitched instruments (trumpet, horn) vs. low-pitched instruments (trombone, tuba) - the accuracy improved to 57%. 5.2 Classification Results Using the Proposed Model Classification Results Using SVM The four-class classification using SVM with landmark data achieved an accuracy of 56%. Analysis of the confusion matrix shown on the left of Figure 3 revealed several key trends. Misclassification occurred most notably between trumpet and trombone, where 34 trumpet samples were correctly classified, but 24 were misclassified as trombone, and 38 trombone samples were correctly identified, while 20 were misclassified as trumpet. In contrast, misclassification between trumpet/trombone and horn/tuba was relatively rare, with only 4 out of 120 samples assigned to the wrong category. A similar pattern was observed within the horn-tuba group, where 44 horn samples were correctly classified, but 14 were misclassified as tuba, and 19 tuba samples were correctly identified, while 27 were misclassified as horn. Meanwhile, misclassification between horn/tuba and trumpet/trombone occurred in 16 out of 120 samples. These results indicate that the trumpet-trombone and horntuba groups can be separated relatively clearly. Overall, SVM demonstrated high accuracy in global classification, effectively distinguishing between the trumpettrombone and horn-tuba groups, but its accuracy in fine-grained classification within each group needs to be further improved. Classification Results Using CNN Building on the four-class classification results from SVM, we implemented a two-stage classification approach using CNN to improve classification accuracy within the same group. Specifically, we categorized trumpet and horn as Class 0 and trombone and tuba as Class 1. Based on the confusion matrix shown in the center of Figure 3, CNN achieved an overall classification accuracy of 79%. For trumpet and horn, classification remained challenging due to their players’ similar lip and jaw structures. However, CNN’s ability to extract subtle image features significantly reduced misclassification compared to SVM. Similarly, in trombone and tuba classification, differences in mouthpiece size and pitch range were reflected in facial features, which CNN successfully captured, leading to improved classification accuracy. Notably, misclassification in the horn-tuba group was significantly reduced, demonstrating CNN’s strength in extracting high-dimensional features. These results indicate that CNN effectively complements SVM by addressing its limitations in fine-grained classification, thereby improving overall classification performance. This demonstrates the effectiveness of the hierarchical approach, where SVM efficiently performs global classification (trumpet/trombone vs. horn/tuba), while CNN handles fine-grained classification within each group. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 502