scieee AI-readable full text Open interactive document viewer

ABOUT ONE METHOD OF SOUND ANALYSIS

Shoista Kasimova

Abstract

First, let's ask the question: who needs to recognize radio ads? This is useful for advertisers, who can track the actual airing of their commercials and spot instances of cuts or interruptions; radio stations can monitor the airing of online ads in regions, and so on. The same recognition problem arises if we want to track the playback of a musical piece or identify a song from a short fragment.

Full text

THE VI INTERNATIONAL SCIENTIFIC CONFERENCE “SCIENTIFIC FOUNDATIONS FOR THE USE OF INFORMATION TECHNOLOGIES OF A NEW LEVEL AND MODERN PROBLEMS OF AUTOMATION”, NOVEMBER 20, 2025 281 ABOUT ONE METHOD OF SOUND ANALYSIS Shoista Kasimova Associate Professor of the Department of Programming Technology at the Tashkent University of Information Technologies named after Muhammad al-Khwarizmi https://doi.org/10.5281/zenodo.17767968 Annotation. First, let's ask the question: who needs to recognize radio ads? This is useful for advertisers, who can track the actual airing of their commercials and spot instances of cuts or interruptions; radio stations can monitor the airing of online ads in regions, and so on. The same recognition problem arises if we want to track the playback of a musical piece or identify a song from a short fragment. A more precise formulation of the problem is as follows: we have a set of reference audio fragments (songs or commercials), and an audio recording of a broadcast in which some of these fragments are presumably played. The goal is to find all the fragments played, determine their start times, and determine their duration. If we're analyzing broadcast recordings, the system as a whole needs to operate faster than real time. Keywords: audio file; audio signal; audio recording; audio playback; audio processing; mixing; acoustics; sampling frequency; analysis; spectral centroid. INTRODUCTION Sounds are time-varying signals in the real world, and indeed, their entire meaning is tied to this temporal variability. Therefore, it is interesting to develop sound analysis methods that capture at least some of the distinctive features of time-varying sounds to facilitate understanding, comparison, modification, and resynthesizes. The development of computer technology has significantly changed many areas of human activity, creating new opportunities and increasing the efficiency of traditional work. One area in which computers have truly revolutionized is audio technology, and in particular, music production in its diverse applications [19, p.232]. Even the first personal computers differed from calculators and mainframes in that they could produce sounds using a small speaker installed in their case. And although PC acoustic capabilities were more than modest, music editors appeared at the dawn of the computer era, allowing one to create a "sound file" for connection to a user-written program. The advent of sound cards opened up new possibilities for users. And it wasn't just the vastly improved sound quality. A new (sound) subsystem emerged. — a complex of software and hardware designed for: • Recording audio signals from external sources, such as a microphone or tape recorder. During the recording process, the input analog audio signals are converted to digital and can then be saved to the PC's hard drive; • playback of recorded previously recorded audio data using an external speaker system or headphones. During playback, the audio signal is read from the storage medium, converted from digital to analog, and sent to the speaker system; • mixing (mixing) when recording or playing signals from multiple sources; • simultaneous recording and playback sound signals; THE VI INTERNATIONAL SCIENTIFIC CONFERENCE “SCIENTIFIC FOUNDATIONS FOR THE USE OF INFORMATION TECHNOLOGIES OF A NEW LEVEL AND MODERN PROBLEMS OF AUTOMATION”, NOVEMBER 20, 2025 282 • processing sound signals: editing, combining or separating signal fragments, filtering, changing its level, etc.; 1. AUDIO ANALYSIS Audio analysis —a field that includes automatic speech recognition (ASR), digital signal processing, as well as music classification, tagging, and generation—is a growing subdomain of deep learning applications. Some of the most popular and widespread machine learning systems, such as virtual assistants Alexa , Siri , and Google Home , are products built on models that extract information from audio signals[16,p.169]. Audio fragments are presented in the format .wav. Sound waves are digitized by sampling them at discrete intervals known as the sampling rate (typically 44.1 kHz for CD-quality audio, or 44,100 samples per second). Each sample represents the amplitude of a wave over a given time interval, where the bit depth (or dynamic range of the signal) determines how detailed the sample will be (usually 16 bits, meaning a sample can range from 65,536 amplitude values). In signal processing, sampling is the conversion of a continuous signal into a series of discrete values. The sampling rate is the number of samples taken over a given fixed period of time. A high sampling rate results in less information loss but higher computational costs. Sound processing applications These include: • Indexing music collections according to their audio features. • Recommendation of music for radio channels. • Similarity search for audio files (Shazam). • Speech processing and synthesis – generating artificial voice for dialog agents. 1.1 Processing Audio Data with Python Sound is represented in the form of an audio signal with parameters such as frequency, bandwidth, decibels, etc. A typical audio signal can be expressed as a function of amplitude and time [2, p.45]. Some devices can capture these sounds and present them in a machine-readable format. Examples of these formats include: • wav (Waveform) Audio File) • mp3 (MPEG-1 Audio Layer 3) • WMA (Windows Media Audio) The audio processing process involves extracting acoustic features relevant to the task, followed by decision-making processes that include detection, classification, and knowledge fusion. Python libraries help solve this problem. 1.2. Python Audio Libraries We will use two libraries to collect and play audio [14, p.130]: - Librosa Librosa can handle any audio signal, but is primarily focused on music. It allows you to create a fully-fledged music information retrieval (MIR) system. The module is well-documented , and numerous user guides are available. You can also install the fmpeg module with many ready-made solutions for converting audio signals . THE VI INTERNATIONAL SCIENTIFIC CONFERENCE “SCIENTIFIC FOUNDATIONS FOR THE USE OF INFORMATION TECHNOLOGIES OF A NEW LEVEL AND MODERN PROBLEMS OF AUTOMATION”, NOVEMBER 20, 2025 283 - IPython.display.Audio IPython.display.Audio allows you to play audio directly in Jupyter Notebook . Sampling rate is the number of samples (vibrations) of sound transmitted per second, measured in Hz or kHz. IPython.display .Audio is used to play audio . Using librosa.display.waveplot , you can visualize an array of audio data. 1.3. Processing of spectograms in Python A spectrogram is a visual representation of the frequency spectrum of sound or other signals changing over time. They are sometimes also called sonograms. In two-dimensional graphs, the first axis represents frequency, and the second axis represents time. To create a spectrogram in Python, we use librosa.display.specshow . Fig.1.1. Spectogram of an audio file The vertical axis represents frequencies (from 0 to 10 kHz), and the horizontal axis represents clip time. Since all significant changes occur at the lower end of the spectrum, the frequency axis can be converted to logarithmic. Extracting entities Each audio signal has many characteristics from which the desired ones must be selected. The process of extracting information for analysis is called feature extraction or entity extraction. Zero crossing frequency Zero crossing frequency ( zero crossing The signal rate is the frequency at which the signal changes sign, i.e., the frequency with which the signal alternates between positive and negative. This function is widely used in both speech recognition and musical information extraction . For metal and rock, this parameter is typically higher than for other genres due to the large number of drums [18-20, p.38]. Let's calculate the zero crossing frequency for our Python example: THE VI INTERNATIONAL SCIENTIFIC CONFERENCE “SCIENTIFIC FOUNDATIONS FOR THE USE OF INFORMATION TECHNOLOGIES OF A NEW LEVEL AND MODERN PROBLEMS OF AUTOMATION”, NOVEMBER 20, 2025 284 Fig. 1.2. Determining the zero-crossing frequency Spectral centroid The spectral centroid indicates where the "center of mass" of a sound is located and is calculated as a weighted average of all frequencies. In blues compositions, frequencies are evenly distributed, and the centroid lies somewhere in the middle of the spectrum. In metal, there is a pronounced frequency shift toward the end of the composition, so the centroid lies closer to the end of the spectrum. Let's calculate the spectral centroid for each frame using librosa.feature.spectral _ centroid: Spectral frequency roll-off It is a measure of the waveform shape, representing the frequency below which a certain percentage of the total spectral energy lies, for example 85%. librosa.feature.spectral_rolloff calculates the frequency rolloff for each frame. Mel-frequency cepstral coefficients Mel-frequency cepstral coefficients ( MFCC ) of a signal are a small set of characteristics (usually around 10-20) that succinctly describe the overall shape of the spectral envelope. This parameter models the characteristics of the human voice [5-11]. We can also perform scaling so that each coefficient dimension has zero mean and unit variance. After getting acquainted with the structure of an acoustic signal and the specifics of the process of extracting musical information, let's look at Python libraries for working with sound. Let's try to model a music genre classifier. It will come in handy if you need to parse a bunch of unknown MP3 files. Music genre classifier Before training a classification model, we need to transform the raw audio data into a more meaningful representation. We'll convert the clips to wav format so that Python can process them using the SoX module[12-14]. Now let's extract all the necessary information from the audio files: • mel-frequency cepstral coefficients, • spectral centroid, • zero crossing frequency, • color frequencies, • spectral frequency roll-off. THE VI INTERNATIONAL SCIENTIFIC CONFERENCE “SCIENTIFIC FOUNDATIONS FOR THE USE OF INFORMATION TECHNOLOGIES OF A NEW LEVEL AND MODERN PROBLEMS OF AUTOMATION”, NOVEMBER 20, 2025 285 We will save all these functions in a .csv file. One can use existing classification algorithms to classify songs into genres, or use spectrograms directly, or extract entities and apply classification models on them. Genre classification is just one of many applied areas of musical information extraction. First, you need to convert the audio files into PNG images (spectrograms). Then, extract the relevant features from them: MFCC, spectral centroid, zero-crossing rate, chromaticity frequencies, and spectral rolloff. After extraction, the features can be added to a CSV file so that the ANN can be used for classification. CONCLUSION Everyone knows that sound (in the narrow sense) is waves of compression and rarefaction propagating through air. A sound recording, for example in a wav file, is a sequence of amplitude values (physically, this corresponds to the degree of compression, or pressure). But we don't perceive these frequency fluctuations directly; instead, we hear sounds of varying frequencies and timbres. Therefore, another method of visualizing sound is often used— a spectrogram , where the horizontal axis represents time, the vertical axis represents frequency, and the color of the dot represents amplitude. The task of searching for a fragment on the air can be broken down into two parts: first, finding candidates among a large number of reference fragments, and then checking whether the candidate actually plays in a given fragment of the air, and if so, at what point the sound begins and ends. Both operations rely on a "fingerprint" of the fragment of the sound. This fingerprint must be noise-resistant and sufficiently compact. This fingerprint is constructed as follows: we divide the spectrogram into short time segments, and in each segment, we search for the frequency with the maximum amplitude (in reality, it is better to search for several maxima in different ranges, but for simplicity, we will take one maximum in the most meaningful range). The set of such frequencies (or frequency indices) constitutes the fingerprint. Very roughly, these are the "notes" sounding at each moment in time. LITERATURE 1. Trevor Cox, The Sound Book: The Science of the Sonic Wonders of the World, AzbukaAttikus Publishing Group, 2018. 2. Rosche s sho Davide. Sound processing, 2012, 244 3. Burger W., Burge MJ Digital Image Processing: An Algorithmic Introduction using Java. – New York: Springer, 2008. –564 p. 4. Gonzalez R., Wood R. Digital image processing. 3rd edition, corrected and supplemented. – Moscow: Tekhnosfera, 2012. – 1104 p. 5. Gashnikov M. V., Glumov N. I., Ilyasova N. Yu., et al. Methods of computer image processing. 2nd edition. Ed. by V. A. Soifer. - Moscow: Fizmatlit , 2003. -784 p. 6. Decree of the President of the Republic of Uzbekistan dated February 7, 2018 No. UP-4947 “On the strategy of actions for the further development of the Republic of Uzbekistan”. 7. Ablameiko S.V., Lagunovsky D.M. Image processing: technology, methods, application. - Minsk: Amalfeya, 2000. 8. Yane B. Digital image processing. - M.: Tekhnosfera , 2007. – 584 p. 9. Gruzman I.S., Kirichuk V.S., Kosykh V.P. et al. Digital image processing in information systems. – Novosibirsk: NSTU, 2002. – 352 p. THE VI INTERNATIONAL SCIENTIFIC CONFERENCE “SCIENTIFIC FOUNDATIONS FOR THE USE OF INFORMATION TECHNOLOGIES OF A NEW LEVEL AND MODERN PROBLEMS OF AUTOMATION”, NOVEMBER 20, 2025 286 10. Vizilter Yu.V. Image processing and analysis in machine vision problems. – M.: Fizmat kniga, 2010. – 672 p. 11. Pratt W.K. Digital Image Processing. – New York: Wiley, 2001. – 792 p 12. Jahne B. Digital Image Processing. - New York: Springer, 2005. – 585 p. 13. Chi Zh., Yan H., Pham T. Fuzzy algorithms: with applications to image processing and pattern recognition. New Jersey: World Scientific, 1996. – 230 p. 14. Gonzalez R., Woods R. Digital Image Processing. 3rd Edition - New Jersey: Prentice-Hall, 2008. – 954 p 15. Zagumennov A. Computer sound processing. – M.: Radio and communication, 2008. – 496 p. 16. Olli Niemitalo. Digital sound processing tutorial for the braindead! -2012,-562 rub. 17. Kintzel T. A Programmer's Guide to Working with Sound. Translated from English. – Moscow: DMK Press, 2000. – 432 p. 18. Loyanich A.A. Recording and processing sound on a computer - M.: Goryachaya Liniya, 2014. - 462 p. 19. Sekunov N.Yu. Sound processing on PC. – SPb.: BHV-Petersburg, 2001. – 1238 p. 20. Furui S. Digital Speech Processing, Synthesis and Recognition. New York: Marcel Dekker, 2001. - 452 p.