Full text
Pitch Estimation in Real Time: Revisiting SWIPE with Causal Windowing Peter Meier1[0000000230941931],SebastianStrahl 1[0009000796547762],Simon Schwär1[000000015780557X],MeinardMüller 1[0000000160627524],andStefan Balke1[0000000313063548] International Audio Laboratories Erlangen, Germany [email protected] Abstract. Pitch estimation in real time is essential for a wide range of Music Information Retrieval (MIR) applications, including intonation monitoring, music education, and interactive systems. Many of these use cases, such as ensemble rehearsal settings, require low-latency, multichannel processing on resource-constrained devices. While recent neural approaches offer high accuracy, they often fall short in real-time performance due to computational demands. In this paper, we revisit the well-established SWIPE algorithm and introduce RT-SWIPE,areal-time variant enabled by using causal windowing. We further propose a delaytolerant evaluation metric that extends Raw Pitch Accuracy (RPA) to account for algorithmic delays. Experimental results on synthetic signals and multi-track ensemble recordings demonstrate that RT-SWIPE provides a practical balance of latency, accuracy, and efficiency. Although our study focuses on wind orchestra scenarios, the method is broadly applicable to similar real-time settings. Keywords: Real-Time ·Multi-Channel ·Pitch Estimation. 1Introduction Real-time pitch estimation plays a crucial role in Music Information Retrieval (MIR), with applications ranging from intonation monitoring [7] and music education [17] to music scene analysis [8] and interactive gaming [13]. Many of these applications operate on resource-constrained platforms, such as smartphones, and demand low-latency, multi-channel processing [6]. In such contexts, lightweight and real-time capable algorithms such as YIN [5] are often favored over more computationally intensive methods like PYIN [10] and CREPE [9], despite potential trade-offs in estimation accuracy. The SWIPE algorithm is a well-established method for pitch estimation, recognized for its robustness under real-world conditions [3]. Although not among the All rights remain with the authors under the Creative Commons Attribution 4.0 International License (CC BY 4.0). Proc. of the 17th Int. Symposium on Computer Music Multidisciplinary Research, London, United Kingdom, 2025 Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 285
P. Meier et al. Future SamplesFuture SamplesFuture Samples (a) (b) (c) Non-Causal (centered) Causal (centered) Causal (right-aligned) Input Waveform Analysis Windows Prediction Point Window Delay Fig. 1. Overview of window positioning strategies for real-time audio processing with SWIPE.(a) Non-causal, center-aligned windows used in the original (offline) SWIPE algorithm. (b) Causal, center-aligned windows with a constant delay for real-time processing. (c) Causal, right-aligned windows. most recent developments, SWIPE remains relevant due to its resilience to noise and crosstalk—a critical factor in the context of wind music ensembles [12]. Its use of the Fast Fourier Transform (FFT) ensures computational efficiency, enabling real-time, multi-channel processing even on resource-constrained hardware. While neural network-based approaches may achieve superior pitch accuracy, their computational complexity and hardware demands often limit their applicability in low-latency, real-time scenarios. In addition, the structure of the spectral kernels in SWIPE offers a flexible foundation for further adaptations. Recent developments such as differentiable SWIPE (dSWIPE), which introduces trainable parameters, demonstrate the algorithm’s ongoing relevance and adaptability [16]. SWIPE was originally developed for offline pitch estimation. Figure 1a illustrates this offline setting, where an input waveform is correlated with a set of analysis windows of varying lengths to estimate the pitch at a given prediction point. This configuration is inherently non-causal, as it requires access to future input samples. A causal adaptation is shown in Figure 1b: By introducing an artificial delay equal to half the length of the longest analysis window, the algorithm can be rendered causal, as it no longer depends on future data. In the illustrated example, the input is a sinusoidal signal with linearly increasing frequency. Due to the uniform shift of all windows, pitch estimates for higher frequencies—which rely on shorter windows—may suffer from reduced accuracy, as the analysis does not capture the most recent signal content. This limitation can introduce a systematic bias toward lower frequencies, particularly problematic in musical contexts involving glissandi or rapid note changes. As shown in Figure 1c, this issue can be mitigated by right-aligning all analysis windows with respect to the prediction point. This strategy ensures that each window captures the most up-to-date input samples. Similar approaches were used before, as demonstrated in [2], which explored frequency dependent latency for the Constant-Q Transform (CQT). While the overall delay—determined by the longest window—remains unchanged, the responsiveness of shorter windows Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 286
Real-Time SWIPE with Causal Windowing improves substantially. Since these windows correspond to higher frequencies, which typically exhibit faster temporal variation, the alignment helps preserve estimation accuracy during fast transitions. In contrast, longer windows associated with lower frequencies are less affected, as low-frequency content tends to vary more slowly. In our research efforts, we aim to apply pitch estimation in ensemble settings, e.g., wind ensembles or orchestras. In such contexts, low latency is essential, as many channels must be processed simultaneously in real time. Furthermore, the presence of variable acoustics and interference from other instruments poses additional challenges, establishing ensemble environments as a demanding yet valuable testbed for real-time pitch estimation. Our previous studies identify SWIPE as the most suitable approach for these scenarios. Compared to CREPE,it is significantly less computationally demanding, and it offers greater robustness to cross-talk than YIN [12]. This paper continues this line of research by adapting SWIPE to operate in acausalmanner,makingitsuitableforreal-timeapplications.Wealsoconduct systematic experiments on ensemble recordings, closely aligned with our target use case. During our evaluation, we observed that Raw Pitch Accuracy (RPA) exhibits limitations, particularly when applied to real-time scenarios. To address this, we propose an extension to the RPA metric that accounts for timing tolerances. While our experiments focus specifically on ensemble music, the underlying concepts and findings are applicable to a broader range of real-time pitch estimation tasks. The remainder of the paper is structured as follows. Section 2 introduces the RT-SWIPE algorithm and analyzes its computational efficiency in multi-channel scenarios. Section 3 presents systematic experiments on wind ensemble recordings and compares RT-SWIPE to state-of-the-art pitch estimation methods. In addition, we examine the influence of algorithmic delays on the standard Raw Pitch Accuracy (RPA) metric and propose an extension to address its limitations in real-time settings. Finally, Section 4 summarizes our findings and outlines directions for future work. Additional materials and resources are available on a supplemental website.1 2Real-TimeSWIPE Our real-time approach builds upon the original SWIPE algorithm [3]. In this section, we outline the key components of SWIPE that are relevant to our real-time modifications. The SWIPE algorithm estimates pitch by examining the harmonic structure of a sound. It utilizes pitch candidate kernels derived from a sawtooth waveform and correlates these with the spectrum of an input signal to identify the best match. The candidate with the highest correlation score is selected as the estimated pitch. The SWIPE algorithm employs different window sizes to accurately estimate the short-time spectrum of an input audio signal. For a given set of predefined 1https://www.audiolabs-erlangen.de/resources/MIR/2025-CMMR-RTSWIPE Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 287
P. Meier et al. pitch candidates fi,eachrepresentedbyafrequency-domainkernel,wedetermine the optimal window sizes Nithat maximize the overlap between these kernels and the spectral lobes. These window sizes are calculated as follows: Ni=8·fs fi ,(1) where fsstands for the sampling rate of the input signal in Hz. However, instead of computing individual spectrograms for each pitch candidate, SWIPE approximates this process by calculating spectrograms for a limited set of window sizes that are powers of two. This approach requires interpolation between window sizes, but significantly reduces computational costs [3]. To generate the spectral representation, SWIPE utilizes window sizes tailored to the pitch range. Let fmin and fmax denote the minimum and maximum expected frequencies in Hz, respectively. We define: a=log2✓8·fs fmax ◆⌫,b=⇠log2✓8·fs fmin ◆⇡.(2) From these, the smallest and largest window sizes are calculated as: Nmin =2 a,N max =2 b.(3) Using these boundaries, we define a set Kthat includes all power-of-two window sizes used for spectrogram computation: K={Nmin,2·Nmin,...,N max}={2a,2a+1, ..., 2b}.(4) For more details on the original method, we refer to [3]. 2.1 Real-Time Audio Processing In real-time audio processing, data is managed block-wise. These blocks, or frames, contain H2Nsamples each and do not overlap. We refer to Has the hop size. When using SWIPE for real-time audio processing, three key considerations arise: (1) Real-time processing follows strict time constraints. The computation time, or processing latency,mustnotexceedasingleframeperiod Tframe,givenby: Tframe =H fs .(5) This constraint is particularly relevant for multi-channel scenarios, where the processing latency may scale with the number of channels. (2) A single audio frame is typically smaller than the maximum window size Nmax 2Nrequired for estimation. To address this, we need to implement audio buffering. (3) In real-time processing, only past data is accessible. Since analysis windows are usually centered, as illustrated in Figure 1b, estimates are delayed by half the maximum window size Nmax,introducingalgorithmic delay. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 288
Real-Time SWIPE with Causal Windowing Fig. 2. Detailed view of adjustable window positioning with delay factor for RT-SWIPE:(a) Centered windows share a common maximum delay. (b) Right-aligned windows use the latest audio data for estimation, each with individual delays. To manage larger window sizes, we use a rolling buffer Bthat efficiently handles continuous audio input by cyclically adding new frames and discarding the oldest ones, defined as follows: B2RC⇥Nmax .(6) C2Nis the number of audio channels, and Nmax is the largest window size (see Equations 2 and 3). For each new frame of audio input, an estimate is generated through a sequence of computations that must be completed within a single frame period Tframe. The following steps are performed: (a) Update buffer Bwith the latest frame of audio input. (b) For each window size Ni2K,extractthecorrespondingdatafromthebuffer according to the window positioning (see Section 2.2). (c) Apply a Hann window to each window. (d) Compute the Discrete Fourier Transform (DFT), vectorized over all Cchannels to ensure computational efficiency. (e) Follow the original method to correlate the resulting spectra with pitch candidate kernels and select the candidate with the highest score as the estimated pitch. 2.2 Adjustable Window Positioning In Figure 2a, we illustrate a more detailed view of causal centered windows as used in the RT-SWIPE method. All windows are not centered around the prediction point (compare to the non-causal case in Figure 1a); instead, each Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 289
P. Meier et al. window extracts data from the center of buffer B,containingonlypastsamples. Consequently, estimates are delayed by half the maximum window size Nmax, even though more recent samples could be used for smaller windows (higher frequencies). To address this, we propose a right-aligned window positioning (Figure 2b), where each window uses the most recent audio samples available in buffer B. This reduces the delay, especially for small window sizes, thereby affecting the expected algorithmic delay for high pitches. To interpolate between centered (offline) and right-aligned (real-time) window positioning, we introduce a normalized delay factor 2[0,1], where =1corresponds to centered and =0 to right-aligned positioning. For each delay factor ,thecorrespondingdelaysizeDin samples is defined as D()=·Nmax 2⌫,(7) which describes the position at which the windows are intended to be centered, i.e., the position of the black dashed line in Figure 2. Depending on the individual window sizes Ni, we further introduce window delays (given in samples) Ni= max ✓Ni 2,D◆,(8) which represent the positions at which the windows are actually centered, considering that only past samples are available. These center positions are illustrated by red lines in Figure 2. For right-aligned window positioning (=0), the window delays Nvary, introducing a pitch-dependent delay. Low pitches corresponding to large windows have large delays, while high pitches with small windows have small delays. In contrast, centered window positioning (=1)providesafrequency-independent delay determined by the maximum window size Nmax. 2.3 Multi-Channel Efficiency For our research and field studies on wind music ensembles, we require not only real-time pitch estimation but also the ability to handle multiple input channels simultaneously. To this end, we investigate the multi-channel efficiency of the RT-SWIPE implementation and test how many input channels can be processed on consumer-level hardware. As illustrated in Figure 3, we measure the real-time factor (RTF) as the number of input channels increases, identifying the maximum number of channels for which the RTF remains 1. An RTF smaller than 1 indicates computations take longer than a frame period, making it unsuitable for real-time processing. In contrast, an RTF greater than 1 means computations are faster than required, making it suitable for real-time applications. For our efficiency experiment, we conduct tests on a Mac Mini 2023 equipped with an Apple M2 Chip and 16 GB of RAM. We compare RT-SWIPE with an offline implementation of SWIPE [15], Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 290
Real-Time SWIPE with Causal Windowing Fig. 3. Efficiency of RT-SWIPE:Real-TimeFactor(RTF)overnumberofchannels. which is not optimized for frame-based processing and, therefore, serves as a lower performance bound. For an upper performance benchmark, we use the YIN implementation from librosa [11], a frame-based algorithm known for its speed and efficiency due to its autocorrelation method. While the YIN implementation can process up to 38 channels on our test computer in real time, the offline SWIPE implementation can handle up to 7. In contrast, RT-SWIPE achieves a solid balance by supporting up to 20 channels in parallel. This highlights the trade-offbetween computational efficiency and robustness in pitch estimation, with SWIPE-based methods generally offering greater reliability in noisy, real-world scenarios compared to YIN [12]. 3 Experiments Dataset. Our experiments are performed on the ChoraleBricks dataset, a compilation of 193 multi-track recordings featuring wind instruments such as brass and woodwinds [1]. The dataset includes performances of 10 different chorale pieces, with soprano, alto, tenor, and bass parts played by different instruments. In addition to the multi-track recordings, ChoraleBricks includes F0 annotations, interactively generated using Sonic Visualiser (v5.0.1) [4] and the pYIN VAMP plugin (v3) [10]. These annotations were verified through sonification methods [14] and manually corrected as needed. The fundamental frequency (f0)ofreferenceannotationsrangefrom38.19Hz(playedbyatuba)to1250.97 Hz (played by a flute). Experimental Setup. In our experiment, we utilize RT-SWIPE with right-aligned window positioning (=0). We compare its performance against the original SWIPE method as implemented by the Python library libf0 [15], and YIN [5], as implemented in librosa.2We also compare it with CREPE [9], using the official implementation available on GitHub.3The experiments are conducted with a 2We specifically used commit ebd878f,whichincludesrecentupdatesandbugfixes for YIN and PYIN. 3https://github.com/marl/crepe. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 291
P. Meier et al. Fig. 4. RPA with a 25-cent tolerance over frequency on the ChoraleBricks dataset for SWIPE,YIN,CREPE,andRT-SWIPE.(a) Instruments without cross-talk. (b) Instruments with cross-talk: Target instrument signals are mixed with a randomly selected interfering instrument from a different voice, using a signal-to-noise ratio (SNR) of 5 dB. Background bars show the pitch distribution in the ChoraleBricks dataset. sample rate fs= 44.1kHz and a hop size H= 512 samples. All estimators operate within a frequency range from note C1 (approximately 32.70 Hz) to note A6 (1760 Hz), effectively covering the entire range of the ChoraleBricks dataset. SWIPE and RT-SWIPE,thecentresolutionisconfiguredto10cents.For the pitch evaluation, we employ RPA with a tolerance of 25 cents. 3.1 Results In Figure 4a, we present the frequency-dependent evaluation results of our experiment. For each estimated frequency, we check if it matches the reference within the cent tolerance. We sort this binary data (1 = correct; 0 = not correct) by annotation frequency and apply a Hanning window-based convolution to smooth the results and achieve RPA values over frequency. The RPA values for SWIPE are illustrated by the black solid line, those for RT-SWIPE are shown by the blue dotted line, CREPE is represented by the green dash-dotted line, and YIN is depicted with the orange dashed line. The step function, with a grey shaded area, indicates the density of f0reference annotations, helping to understand the frequency distribution in the ChoraleBricks dataset. With this, we identify that most notes are around 300 Hz, whereas notes at 50 Hz or 1 kHz are rarely played. Across the dataset, SWIPE achieves an overall RPA of 0.960, while YIN records an RPA of 0.952, and CREPE achieves 0.961, indicating very similar performance among these algorithms. These results align with findings from previous studies [12]. However, RT-SWIPE shows a lower RPA of 0.931. As illustrated in Figure 4a, the decrease in RPA for RT-SWIPE is frequencydependent; all algorithms perform similarly above 500 Hz. For instance, at 200 Hz, the RPAs for SWIPE,YIN,andCREPE are 0.963, 0.948, and 0.959, respectively, compared to 0.922 for RT-SWIPE.At50Hz,thedifferencesbecomemore pronounced: SWIPE,YIN,andCREPE have RPAs of 0.490, 0.492, and 0.486 respectively, while RT-SWIPE falls to 0.434. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 292
Real-Time SWIPE with Causal Windowing Fig. 5. (a) Reference annotation of a vibrato test signal (black line) with estimates for different delay factors .(b) Error of RT-SWIPE (=0)estimatescomparedtothe reference in cents, with corresponding RPA values. Figure 4b shows the results of the different approaches in a cross-talk scenario. Here, target instruments are mixed with randomly selected interfering instruments from a different voice, using a signal-to-noise ratio (SNR) of 5 dB. As observed in our previous study [12], YIN is most susceptible to interference from cross-talk. CREPE achieves the highest overall accuracy, while SWIPE offers afavorablebalancebetweenperformanceandcomputationalefficiency.Notably, all approaches exhibit increased sensitivity to cross-talk in frequency regions above 200 Hz. The observed deviations of RT-SWIPE arise from a mismatch between realtime estimation and the RPA metric, which assumes perfect temporal alignment between estimates and reference annotations. In real-time settings, however, causal windowing introduces algorithmic delays—causing accurate but slightly delayed estimates to be penalized as incorrect. 3.2 Results with Time Tolerance In the following, we illustrate the algorithmic delays introduced through the causal windowing with a synthetic test signal. In Figure 5a, the reference f0 trajectory of the test signal is shown as a black line. This trajectory features sinusoidal oscillations around a base frequency fbase,resemblingavibratopattern. The modulation has a frequency of fmod =1Hz and an amplitude of Amod =0.2·fbase. The test signal includes four segments, each two seconds long, where the base frequency fbase doubles sequentially from 50 Hz to 100 Hz, 200 Hz, and finally 400 Hz. We sonify this f0trajectory using the libsoni [14] toolbox with the function sonify_f0, utilizing 20 partials with individual amplitudes of 1/20,atasamplingratefsof 44.1 kHz to mimic a harmonic signal. Proc. of the 17th International Symposium on CMMR, London, UK, Nov. 3-7, 2025 293