Full text
RESEARCH ARTICLE Amazonian manatee critical habitat revealed by artificial intelligence-based passive acoustic techniques Florence Erbs 1 , Mike van der Schaar 1 , Miriam Marmontel 2 , Marina Gaona 1,2 , Emiliano Ramalho 2 & Michel Andr´ e 1 1 Laboratory of Applied Bioacoustics (LAB), Universitat Polit` ecnica de Catalunya-BarcelonaTech (UPC), Barcelona, Spain 2 Instituto de Desenvolvimento Sustent´ avel Mamirau´ a, Tef´ e, Brazil Keywords bioacoustics, deep learning, habitat, Sirenians, South America, vocal repertoire Correspondence Michel Andr´ e, Laboratory of Applied Bioacoustics (LAB), Universitat Polit` ecnica de Catalunya-BarcelonaTech (UPC), Barcelona, Spain. Tel: (34) 938 967 227; E-mail: michel. [email protected] Funding Information This research is part of Project Providence (http://www.projectprovidence.org/) funded through the Sense of Silence Foundation by the Gordon and Betty Moore Foundation. Additional funding was provided by the Rolex Institute and the Fondation Prince Albert II de Monaco. Editor: Dr. Vincent Lecours Associate Editor: Dr. David Curnick Received: 26 November 2023; Revised: 30 April 2024; Accepted: 4 July 2024 doi: 10.1002/rse2.418 Abstract For many species at risk, monitoring challenges related to low visual detectability and elusive behavior limit the use of traditional visual surveys to collect critical information, hindering the development of sound conservation strategies. Passive acoustics can cost-effectively acquire terrestrial and underwater long-term data. However, to extract valuable information from large datasets, automatic methods need to be developed, tested and applied. Combining passive acoustics with deep learning models, we developed a method to monitor the secretive Amazonian manatee over two consecutive flooded seasons in the Brazilian Amazon floodplains. Subsequently, we investigated the vocal behavior parameters based on vocalization frequencies and temporal characteristics in the context of habitat use. A Convolutional Neural Network model successfully detected Amazonian manatee vocalizations with a 0.98 average precision on training data. Similar classification performance in terms of precision (range: 0.83–1.00) and recall (range: 0.97–1.00) was achieved for each year. Using this model, we evaluated manatee acoustic presence over a total of 226 days comprising recording periods in 2021 and 2022. Manatee vocalizations were consistently detected during both years, reaching 94% daily temporal occurrence in 2021, and up to 11 h a day with detections during peak presence. Manatee calls were characterized by a high emphasized frequency and high repetition rate, being mostly produced in rapid sequences. This vocal behavior strongly indicates an exchange between females and their calves. Combining passive acoustic monitoring with deep learning models, and extending temporal monitoring and increasing species detectability, we demonstrated that the approach can be used to identify manatee core habitats according to seasonality. The combined method represents a reliable, cost-effective, scalable ecological monitoring technique that can be integrated into long-term, standardized survey protocols of aquatic species. It can considerably benefit the monitoring of inaccessible regions, such as the Amazonian freshwater systems, which are facing immediate threats from increased hydropower construction. Introduction Freshwater environments are amongst the most endangered ecosystems (Dudgeon et al., 2006). In particular, freshwater megafauna are disappearing twice as fast as terrestrial and marine megafauna (He et al., 2019; Jenkins, 2003). In the Amazon basin, human activities, including dam construction and mining, are inducing rapid and large-scale degradation of freshwater systems (Castello & Macedo, 2016) with direct and profound consequences for the Amazonian manatee (Trichechus inunguis) populations (Arraut et al., 2017; Arraut & Marmontel, 2016). With their large body, their late sexual maturity, and their low uniparous reproductive rate (one calf every 2 or 3 years, Marmontel, 2009; Rodrigues et al., 2008) sirenians are especially at risk of extinction (Deutsch et al., 2008; Keith Diagne, 2015; Marmontel et al., 2016). Historically hunted throughout their distribution range (Domning, ª2024 The Author(s). Remote Sensing in Ecology and Conservation published by John Wiley & Sons Ltd on behalf of Zoological Society of London. This is an open access article under the terms of the Creative Commons Attribution-NonCommercial-NoDerivs License, which permits use and distribution in any medium, provided the original work is properly cited, the use is non-commercial and no modifications or adaptations are made. 1
1982), Amazonian manatee populations have declined significantly and the species is currently classified as vulnerable (Marmontel et al., 2016). Although manatee hunting was banned in Brazil in 1967, populations have not recovered and hunting for subsistence or illegal trade still occurs (Calvimontes & Marmontel, 2022). The difficulty of detecting manatees through visual methods complicates their conservation status determination and the implementation of adequate conservation strategies. Visual monitoring of manatees is possible under certain conditions, such as when they are resting near the surface in clear shallow waters (e.g., Florida manatees aggregating in springs during winter). Usually, manatees spend most of their time deeper underwater, surfacing for very brief moments with only the tip of the snout usually visible. Due to this inconspicuous behavior, visual observations are difficult, especially in turbid white waters of the Amazon with extremely poor visibility (less than 1 m, Henderson & Crampton, 1997). Amazonian manatees favor mats of floating meadows (Guterres-Pazin, 2014), and are often concealed by the aquatic cover of the aquatic plants, even when surfacing to breathe. Visual surveys (boat, aerial or shore-based) typically achieve low detection rates (Factheu et al., 2023). Underwater imaging techniques such as a side scan sonar have been used recently with some success (Gonzalez-Socoloske & Olivera-G´omez, 2023) and can be particularly useful in murky or turbid waters with limited visibility. A limitation of side scan sonar methods is that animals may actively avoid the research boat, moving outside of the sonar beam detection range. Additionally, this method has a shorter detection range in sediment-filled water with aquatic vegetation. This can introduce noise in the sonar data which makes it difficult to identify the shape of a manatee. All the extant sirenian species, including the three manatee species and the dugong species, produce vocalizations for communication purposes (O’Shea et al., 2022), making them well-suited for studying through passive acoustic monitoring (PAM). A critical advantage of fixed PAM is the ability to continuously survey manatee presence at a location over months or years without interference (i.e., no introduction of additional disturbance and noise associated with boats for visual or sonar surveys). This monitoring technique can reveal valuable information about a species’ temporal occupancy of a habitat, including seasonal variations, at the cost of a lower spatial resolution. To tap into PAM research potential, bioacousticians collect increasing amounts of data that cannot be efficiently processed manually. Technological advances in data acquisition have inaugurated a “big data” era for bioacoustics, demanding automated methods to deal with the analysis of large datasets. In recent years, deep learning has been increasingly used for classification of bioacoustics signals and has been successfully applied to automatic detection of a wide range of soniferous taxa (Stowell, 2022), becoming a major asset for data processing and analysis. Previous work has demonstrated the efficiency of these methods in detecting manatee vocalizations (Factheu et al., 2023). The Amazonian manatee repertoire characterization comes mainly from studies conducted in captivity (Evans & Herald, 1970; Landrau-Giovannetti et al., 2014; Sonoda & Takemura, 1973; Sousa-Lima et al., 2002) with only one study from wild animals describing similar signals (based on a limited sample size; Sousa-Lima et al., 2013). The majority of vocalizations described are short, tonal harmonic complexes that contain one to four notes. Atonal vocalizations have also been noted. In contrast, the acoustic repertoire of West Indian manatees is fairly well described (e.g., Brady et al., 2020; Ramos et al., 2020, Rivera Chavarria et al., 2015; Sousa-Lima et al., 2008) and African manatee vocalizations were recently characterized (Rycyk et al., 2021). The literature demonstrates that manatee vocal repertoire is similar across species (O’Shea et al., 2022; Rycyk et al., 2021). Variability in vocalization frequency and duration have been linked to behavior and group composition (e.g., Brady et al., 2021; Miksis-Olds & Tyack, 2009; O’Shea & Poch´e Jr, 2006). Manatees produce calls across different behavioral contexts (Brady et al., 2021; O’Shea & Poch´e Jr, 2006), making this species a good candidate for PAM studies. Additionally, calling rates may vary widely depending on behavioral states, potentially providing new insights on manatee behavior in the monitored habitat. The highest calling rates are produced during cow-calf interactions and several studies suggested that one of the main functions of acoustic communication in manatees is maintaining contact between a mother and her calf (Bengtson & Fitzgerald, 1985; Bullock et al., 1980; Hartman, 1979; O’Shea & Poch´e Jr, 2006; Reynolds, 1981). Here, we present the results of our ongoing efforts to monitor Amazonian manatees in the Mamirau´aSustainable Development Reserve. The first objective was characterizing manatee presence over two consecutive years at Mamirau´a Lake, a historically important floodplain lake for the population. The second objective was describing the vocal repertoire of the wild Amazonian manatee and analyzing vocal behavior to infer their daily activity patterns. Materials and Methods Study area The study area is located in the Mamirau´a Sustainable Development Reserve (MSDR), in the state of Amazonas, Brazil (Fig. 1). Situated at the confluence of the Solim˜ oes (upper Amazon River) and the Japur´a rivers, with circa 11 000 km 2 2ª2024 The Author(s). Remote Sensing in Ecology and Conservation published by John Wiley & Sons Ltd on behalf of Zoological Society of London. AI-based acoustic monitoring of Amazonian manatees F. Erbs et al. 20563485, 0, Downloaded from https://zslpublications.onlinelibrary.wiley.com/doi/10.1002/rse2.418 by Readcube (Labtiva Inc.), Wiley Online Library on [04/12/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
of floodplains (whitewater river floodplain or v´arzea) habitats, the MSDR is the largest Brazilian protected area dedicated to the conservation of flooded rainforests. Whitewater river floodplains, seasonally inundated by the sedimentand nutrient-laden Amazon River waters and tributaries, are highly productive ecosystems. During high water levels, (semi-)aquatic plants are abundant and constitute the main food source for the Amazonian manatee (GuterresPazin, 2014). Breeding and calving periods seem synchronized with the seasonal availability of food (Best, 1982). Data collection An autonomous SM4 recorder (Wildlife Acoustics, USA) with an HTI-96-Min hydrophone (High Tech Inc., USA, sensitivity: 165 dB re 1 V/uPa) was installed in the Mamirau´a Lake (2°59055.500S, 64°54018.14400W), fixed above the water line to a tree on the lake edge. The hydrophone was placed three meters deep at mid-flood in January with a maximum depth of ca. 12 m at the highest water levels in June. Acoustic recordings were collected on a duty cycle of 5 min on/55 min off, while sampling at 96 kHz. In 2021, 53 days of data were collected from the end of May to mid-July. In 2022, data were collected from mid-January to the end of May and from the beginning of July to mid-August providing 173 recorded days. CNN classification procedure The CNN architecture was initially developed under Project Providence (Zaugg et al., 2023) and was previously applied to underwater data for the classification of Figure 1. Map of the Mamirau´ a Sustainable Development Reserve showing the recording location at Mamirau´ a Lake (red star). The habitat types are modified from Ferreira-Ferreira et al., 2015. ª2024 The Author(s). Remote Sensing in Ecology and Conservation published by John Wiley & Sons Ltd on behalf of Zoological Society of London. 3 F. Erbs et al. AI-based acoustic monitoring of Amazonian manatees 20563485, 0, Downloaded from https://zslpublications.onlinelibrary.wiley.com/doi/10.1002/rse2.418 by Readcube (Labtiva Inc.), Wiley Online Library on [04/12/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Amazon River dolphins (Erbs et al., 2023). The framework was originally created using Tensorflow 1.14 and Python 3.6.8 for backwards compatibility, although for model execution, more recent versions are generally used. The classification procedure is performed in five steps (c.f. Erbs et al., 2023): labeling, augmentation, training, validation and final test evaluation. In the following, label dataset refers to all the data that contains labels used for building the classification model. The training dataset is the portion of the label dataset used to train the CNN model. The model learns patterns and relationships from this data. The validation dataset is another portion of the label dataset used to assess the model’s performance after training, that is, provide performance metrics. The test dataset refers to the data used for a final evaluation of the model’s performance. In our case, it is the totality of the data recorded in 2021 and 2022, for which we separately assess performance per recording period. Four recording periods were defined based on flooding stages. Labeling The datasets from four sources from the Mamirau´a and Aman˜ a Sustainable Development Reserves (label dataset) were annotated to create a database of manatee call labels. To ensure proper soundscape representation in the label dataset, two datasets that did not contain manatee vocalizations, but did contain other sound sources such as dolphin whistles and echolocation clicks, rain, and boat passages, were added. Adding these additional soundscapes resulted in 12 classes in total which allowed to improve the performance of the single manatee class of interest. An overview of the label dataset is presented in Table 1. A custom Python GUI presented 5-s spectrograms for sound labeling. The interface allowed for defining the frequency and time boundaries of the target sounds by drawing boxes on the spectrogram. These labels were then automatically imported into a dedicated database. A total of 5063 segments were labeled, including 726 segments with manatee vocalizations. The label dataset was randomly split into training and testing datasets, with 50% of the labels devoted to training the classifier and 50% to testing. Data augmentation As part of the training procedure, the training dataset was artificially expanded using a data augmentation procedure that included data duplication and transformation through frequency and time shifts, contrast adjustments, and time-warping. All new samples were augmented, the augmentation functions were always applied to the training data at each round; that is each data sample underwent each augmentation. All parameters related to data augmentation were uniformly selected from a defined range. These ranges are generally kept identical once defined for a specific model (specific sound types that need to be classified). The ranges are provided in Table A1. The training set, after data augmentation, comprised 6567 segments, of which 1014 contained manatee vocalizations. Model training A classifier model was trained and validated using a convolutional neural network (CNN). Details of the CNN architecture can be found in Erbs et al. (2023). Following the same procedure, the log-Mel spectrogram was used as Table 1. Overview of the labeling dataset used for training and testing the CNN model. All autonomous recordings were collected at 96 kHz. The boat-based recordings were collected at 576 kHz and resampled to 96 kHz for the analysis. Dataset location Year Habitat Recording equipment Recording method Duration (h) No. of segments with manatee calls No. of segments without manatee calls ASDR Aman˜ a Lake 2018 Entrance of the lake SM4 recorder and HTI-96-Min hydrophone Autonomous 0.67 103 471 MSDR Mamirau´ a Lake 2021 Floodplain lake SM4 recorder and HTI-96-Min hydrophone Autonomous 2.42 623 1740 MSDR 2019 2020 Different floodplain habitats (bay, channel, floodplain lakes, flooded forest) SoundTrap 300 HF recorder Boat-based 2.40 0 1728 MSDR 2019 Flooded forest SM4 recorder and HTI-96-Min hydrophone Autonomous 1.56 0 1124 4ª2024 The Author(s). Remote Sensing in Ecology and Conservation published by John Wiley & Sons Ltd on behalf of Zoological Society of London. AI-based acoustic monitoring of Amazonian manatees F. Erbs et al. 20563485, 0, Downloaded from https://zslpublications.onlinelibrary.wiley.com/doi/10.1002/rse2.418 by Readcube (Labtiva Inc.), Wiley Online Library on [04/12/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
input into the classifier here. The final model was selected with random weight initialization, 300 epochs, 32 batch size, 0.001 learning rate, and 0.005 decay rate. Data were reshuffled for each epoch (see Table A2 for the model dimensions and Figure A1 for the network layout). Model performance on the validation dataset The performance of the final model was evaluated by computing the precision and the recall. The precision (or positive predictive value, PPV) represents the ratio of correct positive predictions to the total positive predictions made by the model. The recall (or true positive rate, TPR) represents the ratio of correct positive predictions out of all actual positive instances in the dataset. These two metrics form the precision-recall curve, with each point of the curve associated to a decision threshold (Fig. 2). The area under the precision-recall curve, also called the average precision, was calculated for the manatee vocalization class. A decision threshold of 0.28, corresponding to a recall of 0.9, was selected from this curve as a starting point for fine tuning the thresholds on the test dataset. Precision PPVðÞ= True positive True positive þFalse positiveðÞ Recall TPRðÞ= True Positive True Positive þFalse negative Average precision APðÞ¼∑t¼T1 t¼0TPR tðÞTPR tþ1ðÞ½ PPV tðÞ with T the number of thresholds and TPR(T) =0, PPV(T) =1. The model performance achieved a 0.98 average precision on the validation dataset. This value corresponds to the area under the precision-recall curve for this model (Fig. 2). Model performance on the test dataset The classification model obtained at the training stage was run over each 5-min file recorded at the Mamirau´ a Lake in 2021 and 2022. A random selection of 100 segments per recording period was then evaluated by a human expert (FE) to assess whether the segment prediction was correct at the decision threshold selected at the training stage. This procedure was repeated to fine-tune the decision threshold in order to achieve similar recall and precision values for each recording period. The decision thresholds and associated precision and recall values are listed in Table A3. The model could confidently identify manatee vocalizations with a recall ranging from 0.97 to 1 and a precision between 0.83 and 1. The lower precision obtained in January–March 2022, when water levels were still low and the hydrophone closer to the surface, was related to the presence of airborne sounds (absent from the training dataset). Manatee occurrence based on CNN classification of vocalizations If the per-segment output of the classifier scored above the decision threshold, the segment was assigned positive for manatee presence. Segments assigned as positive were used to compute the daily number of recording hours with manatee detections. Manatee diel pattern of acoustic activity The number of positive segments (i.e., per-segment output value above the decision threshold) was counted by day and by hour of the day. Time periods between 06:00 and 17:59 (AMT, Amazon time) were assigned to the “daytime” categorical value, and periods between 18:00 and 05:59 were assigned to “night time” based on sunrise and sunset for the recording periods. To test if there was more activity during the day than the night we tested the proportion of daytime activity to the proportion of nighttime activity by performing a proportions z-test for each separate year. Characterization of the Amazonian manatee vocal repertoire To avoid sampling the same individuals that might be present in the area over hours during the same day, only Figure 2. Precision-recall curve of the Amazonian manatee CNN classifier. ª2024 The Author(s). Remote Sensing in Ecology and Conservation published by John Wiley & Sons Ltd on behalf of Zoological Society of London. 5 F. Erbs et al. AI-based acoustic monitoring of Amazonian manatees 20563485, 0, Downloaded from https://zslpublications.onlinelibrary.wiley.com/doi/10.1002/rse2.418 by Readcube (Labtiva Inc.), Wiley Online Library on [04/12/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
one file per day was randomly selected from all the files with classification values above the decision threshold. A total of 85 files were selected, with 48 files from 2021 and 37 from 2022. Each file was manually inspected in Adobe Audition v.5 (Adobe Systems Inc., USA) to confirm the presence of manatee calls. Up to five calls per file were selected with an SNR over 10 dB and measured in Raven Pro 1.5 (Bioacoustics Research Program, 2011) using a 2048-point FFT (Fast Fourier Transform), Hann window with 50% overlap, corresponding to a spectrogram frequency and time resolution of 46.9 Hz and 21.3 ms, respectively. Variables were selected based on previous literature on the manatee acoustic repertoire (Brady et al., 2020; Rycyk et al., 2021; Sousa-Lima et al., 2008). Manatee vocalizations typically present several harmonics. For each manatee call, the following parameters were measured on the emphasized frequency: start frequency, end frequency, start time, end time, duration, peak frequency and number of the emphasized frequency harmonic. The emphasized peak frequency is defined as the peak frequency of the harmonic displaying the highest acoustic energy; the fundamental peak frequency as the peak frequency of the first harmonic. Most of the time, the fundamental frequency was not measurable on the spectrogram due to a lower SNR and/or masking by other biological sounds (amphibians and fish calls). The fundamental peak frequency was computed by dividing the value of the emphasized peak frequency by the number of the corresponding harmonic (following Sousa-Lima et al., 2002). Call frequency contour categories (unmodulated, hill-shaped, multiple inflections and downsweep) were assigned based on similarity with call types described in literature (Brady et al., 2020,2022; O’Shea & Poch´ e Jr, 2006). The number of notes (as described in Sousa-Lima et al., 2002) was counted. To investigate potential changes in vocal behavior between years, a t-test was performed on emphasized frequency and duration. Comparison of Manatee vocalization parameters and acoustic behavior between years Using 64 days of data (randomly selecting 34 files from 2021 and 30 from 2022 to obtain a balanced data set), we compared the emphasized peak frequency and duration between the 2 years (using a t-test). The inter-call interval (ICI) was compared (t-test) between the years based on 10 randomly selected files for each year, on call sequences with similar frequency contour shape. The same files that were used for the ICI were used to compute the vocalization rate. Calls were manually counted and this count was divided by the duration of the file to obtain a value in no. the number of calls per minute. The vocalization rate was computed over the whole file length, including the periods of silence between repetitive call sequences. Manatee individual count per file The 85 files selected for vocalization measurement were manually inspected in Adobe Audition v.5. The number of call contour shapes was counted based on spectral characteristics (especially frequency modulation). As calls of individuals present less variation than inter individual ones (Sousa-Lima et al., 2002), and because most calls occurred in sequences (Figure A2), the number of contour types was used as a proxy for the number of individuals present in one file. Results Manatee occurrence in 2021 and 2022 Manatee vocalizations were detected during the two consecutive years at a time period corresponding to the highest water levels at Mamirau´a Lake (Fig. 3). In 2021, data were collected exclusively during high waters, from the end of May to mid-July. On the 53 recording days, manatees were acoustically present during 50 consecutive days, with a daily presence ranging from one to 11 h. The peak presence occurred during the first half of June. In 2022, the data collection was extended to include the start of the flooded period (January to March) and the beginning of the receding water period (August). Due to technical issues with the recording equipment, data were not collected between mid-May and the beginning of July. Manatees were detected sporadically during the first trimester, with 5 days of presence. Detections increased from April on, and during May manatees were detected on 17 of the 23 recording days, with daily presence ranging from 1 to 5 h, similar to the previous year for the same time period. From mid-July to August, when water levels started to lower, there were only 2 days with one positive hour. Manatee diel acoustic pattern In 2021, manatees appeared to be present and vocally active between 19:00 and 06:00 (all times are local time and in 24 h clock format), that is, during nighttime (Fig. 4A). Manatees were detected in 88% of the nighttime and 44% of the daytime periods. The difference in the proportion of daytime activity to the proportion of nighttime activity was statistically significant (z-value: 4.78; P-value \0.001). The peak detection hour in 2021 occurred at 02:00 and represented 12.8% of the nighttime detection and 11.8% of the total detections. This pattern was not visible in 2022 where manatee detections were 6ª2024 The Author(s). Remote Sensing in Ecology and Conservation published by John Wiley & Sons Ltd on behalf of Zoological Society of London. AI-based acoustic monitoring of Amazonian manatees F. Erbs et al. 20563485, 0, Downloaded from https://zslpublications.onlinelibrary.wiley.com/doi/10.1002/rse2.418 by Readcube (Labtiva Inc.), Wiley Online Library on [04/12/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
similar between daytime and nighttime (40.3 and 38.4%, respectively; z-value: 0.2; P-value =0.84). The highest number of detections occurred during the day, at 06:00 (44.9% of the daytime detections, Fig. 4B). Characterization of Amazonian manatee vocal repertoire From the 85 files selected, a total of 422 calls were selected for measurements. All calls were tonal (without atonal components), and four main call categories were observed based on their frequency contours: unmodulated, hill-shaped, multiple inflections and downsweep. Figure 5provides an example spectrogram of each call category. Nonlinear elements such as subharmonics were observed (Fig. 5A). Multiple inflections of the frequency contour were present in 77.8% of the calls, 11.8% presented a hill-shaped contour, 9.4% had a pronounced downsweep component and less than 1% had a flat frequency contour. The majority of the calls were composed of one (69.0%) or two (27.2%) notes, with only 3.8% containing three notes. In 98% of the calls, the emphasized frequency was centered on the second or higher harmonic. The mean emphasized peak frequency was 10.2 (2.7 SD) kHz. The mean start and end frequencies (measured from the emphasized frequency) were relatively similar (9.98 vs. 9.63 kHz). The fundamental peak frequency was 3.6 (0.4 SD) kHz. The mean duration of a call was 0.294 (0.072 SD) s (Table 2). Comparison of vocalization parameters and acoustic behavior between years Frequency (emphasized frequency) and temporal (duration, inter-call interval) parameters were compared between 2021 and 2022 (Table A4). The mean emphasized frequency for 2021 was higher in 2021 compared to 2022 (t-test, t(327) =2.52, P=0.01) at α=0.05. Duration values (t-test, t(327) =0.69, P=0.49) were not different between years (P[0.05). Inter-call intervals (ICIs) were computed over the whole file (5 min), which often included calls emitted in sequences and silent periods (Figure A2). On average, calls were emitted 2.91 s apart in 2021 and 3.42 s apart in 2022 (t-test, t(1273) =1.78, P=0.08). The maximum ICI values were above 1 min for both years, indicative of calls not emitted in sequences. The average calling rate was 13.9 (2021) and 12.5 (2022) calls per minute, with a maximum of 27 calls per minute (Table A4). Counter calling Files manually analyzed for vocalization measurements were also reviewed for identifying the number of vocalization types in a file. Only 3 out of 64 files presented two different frequency contour shapes (Fig. 6). Two files were recorded during 2021, on June 5 and June 13. The third file was recorded on April 17, 2022. Figure 3. Amazonian Manatee acoustic presence at Mamirau´ a Lake (Brazil), based on CNN classification. The light gray areas indicate an absence of recordings. The dark gray area indicates time periods not monitored due to low water levels. ª2024 The Author(s). Remote Sensing in Ecology and Conservation published by John Wiley & Sons Ltd on behalf of Zoological Society of London. 7 F. Erbs et al. AI-based acoustic monitoring of Amazonian manatees 20563485, 0, Downloaded from https://zslpublications.onlinelibrary.wiley.com/doi/10.1002/rse2.418 by Readcube (Labtiva Inc.), Wiley Online Library on [04/12/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Figure 4. (A) Radial heatmaps of manatee daily acoustic activity for the years 2021 and 2022. The color scale is based on the presence count, that is, the number of segments with detections grouped by hour (see Methods section 2.5). Values on the y-axis are dates in format MM-DD. Gray areas represent absence of data. (B) Histograms of the manatee acoustic detections representing the cumulative number of segments with detections grouped by hour of the day for the years 2021 and 2022. Figure 5. Spectrogram examples of the four categories of Amazonian manatee calls, based on frequency contour. (A) Unmodulated. (B) Hillshaped. (C) Multiple inflections. (D) Downsweep. Note the presence of subharmonics on type A. 8ª2024 The Author(s). Remote Sensing in Ecology and Conservation published by John Wiley & Sons Ltd on behalf of Zoological Society of London. AI-based acoustic monitoring of Amazonian manatees F. Erbs et al. 20563485, 0, Downloaded from https://zslpublications.onlinelibrary.wiley.com/doi/10.1002/rse2.418 by Readcube (Labtiva Inc.), Wiley Online Library on [04/12/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License
Discussion We combined passive acoustic monitoring and deep learning algorithms to monitor the seasonal presence of Amazonian manatees at a floodplain lake where manatees were historically abundant. Our analysis was conducted over two flooded seasons in 2021 and 2022 and show that manatees were present at Mamirau´a Lake during most of the days at peak presence, sometimes up to 11 h a day, demonstrating that this location is likely an important habitat for the local manatee population. Classification performance The CNN classifier developed here presented very high performance for the automatic identification of manatee vocalizations with a 0.98 average precision on training data. On the full Mamirau´a lake dataset, the model Table 2. Summary of the descriptive statistics for the variables measured on wild Amazon manatee vocalizations recorded at Mamirau´ a Lake, Amazonas, Brazil, in 2021 and 2022. Variable Fundamental frequency (Hz) Emphasized frequency (Hz) Start frequency (Hz) End frequency (Hz) Duration (s) N265 423 423 423 423 Mean 3676 10 197 9979 9629 0.294 Standard deviation 442 2728 2631 2611 0.072 Min 2625 2766 2766 3000 0.107 Max 4617 18 000 17 016 17 203 0.499 Median 3656 9891 9656 9281 0.286 Figure 6. Spectrograms of Amazonian manatee “duets” recorded at Mamirau´ a Lake (Brazil) on June 5, 2021 (top), June 13, 2021 (middle), and April 17, 2022 (bottom). The blue and purple marks on each spectrogram indicate the temporal location of two different types of signals (note that the types of signals are not the same between spectrograms). ª2024 The Author(s). Remote Sensing in Ecology and Conservation published by John Wiley & Sons Ltd on behalf of Zoological Society of London. 9 F. Erbs et al. AI-based acoustic monitoring of Amazonian manatees 20563485, 0, Downloaded from https://zslpublications.onlinelibrary.wiley.com/doi/10.1002/rse2.418 by Readcube (Labtiva Inc.), Wiley Online Library on [04/12/2024]. See the Terms and Conditions (https://onlinelibrary.wiley.com/terms-and-conditions) on Wiley Online Library for rules of use; OA articles are governed by the applicable Creative Commons License