Every forest, ocean, and grassland on Earth hums with sound, and for the past two decades scientists have been teaching machines to eavesdrop. A new systematic review published in Applied Intelligence by Eduardo Mora-González, Antonio Canepa-Oneto, José Antonio Barbero-Aparicio, and Álvar Arnaiz-González of the Universidad de Burgos, together with colleagues, has taken stock of that effort, synthesizing 108 studies published between 2002 and 2023 into the most comprehensive map yet of machine learning in bioacoustics. The picture that emerges is one of explosive growth, dramatic technological change, and a field that is finally confronting its own growing pains around reproducibility and rigor.
The review’s starting point is striking: one of the earliest machine learning applications in the field was an automatic frog call monitoring system presented back in 2002. From those modest beginnings, computational bioacoustics has expanded into a discipline that underpins large-scale biodiversity monitoring on every continent. The authors document a marked inflection point around 2016, after which research activity surged. That timing is no coincidence. It follows the deep learning revolution that transformed computer vision after the 2012 ImageNet breakthrough, when convolutional neural networks demonstrated that learned features could outperform hand-crafted ones. Bioacoustics, which treats animal sounds as images by converting recordings into spectrograms, was perfectly positioned to inherit those advances.
Taxonomically, the field has broadened dramatically. Early studies were overwhelmingly bird-focused, reflecting both the richness of avian vocalizations and the availability of recordings from citizen science repositories such as Xeno-Canto. The review shows a steady expansion toward amphibians, insects, marine mammals, fish, bats, rodents, and even livestock, culminating in multispecies soundscape analyses that attempt to characterize entire acoustic communities rather than individual species. Landmark applications along the way include passive acoustic monitoring of northern spotted owls across varied forest conditions, deep learning detection of humpback whale song in long-term datasets, classification of sperm whale bioacoustics, and automated recognition of insect pests such as the red palm weevil and rice weevil. The technology has even been used to distinguish individual chimpanzee infant distress calls and to track song variation in endangered birds like the white-bellied heron in Bhutan.
Methodologically, the review builds a unified comparative framework that standardizes studies across five dimensions: classification objectives, taxonomic groups, preprocessing strategies, learning paradigms, and algorithm families. This kind of harmonization is rare and valuable, because bioacoustics papers have historically been scattered across ecology, computer science, and engineering journals, each with its own vocabulary and conventions. By mapping every study onto a common structure, the authors can identify genuine trends rather than anecdotes. What they find is that supervised learning remains the dominant paradigm, which is unsurprising given that labeled recordings are the fuel of most classifiers, but that unsupervised and semi-supervised approaches have carved out important niches, particularly for exploratory soundscape analysis and for species where labeled data are scarce.
On the algorithmic front, deep learning architectures now reign. Convolutional neural networks, which excel at extracting spatial patterns from spectrograms, and recurrent neural networks, which capture temporal structure in call sequences, have become the most widely adopted approaches for bioacoustic classification. The review traces this shift through the literature, from early support vector machine and random forest classifiers working on hand-engineered features such as mel-frequency cepstral coefficients, to end-to-end deep models, to transfer learning pipelines built on pretrained networks like BirdNET, a deep learning solution for avian diversity monitoring that has become a workhorse for ecologists. Yet the authors are careful not to declare traditional methods obsolete. Support vector machines and random forests continue to deliver competitive performance in many applications, especially when embedded in well-designed analytical workflows, a finding that should give pause to anyone assuming bigger models are always better.
Perhaps the most technically interesting result concerns preprocessing. The comparative analysis indicates that strategies including data augmentation, feature reduction, noise removal, and window segmentation are consistently associated with high-performing classification pipelines. In other words, what happens to the audio before it reaches the model matters enormously. Field recordings are notoriously messy: wind, rain, insects, traffic, and overlapping animal calls all contaminate the signal. Techniques such as spectral filtering, sound separation, careful segmentation of long recordings into analysis windows, and augmentation methods that artificially expand small training sets can make the difference between a classifier that works in the lab and one that works in the rainforest. The review also highlights studies showing that even the parameters used to generate spectrograms measurably affect the accuracy of convolutional neural network classifiers, underscoring how sensitive these systems are to upstream choices.
The authors went beyond simple tallies, applying mixed-effects statistical analyses to ask which factors actually predict reported classification performance. Their conclusion is subtle and important: performance is more strongly associated with specific algorithmic implementations than with broader algorithm families, and significant interactions between methodological components mean that classification outcomes cannot be explained by any single factor considered in isolation. A convolutional neural network is not intrinsically superior to a random forest; rather, success depends on the coherent integration of preprocessing, feature representation, and learning algorithm within a complete pipeline. This holistic view challenges the common practice of comparing algorithms in isolation and suggests that much of the reported variation in the literature reflects pipeline design rather than fundamental differences between model types.
That insight leads directly to the review’s most sobering finding. Despite a decade of methodological progress, considerable heterogeneity persists in datasets, preprocessing procedures, evaluation protocols, and reported performance metrics across the field. Studies differ in how they split data, which metrics they report, how they handle class imbalance, and even how they define a detection. The consequence is that results from different papers often cannot be directly compared, and reproducibility suffers. This mirrors a broader reproducibility crisis documented across machine learning research, and it has practical stakes: conservation managers deciding whether to trust an automated recognizer for monitoring an endangered species need to know how performance estimates were produced and whether they will generalize to new sites and seasons.
To address these problems, the authors propose a set of good-practice recommendations centered on standardized dataset documentation, transparent methodological reporting, consistent evaluation frameworks, and clearly defined research objectives. The spirit of these recommendations is that a bioacoustics study should be legible to outsiders: readers should be able to tell exactly what data were used, how they were processed, how the model was trained and validated, and under what conditions the reported accuracy can be expected to hold. The authors argue that adopting such standards will improve reproducibility, facilitate cross-study comparisons, and strengthen computational bioacoustics as a robust discipline capable of supporting large-scale biodiversity monitoring and conservation.
The timing of this synthesis could hardly be better. Passive acoustic monitoring is rapidly becoming one of the most scalable tools in ecology, with low-cost autonomous recorders now deployed for months at a time in environments from Arctic tundra to tropical rainforest and deep ocean. As foundation models for bioacoustics emerge and recording archives swell into petabytes, the field stands at an inflection point where the lessons of the past twenty years will determine whether the next twenty deliver on the promise of a planet continuously listened to by machines. This review makes clear that the algorithms are largely ready; the challenge now is to make the science around them as rigorous as the technology itself, so that every chirp, click, and song captured by a microphone can be translated into reliable knowledge about the living world.
Subject of Research: Machine learning methods for automated classification of animal vocalizations and soundscapes in computational bioacoustics
Article Title: Two decades of machine learning in bioacoustics: a systematic review
Article References: Mora-González, E., Canepa-Oneto, A., Barbero-Aparicio, J. A., & Arnaiz-González, Á. (2026). Two decades of machine learning in bioacoustics: a systematic review. Applied Intelligence, 56(15), Article 435. https://doi.org/10.1007/s10489-026-07454-0
Image Credits: AI Generated
DOI: 10.1007/s10489-026-07454-0
Keywords: bioacoustics, machine learning, deep learning, convolutional neural networks, passive acoustic monitoring, biodiversity monitoring, systematic review, preprocessing, reproducibility, conservation, soundscape ecology, classification
Tags: bioacousticsbiodiversity monitoringclassificationConservationconvolutional neural networksdeep learningMachine Learningpassive acoustic monitoringpreprocessingReproducibilitysoundscape ecologysystematic review





