Skip to main content
Advertisement
  • Loading metrics

Multi-class, unsupervised detection and classification of biological and anthropogenic sounds in coral reefs

  • Daniel Duane ,

    Roles Conceptualization, Data curation, Formal analysis, Funding acquisition, Investigation, Methodology, Project administration, Resources, Software, Supervision, Validation, Visualization, Writing – original draft, Writing – review & editing

    daniel.m.duane.civ@us.navy.mil

    Affiliation Naval Undersea Warfare Center, Newport, Rhode Island, United States of America

  • Matthew T. Duggan,

    Roles Data curation, Investigation, Validation, Writing – review & editing

    Affiliations Cornell K. Lisa Yang Center for Conservation Bioacoustics, Cornell University, Ithaca, New York, United States of America, Cornell Lab of Ornithology, Cornell University, Ithaca, New York, United States of America

  • Erika Berlik,

    Roles Data curation, Investigation, Validation, Writing – review & editing

    Affiliation FishEye Collaborative, Arlington, Virginia, United States of America

  • Marc S. Dantzker,

    Roles Data curation, Investigation, Project administration, Resources, Visualization, Writing – review & editing

    Affiliation FishEye Collaborative, Arlington, Virginia, United States of America

  • Aaron N. Rice,

    Roles Data curation, Investigation, Project administration, Resources, Writing – original draft, Writing – review & editing

    Affiliations Cornell K. Lisa Yang Center for Conservation Bioacoustics, Cornell University, Ithaca, New York, United States of America, Cornell Lab of Ornithology, Cornell University, Ithaca, New York, United States of America

  • Lauren A. Freeman

    Roles Conceptualization, Data curation, Funding acquisition, Investigation, Project administration, Resources, Supervision, Writing – review & editing

    Affiliation Naval Undersea Warfare Center, Newport, Rhode Island, United States of America

Abstract

Analyzing the complex and diverse soundscapes of ecosystems such as coral reefs remains a challenge for understanding environmental dynamics and processes. While machine learning techniques can significantly improve detection and classification capabilities, applications of traditional supervised learning to underwater acoustics are limited by the size and class-coverage of labeled datasets. Unsupervised machine learning offers the potential to detect and classify sounds without the guidance of human labels, including signals that were unknown to the human analyst. However, the majority of previously developed unsupervised approaches characterize reef soundscapes from correlative metrics without identifying specific sounds, and the few that detect individual signals have been trained on limited data (<10 days), which constrains the potential to generalize across datasets and geographical localities. Here, a convolutional autoencoder was built and trained on year-long acoustic datasets from four Hawaiian coral reefs, and latent embeddings were clustered using Gaussian mixture modeling. A total of 29 classes were automatically generated, and a manual review of samples in each class determined that nine of the classes corresponded to distinct biological and anthropogenic sounds. The classes were identified to be two call types from the damselfish, Dascyllus albisella, parrotfish feeding sounds, holocentrid calls, an unidentified fish sound, three humpback whale song units, and ship noise. The classifier was found to be robust against an independently-collected test dataset with D. albisella calls (AUC = 0.9) with no extra training on the labels. Diel, lunar, and seasonal trends were observed for all nine classes, including previously-unidentified responses of the holocentrid and unknown fish groups to lunar illumination. This work demonstrates the capability of unsupervised algorithms to cluster acoustic signals into identifiable biological and anthropogenic categories in order to examine and characterize ecological trends.

Author’s summary

Detection and classification of biological sounds is essential for monitoring the restoration progress of coral reefs, yet most artificial intelligence (AI) methods require large, hand‑labeled training datasets, which are difficult to obtain in acoustically complex reef environments. Here we present an unsupervised AI algorithm that automatically detects and clusters sounds from acoustic recordings with no manual annotation. We applied the algorithm to four concurrent, year‑long recordings from Hawaiian coral reefs, and it automatically grouped detections into nine distinct signal types: two call types from Domino damselfish, feeding sounds from parrotfish, calls from holocentrids (squirrelfishes and soldierfishes), an unidentified fish sound, three separate humpback whale song units, and ship noise. When tested on an independent dataset with labeled Domino damselfish calls, the model correctly identified more than 82% of damselfish calls while producing false positives on fewer than 15% of non‑damselfish sounds. All nine classes exhibited patterns linked to sunlight or the seasons, and we discovered previously‑unreported responses of the holocentrid and unknown fish groups to moonrise and lunar phase. This work demonstrates that unsupervised AI can detect and cluster sound types and uncover ecological trends without the guidance of a human analyst.

1. Introduction

Coral reefs are acoustically dynamic environments, with hundreds of diverse biological sounds often contained within a single minute of acoustic data [e.g., 1,2] alongside physical and anthropogenic noise [3,4]. These soundscapes serve as important indicators of reef health, with acoustic diversity and activity shown to increase recruitment of pelagic fish, coral, and crustacean larvae [511], and to directly correlate with more resilient reefs [5,12,13]. However, to date, reef ecosystem function as inferred through soundscape metrics has largely been correlative, based on unidentified sound classes [12,14,15], band-limited sound levels [4,1619], or acoustic indices [5,2024], all of which lack taxonomic identification. By parsing multiple individual biological sounds and identifying them to a more precise taxonomic level, coral reef soundscapes can be interpreted to provide improved ecological context and actionable metrics for reef monitoring, conservation, and restoration. This is critical because natural resource management is primarily focused on individual species rather than sound production [2].

The application of machine learning detectors and classifiers allows for the identification of individual signals that would be missed by traditional intensity-based metrics and acoustic indices. For automated detection of fish sounds, supervised machine learning has been typically applied to large, labeled datasets focusing on a single call type, including calls from groupers [25,26], damselfishes [27], and toadfishes [28]. While effective for identifying spatial or temporal call patterns, these techniques are typically not transferable to other environments/locations, other sensors, or other call-types. Other supervised detectors have been trained on fish calls more broadly, without differentiating call-type or species [14,15].

In contrast to supervised learning, unsupervised techniques enable rapid and inexpensive signal labeling without manual annotation. Unsupervised learning reduces inter- and intra-observer variability in event classification by circumventing subjective human interpretation, which can be particularly useful for acoustically diverse datasets where individual sounds are unknown or difficult to disentangle. Previous unsupervised algorithms have been trained to identify biological choruses in long-term spectral averages with minute-scale temporal resolution [2933], or to characterize soundscape recordings without identifying individual signals [23,24]. While a number of studies have used unsupervised clustering to detect and identify sounds associated with marine mammals [3439], fewer studies have used unsupervised techniques to classify individual fish sounds. In a study by Noble et al. [40], handpicked spectral and temporal features were extracted from manually labeled fish calls in order to cluster them into 55 unidentified sub-groups. A study by Ozanich et al. [41] differentiated fish and whale vocalizations using two unsupervised approaches: 1) clustering of handpicked spectral and temporal features similar to Noble et al. [40], and 2) clustering of deep latent features generated by convolutional autoencoders, where the deep learning approach was found to yield significant improvements in classification accuracy. The algorithms developed by Noble et al. [40] and Ozanich et al. [41] only provided broad categorical labels and were both applied to less than ten days of data, which constrained the ability to interpret soundscape composition and resolve diel or seasonal ecological trends.

Here, an unsupervised deep clustering algorithm was trained on four concurrent, year-long acoustic datasets from Hawaiian coral reefs, using an autoencoder-Gaussian mixture model framework similar to Ozanich et al. [41]. A simple manual review of automatically-generated clusters identified nine classes corresponding to distinct acoustic signals, including two call types from the Domino damselfish, Dascyllus albisella, soniferous signatures from holocentrids (squirrelfishes and soldierfishes), an unidentified fish call, scraping from parrotfish grazing, three song units from humpback whales, and noise from ship traffic. The clustering algorithm was verified against an independently collected, human-annotated dataset containing labeled D. albisella calls with synchronized video-audio array verification (AUC = 0.9). Seasonal and diel characteristics of well-studied signals (such as humpback whale and damselfish sounds) are consistent with known behaviors, while the temporal patterns of less-studied signals reveal previously-unidentified ecological trends, including the response of the holocentrid group to lunar illumination. This work demonstrates the capability of unsupervised algorithms to derive biologically meaningful acoustic classes from unlabeled reef soundscapes at an event-level resolution and provide ecological insights through long-term analysis of species- and family-level acoustic signatures.

2. Methods

Long-term passive acoustic measurements were collected at four survey sites with sensor depths ranging from 20-30 m off the western coast of Hawai’i Island (Fig 1). Survey sites were selected based on prior local knowledge to capture a wide range of coastal reef settings. HTI-96-min hydrophones were deployed with Loggerhead LS1 (Survey Sites 1 and 4) or Loggerhead LS1X (Survey Sites 2 and 3) recording packages. Acoustic recorders were set to record at a sampling rate of 96 kHz, recording for 1 minute every 15 minutes for the LS1 recorders and 1 minute every 10 minutes for the LS1X recorders to accommodate the 1-year deployment duration. The sensitivity of the HTI 96-min hydrophones used here was -170 dB re V/μPa, and the frequency range was 2 Hz to 30 kHz. A synchronized video-audio array was independently deployed at a fifth location (“Test Site”, Fig 1) in order to provide ground-truth verification of calls from Dascyllus albisella (Domino damselfish). This is a well-documented vocalizing species in Hawaiian coral reefs [4245], but few of their sounds are publicly available for model training. This passive acoustic camera combines an Insta-360 X2 360° camera within a tetrahedral hydrophone array and is described further in Dantzker et al [2]. The passive acoustic camera was deployed near nests of Dascyllus albisella and allowed for the opportunity to match sounds from specific focal fish species and their associated behaviors with vocalizations.

thumbnail
Fig 1. Map of survey sites on the western coast of Hawai’i Island.

Red, yellow, green, and blue dots correspond to long-term single-sensor survey sites, and the black dot corresponds to the test site with a synchronized video-audio array. Bathymetric data is from the Main Hawaiian Islands Multibeam Bathymetry Synthesis (https://www.soest.hawaii.edu/hmrg/multibeam).

https://doi.org/10.1371/journal.pcbi.1014516.g001

The data analyzed from the four survey sites spans from May 1, 2020 at 00:00 local time to May 1, 2021 at 00:00 local time. Each one-minute audio sample was downsampled to a sampling frequency of 1600 Hz and run through a 4th-order highpass Butterworth filter with critical frequency 30 Hz to reduce low-frequency seismic noise. Spectrograms were generated with 32-sample fast-Fourier transforms with an overlap of 28. Frequencies between 150 and 750 Hz were isolated in order to target a frequency band of interest where diverse biological sounds are present [46].

Detections were made using the Wang & Willett power-law detector, which was chosen for its suitability for transient signals of unknown structure. The power law statistic is defined [47] as

(1)

where is the power spectral density in linear scale at time and frequency , is the background noise level, and =1.7 [47]. Here was calculated for each frequency band as the 3-second median of centered at time . The power law statistic was then smoothed using a time-domain Gaussian filter with a standard deviation of 10. This filter width was chosen to focus the detector on longer-duration sounds (on the order of the 0.36 s sample window used in subsequent analysis) and to help ensure detections remained centered within the sample window. Peaks were then identified at a threshold of one standard deviation above the one-minute mean of the filtered output. This low detection threshold was intentionally chosen so that the automated clustering algorithm would be trained to distinguish relevant signals from background or low signal to noise ratio (SNR) samples. For each detection, a spectrogram sample () centered at the detection time was generated with size 12x144, corresponding to frequency range 150–750 Hz and duration 0.36 s. Spectrogram samples were standardized according to

(2)

where and are the mean and standard deviation of across both time and frequency bins. This standardization step ensures that signals with similar temporal/spectral characteristics but different intensities will be clustered similarly. Normalized samples were then obtained by clipping according to

(3)

effectively setting the lower threshold at one standard deviation above the within-sample mean and the upper threshold at two standard deviations above the within-sample mean. Samples were clipped in this way in order to zero out background fluctuations and emphasize higher-SNR features. The detector was run on the full year of data in all four survey sites. In order to equalize detection rates between sites with different duty cycles, we discarded every third detection at Survey Sites 2 and 3, which had 1.5 times greater temporal coverage than Sites 1 and 4.

A total of 7,767,943 detections were made across the four survey sites (1,514,293 in Site 1, 1,933,200 in Site 2, 1,931,952 in Site 3, and 2,388,498 in Site 4). The detections were randomly split into a training set (90%) and a held-out validation set (10%) and then fed into a convolutional autoencoder which compresses inputs into a 16-dimensional latent space before reconstruction (Fig 2, Table 1). This dimensionality was determined by testing several configurations, selecting the smallest latent space that ensured the reconstructed outputs captured the essential spectral and temporal features of the input spectrograms. Training utilized the Adam optimizer with a learning rate of 0.0001 and a batch size of 64, minimizing the mean squared error (MSE) between input and output tensors. Training concluded automatically when the epoch-to-epoch MSE reduction fell below 1 × 10 ⁻ ⁵, requiring 35 epochs to reach convergence on the training set (Fig 3). On a machine equipped with 128 GB of RAM and two NVIDIA TITAN V GPUs, this training process was completed in less than 26 hours.

thumbnail
Table 1. Autoencoder architecture, ReLu = Rectified linear unit.

https://doi.org/10.1371/journal.pcbi.1014516.t001

thumbnail
Fig 2. Architecture of the convolutional autoencoder.

The encoder compresses a 12x144 input into a 16 element latent embedding (top row). The decoder constructs a recreation of the input using only the latent embedding (bottom row).

https://doi.org/10.1371/journal.pcbi.1014516.g002

thumbnail
Fig 3. Training/validation loss and sample recreation.

Training and validation loss (left) are nearly identical after training ends, indicating that the autoencoder is not overfitting to the training set. Randomly-selected sample input spectrograms (middle) and reconstructions (right) show the autoencoder is retaining the basic spectral and temporal features in the latent embeddings.

https://doi.org/10.1371/journal.pcbi.1014516.g003

After training, the encoder reprocessed the entirety of the dataset to extract latent feature representations. These representations were clustered using Gaussian mixture models (GMMs) with k-means++ initialization and full covariance matrices in order to account for potential correlations between latent features. To ensure clustering stability and mitigate the risk of the expectation-maximization algorithm converging to suboptimal local extrema, we employed a standard multi-start protocol, selecting the final model based on the maximum log-likelihood achieved across ten random restarts. To assess sensitivity to the number of clusters (k), we computed standard model-selection metrics including the Akaike Information Criterion (AIC), Bayesian Information Criterion (BIC), Silhouette score, and Calinski–Harabasz index, across candidate cluster counts ranging from 10 to 45. We observed that these traditional metrics for evaluating cluster count did not converge to an optimal k for this dataset (S1 Fig). This is likely a consequence of a high proportion of noisy or ambiguous samples, which are prevalent in complex coral reef soundscapes. Therefore, we focused our model selection on the portions of the data that were well-modeled by the GMM. For each candidate cluster count k, we identified “well-defined” clusters—those where >10% of assigned samples had likelihoods >0.99—and computed the Akaike Information Criterion exclusively using samples within these well-defined clusters. This approach evaluated the model’s ability to describe well-structured data, while remaining robust to poorly-fit, noisy outliers. The optimal model (k = 29) minimized this modified AIC.

To interpret the acoustic content of these unsupervised groupings, we manually inspected 100 high-likelihood samples (likelihood > 0.99) per cluster by: 1) viewing their normalized spectrograms, and 2) listening to their raw audio. This manual review step is necessary to assign biological and ecological meaning to the automatically generated clusters. We focused this review on high-likelihood samples to establish the dominant acoustic signatures of each soft cluster, acknowledging that low-likelihood samples at the cluster peripheries would likely contain more ambiguous sounds or noise, a natural consequence of the complex acoustic environment. When possible, clusters were associated with source-specific labels (e.g., species or vessel) based on sounds documented in the literature [e.g., 46, 48, 49] or focal observations confirmed with video [2].

3. Results

Manual inspection of automatically generated clusters revealed that nine of the 29 clusters were associated with identifiable manmade and biological signals. These included two distinct sounds produced by Dascyllus albisella (“damselfish1,2”), scraping sounds from parrotfish feeding on coral (“parrotfish”), sounds from holocentrids (“holocentrid”), an unidentified fish sound (“unknown fish”), three distinct humpback whale song units (“humpback1,2,3”), and ship noise (“ship”). Sounds from holocentrid fishes were inferred based on the similarity to sounds produced by Myripristis berndti and other members of the family [5052], but we did not feel confident about assigning species-level labels to the holocentrid class. While the identified humpback song unit clusters represent acoustically distinct signals, we did not assess whether they are functionally distinct within the context of songs they are part of. Example spectrograms from very high-likelihood (>0.999) samples in each of the characterized classes are shown in Fig 4, and example spectrograms with variable thresholds for inclusion ( > 0.9, 0.5, 0.25, and 0) are shown in S2 Fig. While consistent spectral and temporal features are generally preserved as the detection threshold is lowered, this adjustment leads to the inclusion of noisy or misclassified samples.

thumbnail
Fig 4. Randomly selected spectrograms from high-likelihood (>0.999) samples for each of the 9 automatically-generated classes.

Audio files containing sample sounds from each class are in S1 Audio.

https://doi.org/10.1371/journal.pcbi.1014516.g004

Spectrogram samples for the uncharacterized clusters (numbered 1–20) are shown in S3 Fig. Manual review confirmed that many of these uncharacterized clusters contain a mixture of unidentified biological pulses and non-biological transient sounds, or low-SNR signals with few visible features in the spectrograms. One of these uncharacterized clusters (“18”) was found to contain a mixture of humpback whale sounds and ship tonals, indicating some confusion between these signal types.

To evaluate the statistical properties of the 29 clusters automatically generated by the Gaussian mixture model, we analyzed their internal variance, inter-cluster similarity, and classification confidence. The generalized variance for each cluster is shown in S4 Fig, where the comparatively high variance of the damselfish2 and holocentrid classes suggests they encompass a diverse range of vocalizations, and the lower variance of the parrotfish class indicates a narrower range of sound types. The probability distributions for these clusters were found to be mathematically distinct, as the Bhattacharyya coefficients for nearly all cluster pairs were below 0.36, indicating minimal overlap (S5 Fig). An analysis of likelihood scores revealed that the nine characterized sound classes are dominated by high-confidence detections, with prominent modes above a 0.99 likelihood (S6 Fig). In contrast, many of the uncharacterized clusters lack a high-confidence peak, suggesting ambiguous or noisy samples with uncertain assignments.

The distribution of detections for the characterized classes across all survey sites is shown in Fig 5. Significant variations in detection rates for each class are seen across the survey sites. Survey Site 1 saw a majority of damselfish2 call detections (51%) and a plurality of detections across all three humpback classes (40%, 39%, and 52%, respectively), as well as a strikingly low number of unknown fish detections (<0.5%). Site 2 saw a majority of unknown fish detections (67%) and a plurality of damselfish1 detections (46%). Site 3 saw a plurality of holocentrid (48%) and parrotfish (35%) detections. Site 4 was dominated by anthropogenic noise (47% of ship detections) with comparatively few biological detections (3.4% of humpback detections, 5.1% of holocentrid detections, and <15% for all other biological classes). The number of detections per class ranged from 20,000–50,000, with the exception of the damselfish1 class, where there were approximately 5,000.

thumbnail
Fig 5. Number of detections in each class (>0.99) across the 12 month recording period for the four survey sites.

https://doi.org/10.1371/journal.pcbi.1014516.g005

Detection performance for the damselfish1 and 2 classes was evaluated using the labeled dataset of verified Dascyllus albisella calls collected at the Test Site. The passive acoustic camera was deployed near a nest of Dascyllus albisella on July 27, 2022, and manual detections were made over a total of 57 minutes spanning 11:32–11:59 and 12:57–13:27 local time. 85 manual labels of the Dascyllus calls were verified by collocating beamformed acoustic detections with images of the individual from the 360° video imagery (Fig 6a). All of these sounds were associated with Dascyllus courtship/territorial displays, and video samples for four of the calls can be found at https://www.fisheyecollaborative.org/fish-sounds/dascyllus-albisella. The entire 57-minute dataset was run through the detection mechanism outlined in Section 2, where 6,098 initial detections were made, 82 of which overlapped with the manual labels. The 6,098 detections were classified using the pre-trained encoder and GMM clustering algorithms, with no extra training on the labeled dataset. Detections were classified as damselfish when where and are the likelihood scores for the damselfish1 and damselfish2 classes assigned by the Gaussian mixture model, and is an adjustable detection threshold. A ROC curve is generated by varying from 0 to 1 (Fig 6b). True positive rates of >82.5% are possible with false positive rates of <15%, and the area under the curve is 0.90. To validate the choice of this clustering framework, we also implemented standard k-means clustering using the same learned latent representations to serve as a baseline comparison (S7 Fig). The Gaussian mixture model demonstrated improved discrimination performance relative to k-means, achieving a higher ROC AUC (0.90 vs. 0.85) and Precision-Recall AUC (0.32 vs. 0.23). This supports the use of GMM clustering for this application.

thumbnail
Fig 6. Evaluating cluster assignments with an independent, labeled dataset.

Detection performance for the damselfish1 and 2 classes was evaluated using a 1-hour test dataset where a collocated acoustic array and 360° video system allowed for the visual identification of vocalizing fish. A sample video frame is shown in (a), with an overlain visualization of the sound field energy distribution corresponding to the position of a Dascyllus albisella. A Receiver Operating Characteristic (ROC) curve for damselfish detections is shown in (b), with an area under the curve of 0.9. The visualization of the sound field energy is masked to make the individual visible in this illustration.

https://doi.org/10.1371/journal.pcbi.1014516.g006

The diel and seasonal characteristics of the signals in each class are visualized in Fig 7. For classes where diel and seasonal characteristics are known, temporal patterns are consistent with previous observations. Humpback whale vocalizations occurred almost exclusively between the months of December and April, consistent with the known migration patterns of humpback whales to Hawaiian breeding grounds during the winter [53,54]. Damselfish calls and parrotfish scrapes occurred during the day, consistent with previous acoustic observations [27,55]. Ship noise occurred at sporadic intervals mostly during daytime hours, which can be explained by the prevalence of recreational vessels associated with day trips rather than cargo vessels or cruise ships at these survey sites. Significant increases in “ship” detections occurred during the nights of February 5 and 15, and manual inspection of samples confirmed that increased seismic activity that bled into frequencies above 150 Hz was misclassified here as ship noise.

thumbnail
Fig 7. Unique diel, lunar, or seasonal trends are observed for all classes.

Detection rates are shown as a function of time of year (vertical axis) and time of day (horizontal axis) for classes at selected survey sites, where the two damselfish call types and the three humpback song units were merged into singular classes. Solid yellow and cyan lines respectively correspond to sunrise and sunset, dashed yellow and cyan lines respectively correspond to moonrise and moonset, and white dots correspond to full moon. Detections were made here using a likelihood threshold of inclusion > 0.99. Recreations of these plots with variable thresholds ( > 0.9, 0.5, 0.25, and 0) are shown in S8 Fig. Similar plots for all survey sites are shown in S9 Fig.

https://doi.org/10.1371/journal.pcbi.1014516.g007

The output of the classifier also yielded insight into the temporal patterns of signals that have not been extensively studied. In particular, the holocentrid and unidentified fish calls displayed strong diel and lunar patterns. The unknown fish call occurred almost exclusively at night, during hours where the moon was not in the sky (dashed cyan and yellow lines, Fig 7). While changes in sound pressure level associated with moonrise have indicated a biological response to moonlight in this dataset [17], the automated detector and classifier allowed us to isolate the specific signals responsible for the lunar chorus. The holocentrid calls were most prominent in the 1–2 hours after sunset and before sunrise, with increases in activity during the full moon (white dots, Fig 7).

4. Discussion and conclusion

We demonstrated the ability to automatically detect and cluster individual reef sounds with no manual labels, advancing event-level signal detection by deploying established unsupervised learning architectures across multi-site, long-term soundscapes. After training on four concurrent, year-long datasets from Hawaiian coral reefs, our approach successfully partitioned sounds into distinct acoustic groups, which were then manually interpreted to identify ecologically meaningful categories, including five distinct fish call classes, three humpback whale song units, and a ship noise class. The diel and seasonal patterns of damselfish, parrotfish, and humpback whale detections are consistent with previous observations, with damselfish and parrotfish active in the daytime [27,55] and humpback songs detected in the winter months [53]. The holocentrid and unknown fish signals have not been identified previously, likely because they occur during the nighttime when diver or camera surveys are not viable. The holocentrid and unknown signals exhibit distinct responses to sunrise, sunset, and lunar illumination, demonstrating the capability of unsupervised detection and classification algorithms to uncover new signals and trends in ambient soundscapes.

The classifier was found to be robust against a manually annotated Dascyllus albisella dataset (AUC = 0.9), demonstrating the applicability of this model to other survey sites. While this model was specifically tailored to Hawaiian reefs, it can be readily retrained on other acoustic datasets in order to extract commonly occurring signals in those environments. The duration and frequency range of spectrogram samples were chosen to reflect the general characteristics of fish vocalizations in reefs (0.36 s, 150–750 Hz), however the model also detected and classified signals with durations and frequencies larger than the window ranges provided. Parrotfish scrapes have peak frequencies on the order of 2–4 kHz, but are broadband (200–8000 Hz) and commonly occupy the lower frequencies studied here [55,56]. Humpback song units typically have peak frequencies between 0-1.5 kHz, with song unit durations on the order of 1–3 s [e.g., 49,57]. Tonals from passing ships can last minutes or hours, depending on the source level, speed and distance of the ship [e.g., 48,58]. The model architecture can therefore be applied with minimal prior knowledge of the signals that comprise the soundscape and can be used to discover signals that were unknown to the human analyst.

It is important to note that the optimization of the detector and spectrogram parameters for fish vocalizations likely resulted in suboptimal detection performance for signals with different spectral and temporal characteristics. This tradeoff is evident for longer-duration tonal signals, where many of the “ship” detections were mislabeled seismic activity, and one of the ambiguous clusters (“cluster 18”) was found to contain a mixture of ship sounds and humpback sounds (S3 and S10 Figs). In addition, a large number of uncharacterized clusters contain short-duration (<0.02 s) pulses, which were comprised of a mixture of biological and non-biological transient sounds (S3 Fig).

Addressing these limitations presents clear opportunities for future refinement. For instance, the characterization of signals with varying bandwidths and durations could be improved by implementing parallel models with spectrogram parameters tailored to different signal types. Future work could also expand the quantitative testing of classification performance beyond our initial use of labeled damselfish sounds by creating new datasets to calibrate class-specific likelihood thresholds for other sound types. Classification accuracy could be further optimized using “human in the loop” approaches [1], where automatically generated clusters are manually refined or subdivided for supervised training. In this process, growing public libraries of identified fish sounds, such as FishSounds [59], the Global Library of Underwater Biological Sounds (GLUBS) [60], and the Macaulay Library (https://www.macaulaylibrary.org) would serve as valuable reference datasets for annotating and validating these clusters. Further, while D. albisella sounds are well-characterized and readily identifiable [42,43], the ability to extract identified sounds from a combined audio-video sensor array [2] highlights the opportunity to identify sound sources from the field, and then extract a large number of those species-labeled sounds for rapid retraining of machine learning models.

These results demonstrate the significant advantages offered by unsupervised machine learning for soundscape characterization and signal classification at scale. Acoustic signals in a long-term dataset can be ingested, parsed, and automatically clustered without the burden and expense of manual labeling. By successfully scaling an established autoencoder-GMM framework, this study demonstrates its power to extract meaningful, event-level ecological insights from complex, unlabeled acoustic data. In contrast to traditional supervised learning approaches, the detection and clustering architecture introduced here can be easily applied to a variety of unlabeled acoustic datasets. These methods can be particularly useful for datasets where individual sounds are complex or difficult to disentangle, or where the specific signals that comprise the underwater soundscape have not yet been characterized. Such computational efforts will be essential for providing actionable, taxonomically-specific sound identification to enable passive acoustic monitoring for management and conservation of coral reef fishes and ecosystems at global scales [6,61,62].

Supporting information

S1 Audio. Audio samples for each of the nine characterized classes.

Audio files are produced for the nine characterized classes by collating ten randomly-selected one-second sound samples with a very high-likelihood score (>0.999) for that class.

https://doi.org/10.1371/journal.pcbi.1014516.s001

(ZIP)

S1 Fig. Non-convergence of traditional metrics for determining cluster count necessitates a modified AIC approach.

Traditional clustering validation metrics including the Akaike Information Criterion (AIC), Bayesian Information Criterion (BIC), Silhouette score, and Calinski–Harabasz index were evaluated from cluster counts ranging from 10 to 45. To reduce computational costs, a random 10% subsample of the dataset was analyzed for clustering validation. The AIC, BIC, and Calinski–Harabasz scores trended downward with increasing cluster count, despite minor local fluctuations, while the Silhouette Score fluctuated from -0.07 to 0.01 across the surveyed range. The failure of these traditional metrics to converge on a single cluster count reflects the high proportion of noisy or ambiguous samples prevalent in complex coral reef soundscapes. Because these traditional indices struggle to resolve structure in highly noisy datasets, model selection required a targeted approach focused on data well-modeled by the Gaussian mixture model. We implemented a modified AIC, computed exclusively using samples within “well-defined” clusters (defined as clusters where >10% of assigned samples exhibited assignment likelihoods >0.99).

https://doi.org/10.1371/journal.pcbi.1014516.s002

(PDF)

S2 Fig. Recreations of Fig 4, showing the effect of lowering the likelihood threshold for sample inclusion.

For a sample to be included in the visualization for a given class, it must belong to that cluster and exceed a likelihood threshold . The thresholds chosen are >0.9 (A), >0.5 (B), >0.25 (C), >0 (D). While consistent spectral and temporal features are preserved as the detection threshold is lowered, this adjustment leads to the inclusion of lower-SNR samples and an elevated risk of false positives.

https://doi.org/10.1371/journal.pcbi.1014516.s003

(PDF)

S3 Fig. Randomly selected spectrogram samples (>0.9) for the 20 uncharacterized clusters, which lacked consistent, distinct biological or anthropogenic signatures.

Many of these uncharacterized clusters (e.g., 6, 7, 8, 9, 12, 17, 20) contain short-duration (<0.02 s) pulses, which included a mixture of biological pulses and non-biological transient sounds. Other clusters (e.g., 1, 4, 5, 19) contained longer duration (>0.1 s) sounds which were again comprised of a diverse mixture of biological and non-biological transients. Clusters 2 and 13 contained low-SNR signals with few visible features in the spectrograms. Cluster 18 was found to contain a mixture of humpback whale sounds and ship tonals, which is confirmed by the diel/seasonal analysis shown in S10 Fig. A detection threshold of 0.9 was chosen since many of these uncharacterized clusters did not have samples with likelihoods above 0.99.

https://doi.org/10.1371/journal.pcbi.1014516.s004

(PDF)

S4 Fig. The logarithm of the generalized variance for each of the 29 clusters automatically generated by the Gaussian mixture model, where the generalized variance is calculated as the determinant of the covariance matrix .

The 29 clusters show values ranging from approximately 30–60. The damselfish2 and holocentrid classes have the greatest generalized variance values ( > 55), suggesting that these sound classes contain a diverse range of vocalizations. The lowest generalized variance values ( < 35) are associated with the parrotfish class and clusters 13, 17, and 18, suggesting a narrower range of sound types.

https://doi.org/10.1371/journal.pcbi.1014516.s005

(PNG)

S5 Fig. Matrix showing the Bhattacharyya coefficient for each cluster pair within the 29 automatically generated clusters.

The Bhattacharyya coefficient is a measure of the similarity between probability distributions, ranging from 0 to 1. The highest coefficient of 0.50 occurs between the unknown fish class and cluster 14, suggesting moderate overlap. However, as shown in S6 Fig, cluster 14 has a substantially lower proportion of high-likelihood (>0.99) samples compared to the unknown fish class (<0.001% vs. > 10%, respectively), which indicates that the inclusion of cluster 14 would not significantly impact classification performance for high-confidence detections. The Bhattacharyya coefficients for every other cluster pair are less than 0.36, indicating minimal overlap between probability distributions.

https://doi.org/10.1371/journal.pcbi.1014516.s006

(PNG)

S6 Fig. Distribution of likelihood scores for all samples within each of the 29 automatically generated clusters.

The nine characterized clusters have prominent modes above 0.99, indicating a high percentage of high-confidence detections. In contrast, many of the uncharacterized clusters (numbered 1–20) lack a high-confidence peak, with distributions centered around a likelihood of 0.5. This suggests that these clusters may contain a high proportion of ambiguous or noisy samples with uncertain assignments. However, a subset of uncharacterized clusters (e.g., 1, 2, 9, 19, 20) also exhibit high-confidence peaks, suggesting they represent acoustically consistent signals.

https://doi.org/10.1371/journal.pcbi.1014516.s007

(PNG)

S7 Fig. Comparison of detection performance between k-means and Gaussian mixture modeling.

We benchmarked the detection performance of Gaussian mixture modeling (GMM) against k-means clustering, using the labeled damselfish calls as ground truth. Latent representations from the training dataset (7,767,943 samples) were partitioned using k-means with 29 clusters to match the number of GMM clusters and ensure a mathematically fair comparison. Latent representations from the labeled test dataset (6,098 samples) were then clustered using the pre-trained k-means algorithm, with no extra training on the labeled dataset. The k-means clusters were ranked in descending order based on the number of labeled damselfish calls they contained, and test samples were classified as damselfish if they mapped to the top clusters, where serves as an integer detection threshold. ROC curves and Precision-Recall curves were generated for the k-means clusters by varying from 0 to 29. GMM outperformed k-means across both evaluation frameworks, achieving an ROC Area Under the Curve (AUC) of 0.9 compared to 0.85 for k-means (a) and a Precision-Recall AUC of 0.32 compared to 0.23 for k-means (b).

https://doi.org/10.1371/journal.pcbi.1014516.s008

(PNG)

S8 Fig. Recreations of Fig 7, showing the effect of lowering the likelihood threshold for sample inclusion.

For a sample to be included in the visualization for a given class, it must belong to that cluster and exceed a likelihood threshold . The thresholds shown are >0.9 (A), >0.5 (B), >0.25 (C), >0 (D). The diel, lunar, and seasonal trends seen in Fig 7 are mostly preserved as the threshold is lowered. The primary exception is the “ship” class, where a significantly lowered threshold results in elevated nighttime detections, including sunset activity likely associated with biological sound (possible confusion with the “unknown fish” sound).

https://doi.org/10.1371/journal.pcbi.1014516.s009

(PDF)

S9 Fig. Diel and seasonal trends for the nine characterized classes at all survey sites.

Detection rates are shown as a function of time of year (vertical axis) and time of day (horizontal axis). Solid yellow and cyan lines respectively correspond to sunrise and sunset.

https://doi.org/10.1371/journal.pcbi.1014516.s010

(PDF)

S10 Fig. Diel and seasonal trends for cluster 18, which is found to correspond to a mixture of humpback whale vocalizations and ship sounds.

In the survey site where humpback vocalizations are most frequent (Site 1), detections primarily occur between the months of December and April, consistent with the seasonal pattern seen for classes “humpback1,2,3.” In the site where ship sounds are most frequent (Site 4), detections occurred at sporadic intervals mostly during daytime hours, consistent with the diel pattern seen for the “ship” class. Detections in Sites 2 and 3 contained a mixture of the seasonal humpback pattern and the diel ship pattern.

https://doi.org/10.1371/journal.pcbi.1014516.s011

(PNG)

Acknowledgments

We would like to thank Justin Snow at Pacific Solutions, Gavin Key and Katie Key at Pacific Watersports, and David Mann at Loggerhead Instruments.

References

  1. 1. Ferguson SR, Jensen FH, Hyer MD, Noble A, Apprill A, Mooney TA. Ground-truthing daily and lunar patterns of coral reef fish call rates on a US Virgin Island reef. Aquat Biol. 2022;31:77–87.
  2. 2. Dantzker MS, Duggan MT, Berlik E, Delikaris-Manias S, Bountourakis V, Pulkki V, et al. Deciphering complex coral reef soundscapes with spatial audio and 360° video. Methods Ecol Evol. 2025; 16(11): 2622–37.
  3. 3. Ferrier-Pagès C, Leal MC, Calado R, Schmid DW, Bertucci F, Lecchini D, et al. Noise pollution on coral reefs? - A yet underestimated threat to coral reef communities. Mar Pollut Bull. 2021;165:112129. pmid:33588103
  4. 4. Kaplan MB, Lammers MO, Zang E, Aran Mooney T. Acoustic and biological trends on coral reefs off Maui, Hawaii. Coral Reefs. 2017;37(1):121–33.
  5. 5. Kaplan M, Mooney T, Partan J, Solow A. Coral reef species assemblages are associated with ambient soundscapes. Mar Ecol Prog Ser. 2015;533:93–107.
  6. 6. Apprill A, Girdhar Y, Mooney TA, Hansel CM, Long MH, Liu Y, et al. Toward a new era of coral reef monitoring. Environ Sci Technol. 2023;57(13):5117–24. pmid:36930700
  7. 7. Azofeifa-Solano JC, Parsons MJG, Brooker R, McCauley R, Pygas D, Feeney W. Soundscape analysis reveals fine ecological differences among coral reef habitats. Ecol Indic. 2025;171:11.
  8. 8. Gordon TAC, Radford AN, Davidson IK, Barnes K, McCloskey K, Nedelec SL, et al. Acoustic enrichment can enhance fish community development on degraded coral reef habitat. Nat Commun. 2019;10(1):5414. pmid:31784508
  9. 9. Lamont TAC, Williams B, Chapuis L, Prasetya ME, Seraphim MJ, Harding HR, et al. The sound of recovery: coral reef restoration success is detectable in the soundscape. J Appl Ecol. 2022;59(3):742–56.
  10. 10. Suca J, Lillis A, Jones I, Kaplan M, Solow A, Earl A, et al. Variable and spatially explicit response of fish larvae to the playback of local, continuous reef soundscapes. Mar Ecol Prog Ser. 2020;653:131–51.
  11. 11. Tolimieri N, Jeffs A, Montgomery J. Ambient sound as a cue for navigation by the pelagic larvae of reef fishes. Mar Ecol Prog Ser. 2000;207:219–24.
  12. 12. Jarriel SD, Formel N, Ferguson SR, Jensen FH, Apprill A, Mooney TA. Unidentified fish sounds as indicators of coral reef health and comparison to other acoustic methods. Front Remote Sens. 2024;5:1338586.
  13. 13. Freeman L, Freeman S. Rapidly obtained ecosystem indicators from coral reef soundscapes. Mar Ecol Prog Ser. 2016;561:69–82.
  14. 14. McCammon S, Formel N, Jarriel S, Mooney TA. Rapid detection of fish calls within diverse coral reef soundscapes using a convolutional neural networka). J Acoust Soc Am. 2025;157(3):1665–83. pmid:40067342
  15. 15. Mouy X, Archer SK, Dosso S, Dudas S, English P, Foord C, et al. Automatic detection of unidentified fish sounds: a comparison of traditional machine learning with deep learning. Front Remote Sens. 2024;5:1439995.
  16. 16. Cato DH. Marine biological choruses observed in tropical waters near Australia. J Acoust Soc Am. 1978;64(3):736–43.
  17. 17. Duane D, Freeman S, Freeman L. Moonlight-driven biological choruses in Hawaiian coral reefs. PLoS One. 2024;19(3):e0299916. pmid:38507354
  18. 18. Nedelec S, Simpson S, Holderied M, Radford A, Lecellier G, Radford C, et al. Soundscapes and living communities in coral reefs: temporal and spatial variation. Mar Ecol Prog Ser. 2015;524:125–35.
  19. 19. Staaterman E, Rice AN, Mann DA, Paris CB. Soundscapes from a Tropical Eastern Pacific reef and a Caribbean Sea reef. Coral Reefs. 2013;32(2):553–7.
  20. 20. Bertucci F, Parmentier E, Lecellier G, Hawkins AD, Lecchini D. Acoustic indices provide information on the status of coral reefs: an example from Moorea Island in the South Pacific. Sci Rep. 2016;6:33326. pmid:27629650
  21. 21. Dimoff SA, Halliday WD, Pine MK, Tietjen KL, Juanes F, Baum JK. The utility of different acoustic indicators to describe biological sounds of a coral reef soundscape. Ecol Indic. 2021;124:107435.
  22. 22. Elise S, Urbina-Barreto I, Pinel R, Mahamadaly V, Bureau S, Penin L. Assessing key ecosystem functions through soundscapes: a new perspective from coral reefs. Ecol Indic. 2019;107:11.
  23. 23. Williams B, Balvanera SM, Sethi SS, Lamont TAC, Jompa J, Prasetya M, et al. Unlocking the soundscape of coral reefs with artificial intelligence: pretrained networks and unsupervised learning win out. PLoS Comput Biol. 2025;21(4):e1013029. pmid:40294093
  24. 24. Minier L, Rouch J, Sabbagh B, Bertucci F, Parmentier E, Lecchini D, et al. Visualization and quantification of coral reef soundscapes using CoralSoundExplorer software. PLoS Comput Biol. 2025;21(4):e1012050. pmid:40208899
  25. 25. Ibrahim AK, Chérubin LM, Zhuang H, Schärer Umpierre MT, Dalgleish F, Erdol N, et al. An approach for automatic classification of grouper vocalizations with passive acoustic monitoring. J Acoust Soc Am. 2018;143(2):666. pmid:29495690
  26. 26. Ibrahim AK, Zhuang HQ, Schaerer-Umpierre M, Woodward C, Erdol N, Cherubin LM. Fish acoustic detection algorithm research: a deep learning app for Caribbean grouper calls detection and call types classification. Front Mar Sci. 2024;11:1378159.
  27. 27. Munger J, Herrera D, Haver S, Waterhouse L, McKenna M, Dziak R, et al. Machine learning analysis reveals relationship between pomacentrid calls and environmental cues. Mar Ecol Prog Ser. 2022;681:197–210.
  28. 28. Bohnenstiehl DR. Automated cataloging of oyster toadfish (Opsanus tau) boatwhistle calls using template matching and machine learning. Ecol Inform. 2023;77:102268.
  29. 29. Butler J, Pagniello C, Jaffe JS, Parnell PE, Širović A. Diel and seasonal variability in kelp forest soundscapes off the Southern California coast. Front Mar Sci. 2021;8:629643.
  30. 30. Kim EB, Frasier KE, McKenna MF, Kok ACM, Peavey Reeves LE, Oestreich WK, et al. SoundScape learning: an automatic method for separating fish chorus in marine soundscapes. J Acoust Soc Am. 2023;153(3):1710. pmid:37002102
  31. 31. Lin T-H, Akamatsu T, Tsao Y. Sensing ecosystem dynamics via audio source separation: a case study of marine soundscapes off northeastern Taiwan. PLoS Comput Biol. 2021;17(2):e1008698. pmid:33600436
  32. 32. Lin T-H, Fang S-H, Tsao Y. Improving biodiversity assessment via unsupervised separation of biological sounds from long-duration recordings. Sci Rep. 2017;7(1):4547. pmid:28674439
  33. 33. Mahale VP, Chanda K, Chakraborty B, Salkar T, Sreekanth GB. Biodiversity assessment using passive acoustic recordings from off-reef location-Unsupervised learning to classify fish vocalization. J Acoust Soc Am. 2023;153(3):1534. pmid:37002105
  34. 34. Buchan SJ, Mahú R, Wuth J, Balcazar-Cabrera N, Gutierrez L, Neira S, et al. An unsupervised Hidden Markov Model-based system for the detection and classification of blue whale vocalizations off Chile. Bioacoustics. 2019;29(2):140–67.
  35. 35. Frasier KE. A machine learning pipeline for classification of cetacean echolocation clicks in large underwater acoustic datasets. PLoS Comput Biol. 2021;17(12):e1009613. pmid:34860825
  36. 36. Frasier KE, Roch MA, Soldevilla MS, Wiggins SM, Garrison LP, Hildebrand JA. Automated classification of dolphin echolocation click types from the Gulf of Mexico. PLoS Comput Biol. 2017;13(12):e1005823. pmid:29216184
  37. 37. Li K, Sidorovskaia NA, Tiemann CO. Model-based unsupervised clustering for distinguishing Cuvier’s and Gervais’ beaked whales in acoustic data. Ecol Inform. 2020;58:101094.
  38. 38. Murray SO, Mercado E, Roitblat HL. The neural network classification of false killer whale (Pseudorca crassidens) vocalizations. J Acoust Soc Am. 1998;104(6):3626–33. pmid:9857520
  39. 39. Cohen RE, Frasier KE, Baumann-Pickering S, Wiggins SM, Rafter MA, Baggett LM, et al. Identification of western North Atlantic odontocete echolocation click types using machine learning and spatiotemporal correlates. PLoS One. 2022;17(3):e0264988. pmid:35324943
  40. 40. Noble AE, Jensen FH, Jarriel SD, Aoki N, Ferguson SR, Hyer MD, et al. Unsupervised clustering reveals acoustic diversity and niche differentiation in pulsed calls from a coral reef ecosystem. Front Remote Sens. 2024;5:1429227.
  41. 41. Ozanich E, Thode A, Gerstoft P, Freeman LA, Freeman S. Deep embedded clustering of coral reef bioacoustics. J Acoust Soc Am. 2021;149(4):2587. pmid:33940892
  42. 42. Lobel PS, Mann DA. Spawning sounds of the damselfish, Dascyllus albisella(Pomacentridae), and relationship to male size. Bioacoustics. 1995;6(3):187–98.
  43. 43. Mann DA, Lobel PS. Passive acoustic detection of sounds produced by the damselfish,Dascyllus albisella(Pomacentridae). Bioacoustics. 1995;6(3):199–213.
  44. 44. Mann DA, Lobel PS. Acoustic behavior of the damselfish Dascyllus albisella: behavioral and geographic variation. Environ Biol Fish. 1998;51(4):421–8.
  45. 45. Laboury S, Parmentier E, Lobel PS. Are there individual acoustic signatures in the damselfish Dascyllus albisella?. J Acoust Soc Am. 2025;157(1):48–56. pmid:39775844
  46. 46. Tricas T, Boyle K. Acoustic behaviors in Hawaiian coral reef fish communities. Mar Ecol Prog Ser. 2014;511:1–16.
  47. 47. Wang Z, Willett PK. All-purpose and plug-in power-law detectors for transient signals. IEEE Trans Signal Process. 2001;49(11):2454–66.
  48. 48. McKenna MF, Rowell TJ, Margolina T, Baumann-Pickering S, Solsona-Berga A, Adams JD, et al. Understanding vessel noise across a network of marine protected areas. Environ Monit Assess. 2024;196(4):369. pmid:38489113
  49. 49. Au WWL, Pack AA, Lammers MO, Herman LM, Deakos MH, Andrews K. Acoustic properties of humpback whale songs. J Acoust Soc Am. 2006;120(2):1103–10. pmid:16938996
  50. 50. Banse M, Bertimes E, Lecchini D, Donaldson TJ, Bertucci F, Parmentier E. Sounds as taxonomic indicators in holocentrid fishes. npj Biodivers. 2024;3(1):33. pmid:39501023
  51. 51. Banse M, Hanssen N, Sabbe J, Lecchini D, Donaldson TJ, Iwankow G, et al. Same calls, different meanings: acoustic communication of Holocentridae. PLoS One. 2024;19(11):e0312191. pmid:39570950
  52. 52. Salmon M. Acoustical behavior of the menpachi, Myripristis berndti, in Hawaii. Pacific Science. 1967;21:364–81.
  53. 53. Baker C, Herman L, Perry A, Lawton W, Straley J, Wolman A, et al. Migratory movement and population structure of humpback whales (Megaptera novaeanglieae) in the central and eastern North Pacific. Mar Ecol Prog Ser. 1986;31:105–19.
  54. 54. Lammers MO, Goodwin B, Kügler A, Zang EJ, Harvey M, Margolina T. The occurrence of humpback whales across the Hawaiian archipelago revealed by fixed and mobile acoustic monitoring. Front Mar Sci. 2023;10:1083583.
  55. 55. Tricas T, Boyle K. Parrotfish soundscapes: implications for coral reef management. Mar Ecol Prog Ser. 2021;666:149–69.
  56. 56. Sartori JD, Bright TJ. Hydrophonic study of the feeding activities of certain Bahamian parrot fishes, family Scaridae. Hydro-Lab J. 1973;2(1):25–56.
  57. 57. Kügler A, Lammers MO, Zang EJ, Pack AA. Male humpback whale chorusing in Hawai’i and its relationship with whale abundance and density. Front Mar Sci. 2021;8:735664.
  58. 58. Kline LR, DeAngelis AI, McBride C, Rodgers GG, Rowell TJ, Smith J. Sleuthing with sound: understanding vessel activity in marine protected areas using passive acoustic monitoring. Mar Policy. 2020;120:104138.
  59. 59. Looby A, Vela S, Rice AN, Bravo S, Davies HL, Murchy KA, et al. FishSounds versions 2 and 3: achieving the largest global database of fish sound production. Global Ecol Biogeogr. 2025;34(11).
  60. 60. Parsons MJ, Looby A, Chanda K, Di Iorio L, Erbe C, Frazao F, et al. A global library of underwater biological sounds (GLUBS): an online platform with multiple passive acoustic monitoring applications. The effects of noise on aquatic life: principles and practical considerations. Cham: Springer International Publishing; 2024. 2149–73.
  61. 61. Mooney TA, Di Iorio L, Lammers M, Lin T-H, Nedelec SL, Parsons M, et al. Listening forward: approaching marine biodiversity assessments using acoustic methods. R Soc Open Sci. 2020;7(8):201287. pmid:32968541
  62. 62. Hodson EJ, Cox K, Juanes F, Looby A. Actively soniferous tropical reef fishes are diverse, vulnerable, and valuable. J Fish Biol. 2025;106(4):990–5. pmid:39681114