
 Research journal  First study: plant promoters  Reproduction package 

STUDY 02 · COMPUTATIONAL BIOACOUSTICS · 19 SEPTEMBER 2026
Non-overlapping bird calls suppress detections in BirdNET v2.4
A controlled context intervention with distinct focal recordings, natural backgrounds, and field-recording ablations.
Executed research note · novelty provisionally supported by two literature reviews · not peer reviewed


The finding
Adding another bird in a separate part of a three-second window made 54.5% of otherwise detectable target appearances fall below a score of 0.5. The target audio itself was unchanged. The larger experiment used 56 passages from 16 species, recorded in 56 distinct files by 52 credited recordists.
This is a measured failure under a controlled intervention. It is not an estimate that BirdNET misses this fraction of birds in the wild.


Abstract
Multi-species acoustic classifiers can introduce dependence between detections. To isolate this dependence from simultaneous acoustic overlap, we placed annotated 1.2-second single-species passages in disjoint intervals of a three-second input, separated by 0.3 seconds. Target components had identical gain and position in baseline and combined inputs. A focal-recording extension prospectively specified in the journal produced 3,198 losses among 5,872 baseline-eligible target appearances (54.5%; source-file bootstrap interval 49.0%–59.5%). Both orders were tested, and every included species was affected. The result survived common waveform normalization and continuous natural background. A secondary temporal ablation of 364 unmodified field windows found conditional losses of 26.4%. Exact enumeration of independent calling states generated a negative detected association (φ = -0.472) despite true independence. These results motivate auditing joint detectability before interpreting fine-scale classifier co-detections as bird interactions.

All bars describe conditional detection loss at score 0.5. Samples and eligibility differ between panels: pilot 9 species; held-out soundscapes 4; focal extension 16. Fixed-normalization and ambient controls use fixed subsets of focal pairs. The natural-window ablation is secondary. Bars are not independent population estimates.


Question and experimental logic
In a detector with reliable multi-label recognition, adding another identifiable call need not eliminate evidence already present for a target. We tested whether it nevertheless does. The primary question concerns the classifier’s observation process, not whether birds compete or avoid each other.

Baseline A  A · 0.15–1.35 s  silence 
Baseline B  silence  B · 1.65–2.85 s 
Combined  identical A  identical B 
All source excerpts were mono, resampled to 48 kHz, mean-centered and faded over 5 ms at their edges. They were normalized to a common RMS, followed by one global clipping-safe scaling factor and a factor of 0.5 applied equally to all components. The combined waveform never clipped. We tested both temporal orders and simultaneous mixtures. Scores use the model’s 6,522 outputs with the source implementation’s clipped sigmoid; no geographic filter, custom classifier, top-k truncation, or retraining was used.
Primary eligibility required score ≥0.5 in both time positions. This deliberately restricts inference to calls the model can recognize alone. Candidate passages, recording limits and advancement rules were recorded before their mixture outcomes. The score is a model output, not a calibrated probability of species presence.


Discovery history and frozen gates
 Stage  Sample  Outcome 
 015: isolated-note feasibility  37 candidates; only 1 species qualified  Failed before mixture testing 
 018: single-species passages  98 candidates → 18 passages/9 species  Non-overlap losses 53.6% and 57.5%; pilot gate passed 
 018: untouched soundscape files  178 candidates → 16 passages/4 species  Effect replicated at 43.4%; failed the frozen 8-species diversity gate 
 019: separate focal source  360 candidates → 56 passages/16 species  Passed ≥8 species, ≥24 passages, ≥20% losses and ≥4 affected species 
The focal extension was separately specified after the soundscape diversity shortfall. It was not represented as success of that failed gate. Its clips came from the CEB focal subset, while the earlier clips came from soundscape recordings. Xeno-Canto recordings may have been used during BirdNET training, so this is not a held-out model-accuracy benchmark. The counterfactual manipulation and source extension are the units of validation.
The journal also retains two closed plant hypotheses from this round. A copied-output-path error during the bird pilots was recovered from git; separated reruns reproduced the same outcomes.  Recovery record . Every artifact of the first completed plant study remains unchanged.


Results
The focal experiment comprised 1,468 heterospecific recording pairs, both orders, and two target appearances per order: 5,872 target appearances, not 5,872 independent birds. Loss was 54.6% in one order and 54.3% in the reverse order. Simultaneous loss was 55.7%. Equal weighting of species gave 53.1% non-overlap loss. A two-endpoint source-file bootstrap with 2,000 resamples gave the interval above; it is conditional on this set of selected species and detectable passages.
Independent passages from the same species produced 0 losses among 288 non-overlapping target appearances. This control shows that adding another passage or more total sound does not inevitably reduce the target score. It does not prove which learned features cause heterospecific suppression.

Rows are the species whose detection is evaluated; columns are the species added elsewhere in the same window. Each cell pools multiple source-recording combinations and both orders. Blank diagonal cells are the separate conspecific control. This is a detector-response matrix, not a map of ecological competition.


Alternative explanations and controls
Waveform normalization
The pinned model begins with per-window min/max waveform normalization. We exposed the tensor immediately after its seven normalization operations and retained all 549 model buffers and every later operation. Feeding the original normalization into this modified graph reproduced all tested baseline logits exactly. For 300 fixed-hash focal pairs, we then used the same original-mixture extrema to normalize both baselines and the mixture. Target samples after normalization were bit-identical. Loss remained 57.1% across 1,196 eligible target appearances. The original and modified models’ mixture scores agreed within 7.2e-08. Thus, waveform min/max rescaling alone cannot explain the result; we do not attribute it to a specific later layer.
Continuous natural background
Three annotation-free three-second background passages were selected from distinct soundscape files before mixture scores. Each was scaled to RMS 0.0015 and included unchanged in background-only, single-bird and two-bird inputs. Among 2,400 target appearances from 200 fixed pairs and both orders, 2,139 remained detectable alone and absent in the background-only prediction. Conditional loss was 41.1%, with 40.0–42.1% across backgrounds. This argues against zero padding alone as an explanation. Annotation-free does not guarantee complete acoustic silence.
Natural field-window ablation
The metadata screen found 366 windows; two involved a species absent from the model label list, leaving every model-covered three-second grid window with exactly two annotated species whose bouts were separated by ≥0.1 s: 364 windows in 88 files. Complementary masks transitioned only inside the annotated gap; their components sum to the original recording and preserve the target’s annotated samples. Among 140 individually eligible appearances, loss in the original full context was 26.4% (file-bootstrap interval 18.3%–35.0%). With shared original-window normalization it was 25.4%. The stricter subset where both species were independently detectable contained 18 appearances and had 44.4% loss.
Across all 728 target appearances, there were 37 threshold losses and 29 gains. Median score change across all appearances was +0.0033. This secondary test reuses soundscape data and removes background together with the neighboring passage; it supports context dependence without isolating recognized bird identity as the sole cause.

Every selected natural-window target is shown, including gains. Points below the diagonal receive a lower score in the original field context than after the disjoint context is removed. Dashed lines mark 0.5; this is not a field-prevalence estimate.


What it means for co-detection
For each eligible pair and order, we enumerated four states—neither caller, A alone, B alone, both—each with probability 0.25. Calling is therefore exactly independent. Here, φ is the ordinary correlation between two binary detection indicators: negative values mean they are detected together less often than independence predicts. All 2,936 focal pair-orders also passed the specificity check: neither species was detected in silence or in the other’s single-source input. The classifier nevertheless produced a pooled detected φ of -0.472; 90.8% of pair-orders had negative detected association. No random simulation or significance test is needed for this finite enumeration.
The size of this artificial association depends on calling frequency and analysis scale. In an explicitly hypothetical independent-window repeat model, reducing per-window calling probability to 0.05 gives φ≈−0.046 at 3 s, and aggregating those detections to 5-minute presence gives φ≈−0.0015. The experiment therefore does not invalidate long-duration occupancy studies. It identifies an observation-process hazard most directly relevant to short-window vocal associations.

Left: independent true calling states. Middle: the classifier’s pooled detected states at calling probability 0.5. Right: analytic sensitivity under an independent-window toy model; these calling rates and aggregation assumptions were not estimated from wild birds.


Novelty and practical value
Previous work establishes environmental interference, overlapping-source separation, multi-label recognition, heterogeneous classifier errors and detection-aware occupancy models. None of those ideas is claimed as new. Two documented searches, repeated after the larger test, did not locate a directly equivalent combination of disjoint-time bird counterfactuals, matched waveform-normalization control and explicit independent-presence association enumeration for BirdNET v2.4.
The original contribution is a reproducible, narrowly scoped empirical result and stress-test resource. It can be used to audit a detector before interpreting fine-scale co-detections. Novelty remains provisional pending expert review; the note has not been submitted or peer reviewed. No deployable correction or new separation algorithm is claimed. Simply taking maxima over extra crops can create false positives and was not validated as a remedy here.
 First novelty review and exact queries  ·  Second novelty review 


Limitations
One pinned BirdNET v2.4 checkpoint, a selected European species set, short excerpts, and a fixed primary threshold. Results do not establish behavior of newer BirdNET models or other architectures.
Baseline eligibility is deliberate conditioning. The quoted loss percentages are not unconditional recall, false-negative rates in nature, or estimates of missed populations.
Files, passages, species and pair combinations are distinct units. Individual birds and recording sites are not independently identified; combinations repeatedly reuse recordings.
Annotations can miss sounds. Focal recordings may overlap model training; natural-window ablation was secondary and reused the CEB soundscapes.
The common-normalization control rules out one explanation. It does not identify a neural circuit or exclude all forms of spectral/context dependence.
Observed classifier dependence cannot establish actual bird competition, temporal avoidance, or bias in every ecological study.


Data, credits and reproducibility
Audio and annotations:  Martin et al., CEB v1.1 , published 6 August 2026. Focal audio retains its individual Xeno-Canto licenses and recordist credits; it must not be relabeled under the annotation license. Model:  BirdNET-Analyzer v1.5.1, v2.4 FP32 checkpoint  (SHA256 55f3e405…165e4c), with source sigmoid convention. The modified diagnostic checkpoint is generated locally, not redistributed.
 All selected recordings, original links, recordists and licenses  ·  Exact input hashes  ·  Download reproduction package Included focal species and source counts
  
    

       species 
       passages 
       recordists 
    
  
  
    

       Carduelis carduelis 
       3 
       3 
    
    

       Certhia brachydactyla 
       4 
       3 
    
    

       Certhia familiaris 
       4 
       4 
    
    

       Chloris chloris 
       3 
       3 
    
    

       Coccothraustes coccothraustes 
       4 
       4 
    
    

       Corvus corax 
       4 
       4 
    
    

       Cuculus canorus 
       3 
       3 
    
    

       Cyanistes caeruleus 
       4 
       4 
    
    

       Erithacus rubecula 
       3 
       3 
    
    

       Fringilla coelebs 
       3 
       3 
    
    

       Fringilla montifringilla 
       4 
       4 
    
    

       Lophophanes cristatus 
       3 
       3 
    
    

       Phylloscopus trochilus 
       3 
       3 
    
    

       Picus viridis 
       3 
       3 
    
    

       Poecile palustris 
       4 
       4 
    
    

       Sylvia atricapilla 
       4 
       4 
    
  

Run instructions, dependencies, scripts, frozen plans, full result tables and vector figures are included in the package. A  full isolated CPU rerun  reproduced all 23 result tables byte for byte and matched all 10 scientific summaries. It used the original installed interpreter with a separate output tree and cache; this is a computational reproduction, not an external biological replication. It downloads about 3.2 GB of public inputs and runs on CPU. The completed analyses used single-thread inference on this server; no large-compute handoff was required. The  append-only journal  retains failed leads and the first completed study.
 All focal mixture outcomes  ·  Primary statistics  ·  Normalization intervention  ·  Ambient control  ·  Natural-window scores  ·  Scale sensitivity 


Closest prior work
 Clark et al. 2023, soundscape composition  — Observed soundscape classes and errors; broad biophony class did not resolve interference among animal vocalizations. Uses sigmoid multi-label outputs and discusses overlapping sounds.
 Denton, Wisdom and Hershey, ICASSP2022  — Unsupervised separation of overlapping songs/noise improves classification; pooling original and separated channels is prior art.
 Sasek et al. 2024  — Synthetic overlapping background/target mixtures; site-specific source separation and BirdNET robustness tests for Golden-cheeked Warbler.
 Separating Overlapping Birdsongs Enhances the Reliability of Avian Vocal Activity Analysis, 2026  — 30-species synthetic overlapping mixtures, source separation and ecological vocal-activity analyses; full text examined.
 Multi-species Mixing for Weakly Supervised SED Under Domain Shift  — Synthetic multi-species training augmentation with BirdNET embeddings and linear classification, evaluated on five anurans.
 How to analyse overlapping sounds in the marine environment using supervised multi-label classification, 2026  — Naturally co-occurring marine sounds; training co-occurrence structure and architectural interference.
 Richmond, Hines and Beissinger2010  — Conditional two-species occupancy models already account for false absences and detection dependence. Do not claim a new ecological statistical principle.
 Tobler et al.2019  — JSDMs with imperfect detection; detection heterogeneity can distort inferred species correlations.
 Metcalf et al.2022  — Acoustic classification error heterogeneity and contextual correction; recommends validation for ecological predictor analyses.
 Chambert et al.2018  — Two-species occupancy models accommodating misidentification and non-detection.
 Briggs et al., Acoustic classification of multiple simultaneous bird species  — Segmentation and multi-instance/multi-label recognition are longstanding; no novel segmentation algorithm claimed.
 Stowell et al., first Bird Audio Detection challenge  — Maskers/distractors and background-domain robustness were already reported for binary bird detection.
 Two-stage HuBERT fine-tuning,2026  — Synthetic overlapping bird mixtures for training and recognition; does not establish this fixed-model disjoint-time counterfactual.
 Cole et al.2022  — Longer acoustic sampling can yield occupancy results similar to manual annotation; the new short-window result must not be extrapolated to invalidate occupancy monitoring.
 Chiatante and Canestrelli 2026, Temporal dynamics of animal vocal behavior affect semi-automatic identifications  — Primary study of month-specific BirdNET confidence calibration and vocal phenology. Full methods examined; does not manipulate disjoint passages within a fixed three-second input.
 Cornell Lab Bird Academy: BirdNET developer discussion of multiple simultaneous species  — The developer explicitly acknowledges difficulty when many birds overlap. General multi-bird interference is not a discovery here.
 Fine-Tuning BirdNET for the Automatic Ecoacoustic Monitoring of Bird Species in the Italian Alpine Forests, Information 2025, 16, 628  — Fine-tuned BirdNET: observational spectrogram examples of confidence degradation with environmental interference and simultaneous species. Full text and figures 6–7 examined; figure 7 shows temporal overlap, without same-component disjoint-time baselines, common normalization or independent-presence enumeration.Study 02 · Completed computational analysis, provisional novelty assessment. Original source authors receive credit for all recordings, annotations and model development.
