STUDY 02 · COMPUTATIONAL BIOACOUSTICS · 19 SEPTEMBER 2026

Non-overlapping bird calls suppress detections in BirdNET v2.4

A controlled context intervention with distinct focal recordings, natural backgrounds, and field-recording ablations.

Executed research note · novelty provisionally supported by two literature reviews · not peer reviewed

The finding

Adding another bird in a separate part of a three-second window made 54.5% of otherwise detectable target appearances fall below a score of 0.5. The target audio itself was unchanged. The larger experiment used 56 passages from 16 species, recorded in 56 distinct files by 52 credited recordists.

This is a measured failure under a controlled intervention. It is not an estimate that BirdNET misses this fraction of birds in the wild.

Abstract

Multi-species acoustic classifiers can introduce dependence between detections. To isolate this dependence from simultaneous acoustic overlap, we placed annotated 1.2-second single-species passages in disjoint intervals of a three-second input, separated by 0.3 seconds. Target components had identical gain and position in baseline and combined inputs. A focal-recording extension prospectively specified in the journal produced 3,198 losses among 5,872 baseline-eligible target appearances (54.5%; source-file bootstrap interval 49.0%–59.5%). Both orders were tested, and every included species was affected. The result survived common waveform normalization and continuous natural background. A secondary temporal ablation of 364 unmodified field windows found conditional losses of 26.4%. Exact enumeration of independent calling states generated a negative detected association (φ = -0.472) despite true independence. These results motivate auditing joint detectability before interpreting fine-scale classifier co-detections as bird interactions.

2026-09-19T08:07:08.711541 image/svg+xml Matplotlib v3.6.3, https://matplotlib.org/ 0 10 20 30 40 50 60 70 Loss of detection among eligible target appearances (%) Soundscape pilot 9 species Soundscape validation 4 species Focal extension 16 species Fixed normalization 300 focal pairs Continuous background 200 focal pairs × 3 Natural field context secondary ablation Same species focal control 55.5% 43.4% 54.5% 57.1% 41.1% 26.4% 0.0% A second, disjoint-time passage can suppress a detection
All bars describe conditional detection loss at score 0.5. Samples and eligibility differ between panels: pilot 9 species; held-out soundscapes 4; focal extension 16. Fixed-normalization and ambient controls use fixed subsets of focal pairs. The natural-window ablation is secondary. Bars are not independent population estimates.

Question and experimental logic

In a detector with reliable multi-label recognition, adding another identifiable call need not eliminate evidence already present for a target. We tested whether it nevertheless does. The primary question concerns the classifier’s observation process, not whether birds compete or avoid each other.

Baseline A A · 0.15–1.35 ssilence
Baseline B silenceB · 1.65–2.85 s
Combined identical Aidentical B

All source excerpts were mono, resampled to 48 kHz, mean-centered and faded over 5 ms at their edges. They were normalized to a common RMS, followed by one global clipping-safe scaling factor and a factor of 0.5 applied equally to all components. The combined waveform never clipped. We tested both temporal orders and simultaneous mixtures. Scores use the model’s 6,522 outputs with the source implementation’s clipped sigmoid; no geographic filter, custom classifier, top-k truncation, or retraining was used.

Primary eligibility required score ≥0.5 in both time positions. This deliberately restricts inference to calls the model can recognize alone. Candidate passages, recording limits and advancement rules were recorded before their mixture outcomes. The score is a model output, not a calibrated probability of species presence.

Discovery history and frozen gates

StageSampleOutcome
015: isolated-note feasibility37 candidates; only 1 species qualifiedFailed before mixture testing
018: single-species passages98 candidates → 18 passages/9 speciesNon-overlap losses 53.6% and 57.5%; pilot gate passed
018: untouched soundscape files178 candidates → 16 passages/4 speciesEffect replicated at 43.4%; failed the frozen 8-species diversity gate
019: separate focal source360 candidates → 56 passages/16 speciesPassed ≥8 species, ≥24 passages, ≥20% losses and ≥4 affected species

The focal extension was separately specified after the soundscape diversity shortfall. It was not represented as success of that failed gate. Its clips came from the CEB focal subset, while the earlier clips came from soundscape recordings. Xeno-Canto recordings may have been used during BirdNET training, so this is not a held-out model-accuracy benchmark. The counterfactual manipulation and source extension are the units of validation.

The journal also retains two closed plant hypotheses from this round. A copied-output-path error during the bird pilots was recovered from git; separated reruns reproduced the same outcomes. Recovery record. Every artifact of the first completed plant study remains unchanged.

Results

The focal experiment comprised 1,468 heterospecific recording pairs, both orders, and two target appearances per order: 5,872 target appearances, not 5,872 independent birds. Loss was 54.6% in one order and 54.3% in the reverse order. Simultaneous loss was 55.7%. Equal weighting of species gave 53.1% non-overlap loss. A two-endpoint source-file bootstrap with 2,000 resamples gave the interval above; it is conditional on this set of selected species and detectable passages.

Independent passages from the same species produced 0 losses among 288 non-overlapping target appearances. This control shows that adding another passage or more total sound does not inevitably reduce the target score. It does not prove which learned features cause heterospecific suppression.

2026-09-19T08:07:09.035808 image/svg+xml Matplotlib v3.6.3, https://matplotlib.org/ Carduelis carduelis Certhia brachydactyla Certhia familiaris Chloris chloris Coccothraustes coccothraustes Corvus corax Cuculus canorus Cyanistes caeruleus Erithacus rubecula Fringilla coelebs Fringilla montifringilla Lophophanes cristatus Phylloscopus trochilus Picus viridis Poecile palustris Sylvia atricapilla Added species (in the other time interval) Carduelis carduelis Certhia brachydactyla Certhia familiaris Chloris chloris Coccothraustes coccothraustes Corvus corax Cuculus canorus Cyanistes caeruleus Erithacus rubecula Fringilla coelebs Fringilla montifringilla Lophophanes cristatus Phylloscopus trochilus Picus viridis Poecile palustris Sylvia atricapilla Target species Focal recordings: directed detection loss (%) 0 20 40 60 80 100 Conditional loss (%)
Rows are the species whose detection is evaluated; columns are the species added elsewhere in the same window. Each cell pools multiple source-recording combinations and both orders. Blank diagonal cells are the separate conspecific control. This is a detector-response matrix, not a map of ecological competition.

Alternative explanations and controls

Waveform normalization

The pinned model begins with per-window min/max waveform normalization. We exposed the tensor immediately after its seven normalization operations and retained all 549 model buffers and every later operation. Feeding the original normalization into this modified graph reproduced all tested baseline logits exactly. For 300 fixed-hash focal pairs, we then used the same original-mixture extrema to normalize both baselines and the mixture. Target samples after normalization were bit-identical. Loss remained 57.1% across 1,196 eligible target appearances. The original and modified models’ mixture scores agreed within 7.2e-08. Thus, waveform min/max rescaling alone cannot explain the result; we do not attribute it to a specific later layer.

Continuous natural background

Three annotation-free three-second background passages were selected from distinct soundscape files before mixture scores. Each was scaled to RMS 0.0015 and included unchanged in background-only, single-bird and two-bird inputs. Among 2,400 target appearances from 200 fixed pairs and both orders, 2,139 remained detectable alone and absent in the background-only prediction. Conditional loss was 41.1%, with 40.0–42.1% across backgrounds. This argues against zero padding alone as an explanation. Annotation-free does not guarantee complete acoustic silence.

Natural field-window ablation

The metadata screen found 366 windows; two involved a species absent from the model label list, leaving every model-covered three-second grid window with exactly two annotated species whose bouts were separated by ≥0.1 s: 364 windows in 88 files. Complementary masks transitioned only inside the annotated gap; their components sum to the original recording and preserve the target’s annotated samples. Among 140 individually eligible appearances, loss in the original full context was 26.4% (file-bootstrap interval 18.3%–35.0%). With shared original-window normalization it was 25.4%. The stricter subset where both species were independently detectable contained 18 appearances and had 44.4% loss.

Across all 728 target appearances, there were 37 threshold losses and 29 gains. Median score change across all appearances was +0.0033. This secondary test reuses soundscape data and removes background together with the neighboring passage; it supports context dependence without isolating recognized bird identity as the sole cause.

2026-09-19T08:07:09.625221 image/svg+xml Matplotlib v3.6.3, https://matplotlib.org/ 0.0 0.2 0.4 0.6 0.8 1.0 Target score after neighboring context is removed 0.0 0.2 0.4 0.6 0.8 1.0 Target score in the original field window Natural recordings: 364 annotated windows
Every selected natural-window target is shown, including gains. Points below the diagonal receive a lower score in the original field context than after the disjoint context is removed. Dashed lines mark 0.5; this is not a field-prevalence estimate.

What it means for co-detection

For each eligible pair and order, we enumerated four states—neither caller, A alone, B alone, both—each with probability 0.25. Calling is therefore exactly independent. Here, φ is the ordinary correlation between two binary detection indicators: negative values mean they are detected together less often than independence predicts. All 2,936 focal pair-orders also passed the specificity check: neither species was detected in silence or in the other’s single-source input. The classifier nevertheless produced a pooled detected φ of -0.472; 90.8% of pair-orders had negative detected association. No random simulation or significance test is needed for this finite enumeration.

The size of this artificial association depends on calling frequency and analysis scale. In an explicitly hypothetical independent-window repeat model, reducing per-window calling probability to 0.05 gives φ≈−0.046 at 3 s, and aggregating those detections to 5-minute presence gives φ≈−0.0015. The experiment therefore does not invalidate long-duration occupancy studies. It identifies an observation-process hazard most directly relevant to short-window vocal associations.

2026-09-19T08:07:09.398232 image/svg+xml Matplotlib v3.6.3, https://matplotlib.org/ B absent B present A absent A present 25.0% 25.0% 25.0% 25.0% True presence: φ = 0 B absent B present A absent A present 29.5% 33.1% 35.1% 2.3% Detected: φ = -0.472 0.0 0.1 0.2 0.3 0.4 0.5 True calling probability / species / 3 s −0.4 −0.3 −0.2 −0.1 0.0 Detected association φ Toy model: bias depends on event rate 3 s bin 15 s bin 60 s bin
Left: independent true calling states. Middle: the classifier’s pooled detected states at calling probability 0.5. Right: analytic sensitivity under an independent-window toy model; these calling rates and aggregation assumptions were not estimated from wild birds.

Novelty and practical value

Previous work establishes environmental interference, overlapping-source separation, multi-label recognition, heterogeneous classifier errors and detection-aware occupancy models. None of those ideas is claimed as new. Two documented searches, repeated after the larger test, did not locate a directly equivalent combination of disjoint-time bird counterfactuals, matched waveform-normalization control and explicit independent-presence association enumeration for BirdNET v2.4.

The original contribution is a reproducible, narrowly scoped empirical result and stress-test resource. It can be used to audit a detector before interpreting fine-scale co-detections. Novelty remains provisional pending expert review; the note has not been submitted or peer reviewed. No deployable correction or new separation algorithm is claimed. Simply taking maxima over extra crops can create false positives and was not validated as a remedy here.

First novelty review and exact queries · Second novelty review

Limitations

Data, credits and reproducibility

Audio and annotations: Martin et al., CEB v1.1, published 6 August 2026. Focal audio retains its individual Xeno-Canto licenses and recordist credits; it must not be relabeled under the annotation license. Model: BirdNET-Analyzer v1.5.1, v2.4 FP32 checkpoint (SHA256 55f3e405…165e4c), with source sigmoid convention. The modified diagnostic checkpoint is generated locally, not redistributed.

All selected recordings, original links, recordists and licenses · Exact input hashes · Download reproduction package

Included focal species and source counts
species passages recordists
Carduelis carduelis 3 3
Certhia brachydactyla 4 3
Certhia familiaris 4 4
Chloris chloris 3 3
Coccothraustes coccothraustes 4 4
Corvus corax 4 4
Cuculus canorus 3 3
Cyanistes caeruleus 4 4
Erithacus rubecula 3 3
Fringilla coelebs 3 3
Fringilla montifringilla 4 4
Lophophanes cristatus 3 3
Phylloscopus trochilus 3 3
Picus viridis 3 3
Poecile palustris 4 4
Sylvia atricapilla 4 4

Run instructions, dependencies, scripts, frozen plans, full result tables and vector figures are included in the package. A full isolated CPU rerun reproduced all 23 result tables byte for byte and matched all 10 scientific summaries. It used the original installed interpreter with a separate output tree and cache; this is a computational reproduction, not an external biological replication. It downloads about 3.2 GB of public inputs and runs on CPU. The completed analyses used single-thread inference on this server; no large-compute handoff was required. The append-only journal retains failed leads and the first completed study.

All focal mixture outcomes · Primary statistics · Normalization intervention · Ambient control · Natural-window scores · Scale sensitivity

Closest prior work