FIELD NOTES / BIOINFORMATICS DISCOVERY LOOP

Finding a contribution in biology + computing

Find and execute a valuable, plausibly original research contribution using public data and programming. Interests: genetics, plant design, birds and birdwatching. Failed and overlapping ideas remain visible.

Discovery resumed: seeking a second original study. First completed study and all original results preserved.

Started 2026-09-19 · Updated 2026-09-19 04:16 UTC

What counts as success

Read the completed plant promoter robustness study

ATTEMPT 001

Are source recordings independent in a nocturnal bird-call benchmark?

Closed — no direct train/test overlap detected

Question

Do publicly identifiable shared source recordings cross the NBM train/test split, and could this invalidate an otherwise independent comparison?

Why it might matter

A direct audit is cheap and could yield a useful correction if genuine overlap exists. The authors already address sampling bias and deliberately select later uploads for testing, so general concerns about leakage are not new.

Minimal test

Retrieve metadata; compare recording identifiers, filenames, source origin and available contributor/session fields. If no direct overlap is found, reject this narrow lead unless independent evidence supports audio duplication.

First novelty check

Initial reading: source paper explicitly discusses class imbalance, contributor imbalance, train/test upload-date separation, possible BirdNET pretraining overlap, and focal-versus-passive domain shift. These are established issues, not our contributions.

Decision and next step

Reject the narrow direct-leakage hypothesis. This test cannot exclude cropped, re-encoded or same-session recordings; no broad claim of independence is made.

Observed results

Evidence and outputs

Sources and closest prior work

Exact literature search queries
  • "Nocturnal Bird Migration" dataset leakage duplicates evaluation
  • "NBM" "Airale" dataset bias annotation
ATTEMPT 002

Can duplicate audio reveal annotation uncertainty?

Closed — annotation provenance insufficient

Question

Can multiple annotations of repeated NBM audio quantify uncertainty in detection benchmarks?

Why it might matter

The directory audit exposed repeated audio candidates with differently named annotation files.

Minimal test

Read the annotation files of all ten candidate groups and distinguish mere formatting, revised labels, missing files and independently drawn boundaries.

First novelty check

Inter-annotator agreement and boundary uncertainty in bioacoustics are already established topics. Accidental repeated files would only support a new result if their annotation provenance and independence were known.

Limitations

No annotator identity or independent repeated-labeling design is provided. CRC identity is only a screen. Ten convenience-sampled groups cannot represent the full dataset.

Decision and next step

Retain as a limited data-quality finding. No general uncertainty estimate and no novelty claim. Large WAV pairs were not downloaded merely to confirm this unsuitable lead.

Observed results

Evidence and outputs

Sources and closest prior work

Exact literature search queries
  • bioacoustic bounding box annotation inter annotator IoU repeatability bird detection
  • "NBM" "duplicate" birds dataset
  • "bioacoustic" "inter-annotator" bounding boxes
ATTEMPT 003

Does changing song-versus-call composition distort automated activity estimates?

Closed — pilot failed advancement gates

Question

Within the same bird species and soundscape, does BirdNET recall differ between songs and contact calls enough to distort activity estimates when their mixture changes?

Why it might matter

A detector can recognize a species well overall yet undersample one behavior. A useful contribution would quantify the within-species, within-recording effect across many species and determine how much duration/context explains it.

Minimal test

Use the new CEB soundscape annotations to count species with both songs and contact calls. Run a frozen BirdNET v2.4 model on a bounded stratified sample of eligible three-second windows, with no geographic prior. Compare per-species recall at fixed thresholds 0.1 and 0.5. Exclude windows containing both song and contact-call annotations of the focal species. Frozen implementation: SHA256 seed 20260919 assigns whole files to pilot (hash mod 3 == 0) or validation. Keep species with at least three paired files overall and at least one in each phase; at most three pilot files/species. Select one window per behavior/species/file using event-ID hashes. Exclude other annotated behaviors of the focal species from each window. No padding; source windows remain inside the three-minute file. Primary threshold 0.1; 0.5 is sensitivity only. These paired source files are not independent identified birds.

First novelty check

Broad call-type evaluation is not new: Weldy et al. (2025) compare sonotypes, and call-type classifiers exist. The narrower candidate is a comparative, same-recording audit of behavior-mixture bias and a practical calibration/abstention diagnostic. Need determine whether this distinction is substantive after the pilot. Clapp et al. (2026) explicitly identify time-varying recall as an open gap while studying precision; a 2022 researcher discussion already connects song/call mixtures to phenology. The broad problem is established. Our possible contribution is empirical estimation and validation of within-recording behavior-specific recall, not originating the concern.

Larger test / frozen decision gates

Advance only if at least ten species have usable paired data and the median species-level recall gap is at least 0.15 at a prespecified threshold. Larger validation must use disjoint recording groups, resample species and recordings rather than calls, examine duration/context and source effects, and seek external replication. Do not call a threshold-free score difference a recall effect.

Limitations

No seasonal or population-abundance inference is available from these labels alone. CEB test data are a convenience sample; independence from original BirdNET training needs verification. Annotation type is expert-assigned behavior, not individual sex/age.

Decision and next step

Do not advance to validation or claim a general song advantage. Held-out recordings have not been scored. A future substantially different hypothesis must be logged explicitly and cannot reuse these pilot results as confirmation.

Observed results

Evidence and outputs

Sources and closest prior work

Exact literature search queries
  • BirdNET accuracy "begging" calls song
  • BirdNET vocalization type performance call song bias evaluation
  • "BirdNET" "vocalization type" performance bias phenology
  • "Central European Birds" vocalization types benchmark
  • "BirdNET" "songs and calls" "recall"
  • "BirdNET" "song" "contact call" performance
  • "Passive Acoustic Data as Phenological Distributions" data call song
ATTEMPT 004

Do bird mitochondrial coding annotations reproduce their deposited proteins?

Closed — control passes; closely covered by prior work

Question

Can a small annotation audit distinguish legitimate mitochondrial recoding from incorrect coding annotations in public bird genomes?

Why it might matter

Bird mitochondrial ND3 sometimes contains an extra nucleotide handled by recoding. Incorrect treatment can create misleading proteins in comparative analyses. A useful project would need a specific new failure mode beyond established mitogenome quality audits.

Minimal test

Download a well-characterized chicken mitogenome from ENA. Parse CDS features, translation tables and exceptions; compare extracted coding sequences to deposited protein translations. Distinguish incomplete terminal stops and programmed frameshifts from defects.

First novelty check

Andreu-Sánchez et al. (2021) analyzed tens of thousands of ND3 sequences, including recoding-site evolution and annotation issues. Sangster & Luksenburg (2021) systematically investigated avian mitogenome errors and downstream reuse. This proposed general audit lacks a sufficiently distinct question.

Decision and next step

Stop this candidate after the control and first novelty check; do not perform a larger scan merely to rediscover known problems.

Observed results

Evidence and outputs

Sources and closest prior work

Exact literature search queries
  • avian mitochondrial genome ND3 frameshift annotation errors 2024 2025 2026
  • "Sharp Increase of Problematic Mitogenomes" data github
ATTEMPT 005

Can seasonal species filters censor early migrants?

Closed — direct prior study found

Question

Does a date-dependent species allowlist introduce an artificial detection boundary for early or late migrating birds even when the acoustic evidence is unchanged?

Why it might matter

A seasonal prior can improve identification, but phenology studies target deviations from expected timing. A substantive contribution would quantify documented true-positive losses and resulting timing bias, beyond the familiar warning that priors can miss rare birds.

Minimal test

Run the official BirdNET v2.4 geographic model at a fixed European location for all 48 modeled weeks. Record default-threshold allowlist crossings for common migratory birds and compare with the year-round mode. This is a counterfactual model diagnostic, not evidence that a bird was present on any particular date.

First novelty check

A June 2026 preprint by Pérez-Granados et al. directly compares no filter, spatial filter and spatiotemporal filter across worldwide annotated soundscapes. Its discussion specifically warns about missing early and late migrants. This materially overlaps the question.

Larger test / frozen decision gates

Only advance if meaningful week-dependent exclusions occur and the exact research question is not already answered. Larger validation needs dated, geolocated, independently annotated recordings and must distinguish removing false positives from losing genuine early/late detections. Model-only calendar curves cannot establish real ecological bias.

Decision and next step

Reject the broad seasonal-censoring study as insufficiently new; no expensive soundscape analysis.

Observed results

Evidence and outputs

Sources and closest prior work

Exact literature search queries
  • BirdNET seasonal geographic filter early arrival phenology bias
  • BirdNET range filter climate change migration detection prior circularity
  • "BirdNET" "week" "bias" phenology
ATTEMPT 006

Is a year-round species list guaranteed to contain every weekly list?

Closed — related behavior already documented; small pilot anomaly

Question

Does the BirdNET year-round filter obey the expected nesting of all weekly allowlists?

Why it might matter

A year-round filter is trained to predict peak weekly occurrence, yet unconstrained neural predictions might violate that relationship. This would be a sharper question than the seasonal-filter warning.

Minimal test

Compare official week=-1 outputs with the maximum of weeks 1–48 at five fixed locations, using the default 0.03 cutoff.

First novelty check

BirdNET-Analyzer issue #211 already reports year-round lists differing from the union of weekly lists. The maintainer explains that the -1 mode is trained separately against maximum weekly checklist frequencies. Our reverse-direction exceptions are a narrow consequence of the same unconstrained approximation; no material ecological impact has been established.

Limitations

Five model coordinates, no independently confirmed bird occurrences, and two small cutoff crossings. No evidence of missed real recordings or precision improvement from a proposed fix.

Decision and next step

Retain the reproducible observation but do not call it a new research contribution or enlarge it solely to accumulate discrepancies.

Observed results

Evidence and outputs

Sources and closest prior work

Exact literature search queries
  • "BirdNET" "year-round" "maximum"
  • "BirdNET" "year-round" "missing" species
  • "BirdNET" "species lists" "nested"
  • GitHub issues search: repo:birdnet-team/BirdNET-Analyzer "year round"
ATTEMPT 007

Do plant promoter benchmarks separate independent DNA sequences?

Closed — exact overlap too small to support lead

Question

Do public plant promoter benchmarks contain exact or reverse-complement-identical sequences across train/test splits, or duplicated parent constructs that permit target memorization?

Why it might matter

Plant design increasingly depends on sequence-to-function models. A concrete, consequential benchmark flaw could be valuable to document, but generic concerns about genomic train/test leakage are established and do not themselves constitute novelty.

Minimal test

Download the public PDLLMs core-promoter classification splits and the PGB tobacco-leaf promoter-strength splits. Audit exact sequences, reverse complements, duplicated labels and available sequence IDs. If there is a material overlap, test a training-only lookup baseline before any expensive model training.

First novelty check

Initial source reading: AgroNT uses held-out chromosomes for several genomic tasks and adopts the published Jores promoter-strength splits. Respect each benchmark’s intended generalization target; related sequences alone are not proof of inappropriate leakage. Generic homology-induced evaluation leakage is studied directly by Rafi et al. (2025 preprint). The original Jores study explicitly used random promoter holdouts; finding related promoters would not alone contradict that design.

Larger test / frozen decision gates

Advance only if an overlap or shortcut has a measurable predictive consequence and contradicts a stated evaluation target; require new holdouts or group-aware splits and re-check whether the specific issue has already been reported.

Limitations

Reverse-complement identity does not guarantee the same directional promoter function. Exact overlap counts alone are not a measure of model optimism.

Decision and next step

Close this narrow exact-overlap hypothesis. No substantive overall evaluation effect is established. Homology and nearest-neighbor leakage require a different analysis; absence of exact copies does not exclude them.

Observed results

Evidence and outputs

Sources and closest prior work

Exact literature search queries
  • plant genomic benchmark sequence data leakage AgroNT PlantCaduceus 2026
  • "AgroNT" "leakage"
  • "plant-genomic-benchmark" "duplicate"
  • "plant-multi-species-core-promoters" leakage
  • "PDLLMs" leakage duplicate benchmark
ATTEMPT 008

Does arbitrary window alignment change recognition of a complete short bird call?

Closed — established overlap/context problem

Question

How sensitive are BirdNET scores to temporal translation of a short vocalization that remains fully inside its three-second input?

Why it might matter

If small changes in the analysis grid alter complete-event detection, comparing counts from differently aligned files can add avoidable uncertainty. A new contribution would need a measured practical consequence and a validated correction beyond standard overlapping windows.

Minimal test

Reuse the previously frozen pilot sample, selecting only events at most 1.2 seconds long. Circularly shift each identical three-second waveform by −0.6, −0.3, 0, +0.3 and +0.6 seconds. The centered target event remains intact. Also shift real source-window starts by those offsets. Compare threshold-crossing fractions at 0.1 (primary) and 0.5 (secondary). Advance to novelty review if at least 20% of eligible events cross 0.1 across shifts.

First novelty check

Prior BirdNET/Chirpity implementations already address segmentation-dependent false positives through overlapping-window consensus. Translation sensitivity of CNNs and bird-audio window-size effects are established. Need a distinct practical contribution before larger validation. The 2024 Chirpity discussion describes the boundary-confusion mechanism and an implemented adjacent-window rule; the global BirdNET assessment cites tests of overlap and vocalization duration. Our observations are too close to these established concerns.

Limitations

Circular shifts can cut background events and create a boundary seam. Real source shifts change contextual audio. These are diagnostics, not yet a pure causal attribution or field error rate; pilot selection is inherited from an earlier question.

Decision and next step

Do not enlarge this generic translation diagnostic. The known overlap/consensus approaches already address the practical concern, and our pilot cannot isolate a new failure mechanism. Preserve scores and move to a different question.

Observed results

Evidence and outputs

Sources and closest prior work

Exact literature search queries
  • BirdNET "time shift" prediction
  • BirdNET "window" "position" detection confidence overlap
  • bird sound classification temporal translation invariance window alignment
ATTEMPT 009

How reliable are the expression targets in plant promoter benchmarks?

Closed broad noise question — follow-up on coding barcodes

Question

Do low-precision expression measurements change evaluation or design rankings in reused plant promoter benchmarks?

Why it might matter

Repeated floating-point target values in the downloaded promoter data motivate checking their experimental origin. Count-based reporter measurements have unequal precision; a useful study would connect that uncertainty to design decisions or model evaluation, beyond merely observing measurement noise.

Minimal test

Read original measurement-processing code and compare benchmark labels with source measurements and biological replicates. First determine whether repeated labels reflect low counts, normalization or something else; do not infer low coverage from ties alone. Quantify replicate disagreement across a fixed Arabidopsis leaf pilot before selecting a larger test.

First novelty check

Heteroscedastic MPRA measurements, count-ratio bias and barcode-related uncertainty are established. MTSA (Lee et al., 2021) directly models sequence-specific barcode effects. The source study already reports replicate correlations. No general novelty claim from re-quantifying these facts.

Limitations

Noise-aware MPRA analysis already exists. Novelty would require a distinct demonstrated consequence in plant design benchmarks. Expression differences between leaf and protoplast assays confound species, tissue and assay system.

Decision and next step

Do not claim generic measurement-noise discovery. The source construct places a 12-base random barcode in the first four codons after the initiation codon, motivating a separate specific hypothesis about positional barcode sequence effects.

Observed results

Evidence and outputs

Sources and closest prior work

Exact literature search queries
  • "plant" "promoter" "benchmark" "measurement noise"
  • "Jores" "promoter" "low" "counts"
ATTEMPT 010

Do short coding barcodes systematically alter plant promoter readouts?

Closed — below frozen pilot thresholds

Question

In Plant STARR-seq, does the identity and position of bases in a four-codon barcode predict RNA/DNA enrichment among barcodes attached to the same promoter?

Why it might matter

The barcode follows the start codon rather than lying in a 3′ UTR. A position-specific, reproducible effect could implicate assay architecture and guide barcode design. We cannot infer a translation or RNA-decay mechanism solely from count associations.

Minimal test

Use Arabidopsis library, tobacco leaf, enhancer present, dark, replicate 1. Reproduce original RNA/DNA count cutoff ≥5 and WT/full-length mapping filter. Fit within-promoter effects of one-hot bases at all 12 barcode positions, holding out complete promoter groups by fixed SHA256 hash. Compare held-out residual prediction with zero. Advance if held-out R² is at least 0.01 or the fitted first-barcode-base contrast spans at least 0.1 log2 units; validate in an independent biological replicate and another assay system if novelty survives.

First novelty check

General sequence-specific barcode bias and correction are known from MTSA. The candidate is the specific positional behavior of short translated barcodes in plant assays and its reproducible practical impact. This distinction must be tested against the full prior literature, not inferred from the different organism alone.

Limitations

Original count truncation, promoter sequence variants, barcode misassignment and sequence-dependent sequencing/RT efficiencies are alternative explanations. Shared barcodes across replicates are not independent randomized interventions. Source code clarifies that the metadata mutations field lists intentional restriction-site-removal edits, while variant=WT identifies a match to the designed construct; nonempty mutations is not evidence of sequencing error.

Decision and next step

Do not enlarge the study by adjusting the gate after seeing a near miss. Small sequence effects may exist, but this pilot does not establish a practically consequential new bias or a translation mechanism.

Observed results

Evidence and outputs

Sources and closest prior work

Exact literature search queries
  • plant STARR seq barcode sequence bias promoter Jores
  • "barcode" "sequence" "bias" MPRA RNA stability
ATTEMPT 011

Are bird-audio predictions invariant to computational batching?

Closed — batching invariance passes

Question

Does an identical waveform receive materially different BirdNET scores when batched with silence, a different recording, or an amplified recording?

Why it might matter

Batching should alter throughput, not the meaning of individual samples. A failure could reveal an actionable normalization issue rather than a generic ecological limitation.

Minimal test

Check dynamic-batch support in the frozen official TFLite model. Compare identical waveform logits alone and in batches of two with zero, independent and 10×-amplitude distractors. A material failure requires a score difference greater than 0.01 or a change at score threshold 0.1, beyond numerical roundoff.

Limitations

This tests one frozen acoustic model and inference backend. It is not a comparison of model versions or a biological effect.

Decision and next step

Reject the batch-dependence hypothesis for this model/backend. No reason for a larger experiment.

Observed results

Evidence and outputs

Sources and closest prior work

Exact literature search queries
ATTEMPT 012

Does barcode count filtering make promoter ranks depend on sequencing depth?

Executed — validated resource-specific robustness result

Question

When the same RNA count table is computationally thinned, does the published per-barcode RNA≥5 filter and median-ratio estimator systematically change promoter rankings?

Why it might matter

This is a specific ascertainment question prompted by source-code inspection, distinct from ordinary replicate noise. Low-RNA barcodes disappear as depth falls; a retained-only median may shift as the surviving barcode set changes. A useful contribution requires a material and avoidable effect on current plant-design benchmark targets.

Minimal test

Use Arabidopsis leaf replicate 1 raw counts. Keep the original WT/full-length associations and original DNA≥5 rule. Compare full counts with seeded binomial RNA thinning to 50% and 25%; recompute the original median of retained log2 RNA/DNA ratios and correct the known dilution. Contrast with a fixed-barcode pooled-count ratio retaining RNA zeros. Analyze common promoters with ≥10 DNA-supported barcodes, report retention and all rank shifts. Advance only if median absolute activity shift exceeds 0.25 log2 units or rank correlation drops below 0.95 at 25% depth.

First novelty check

mpralm (2019) derives lower bias for aggregate ratios than mean log ratios; a 2025 human-assay comparison recommends DNA-based filtering. These establish the statistical rationale. Searches so far did not locate a direct depth-intervention audit of the Jores/PGB plant promoter targets. This is provisional novelty of an empirical benchmark consequence, not a new estimator.

Larger test / frozen decision gates

Before inspecting further outcomes: reconstruct normalized original labels exactly using barcode controls and two-replicate means; verify against deposited source data. Validate with At leaf replicate 2 and Zm leaf replicate 1, with Zm leaf replicate 2 for complete benchmark reproduction. Use five seeded thinnings at 50% and 25% RNA depth. Compare retained median, all-barcode median with 0.5 pseudocount, retained pooled counts and all-barcode pooled counts. Primary identical promoter set: >=10 DNA-supported barcodes, full-depth observed in all methods, and retained in the thin original estimator; separately report dropout. Preserve source cutoff DNA>=5; sensitivity DNA>=20. Advancement requires original-estimator quarter-depth Spearman<0.95 in both new validation libraries and pooled-all improvement >=0.05; test pooled-retained to isolate censorship from aggregation. Evaluate actual independent-replicate agreement before recommending any corrected label set. Seek published-model consequences using original held-out predictions. Clarification before validation outcomes: exact historical reconstruction is established by attempt 013; depth tests use robust whitespace counts and independently reconstructed corrected values. PGB targets are a distinct snapshot and cannot yet be called corrected by these values.

Second novelty check

MPRAsnakeflow (2025) already evaluates count/barcode downsampling and replicate agreement in human libraries. mpralm (2019) already explains advantages of pooling counts. A general downsampling tool or new pooling estimator would not be novel. The remaining contribution is a reproducible, resource-specific audit of plant promoter target robustness, coupled to verified historical parser effects and separated target snapshots. Searches did not locate that applied result; novelty remains provisional pending expert review.

Limitations

Computational thinning only probes read-sampling depth, not transfection noise or between-library reproducibility. Fixed observed counts are not biological ground truth.

Decision and next step

Completed full execution as study 014, using all twelve model-relevant libraries and six replicate pairs; the result supports a computational resource note, with provisional novelty and explicit limits.

Observed results

Evidence and outputs

Sources and closest prior work

Exact literature search queries
  • "MPRA" "filtering" "RNA" "bias" threshold
  • "Jores" "downsampling" promoter
  • "STARR-seq" "barcode" "censoring"
  • "promoter" "sequencing depth" "median" barcode
  • "Jores" "RNA" "filtering" "bias"
  • "Jores" "sequencing depth" promoter
  • "plant" "promoter" "barcode" "downsampling"
  • "MPRA" "RNA" "filtering" "censoring"
  • "MPRAsnakeflow" "downsampling"
  • "Jores" "mpralm"
ATTEMPT 013

Did historical fixed-width count parsing alter inherited plant promoter labels?

Verified archive discrepancy — correctly parsed values already public

Question

Did the original readr fixed-width parser truncate high barcode counts, and are the resulting altered promoter values present in later plant-model benchmarks?

Why it might matter

Exact reconstruction exposed a deterministic discrepancy: 1137→137 for one RNA count reproduces an archived promoter value. The old parser inferred field boundaries from the first 1000 rows. This can be tested against the full raw files and archived downstream artifacts.

Minimal test

Compare robust whitespace count parsing to the fixed-width boundaries inferred from the first 1000 rows, using the original readr v1.4.0 source as reference. Confirm that all initially mismatched Arabidopsis values are reproduced without per-gene fitting; use maize libraries as negative controls.

First novelty check

The generic readr fixed-width limitation is documented. Novelty would be a previously unreported, verified impact on this experimental resource and its downstream model targets, with explicit corrected records. No claim of a new parser bug or broad invalidity of the biological study.

Larger test / frozen decision gates

Before scope outcomes: audit every native-promoter leaf/protoplast count file from the pinned source repository, infer widths from file layout rather than selecting moduli from outcomes, reconstruct all deposited per-replicate values with legacy-compatible parsing, and compute corrected labels with whitespace parsing. Verify inherited values in PGB splits and original CNN inputs. Quantify affected records, practical size, training/test membership and model-metric consequences. Create an executable historical-R reproduction or clearly delimit any remaining emulator uncertainty. Recheck closest prior work after results.

Second novelty check

The original repository has one public issue, requesting replicate data, whose author response supplies the correctly parsed measurements in 2023. PGB discussions concern unrelated viewer/gene-expression labels. A newly described discrepancy between snapshots may be useful provenance documentation, but producing the corrected values does not satisfy novelty on its own.

Limitations

Exact downstream agreement strongly supports the parser explanation but is not a record of the authors’ actual execution environment. A small number of changed labels may have negligible aggregate model impact; do not inflate severity.

Decision and next step

Close as a standalone discovery claim: retain the verified provenance audit and attachment comparison as support for the measurement-robustness study. Do not report PGB parser corruption or claim first correction of the source values.

Observed results

Evidence and outputs

Sources and closest prior work

Exact literature search queries
  • readr read_table truncates first digit counts fixed width
  • read_table readr digits cut off leading whitespace bug
  • "Jores" "read_table"
  • "Jores" "readr" truncation
  • "plant genomic benchmark" "parsing"
  • "promoter" "read_table" "bug"
  • "Jores" "promoter" "labels" benchmark discrepancy
  • "Synthetic promoter" "readr"
  • "plant genomic benchmark" promoter preprocessing Jores
ATTEMPT 014

Completed study: RNA-count filtering destabilizes plant promoter ranks

Executed — figures, data, reproducible workflow and research note

Question

Does the validated depth effect persist across all three promoter-origin species and both assay hosts, and does it affect evaluation of frozen published predictions?

Why it might matter

This executes the lead from 012. The applied contribution is measured resource-specific robustness, not a new estimator or biological regulatory mechanism.

Minimal test

See 012 for pilot and validation gates. No new thresholds were selected for this extension.

First novelty check

Pooling, MPRA downsampling and processing-bias studies precede this work. Targeted searches found no report of the exact Jores-resource comparison. The original author already supplied correct original-estimator values in 2023, so the parser correction is excluded from novelty.

Larger test / frozen decision gates

The extension protocol recorded under 012 was executed unchanged for Sb leaf replicates and all six protoplast libraries. Five seeds at 50%/25% RNA depth; DNA cutoffs 5/20; all four estimator variants, shared cohorts and dropout. Frozen CNN test predictions were mapped only where species/target matches uniquely identified a gene.

Second novelty check

Reviewed mpralm (2019), MPRAsnakeflow (2025), human MPRA processing comparisons (2025), the source study and its full public issue history, plus PGB discussions and targeted Jores/promoter/filtering/downsampling searches. No exact prior applied result found. Novelty remains provisional until expert review; no guarantee of publication.

Limitations

One experimental resource; two biological replicates per library/host combination; RNA-only thinning; high-coverage selected cohorts; pooled weighting is not proven biological truth; frozen-CNN subset evaluation; unresolved CNN-export provenance. Monte Carlo seeds are not biological replication.

Decision and next step

The local computational study is executed and yields a useful original empirical result under the searched literature scope. Package it for review as a short computational resource/methods note. Do not submit or contact authors without user instruction.

Observed results

Evidence and outputs

Sources and closest prior work

Exact literature search queries
  • "plant promoter" "RNA filtering" sequencing
  • "Jores" "barcode" "pooling"
  • "plant promoter" "downsampling" "median"
  • "Jores" "replicate" "reanalysis"
  • "Jores" "promoter" "correction" data counts
  • "Synthetic promoter designs" "error" data
ATTEMPT 015

Can automated bird detection manufacture apparent avoidance between species?

Closed — insufficient isolated, detectable source events

Question

If the occurrence of two bird vocalizations is experimentally independent, does a fixed-window audio classifier create a negative association between their detections, including when their calls do not overlap in time?

Why it might matter

Species co-occurrence and temporal partitioning are ecological quantities. Dependence of one species detection on another species vocalizing could manufacture evidence of competition or avoidance. A controlled mixture establishes presence without requiring inferred field interactions.

Minimal test

Select a fixed, deterministic sample of annotated short single-species events from the cached CEB test recordings, with no other annotated bird in the source excerpt. Require baseline target scores >=0.5 and use at most two excerpts per species from distinct source files. Combine pairs at matched event RMS, with simultaneous and non-overlapping event timing and matched attenuation controls. No clipping; hold baseline gain and placement explicit. Primary advancement gate: in non-overlapping mixtures at least 20% of baseline-detectable target appearances fall below 0.5, across at least four source species and two distinct recording pairs per affected species. Also report continuous score changes, simultaneous-mixture results and failed eligibility. Pilot file selection must not be reused as held-out validation.

First novelty check

Candidate generation found broad prior work on soundscape interference and synthetic multi-species mixing. The specific novelty question, after the minimal test, is whether controlled independent vocalizations produce spurious ecological association even without acoustic event overlap, and whether an explicit calibration can correct it.

Limitations

Annotation-based isolation does not guarantee absence of every unannotated sound. Audio mixing is an intervention on detectability, not evidence of real interactions among wild birds.

Decision and next step

Close this dataset-specific pilot without relaxing selection or confidence gates. Continue to another biological question.

Observed results

Evidence and outputs

Sources and closest prior work

Exact literature search queries
  • bird vocalization classifier synthetic mixtures co-occurrence bias
  • bird bioacoustics classifier mixed species false positive co occurrence networks BirdNET
ATTEMPT 016

Does strand-specific G-rich promoter architecture predict host portability?

Closed — predeclared motif-carrier gate failed

Question

Do canonical G-quadruplex-like promoter motifs show a strand-specific association with changes in activity between tobacco leaves and maize protoplasts, after accounting for ordinary sequence composition?

Why it might matter

A sequence feature associated with host portability could help design promoters that transfer between assay systems. G-rich secondary-structure motifs are plausible regulatory features, but their apparent effects can simply reflect GC or ordinary transcription-factor motifs. The pilot tests a narrow residual association and does not presume DNA folding.

Minimal test

Use maize-origin native promoter sequences and paired host activity estimates, requiring >=10 DNA-supported barcodes in both biological replicates of both hosts. Scan a fixed canonical G3 motif (four runs of >=3 G separated by 1–7 bases) and its reverse-complement C-rich counterpart. Before outcome testing require >=50 promoters for each orientation. Fit log2(leaf activity)-log2(protoplast activity) on motif indicators, mononucleotide and dinucleotide composition, regional GC content, quadratic/cubic overall GC and a canonical core TATA indicator. Primary contrast: G-motif coefficient minus C-motif coefficient; advance only if absolute contrast >=0.25 log2 and HC3 robust |t|>=3, with consistent sign using both original-median and pooled activity estimates. Hold sorghum and Arabidopsis outcomes out of this pilot.

First novelty check

Plant G4 sequence surveys and regulatory mechanisms already exist. A new contribution would require a specific, reproducible host-portability relationship that survives composition controls and independent validation, not simply enrichment of G-rich sequence in promoters.

Limitations

Motif matching does not demonstrate G4 formation, and transient reporter assays are not native genomic expression. This is exploratory mechanistic prioritization until replicated and independently validated.

Decision and next step

Close without lowering the minimum carrier count or inspecting outcome contrasts. Continue discovery.

Observed results

Evidence and outputs

Sources and closest prior work

Exact literature search queries
  • plant core promoter G quadruplex Jores promoter strength
ATTEMPT 017

Are the two ends of a plant gene functionally coordinated?

Closed — no meaningful coordination in the pilot

Question

Do separately measured native core-promoter and terminator activities covary for the same gene beyond shared sequence composition, and is this relationship consistent across assay hosts?

Why it might matter

A natural gene pairs its promoter and terminator, but synthetic designs often choose these parts independently. A reproducible residual association could test proposed coordination between transcription initiation and termination; it would not establish that a particular pair physically interacts.

Minimal test

Pilot only Arabidopsis chromosome 1 genes. Join native 170-base terminator activities from Gorjifard et al. 2024 to the 2021 promoter sequences and original-median / pooled promoter estimates, requiring >=10 DNA-supported promoter barcodes in both replicates of both hosts. Use gene-level averages when there are multiple terminator entries. Compute Spearman correlation and partial rank correlation after adjusting both activities for mono- and dinucleotide composition of both promoter and terminator, regional GC and overall GC squared/cubed. Advance only with >=500 matched genes and a positive partial correlation >=0.15 in both hosts using both promoter estimators. Hold other Arabidopsis chromosomes and maize out of this pilot.

First novelty check

Gene harmony is a prior hypothesis. An author dissertation also compares promoter, terminator and enhancer strengths with gene expression; this is a potential direct overlap to inspect before advancing. Merely joining two tables is not novel.

Limitations

Different reporter assays measure each end independently; these data cannot establish promoter–terminator interaction or causal evolutionary coadaptation. Chromosomes are partitions, not independent species.

Decision and next step

The frozen effect gate failed. Do not inspect held-out chromosomes or maize for this hypothesis.

Observed results

Evidence and outputs

Sources and closest prior work

Exact literature search queries
  • plant promoters terminators strengths same gene correlation Jores terminators 2022
  • "Arabidopsis and maize terminator strength" "promoter" correlation
ATTEMPT 018

Bird co-detection bias using single-species passages

Validation effect replicated; diversity gate failed

Question

Can two separately detectable bird vocalizations suppress one another in an automated detector even when they occupy disjoint parts of its input window?

Why it might matter

Entry 015 required exactly one annotation per excerpt. That excludes a bird producing several notes, although the excerpt can still contain only one species. This follow-up changes source eligibility to single-species passages and retains the original causal comparison and score/effect gates.

Minimal test

Use only the entry-015 pilot recording partition (SHA256 of 015:20260919:filepath, first 8 hex digits modulo 3 equals 0), keeping the other recordings uninspected for validation. For each annotated event center, take 1.2 seconds and require every intersecting valid annotation to belong to the target species, allowing multiple same-species events and excerpts of longer vocalizations. Select at most 8 candidate passages per species from distinct files, and 240 total, deterministically before model scores. Scale excerpts to a common RMS with a single clipping-safe scale. Require target score >=0.5 at both time positions, at least two source files per species; retain at most two passages per species. Compare exact-component baselines to non-overlapping mixtures in both orders and simultaneous mixtures. Advance if >=20% of non-overlapping target appearances lose detection at 0.5, across >=4 species with >=2 distinct recording pairs per affected species. The baseline and advancement thresholds are unchanged from entry 015.

First novelty check

Noise and soundscape interference are known. If the pilot succeeds, investigate prior demonstrations of spurious ecological association from temporal-context competition, and existing mitigations, before validating.

Limitations

Single-species annotation does not prove acoustic purity. Model conditioning on an eligible detectable subset limits population generalization; any result describes this controlled counterfactual experiment.

Decision and next step

Do not call this a completed study. Record the shortfall and pursue a separately specified focal-recording extension (entry019), retaining the existing negative feasibility outcomes.

Observed results

Evidence and outputs

Sources and closest prior work

Exact literature search queries
  • "BirdNET" "co-occurrence" bias mixtures