Statistical validation of a simple cross-correlation statistic on LIGO O1/O2 open data — reproducible baseline and audit trail

Hi everyone,

Following my earlier posts on this forum, I have completed a full statistical validation of a deliberately simple template-free detection statistic for short GW transients. I’m sharing it here before journal submission to invite technical feedback, particularly on the null construction and injection methodology.

The statistic: maximum inter-detector cross-correlation (max_corr) between whitened H1 and L1 strain in a 0.25 s window, evaluated over a 9-point offset grid with a continuous physical delay taper. No templates, no training, no tunable parameters after freezing.

Global null: 10,200 background realizations stratified across 3,967.8 h of coincident CAT2-clean O1+O2 livetime, with hardware-injection and catalog-event vetoes and a look-elsewhere-matched procedure (events and null both scored as best-of-9 offsets).

Catalog results: Five GWTC-1 events significant after Holm–Bonferroni correction over eleven tests: GW150914 and GW170814 (p < 1×10⁻⁴), GW170608 (p = 6×10⁻⁴), GW170104 (p = 2.4×10⁻³), GW170729 (p = 6.0×10⁻³).

Injection recovery (real detector noise, rule-matched to null):

  • 400 BBH injections (SNR 5–30): AUC = 0.815 [95% CI 0.789–0.839]
  • 100 sine-Gaussian bursts: AUC = 0.870 [0.824–0.911]
  • 100 white-noise bursts: AUC = 0.793 [0.734–0.857]

Detection threshold is sharply at network SNR ≈ 12–15. Below SNR 8 the method is effectively blind.

Audit trail: Six pipeline errors documented in the paper, each caught by a validation test — including a look-elsewhere mismatch, a silent FFT truncation in SNR scaling, and an offset-selection rule asymmetry between events and injections. The corrected BBH AUC dropped from 0.901 to 0.815; the catalog table was unaffected.

Event-twin injections (n=100 per event, bootstrap CI):

  • GW170608: 46th percentile [36th–56th] of twins — typical detection
  • GW151226: 1st percentile [0th–3rd] — borderline population (twin median 0.285 coincides with null p99 threshold 0.284; 53% twin efficiency)
  • GW150914: 47th percentile [37th–57th] under tie-free raw-max accounting — unremarkable

The method is not competitive with matched filtering or cWB and is not intended to be. Its contributions are: (i) a fully interpretable, exactly reproducible floor against which more complex template-free methods can report what their complexity buys; (ii) a portable validation protocol (stratified global null, matched look-elsewhere, local null, rule-matched injection recovery, event-twin comparisons) that transfers to any two-sensor coincident-transient problem with a physical delay bound.

All code, seeds, and sampled epochs are released for exact reproduction on a consumer laptop:
Zenodo deposit: A Transparent, Training-Free Inter-Detector Cross-Correlation Statistic for Short Gravitational-Wave Transients: Statistical Validation on LIGO O1/O2 Open Data (DOI: 10.5281/zenodo.21435218)
— all scripts, seeds, null realizations, injection sets, and pre/post-fix
data products for exact reproductionSpecific questions for the community:**

  1. Is the look-elsewhere matching (best-of-9 for both events and null) correctly implemented for empirical p-values, or is there a subtlety I’m missing?
  2. The correlated-noise confounder (Schumann resonances in-band) is mitigated empirically by building the null from simultaneous H1/L1 data — is a time-slide null considered necessary for journal submission, or is the empirical approach acceptable?
  3. Which venue would you recommend — CQG or PRD — for a methodological contribution of this type?

Thanks in advance for any feedback.

Dimitar Kretski
Independent researcher, Varna, Bulgaria
ORCID: 0000-0001-5108-2243