Free for humans

SPBench: A Multi-Task Evaluation Benchmark for Exploration Seismic Processing

A public seismic-processing benchmark re-runs 24 learning methods on shared data and shows that winning on synthetics does not reliably mean winning on field records.

arXiv:2609.289255 min readScore 65/100 · editorial triage · not peer reviewPaper hub2026-W40

The 30-second take

  • What: SPBench covers six processing tasks, reproduces 24 supervised methods on 10 datasets under 43 settings, and adds SCoRE and a ridge-curvature pick score — finding synthetic rankings fail to predict field rankings and that method order shuffles as degradations get harder.
  • Abundance angle: today, honest comparison of seismic ML is a scarce, private-data luxury. A reproducible public arena is a step toward default, cheaper evaluation — not cheaper oil (mid-horizon: data access and criteria still decide).
  • Who should care: Exploration geophysicists, seismic-ML authors tired of irreproducible tables, and buyers who need to know when a paper’s ranking is an artifact of the test.

What the paper actually did

Exploration seismic processing is the backbone of subsurface imaging, but learning-based methods are hard to compare. The authors surveyed 368 papers and found widespread private or hard-to-reproduce data, with only 25 providing public code — so reported gains may be model design or just the experimental setting.

They introduce the Seismic Processing Benchmark (SPBench) with six tasks: random noise attenuation, trace interpolation, ground-roll suppression, multiple suppression, deblending, and first-arrival picking. They reproduce 24 supervised methods on 10 datasets under 43 standardized settings and release datasets, implementations, configs, evaluation scripts, and results.

Besides global scores and per-trace pick errors they add signal-component-resolved evaluation (SCoRE) for reconstruction and a reference-free ridge-curvature score (RC_norm) for first-arrival picking. Analyses: synthetic rankings do not reliably predict field rankings (agreement is task-dependent when models train inside each setting). As degradation strengthens, rankings reorder more under coherent ground roll than under random-like interference. RC_norm agrees with MAE-based model rankings in the evaluated settings (mean Kendall correlation 0.881 across three field surveys), while SCoRE shows frequency- and energy-dependent differences that global scores hide.

What makes this disruptive

The scarce capability is a fair, public seismic-ML leaderboard. Private data plus missing code made the last few hundred papers hard to trust as a cumulative science.

Reproducing 24 methods under 43 settings and then showing that synthetic rank ≠ field rank is the punch. SCoRE and RC_norm attack two evaluation failure modes: global scores that hide which frequencies were fixed, and pick metrics that need a reference.

This is a benchmark paper. Treat Kendall 0.881 and the synthetic-vs-field warning as their measurements on this suite, not a law of geophysics.

Why it matters (outside the lab)

Abundance lens: high-quality seismic processing know-how is still concentrated in shops with private benchmarks. If SPBench becomes a default comparison layer, more of that evaluation can be ordinary and cheaper — which is a scientific abundance, not a promise of cheaper hydrocarbons.

Near-term, use it to distrust single-dataset tables. Medium-term, community uptake, more field surveys, and whether unsupervised methods get added decide if it becomes infrastructure.

No date. Better ML eval does not by itself change drilling.

Limitations & open questions

Preprint benchmark. Twenty-four supervised methods will age; unsupervised and industry in-house stacks are outside the 43 settings. We have not rerun the suite. “Only 25 of 368 papers had public code” is their survey, not independently audited here.

Synthetic-versus-field disagreement is task-dependent — not a blanket “synthetics are useless.” RC_norm’s Kendall 0.881 is on three field surveys. Released data still may not match every company’s geology.

Abundance is not automatic: open eval does not open proprietary seismic libraries.

Explain ladder

Default article depth

Machine-learning papers that clean seismic records often train and test on private shots, so nobody can tell if a new network is actually better. The authors counted hundreds of papers and almost no public code, then built SPBench: six jobs (denoise, interpolate, kill ground roll and multiples, unblend, pick first breaks) and 24 reproduced methods on shared data.

The uncomfortable result: the model that wins on synthetic damage is often not the one that wins on field data. Rankings also reshuffle more when the junk is coherent ground roll than when it is random-ish noise. A new component-wise score (SCoRE) shows who fixed which frequencies; a ridge-shape score for picks mostly agrees with error-against-human rankings without needing those humans at test time.

If you read seismic ML, demand the SPBench setting before believing a gold medal.

Key terms

Ground roll
High-amplitude, low-velocity surface waves that coherent-noise suppressors try to remove from seismic records.
Deblending
Separating overlapping shots when surveys fire sources in a blended, efficient pattern.
SCoRE
The authors’ signal-component-resolved evaluation for reconstruction quality beyond a single global error.
Democratization of abundance
Editorial lens: scarce reproducible evaluation can become a cheaper default scientific layer — no promised year.

Sources

Related explainers

Same topic and week first — keep exploring the scarcity → abundance map.

Editorial explainer · not peer review · always read the primary paper.

Byline: Disruptive Concepts editorial.