Impact of High Signal-to-Noise Ratio Screening on Seismic Deep Learning Training Data

arXiv Physics · · 3 min read · Natural Sciences

Read research and analysis on Impact of High Signal-to-Noise Ratio Screening on Seismic Deep Learning Training Data published by ICANEWS, a global research journal for emerging researchers.

Key Takeaways

  • SNR screening for phase picking showed no consistent benefit under fine-tuning or random initialization, even with distance and phase composition controls.
  • Dispersion screening initially increased error at five epochs but improved performance against an equally sized unscreened set after longer training.
  • Utilizing all 29,788 eligible fit paths for dispersion reduced unfiltered-test error by 3.0% compared to 9,929 randomly sampled paths.
  • Full-available and strict-SNR training for dispersion resulted in similar mean errors (0.0405 km s$^{-1}$ vs 0.0406 km s$^{-1}$), but strict training had a 9.9% advantage on SNR-selected tests.
  • High SNR did not reliably identify better training data, with gains dependent on task, training schedule, and test domain.

Why This Matters

This research suggests that automatically discarding lower signal-to-noise ratio seismic data may not always optimize deep learning performance. It highlights the need for task-specific evaluation of data screening strategies, balancing data quality with the benefits of larger training datasets.

Overview

This research investigates whether screening seismic records for a high signal-to-noise ratio (SNR) yields better training data for deep learning applications, specifically in seismic phase picking and ambient-noise dispersion analysis. The study explored whether the benefit of clearer waveforms outweighs the loss of discarded data when training counts are matched, and how retaining a larger eligible data pool, rather than a screened subset, affects learning outcomes.

Research Context

Deep learning models for seismic analysis frequently utilize large datasets. A common practice in data preparation involves hard SNR screening, which removes seismic records deemed to have weak signals. The core question addressed here is whether this removal process, intended to enhance data quality, genuinely improves the learning process sufficiently to compensate for the reduction in available training examples.

Approach

The investigation was structured around two primary seismic tasks: phase picking and ambient-noise dispersion. The researchers conducted tests in two main configurations:

  • Matched Training Counts: Initial comparisons involved ensuring an equal number of training examples between screened and unscreened datasets.
  • Full Eligible Pool: Subsequent tests focused on utilizing the entire pool of eligible dispersion data, as opposed to a randomly sampled subset or a strictly SNR-screened subset, while maintaining a fixed computational budget.

For phase picking, the screening effects were assessed under both fine-tuning and random initialization conditions. A control group, matched on distance and phase composition, was also included in these phase-picking evaluations. For ambient-noise dispersion, the impact of screening was observed at a five-epoch endpoint and after longer training periods, comparing performance against an equally sized unscreened set.

Findings

Phase Picking

  • SNR screening for phase picking yielded no consistent benefit to unfiltered test performance. This held true for models undergoing fine-tuning or random initialization.
  • The lack of consistent benefit also applied to a control dataset specifically matched on distance and phase composition.

Ambient-Noise Dispersion

  • At the original five-epoch endpoint, dispersion screening led to an increase in error.
  • After extended training, however, the screened dispersion dataset improved performance when compared to an equally sized unscreened set.
  • Retaining all 29,788 eligible fit paths, instead of a subset of 9,929 randomly sampled paths, reduced unfiltered-test error by 3.0% across three distinct seeds, using the same update budget.
  • When comparing full-available training (using all eligible data) with strict-SNR training, both approaches exhibited similar mean errors of 0.0405 km s$^{-1}$ and 0.0406 km s$^{-1}$, respectively. The seed-paired ordering of their performance was mixed.
  • Strict SNR training maintained a 9.9% advantage when evaluated on the SNR-selected test dataset.

Monitoring Applications

  • In a two-day monitoring scenario, SNR regulated pick streams.
  • Event recovery, however, was observed to be dependent on how the SNR measurement was defined.

Overall Conclusion on SNR Screening

High SNR did not reliably identify better training data in a universal manner. The observed gains were contingent upon several factors, including the specific task being addressed, the training schedule employed, and the domain of the test data. Comparisons made using fixed data counts were identified as potentially omitting the benefits associated with retaining a larger number of additional valid training examples.

Why This Matters

The findings indicate that the efficacy of signal-to-noise ratio-based data screening for deep learning in seismology is not a universal constant. Its utility is task-dependent and influenced by training methodologies, suggesting that researchers and practitioners need to carefully evaluate the trade-offs between data quality (high SNR) and data quantity for specific applications rather than assuming a default benefit.

Research Information

Institution
arXiv Physics
Original Study
View Publication
Source
arXiv Physics

About ICANEWS

ICANEWS is a global research journal for emerging researchers, publishing student and emerging researcher work across all fields.