Skip to content

API reference

Generated from the source docstrings. Every module is importable on its own, e.g. from sweep_io.seismic_plan import PlanReader.

SEG-Y

sweep_io.segy.SEGYReader

SEGYReader(path: str | PathLike, *, mmap_mode: bool = True)

Single-file SEG-Y reader keyed by absolute byte offsets.

On construction, parses the 400-byte binary header to extract n_samples, dt, and sample_format. Trace data is then read on demand by absolute byte offset — no header re-parsing per trace, no Python loop per sample.

Designed to be the fast lane behind a pre-built shot index. Pair it with :class:MultiFileSEGYReader when shots live across many files.

Parameters:

  • path (str | PathLike) –

    File path.

  • mmap_mode (bool, default: True ) –

    If True (default), memory-map the file (zero-copy reads, kernel page cache reuse). If False, use pread / seek+read. mmap can misbehave on certain network filesystems — try False if you see SIGBUS or weird truncation.

Concurrency

A single SEGYReader instance is thread-safe for reads when mmap_mode=True (each read_trace_data call takes a numpy view of the mmap; no shared file pointer). With mmap_mode=False reads use os.pread, which is also thread-safe. So either way you can share one reader across all your prefetch workers.

n_traces property

n_traces

Total number of traces in the file (computed from file size).

read_trace_data

read_trace_data(
    byte_offsets: Sequence[int] | ndarray, *, coalesce_gap: int = 0
) -> np.ndarray

Read trace data at given absolute byte offsets.

Each offset must point at a trace header (not at the data block — the 240-byte header is skipped internally).

Parameters:

  • byte_offsets (Sequence[int] | ndarray) –

    1-D array of absolute byte offsets, one per trace to read. Order is preserved in the output, but the actual disk reads are issued in ascending order (and coalesced where possible) for sequential bandwidth.

  • coalesce_gap (int, default: 0 ) –

    Adjacent reads whose gap (between the end of one and the start of the next) is ≤ this many bytes are merged into a single pread. 0 disables coalescing; sensible values on Lustre / NVMe are 0 (already sorted-sequential is fast enough) up through one filesystem block (~1 MB).

Returns:

  • ndarray –

    (n_traces, n_samples) float32 array, ordered the same as byte_offsets.

read_trace_header

read_trace_header(byte_offset: int) -> bytes

Return the 240-byte trace header at byte_offset.

read_text_header

read_text_header() -> bytes

First 3200 bytes (the EBCDIC text header).

read_binary_header

read_binary_header() -> bytes

Bytes 3200–3600 (the binary header).

sweep_io.segy.MultiFileSEGYReader

MultiFileSEGYReader(paths: Sequence[str | PathLike], *, mmap_mode: bool = True)

Dispatch byte-offset reads across many SEG-Y files by file_id.

Holds one open :class:SEGYReader per path. All readers must agree on n_samples, dt, and sample_format (the usual case for a single survey split across many files). A mismatch raises at construction.

Parameters:

  • paths (Sequence[str | PathLike]) –

    List of file paths. file_id = 0 corresponds to paths[0], etc.

  • mmap_mode (bool, default: True ) –

    Passed through to each underlying :class:SEGYReader.

read_traces

read_traces(
    file_ids: Sequence[int] | ndarray,
    byte_offsets: Sequence[int] | ndarray,
    *,
    coalesce_gap: int = 0
) -> np.ndarray

Read traces from heterogeneous files, preserving caller order.

Parameters:

  • file_ids (Sequence[int] | ndarray) –

    1-D array, same length as byte_offsets. Index into self.readers.

  • byte_offsets (Sequence[int] | ndarray) –

    1-D array of trace-header offsets within each file.

  • coalesce_gap (int, default: 0 ) –

    Forwarded to each per-file read.

Returns:

  • ndarray –

    (n, n_samples) float32, ordered by the caller.

sweep_io.segy.read_segy

read_segy(
    path: str | PathLike, *, iline: int = 189, xline: int = 193, strict: bool = False
) -> Tuple[np.ndarray, dict]

Read an entire SEG-Y volume into memory (via the optional segyio dep).

sweep_io.segy.write_segy

write_segy(
    path: str | PathLike, data: ndarray, *, dt: float, delay_recording_time: float = 0.0
) -> None

Write a 2-D (n_traces, n_samples) SEG-Y via segyio.

sweep_io.segy.write_segy_minimal

write_segy_minimal(
    path: str | PathLike, data: ndarray, *, dt: float, sample_format: int = 5
) -> None

Write a SEG-Y file with empty text/trace headers — useful for fixtures.

No segyio required. Produces a valid layout that :class:SEGYReader can parse: - 3200-byte zeroed text header - 400-byte binary header with dt / n_samples / sample_format set - one (zeroed trace header + sample data) per row

Only formats 1 (IBM) and 5 (IEEE32) are supported here; we don't write integers from this helper.

Header index

sweep_io.segy_index.build_segy_index

build_segy_index(
    paths: Sequence[str | Path],
    *,
    byte_map: dict | None = None,
    source_depth_m_override: float | None = None,
    receiver_depth_m_override: float | None = None,
    coord_scalar_override: float | None = None,
    num_workers: int = 4,
    show_progress: bool = False
) -> SEGYIndex

Scan headers of paths in parallel and build a unified :class:SEGYIndex.

Parameters:

  • paths (Sequence[str | Path]) –

    SEG-Y file paths in the order you want them assigned file_id.

  • byte_map (dict | None, default: None ) –

    Trace-header byte-offset overrides. Defaults to SEG-Y rev1 standard (:data:SEGY_REV1_BYTES). Override per-key, e.g. {"source_depth": 44} for an OBN airgun depth.

  • source_depth_m_override (float | None, default: None ) –

    If the SEG-Y file's depth bytes are zero (or unreliable), use these instead. Common case for marine streamers.

  • receiver_depth_m_override (float | None, default: None ) –

    If the SEG-Y file's depth bytes are zero (or unreliable), use these instead. Common case for marine streamers.

  • coord_scalar_override (float | None, default: None ) –

    If set, ignore the per-trace coord_scalar field and apply this constant multiplier to (sx, sy, rx, ry). Use when you know the files have inconsistent or wrong scalar values.

  • num_workers (int, default: 4 ) –

    Parallel header-scan workers. find -name "*.sgy" | wc -l is a good upper bound; 4-8 is typical for SSD / Lustre.

  • show_progress (bool, default: False ) –

    Print one line per finished file. Useful when scanning hundreds.

Returns:

  • SEGYIndex –

    With len(paths) distinct file_id values.

sweep_io.segy_index.SEGYIndex

SEGYIndex(file_paths: list[str], file_id: ndarray, byte_offset: ndarray, shot_id: ndarray, receiver_id: ndarray, sx_m: ndarray, sy_m: ndarray, sz_m: ndarray, rx_m: ndarray, ry_m: ndarray, rz_m: ndarray, dt_s: float, n_samples: int, sample_format: int = 5, meta: dict = <factory>, SCHEMA_VERSION: int = 1)

Per-trace catalog backed by 1-D numpy columns.

All trace-level arrays have the same length = total trace count across all files. file_id[i] indexes :attr:file_paths. shot_id[i] is the group key (usually FFID); the same value across all traces of one shot. receiver_id[i] is the trace's index within its shot (0-based after sort).

Per-shot fast lookup is built on construction via :meth:_build_shot_lookup. lookup_shot(shot_id) returns the slice in O(log n).

SCHEMA_VERSION class-attribute

SCHEMA_VERSION = 1

int([x]) -> integer int(x, base=10) -> integer

Convert a number or string to an integer, or return 0 if no arguments are given. If x is a number, return x.int(). For floating point numbers, this truncates towards zero.

If x is not a number or if base is given, then x must be a string, bytes, or bytearray instance representing an integer literal in the given base. The literal can be preceded by '+' or '-' and be surrounded by whitespace. The base defaults to 10. Valid bases are 0 and 2-36. Base 0 means to interpret the base from the string as an integer literal.

int('0b100', base=0) 4

sample_format class-attribute

sample_format = 5

int([x]) -> integer int(x, base=10) -> integer

Convert a number or string to an integer, or return 0 if no arguments are given. If x is a number, return x.int(). For floating point numbers, this truncates towards zero.

If x is not a number or if base is given, then x must be a string, bytes, or bytearray instance representing an integer literal in the given base. The literal can be preceded by '+' or '-' and be surrounded by whitespace. The base defaults to 10. Valid bases are 0 and 2-36. Base 0 means to interpret the base from the string as an integer literal.

int('0b100', base=0) 4

lookup_shot

lookup_shot(shot_id: int) -> dict

Return all per-trace + source info for one shot. O(log n_shots).

save

save(path: str | Path) -> Path

Save as .npz (compressed). Loads back via :meth:load.

to_physical_geometry

to_physical_geometry(*, shot_ids: Sequence[int] | None = None)

Materialise a :class:sweep_io.geometry.PhysicalGeometry.

Expects every shot to have the same receiver count (typical for streamers). For OBN / variable-receiver-count layouts, slice the index by shot_id first.

to_seismic_plan

to_seismic_plan(**kwargs)

Convenience: build a :class:sweep_io.seismic_plan.SeismicPlan.

All kwargs forward to :func:sweep_io.seismic_plan.build_seismic_plan. Use grouping='csg' (default) for shot-gather organisation (typical 2-D / 3-D streamer) or grouping='crg' (plus receiver_quantize_m) for common-receiver gathers (OBN).

Examples:

>>> idx = build_segy_index(["viking.sgy"])
>>> plan_csg = idx.to_seismic_plan(grouping="csg")
>>> plan_csg.save("viking_csg_plan.npz")

sweep_io.segy_index.IndexedShotGatherDataset

IndexedShotGatherDataset(
    index: SEGYIndex,
    reader: MultiFileSEGYReader | None = None,
    *,
    shot_ids: Sequence[int] | None = None,
    coalesce_gap: int = 0
)

Per-shot dataset backed by :class:SEGYIndex + :class:MultiFileSEGYReader.

Each __getitem__ issues one coalesced byte-offset read for the selected shot's traces (sorted ascending, runs merged). Output shape is (n_recv, n_samples) float32.

Pair with :class:sweep_io.prefetch.Prefetcher or :class:sweep_io.cuda_prefetch.CUDAPrefetcher for I/O / compute overlap.

Parameters:

  • index (SEGYIndex) –

    Pre-built SEGYIndex.

  • reader (MultiFileSEGYReader | None, default: None ) –

    Optional :class:MultiFileSEGYReader. Built from index.file_paths if omitted.

  • shot_ids (Sequence[int] | None, default: None ) –

    Subset (and order) of shot IDs to expose. Default: all shots, sorted.

  • coalesce_gap (int, default: 0 ) –

    Forwarded to reader.read_traces; merges adjacent reads with at most this many bytes of gap.

Plans and readers

sweep_io.seismic_plan.build_seismic_plan

build_seismic_plan(
    index,
    *,
    grouping: str = "csg",
    shot_ids: Sequence[int] | None = None,
    receiver_quantize_m: float | None = None,
    offset_min_m: float | None = None,
    offset_max_m: float | None = None,
    max_traces_per_group: int | None = None,
    seed: int = 0,
    build_label: str | None = None
) -> SeismicPlan

Build a :class:SeismicPlan from a :class:SEGYIndex.

Parameters:

  • index –

    A :class:sweep_io.segy_index.SEGYIndex — the full per-trace catalog. Not mutated.

  • grouping (str, default: 'csg' ) –

    "csg" (default) → one group per unique shot_id (FFID). "crg" → one group per receiver cell (use receiver_quantize_m to define the cell size).

  • shot_ids (Sequence[int] | None, default: None ) –

    Optional whitelist of shot IDs to keep (intersected with what the index has). None = all shots in the index.

  • receiver_quantize_m (float | None, default: None ) –

    For grouping="crg": quantization tolerance for collapsing per-trace receiver positions to a finite set of cells. Required for CRG; ignored for CSG.

  • offset_min_m (float | None, default: None ) –

    Per-trace source-receiver horizontal-offset filter (meters). Traces falling outside [offset_min_m, offset_max_m] are dropped pre-grouping.

  • offset_max_m (float | None, default: None ) –

    Per-trace source-receiver horizontal-offset filter (meters). Traces falling outside [offset_min_m, offset_max_m] are dropped pre-grouping.

  • max_traces_per_group (int | None, default: None ) –

    Hard cap on rows per group after grouping. Excess rows are randomly sub-sampled (seeded via seed).

  • seed (int, default: 0 ) –

    RNG seed for the per-group cap sub-sampling.

  • build_label (str | None, default: None ) –

    Optional human-readable tag stored in build_meta for provenance.

sweep_io.seismic_plan.SeismicPlan

SeismicPlan(grouping: str, files: list[Path], trace_size_per_file: ndarray, sample_format: int, samples_per_trace: int, dt_s: float, row_file_id: ndarray, row_trace_offset: ndarray, row_source_xyz: ndarray, row_receiver_xyz: ndarray, group_id: ndarray, group_xyz: ndarray, group_offsets: ndarray, build_meta: dict = <factory>)

A filter + grouping view over a :class:SEGYIndex, persisted to npz.

Carries enough metadata to read the underlying SEG-Y traces lazily (via :class:PlanReader) but holds NO sample data itself — on-disk size is O(n_rows × 80 bytes), not O(n_rows × nt × 4).

Construct via :func:build_seismic_plan or load via :meth:load.

save

save(path: str | Path) -> Path

Save as a single npz; loads back via :meth:load.

load classmethod

load(path: str | Path) -> SeismicPlan

Load a seismic_plan_v1 npz.

filter_rows

filter_rows(row_mask: ndarray) -> SeismicPlan

Drop rows where row_mask is False; recompute group_offsets.

Groups that lose every row survive with a zero-length slice; compose with :meth:drop_empty_groups to remove them.

filter_groups

filter_groups(group_mask: ndarray) -> SeismicPlan

Drop groups where group_mask is False; recompute offsets + row arrays.

drop_empty_groups

drop_empty_groups() -> SeismicPlan

Drop groups with zero rows; preserves order of surviving groups.

sweep_io.seismic_plan.PlanReader

PlanReader(
    plan: SeismicPlan,
    *,
    mmap: bool = True,
    cache_all: bool = False,
    trace_cache_bytes: int = 0,
    coalesce_gap: int = 0
)

Lazy SEG-Y trace reader bound to a :class:SeismicPlan.

The plan tells what to read (file_ids + byte offsets, grouped); this reader does the actual IBM-float decoding via :class:MultiFileSEGYReader. Trace order within a group is preserved (matches plan.row_*[group_slice(g)]).

Parameters:

  • plan (SeismicPlan) –

    A :class:SeismicPlan (typically loaded from disk).

  • mmap (bool, default: True ) –

    Whether the underlying SEG-Y reader should mmap files. Default True — the OS page cache then shares pages across processes in distributed runs.

  • cache_all (bool, default: False ) –

    When True, read every plan row into a single (n_rows, nt) float32 tensor at construction and serve subsequent read_group(g) calls from the cache. Useful for small datasets (e.g. 2-D Viking — ~700 MB fits in RAM trivially). For large 3-D surveys leave it False so memory stays bounded.

Parameters:

  • plan (SeismicPlan) –

    A :class:SeismicPlan (typically loaded from disk).

  • mmap (bool, default: True ) –

    Whether the underlying SEG-Y reader should mmap files.

  • cache_all (bool, default: False ) –

    When True, eagerly read every plan row into a (n_rows, nt) float32 tensor at construction.

  • trace_cache_bytes (int, default: 0 ) –

    Optional in-RAM LRU cache for decoded traces, keyed by (file_id, byte_offset). Used by :meth:read_rows to absorb cold-Lustre re-reads when the same physical traces are sampled across multiple iterations (the OBN multisource supershot pattern). 0 disables. -1 means unbounded (matches the legacy --trace-cache-bytes -1 default for large surveys). Any positive value is a byte budget. Ignored when cache_all=True (the full cache is already materialised).

  • coalesce_gap (int, default: 0 ) –

    Bytes-within-the-same-file gap below which two miss-traces get merged into a single pread call. 0 disables. Sensible value: 4× the trace stride (~240 + nt*4 bytes) so any two sequentially-indexed traces in the same file get coalesced. Critical for Lustre where per-call latency dominates small random reads. Only used by :meth:read_rows (the cached path); read_group always reads the full group slice which is already contiguous.

trace_cache_stats property

trace_cache_stats

Cache hit/miss counters + current byte usage, or None when trace_cache_bytes=0 was set at construction.

read_group

read_group(g: int) -> np.ndarray

Return (n_in_group, nt) float32 for group g.

read_groups

read_groups(gs: Iterable[int]) -> list[np.ndarray]

Return one (n_in_g, nt) array per group id in gs.

read_rows

read_rows(row_idx: ndarray) -> np.ndarray

Return (n_picked, nt) float32 for an arbitrary set of plan rows.

Used by sampler-driven loaders (e.g. CRG shared-shot batches) that pick row indices across multiple groups rather than reading whole groups. Row order in the output matches row_idx.

Three code paths, in priority order:

  1. cache_all=True at construction → fancy-index slice on the in-RAM (n_rows, nt) tensor (single numpy op, instant).
  2. trace_cache_bytes set (per-trace LRU + coalesce_gap) → :func:sweep_io.crg_dataset.read_traces_cached. This is the legacy OBN supershot IO path: hot traces (re-sampled across iters) come from RAM, cold traces are coalesced into bigger pread calls. Cuts wait_io about five-fold on a large survey.
  3. Otherwise → raw :meth:MultiFileSEGYReader.read_traces (one pread per miss, no caching). Slow on Lustre with random access patterns; only useful for one-shot reads.

read_all

read_all() -> np.ndarray

Return (n_rows, nt) float32 of every plan row, in plan order.

sweep_io.seismic_plan.sample_shared_shots_from_plan

sample_shared_shots_from_plan(
    plan: SeismicPlan,
    rng: Generator,
    *,
    batch_size: int,
    source_lines_per_group: int = 0,
    max_traces_per_sourceline: int = 0,
    min_coverage: int = 0,
    eligible_groups: ndarray | None = None,
    max_retries: int = 10,
    precomputed_group_unique_keys: list[ndarray] | None = None
) -> SharedShotBatch

Pick batch_size groups whose physical-shot sets intersect.

Operates on a grouping='crg' :class:SeismicPlan: each group is a virtual source (e.g. one OBN node, or a quantised receiver cell), and its rows are the physical shots that group recorded. The intersection across the chosen groups gives the shots EVERY picked group recorded — the invariant needed for source-encoded supershot FWI.

This is the unified-plan port of :func:sweep_io.crg_plan.sample_crg_shared_shots; behaviour is deliberately byte-equivalent so existing runner code can swap one for the other with no FP drift (assuming the same RNG state).

Optional sub-sampling within the intersection:

  • source_lines_per_group > 0 — keep that many random source-line file_ids (= source-line SEG-Y files).
  • max_traces_per_sourceline > 0 — keep that many random shots per kept source line.
  • 0 for either disables that level.

Parameters:

  • plan (SeismicPlan) –

    A grouping='crg' :class:SeismicPlan (CSG plans cannot be sampled this way — each CSG group has a single source, so the intersection is trivially that one source).

  • rng (Generator) –

    Caller-owned np.random.Generator (advanced in-place).

  • batch_size (int) –

    Number of groups (= virtual sources) per batch.

  • min_coverage (int, default: 0 ) –

    Drop groups whose row count is below this. 0 disables.

  • eligible_groups (ndarray | None, default: None ) –

    Optional pre-curated set of group indices to sample from. When None and min_coverage > 0, computed from :meth:SeismicPlan.per_group_row_counts.

  • max_retries (int, default: 10 ) –

    Re-roll group selection up to this many times when the picked groups' shot intersection is empty (rare with min_coverage on but possible for partial-coverage edge nodes).

Raises:

  • ValueError –

    When plan.grouping != 'crg', or when the shot intersection is empty after max_retries attempts, or when sub-sampling empties the intersection.

sweep_io.seismic_plan.sample_percrg_independent

sample_percrg_independent(
    plan: SeismicPlan,
    rng: Generator,
    *,
    batch_size: int,
    source_lines_per_group: int = 0,
    max_traces_per_sourceline: int = 0,
    min_coverage: int = 0,
    eligible_groups: ndarray | None = None
) -> PerCRGBatch

Pick batch_size CRG groups; each keeps its OWN sub-sampled rows.

Unlike :func:sample_shared_shots_from_plan, this does NOT intersect the groups' shot sets — there is no shared receiver geometry. The sub-sampling knobs (source_lines_per_group → keep that many random source-line file_ids; max_traces_per_sourceline → that many random shots per kept line) are applied PER GROUP on that group's own recorded shots. Use for the per-shot (non-encoded) FWI path so each CRG node inverts its full aperture; over iterations the stochastic per-node draw sweeps each node's entire coverage.

Grid-cell dedup of the receivers is deliberately left to the caller — it needs the model frame / grid origin this plan-space sampler is agnostic to.

Data and model windows

sweep_io.plan.DataPlan

DataPlan(
    shot_start: int = 0,
    shot_stop: int | None = None,
    shot_stride: int = 1,
    shot_indices: Sequence[int] | ndarray | None = None,
    receiver_stride: int = 1,
    offset_min_m: float | None = None,
    offset_max_m: float | None = None,
    abs_offset: bool = True,
    dt_target_s: float | None = None,
    time_decimate: int | None = None,
    t_start_s: float | None = None,
    t_end_s: float | None = None,
)

Which shots / receivers / time samples to feed into FWI.

Attributes:

  • shot_start, shot_stop, shot_stride –

    Standard range-style shot selection. Ignored when shot_indices is given.

  • shot_indices –

    Explicit list of shot indices to keep. Overrides the range fields.

  • receiver_stride –

    Keep every N-th receiver per shot (uniform across shots).

  • offset_min_m, offset_max_m –

    Receiver-offset window in meters (signed if abs_offset=False, else compared to |offset|). Offsets are computed in the horizontal plane only — ndim=2 uses |rx - sx|, ndim=3 uses sqrt((rx-sx)^2 + (ry-sy)^2).

  • abs_offset –

    Treat offsets as absolute values when comparing to the window.

  • dt_target_s –

    Time-resample the observed data to this dt (delegates to sweep_io._resample.resample_time). Mutually exclusive with time_decimate.

  • time_decimate –

    Keep every K-th time sample (no anti-alias filtering — that's on the caller). Mutually exclusive with dt_target_s.

  • t_start_s, t_end_s –

    Trim the time axis to this window (inclusive of start, exclusive of end), applied after any resample / decimate.

abs_offset class-attribute

abs_offset = True

bool(x) -> bool

Returns True when the argument x is true, False otherwise. The builtins True and False are the only two instances of the class bool. The class bool is a subclass of the class int, and cannot be subclassed.

receiver_stride class-attribute

receiver_stride = 1

int([x]) -> integer int(x, base=10) -> integer

Convert a number or string to an integer, or return 0 if no arguments are given. If x is a number, return x.int(). For floating point numbers, this truncates towards zero.

If x is not a number or if base is given, then x must be a string, bytes, or bytearray instance representing an integer literal in the given base. The literal can be preceded by '+' or '-' and be surrounded by whitespace. The base defaults to 10. Valid bases are 0 and 2-36. Base 0 means to interpret the base from the string as an integer literal.

int('0b100', base=0) 4

shot_start class-attribute

shot_start = 0

int([x]) -> integer int(x, base=10) -> integer

Convert a number or string to an integer, or return 0 if no arguments are given. If x is a number, return x.int(). For floating point numbers, this truncates towards zero.

If x is not a number or if base is given, then x must be a string, bytes, or bytearray instance representing an integer literal in the given base. The literal can be preceded by '+' or '-' and be surrounded by whitespace. The base defaults to 10. Valid bases are 0 and 2-36. Base 0 means to interpret the base from the string as an integer literal.

int('0b100', base=0) 4

shot_stride class-attribute

shot_stride = 1

int([x]) -> integer int(x, base=10) -> integer

Convert a number or string to an integer, or return 0 if no arguments are given. If x is a number, return x.int(). For floating point numbers, this truncates towards zero.

If x is not a number or if base is given, then x must be a string, bytes, or bytearray instance representing an integer literal in the given base. The literal can be preceded by '+' or '-' and be surrounded by whitespace. The base defaults to 10. Valid bases are 0 and 2-36. Base 0 means to interpret the base from the string as an integer literal.

int('0b100', base=0) 4

sweep_io.plan.apply_data_plan

apply_data_plan(
    plan: DataPlan,
    geom: PhysicalGeometry,
    obs: ndarray,
    *,
    time_axis: int = -2,
    receiver_axis: int = -1
) -> tuple[PhysicalGeometry, np.ndarray, np.ndarray]

Apply the data plan to a (geometry, obs) pair.

Parameters:

  • plan (DataPlan) –

    :class:DataPlan to apply.

  • geom (PhysicalGeometry) –

    Source-of-truth geometry in physical units.

  • obs (ndarray) –

    Observed data tensor. Shape must include shot axis at 0, a time axis at time_axis, and a receiver axis at receiver_axis. The default layout matches sweep's binding output: (nshots, nt, nrec) (or (nshots, nt, nrec, nchannel)).

  • time_axis (int, default: -2 ) –

    Position of the time axis (default -2).

  • receiver_axis (int, default: -1 ) –

    Position of the receiver axis (default -1).

Returns:

  • geom_out ( PhysicalGeometry ) –

    Geometry restricted to the selected shots (receiver dim untouched because the receiver mask is per-shot non-uniform when offset filters are active).

  • obs_out ( ndarray ) –

    Observed data after shot subset + time resample.

  • receiver_keep_mask ( ndarray ) –

    (nshots_kept, nreceivers) boolean mask. The caller applies it per-shot to obs_out along receiver_axis because the kept count may vary per shot (offset window depends on the source). True entries indicate "use this receiver".

sweep_io.plan.ModelPlan

ModelPlan(
    x_window_m: tuple[float, float] | None = None,
    y_window_m: tuple[float, float] | None = None,
    z_window_m: tuple[float, float] | None = None,
    drop_outside_sources: bool = True,
    drop_outside_receivers: bool = True,
)

Crop a velocity model to a region of interest (axes are zero-indexed in meters).

Coordinate convention matches sweep: the model array has shape (nz, nx) for 2-D or (nz, ny, nx) for 3-D, and the physical origin of the array (index 0) is at origin_xyz_m (defaults to zero on every axis).

Attributes:

  • x_window_m, y_window_m, z_window_m –

    (low, high) inclusive windows in meters. None keeps the full axis.

  • drop_outside_sources, drop_outside_receivers –

    When True (default) acquisition geometry entries whose physical position falls outside the kept window are dropped.

drop_outside_receivers class-attribute

drop_outside_receivers = True

bool(x) -> bool

Returns True when the argument x is true, False otherwise. The builtins True and False are the only two instances of the class bool. The class bool is a subclass of the class int, and cannot be subclassed.

drop_outside_sources class-attribute

drop_outside_sources = True

bool(x) -> bool

Returns True when the argument x is true, False otherwise. The builtins True and False are the only two instances of the class bool. The class bool is a subclass of the class int, and cannot be subclassed.

sweep_io.plan.apply_model_plan

apply_model_plan(
    plan: ModelPlan,
    vp: ndarray,
    dh: tuple[float, ...],
    geom: PhysicalGeometry | None = None,
    *,
    origin_xyz_m: tuple[float, ...] | None = None
) -> tuple[np.ndarray, PhysicalGeometry | None, np.ndarray | None]

Crop a velocity model and (optionally) drop out-of-window acquisition.

Parameters:

  • plan (ModelPlan) –

    :class:ModelPlan.

  • vp (ndarray) –

    Model array, shape (nz, nx) (2-D) or (nz, ny, nx) (3-D).

  • dh (tuple[float, ...]) –

    Grid spacing tuple matching the model axes: (dz, dx) for 2-D or (dz, dy, dx) for 3-D.

  • geom (PhysicalGeometry | None, default: None ) –

    Optional acquisition. If given, sources / receivers outside the cropped window are dropped per the plan's flags. The returned geometry has its physical positions rebased so that the new cropped origin sits at (0, 0).

  • origin_xyz_m (tuple[float, ...] | None, default: None ) –

    Coordinate origin of the input model. Defaults to zero on every axis. Note that axis order matches vp.shape so it's (z_origin, x_origin) for 2-D.

Returns:

  • vp_out ( ndarray ) –

    Cropped model.

  • geom_out ( PhysicalGeometry | None ) –

    Cropped + rebased geometry, or None if no geom was given.

  • keep_shots_mask ( ndarray | None ) –

    (nshots,) boolean — surviving shots. None if no geom given or drop_outside_sources=False. The corresponding obs slice is obs[keep_shots_mask].

Geometry

sweep_io.geometry.Geometry

Geometry(sources: ndarray, receivers: ndarray, dt: float | None = None, nt: int | None = None, dh: Tuple[float, ...] | None = None, meta: dict[str, Any] = <factory>)

Source / receiver acquisition geometry.

Attributes:

  • sources –

    (nshots, ndim) grid indices. ndim=2 for 2-D (x, z), ndim=3 for 3-D (x, y, z).

  • receivers –

    (nshots, nreceivers, ndim) grid indices.

  • dt –

    Time sampling interval (s). Optional.

  • nt –

    Number of time samples. Optional.

  • dh –

    Spatial sampling per axis (same length as ndim). Optional.

  • meta –

    Arbitrary user metadata that gets serialized verbatim.

save

save(path: str | Path) -> None

Save as JSON (always) or .npz (binary, more compact).

sweep_io.geometry.PhysicalGeometry

PhysicalGeometry(sources_xyz_m: ndarray, receivers_xyz_m: ndarray, dt: float, nt: int, meta: dict[str, Any] = <factory>)

Source/receiver positions in meters — independent of any grid.

For multi-scale FWI the same geometry is reused across many dh; holding the physical form and calling :meth:to_grid per stage is the canonical pattern.

Attributes:

  • sources_xyz_m –

    (nshots, ndim) source coordinates in meters. ndim=2 is (x, z), ndim=3 is (x, y, z).

  • receivers_xyz_m –

    (nshots, nreceivers, ndim) receiver coordinates in meters.

  • dt –

    Time sampling interval (s).

  • nt –

    Number of time samples.

  • meta –

    Optional metadata dict.

to_grid

to_grid(
    dh: Tuple[float, ...] | float,
    *,
    origin_xyz_m: Tuple[float, ...] | None = None,
    dedupe: bool = True,
    dedup_method: Literal["nearest", "first"] = "nearest"
) -> Tuple[Geometry, np.ndarray]

Snap to nearest grid cell for the given dh.

Parameters:

  • dh (Tuple[float, ...] | float) –

    Grid spacing in meters. Scalar broadcasts to all axes; otherwise a tuple of length ndim.

  • origin_xyz_m (Tuple[float, ...] | None, default: None ) –

    Coordinate origin (the position of grid index 0). Defaults to zero on every axis.

  • dedupe (bool, default: True ) –

    When two or more receivers in the same shot snap to the same grid cell, drop the redundant ones. Sources are not deduped (each source already targets a distinct shot).

  • dedup_method (Literal['nearest', 'first'], default: 'nearest' ) –

    "nearest" — keep the receiver whose true position is closest to the grid-cell center. "first" — keep the first receiver in input order.

Returns:

  • geom ( Geometry ) –

    Grid-index geometry with dh populated.

  • keep_mask ( ndarray ) –

    Boolean array of shape (nshots, nreceivers). True where the receiver survived dedupe. Use this to slice observed-data tensors along the receiver axis so they line up with the surviving receivers per shot::

    gg, mask = pg.to_grid(dh=(25.0, 25.0))
    obs_aligned = [obs[s][..., mask[s], :] for s in range(pg.nshots)]
    

apply_rotation

apply_rotation(frame: RotatedFrame) -> PhysicalGeometry

Return a copy with the horizontal (x, y) coordinates rotated from UTM into the model frame defined by frame.

Only valid for ndim == 3. The vertical z coordinate is passed through unchanged. The returned geometry shares dt / nt and carries a meta['rotation'] payload recording the frame parameters so downstream consumers can round-trip back to UTM via :meth:RotatedFrame.to_utm.

sweep_io.geometry.RotatedFrame

RotatedFrame(
    origin_xy_utm: ndarray,
    rotation_matrix: ndarray,
    target_axis: Literal["x", "y"] = "x",
    inline_shift: float = 0.0,
    crossline_shift: float = 0.0,
)

UTM ↔ model-frame 2-D rotation.

Maps model_xy = (utm_xy - origin_xy_utm) @ R^T + shifts where R is a 2×2 rotation matrix. The model frame is what the propagator's axis-aligned grid consumes; the UTM frame is what field-data SEG-Y headers (and the OBN survey metadata) live in.

Sign convention matches the legacy rotation_metadata.json schema produced by fwi_workflow-dev: rotation_matrix is the matrix that rotates a UTM displacement vector (relative to origin_xy_utm) onto the model-frame axes. target_axis="x" means the inline azimuth points along model +x; "y" means inline points along model +y.

Parameters:

  • origin_xy_utm (ndarray) –

    (2,) UTM origin that maps to model-frame (0, 0) before any inline / crossline shift.

  • rotation_matrix (ndarray) –

    (2, 2) rotation matrix. If only a rotation angle is known, construct with :meth:from_rotation_deg.

  • target_axis (Literal['x', 'y'], default: 'x' ) –

    "x" (default) or "y" — which model axis is the inline.

  • inline_shift (float, default: 0.0 ) –

    Optional rigid offsets along the rotated inline / crossline axes (meters). Default 0.0.

  • crossline_shift (float, default: 0.0 ) –

    Optional rigid offsets along the rotated inline / crossline axes (meters). Default 0.0.

crossline_shift class-attribute

crossline_shift = 0.0

Convert a string or number to a floating point number, if possible.

inline_shift class-attribute

inline_shift = 0.0

Convert a string or number to a floating point number, if possible.

target_axis class-attribute

target_axis = 'x'

str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.str() (if defined) or repr(object). encoding defaults to sys.getdefaultencoding(). errors defaults to 'strict'.

from_rotation_deg classmethod

from_rotation_deg(
    *,
    origin_xy_utm: ndarray | Tuple[float, float],
    rotation_deg: float,
    target_axis: Literal["x", "y"] = "x",
    inline_shift: float = 0.0,
    crossline_shift: float = 0.0
) -> RotatedFrame

Build a frame from a single rotation angle (degrees, CCW).

from_metadata classmethod

from_metadata(metadata: str | Path | Mapping[str, object]) -> RotatedFrame

Build a frame from a rotation_metadata.json path or dict.

Compatible with the legacy fwi_workflow-dev schema: origin_xy, rotation_matrix, optional inline_shift, crossline_shift, target_axis.

Missing shifts default to 0.0; missing target_axis defaults to "x".

fit classmethod

fit(
    points_xy: ndarray,
    *,
    target_axis: Literal["x", "y"] = "x",
    shift_inline_to_zero: bool = False,
    shift_crossline_to_zero: bool = False,
    shift_points_xy: ndarray | None = None
) -> RotatedFrame

Derive a frame from scattered UTM points via a 2-D PCA line fit.

This is the companion constructor to :meth:from_rotation_deg / :meth:from_metadata: those consume a known rotation, while fit derives one from the acquisition geometry. It estimates the dominant survey-line azimuth from points_xy with :func:principal_direction and builds the frame that rotates that azimuth onto the model target_axis.

Ports the rotation estimation of fwi_workflow.geometry.line2d.rotate_csg_index: identical SVD principal direction, origin = mean(points), and [[cos, -sin], [sin, cos]] matrix convention. The resulting frame therefore serialises (via :meth:to_dict) to the same rotation_metadata.json schema and reproduces the legacy inline / crossline coordinates bit-for-bit.

Parameters:

  • points_xy (ndarray) –

    (npoints, 2) UTM xy used to estimate the line azimuth (e.g. one point per shot for a 2-D streamer line).

  • target_axis (Literal['x', 'y'], default: 'x' ) –

    Model axis the inline direction is rotated onto: "x" (default, 2-D streamer convention) or "y".

  • shift_inline_to_zero (bool, default: False ) –

    Shift the inline coordinate so its minimum is 0 (line starts at the grid edge). Matches the legacy 2-D streamer default.

  • shift_crossline_to_zero (bool, default: False ) –

    Shift the crossline coordinate so its minimum is 0. Combine with shift_inline_to_zero for the 3-D OBN first-quadrant layout.

  • shift_points_xy (ndarray | None, default: None ) –

    Points used to compute the shift minima; defaults to points_xy. Pass the union of sources + receivers to reproduce the legacy rotate_csg_index shift, which always minimised over both even when the rotation was fit on shots only.

Returns:

  • RotatedFrame –

    Frame whose :meth:to_model aligns the fitted line to target_axis and applies the requested shifts.

to_model

to_model(xy_utm: ndarray) -> np.ndarray

Project (..., 2) UTM xy to model-frame xy.

to_utm

to_utm(xy_model: ndarray) -> np.ndarray

Inverse of :meth:to_model — model-frame xy back to UTM xy.

sweep_io.geometry.load_rotation_metadata

load_rotation_metadata(path: str | Path) -> RotatedFrame

Build a :class:RotatedFrame from a JSON metadata file.

Thin wrapper around :meth:RotatedFrame.from_metadata for the file-path call style; matches the legacy fwi_workflow.geometry.line2d.load_rotation_transform.

Velocity models

sweep_io.models.load_velocity

load_velocity(
    path: str | Path,
    *,
    shape: Tuple[int, ...] | None = None,
    dtype: str | dtype = "float32",
    order: str = "C"
) -> np.ndarray

Load a velocity/parameter model.

Parameters:

  • path (str | Path) –

    File path. Format is dispatched by extension.

  • shape (Tuple[int, ...] | None, default: None ) –

    Used only for raw-binary loads (e.g. .bin, .raw, or unknown extensions). Ignored for self-describing formats.

  • dtype (Tuple[int, ...] | None, default: None ) –

    Used only for raw-binary loads (e.g. .bin, .raw, or unknown extensions). Ignored for self-describing formats.

  • order (Tuple[int, ...] | None, default: None ) –

    Used only for raw-binary loads (e.g. .bin, .raw, or unknown extensions). Ignored for self-describing formats.

Returns:

  • ndarray –

    Array as stored on disk. The caller is responsible for reshape / transpose — sweep-io does not impose (nz, nx) vs (nx, nz).

sweep_io.models.save_velocity

save_velocity(
    path: str | Path, arr: ndarray, *, header: dict[str, Any] | None = None
) -> None

Save a velocity/parameter model. Format dispatched by extension.

header is an optional metadata dict; honored only by formats that can store it (HDF5). Otherwise silently ignored.

Prefetching and datasets

sweep_io.prefetch.Prefetcher

Prefetcher(
    source: Iterable[T],
    *,
    queue_depth: int = 2,
    name: str | None = None,
    daemon: bool = True
)

Bases: typing.Generic

Yield items from source while a background thread reads ahead.

Parameters:

  • source (Iterable[T]) –

    Any iterable. Must be safe to consume from a single background thread (most generators are; only iterators with hard thread-affinity requirements are not).

  • queue_depth (int, default: 2 ) –

    Maximum number of items the worker may hold ahead of the consumer. 2 is the canonical double-buffered default. Higher values trade peak RAM for more slack against bursty compute / I/O.

  • name (str | None, default: None ) –

    Optional thread name (helps in py-spy / gdb).

  • daemon (bool, default: True ) –

    Whether the worker is a daemon thread. Default True so Python doesn't wait for it on interpreter exit.

Lifecycle

The worker starts in __init__ and runs until the source is exhausted, an exception is raised, or close() is called. Reusing a consumed instance is not supported; build a new one.

Examples:

>>> def load(i):
...     # imagine slow disk I/O here
...     return i * 2
>>> with Prefetcher((load(i) for i in range(5))) as pf:
...     out = list(pf)
>>> out
[0, 2, 4, 6, 8]

close

close(*, timeout: float = 1.0) -> None

Signal the worker to stop, drain the queue, join the thread.

sweep_io.prefetch.ThreadPoolPrefetcher

ThreadPoolPrefetcher(
    load_fn: Callable[[int], T],
    indices: Sequence[int] | Iterable[int],
    *,
    num_workers: int = 2,
    queue_depth: int = 2,
    name: str | None = None
)

Bases: typing.Generic

Order-preserving prefetcher with num_workers parallel loaders.

Use this when each item is loaded by an independent call to a function load(idx), the calls are I/O-bound (so threading helps), and there's enough parallel bandwidth to benefit from multiple in-flight reads.

Parameters:

  • load_fn (Callable[[int], T]) –

    Callable (idx) -> item. Must be thread-safe; if it mutates shared state, you must serialize access yourself.

  • indices (Sequence[int] | Iterable[int]) –

    Sequence of indices to feed load_fn. Items are yielded in the order of this sequence (not the order workers finish).

  • num_workers (int, default: 2 ) –

    Worker thread count. 2 is a sensible default for OS-cached spinning disk; 4-8 for NVMe or Lustre.

  • queue_depth (int, default: 2 ) –

    How far ahead of the consumer the worker pool may run. Roughly: peak RAM ≈ queue_depth × item_size.

Examples:

>>> def load(i): return i * i
>>> with ThreadPoolPrefetcher(load, range(5), num_workers=2) as pf:
...     out = list(pf)
>>> out
[0, 1, 4, 9, 16]

close

close() -> None

Shut down the executor, cancelling any pending futures.

sweep_io.prefetch.TimingPrefetcher

TimingPrefetcher(source: Iterable[T], *, queue_depth: int = 2)

Bases: sweep_io.prefetch.Prefetcher

Variant that also records per-item wait times.

Records, on the consumer side, how long __next__ blocked waiting for the worker. wait_times accumulates one entry per consumed item. Useful for benchmarking the prefetch / compute overlap:

  • all-zero wait times → consumer is the bottleneck (compute-bound)
  • large wait times → worker can't keep up (I/O-bound)

sweep_io.cuda_prefetch.CUDAPrefetcher

CUDAPrefetcher(
    source: Iterable[Any],
    *,
    device: str | device = "cuda",
    queue_depth: int = 2,
    stream: Stream | None = None
)

Wrap an iterable so each yielded item lands on GPU before compute uses it.

Parameters:

  • source (Iterable[Any]) –

    Any iterable. Items can be numpy arrays, torch tensors, or nested dict/list/tuple of those. Non-tensor leaves pass through.

  • device (str | device, default: 'cuda' ) –

    Target CUDA device ("cuda", "cuda:0", torch.device).

  • queue_depth (int, default: 2 ) –

    How many items the background thread keeps ready. 2 for double buffering; 3+ if compute is very fast / bursty.

  • stream (Stream | None, default: None ) –

    CUDA stream to issue H2D on. Defaults to a fresh stream on device.

Notes
  • The consumer must use the returned tensors on the current stream. We insert a GPU-side wait so the current stream sees the copy as completed; the host returns immediately.
  • For maximum throughput, pair this with torch.cuda.set_per_process_memory_fraction or a persistent pool to avoid pinned-allocator pressure under high queue depth.

Examples:

>>> import numpy as np
>>> def load(i): return np.random.randn(2048, 256).astype("float32")
>>> pf = CUDAPrefetcher((load(i) for i in range(N)),
...                     device="cuda:0", queue_depth=2)
>>> for shot in pf:
...     pred = solver(shot)

sweep_io.datasets.ShotGatherDataset

ShotGatherDataset(
    geometry: Geometry,
    obs: ndarray | Tensor | Callable[[int], Any],
    *,
    transform: Callable[[dict], dict] | None = None
)

Bases: torch.utils.data.dataset.Dataset

Per-shot dataset: each item is (sources_i, receivers_i, obs_i).

Parameters:

  • geometry (Geometry) –

    Acquisition geometry. Provides sources and receivers slices.

  • obs (ndarray | Tensor | Callable[[int], Any]) –

    Observed data, either: - a numpy array / torch tensor of shape (nshots, ...), or - a callable obs(shot_index) -> array_like for out-of-core access (e.g. one SEG-Y per shot).

  • transform (Callable[[dict], dict] | None, default: None ) –

    Optional transform(sample) -> sample applied lazily.

iter_prefetched

iter_prefetched(
    indices: Sequence[int] | Iterable[int] | None = None,
    *,
    queue_depth: int = 2,
    num_workers: int = 0
) -> Iterator[dict[str, Any]]

Iterate samples with a background prefetch thread.

Parameters:

  • indices (Sequence[int] | Iterable[int] | None, default: None ) –

    Sequence of shot indices to iterate. Defaults to range(len(self)).

  • queue_depth (int, default: 2 ) –

    How many samples to keep ready ahead of the consumer.

  • num_workers (int, default: 0 ) –

    0 → single background thread (good when the underlying __getitem__ does one big I/O call). >0 → a thread pool of that size (good when __getitem__ can run in parallel, e.g. distinct SEG-Y files per shot).

Notes

This is not torch.utils.data.DataLoader. There's no batch collation here — one item per next. Wrap with your own batching if you want it. For the typical FWI loop you want one shot at a time anyway, so the loop is:

for sample in ds.iter_prefetched():
    pred = solver(sample["sources"], sample["receivers"], ...)
    ...

CRG plan cache

sweep_io.crg_build.build_crg_plan_from_segy

build_crg_plan_from_segy(
    paths: Sequence[str | Path],
    *,
    byte_map: dict | None = None,
    source_depth_m_override: float | None = None,
    receiver_depth_m_override: float | None = None,
    coord_scalar_override: float | None = None,
    num_workers: int = 16,
    use_mpi: bool | None = None,
    mpi_stage_dir: str | Path | None = None,
    receiver_quantize_m: float = 0.5,
    receiver_stride: int = 1,
    receiver_max: int | None = None,
    receiver_ids: Sequence[int] | None = None,
    shots_per_source_line: int | None = None,
    source_line_stride: int = 1,
    rng_seed: int = 20260509,
    show_progress: bool = False
) -> CRGPlan | None

End-to-end: scan raw SEG-Y headers, then group into a :class:CRGPlan.

Convenience wrapper around :func:sweep_io.segy_index.build_segy_index (parallel header scan) followed by :func:build_crg_plan_from_index (receiver grouping + sub-sampling).

Parameters mirror both helpers. See :func:sweep_io.segy_index.build_segy_index for the SEG-Y header knobs (byte_map, source_depth_m_override etc.) and the arguments of :func:build_crg_plan_from_index for the receiver-side knobs.

Parallelism modes (mutually exclusive):

  • MPI (set use_mpi=True or run under mpiexec with use_mpi=None auto-detect): rank 0 partitions the file list, every rank scans a contiguous slice single-threaded, rank 0 gathers + merges. Non-root ranks return None.
  • ThreadPool (default, use_mpi=False): one process, num_workers background threads for the per-file scan.

The num_workers knob is ignored under MPI (MPI ranks ARE the parallelism; nesting threads inside MPI workers competes for file-descriptor cache without helping).

sweep_io.crg_build.build_crg_plan_from_index

build_crg_plan_from_index(
    index: SEGYIndex,
    *,
    receiver_quantize_m: float = 0.5,
    receiver_stride: int = 1,
    receiver_max: int | None = None,
    receiver_ids: Sequence[int] | None = None,
    shots_per_source_line: int | None = None,
    source_line_stride: int = 1,
    rng_seed: int = 20260509
) -> CRGPlan

Group an :class:SEGYIndex by quantised receiver position, then apply optional receiver / shot sub-sampling, into a :class:CRGPlan.

Quantisation: traces whose (rx, ry, rz) collapse to the same cell at receiver_quantize_m tolerance are considered the same OBN node. A regular node grid can use e.g. 0.5 m.

Parameters:

  • index (SEGYIndex) –

    Loaded :class:SEGYIndex over all source-line SEG-Y files. Built by :func:sweep_io.segy_index.build_segy_index.

  • receiver_quantize_m (float, default: 0.5 ) –

    Cell size (m) for receiver position quantisation.

  • receiver_stride (int, default: 1 ) –

    Keep every N-th unique receiver (after quantisation). 1 = all.

  • receiver_max (int | None, default: None ) –

    Cap the kept-receiver count after striding.

  • receiver_ids (Sequence[int] | None, default: None ) –

    Explicit receiver indices to keep (overrides stride / max). Indices refer to the receiver-id ordering AFTER quantisation + sort (same order CRGPlan.receiver_xyz_used is written in).

  • shots_per_source_line (int | None, default: None ) –

    Per-receiver cap on shots picked from each source-line file. None keeps all; e.g. 4 keeps at most four shots per line when building the canonical plan from CRG.

  • source_line_stride (int, default: 1 ) –

    Drop every Nth source line per receiver before the per-line cap. 1 keeps all lines.

  • rng_seed (int, default: 20260509 ) –

    Deterministic RNG seed for the per-line shot sub-sampler.

Returns:

  • CRGPlan –

    Ready to feed obs.crg_plan.plan_path in a sweep-tasks FWI YAML, or to save_crg_plan for the on-disk cache.

sweep_io.crg_plan.CRGPlan

CRGPlan(
    files: list[Path],
    trace_size_per_file: ndarray,
    sample_format: int,
    samples_per_trace: int,
    dt_s: float,
    receiver_indices: ndarray,
    receiver_xyz_used: ndarray,
    plan_file_ids: ndarray,
    plan_trace_offsets: ndarray,
    plan_sx: ndarray,
    plan_sy: ndarray,
    plan_sz: ndarray,
    plan_offsets: ndarray,
)

Slim per-virtual-source CRG plan loaded from a crg_fwi_plan_v1 cache.

Each i ∈ [0, n_receivers_used) is one virtual source. Its physical-shot rows are plan_*[slot_slice(i)].

Attributes:

  • files –

    SEG-Y source-line files (per-file read_traces dispatch keys are plan_file_ids).

  • trace_size_per_file –

    (n_files,) byte stride per trace (header + samples).

  • sample_format, samples_per_trace, dt_s –

    SEG-Y trace parameters; identical across files in a plan.

  • n_receivers_used –

    Number of virtual sources (OBN nodes) in the plan.

  • receiver_xyz_used –

    (n_receivers_used, 3) virtual-source positions in UTM (m). These become the FWI sources.

  • plan_file_ids, plan_trace_offsets –

    Flat per-row arrays of (file index, byte offset) tuples handed to :meth:sweep_io.segy.MultiFileSEGYReader.read_traces.

  • plan_sx, plan_sy, plan_sz –

    Flat per-row physical shot positions in UTM (m). These become the FWI receivers (one per recorded trace).

  • plan_offsets –

    (n_receivers_used + 1,) CSR-style row pointer: plan_offsets[i] : plan_offsets[i+1] are the row indices of slot i's gather.

virtual_source_xy property

virtual_source_xy

(n_receivers_used, 2) virtual-source horizontal coordinates (m).

slot_slice

slot_slice(slot: int) -> slice

Flat-row slice for the slot-th virtual source.

slot_row_count

slot_row_count(slot: int) -> int

Number of physical shots that hit virtual source slot.

per_slot_row_counts

per_slot_row_counts() -> np.ndarray

(n_receivers_used,) shot count per virtual source.

slot_receiver_xyz

slot_receiver_xyz(slot: int) -> np.ndarray

Position of the slot-th OBN node (= virtual source position).

slot_shot_xyz

slot_shot_xyz(slot: int) -> np.ndarray

(n_shots_for_slot, 3) physical shot positions (= virtual recv).

filter_by_row_mask

filter_by_row_mask(row_mask: ndarray) -> CRGPlan

Return a new plan keeping only flat rows where row_mask is True.

The per-slot plan_offsets are recomputed so each slot still addresses a contiguous run of its surviving rows. Slots that lose ALL their rows survive with a zero-length slice — caller should compose with :meth:filter_by_coverage to drop them entirely.

Used by the OBN CRG path to drop physical shots whose model-frame position falls outside the (optionally cropped) inversion grid.

filter_by_coverage

filter_by_coverage(min_shots: int) -> CRGPlan

Return a new plan keeping only slots with >= min_shots shots.

Mirrors the legacy min_coverage filter — drops edge / partial- coverage OBN nodes that record so few shots they would hurt source-encoded supershot inversions (per crg_data._sample_crg_iter_shared_shots).

sweep_io.crg_plan.load_crg_shot_plan_cache

load_crg_shot_plan_cache(cache_path: str | Path) -> CRGPlan

Load a crg_fwi_plan_v1 npz cache into a :class:CRGPlan.

SEG-Y paths recorded in the cache are remapped via :func:_remap_segy_path so the same cache works across hosts when FWI_SEGY_ROOT or FWI_SEGY_REMAP is exported.

sweep_io.crg_dataset.CRGBatchDataset

CRGBatchDataset(
    plan: CRGPlan,
    *,
    slot_indices: ndarray | None = None,
    max_shots_per_slot: int | None = None
)

Dataset yielding one virtual-source gather per __getitem__ call.

Designed for torch.utils.data.DataLoader with num_workers > 0 + :func:pad_collate_fn + :func:worker_init_fn.

Parameters:

  • plan (CRGPlan) –

    Loaded :class:CRGPlan.

  • slot_indices (ndarray | None, default: None ) –

    Subset of slot indices to expose (default: all). Use this to shard slots across distributed ranks before the DataLoader.

  • max_shots_per_slot (int | None, default: None ) –

    Truncate each slot's shot list to this length (or None to keep them all). Truncation is deterministic (first max_shots); per-iter random sub-sampling belongs upstream in the runner.

n_shots_max property

n_shots_max

Max gather length across exposed slots — for n_recv_max sizing.

n_shots_per_slot property

n_shots_per_slot

(len(self),) shot count after max_shots_per_slot cap.

Wavelets

sweep_io.wavelet.load_wavelet_npz

load_wavelet_npz(
    path: str | Path,
    *,
    prefer_keys: Sequence[str] = (
        "wavelet",
        "optimized_siren_wavelet",
        "direct_causal_wavelet",
        "initial_causal_wavelet",
    ),
    explicit_key: str | None = None
) -> WaveletNPZ

Read a wavelet from an .npz file.

The selected array key is, in order: explicit_key (if non-None); otherwise the first key in prefer_keys present in the npz. Raises KeyError if none match.

The sample interval is taken from dt_s (scalar) when present; otherwise inferred from time_s via the median diff (matches the legacy fwi_workflow-dev behaviour). If neither is present a KeyError is raised — sweep-stack will not invent a dt.

Parameters:

  • path (str | Path) –

    Path to the .npz file.

  • prefer_keys (Sequence[str], default: ('wavelet', 'optimized_siren_wavelet', 'direct_causal_wavelet', 'initial_causal_wavelet') ) –

    Ordered fallback list of npz keys for the wavelet array.

  • explicit_key (str | None, default: None ) –

    Force-select this key; bypasses prefer_keys.

sweep_io.wavelet.WaveletNPZ

WaveletNPZ(
    samples: ndarray, dt_s: float, source_delay_s: float | None = None, key: str = ""
)

Loaded source-wavelet payload.

Attributes:

  • samples –

    1-D float32 array of length nt with the wavelet amplitudes.

  • dt_s –

    Sample interval in seconds.

  • source_delay_s –

    Optional zero-prepad in front of the wavelet (s). The SIREN wavelet-fit pipeline writes this so downstream FWI can left-pad observed traces by the same amount; None when absent.

  • key –

    Which npz key the samples were read from (helpful for logs).

key class-attribute

key = ''

str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.str() (if defined) or repr(object). encoding defaults to sys.getdefaultencoding(). errors defaults to 'strict'.