API reference¶
Generated from the source docstrings. Every module is importable on its own, e.g.
from sweep_io.seismic_plan import PlanReader.
SEG-Y¶
sweep_io.segy.SEGYReader ¶
Single-file SEG-Y reader keyed by absolute byte offsets.
On construction, parses the 400-byte binary header to extract
n_samples, dt, and sample_format. Trace data is then
read on demand by absolute byte offset — no header re-parsing
per trace, no Python loop per sample.
Designed to be the fast lane behind a pre-built shot index. Pair it
with :class:MultiFileSEGYReader when shots live across many files.
Parameters:
-
path(str | PathLike) –File path.
-
mmap_mode(bool, default:True) –If
True(default), memory-map the file (zero-copy reads, kernel page cache reuse). IfFalse, usepread/seek+read. mmap can misbehave on certain network filesystems — tryFalseif you seeSIGBUSor weird truncation.
Concurrency
A single SEGYReader instance is thread-safe for reads when
mmap_mode=True (each read_trace_data call takes a numpy view
of the mmap; no shared file pointer). With mmap_mode=False reads
use os.pread, which is also thread-safe. So either way you can
share one reader across all your prefetch workers.
read_trace_data ¶
Read trace data at given absolute byte offsets.
Each offset must point at a trace header (not at the data block — the 240-byte header is skipped internally).
Parameters:
-
byte_offsets(Sequence[int] | ndarray) –1-D array of absolute byte offsets, one per trace to read. Order is preserved in the output, but the actual disk reads are issued in ascending order (and coalesced where possible) for sequential bandwidth.
-
coalesce_gap(int, default:0) –Adjacent reads whose gap (between the end of one and the start of the next) is ≤ this many bytes are merged into a single
pread.0disables coalescing; sensible values on Lustre / NVMe are0(already sorted-sequential is fast enough) up through one filesystem block (~1 MB).
Returns:
-
ndarray–(n_traces, n_samples)float32 array, ordered the same asbyte_offsets.
read_trace_header ¶
Return the 240-byte trace header at byte_offset.
sweep_io.segy.MultiFileSEGYReader ¶
Dispatch byte-offset reads across many SEG-Y files by file_id.
Holds one open :class:SEGYReader per path. All readers must agree
on n_samples, dt, and sample_format (the usual case for a
single survey split across many files). A mismatch raises at construction.
Parameters:
-
paths(Sequence[str | PathLike]) –List of file paths.
file_id = 0corresponds topaths[0], etc. -
mmap_mode(bool, default:True) –Passed through to each underlying :class:
SEGYReader.
read_traces ¶
read_traces(
file_ids: Sequence[int] | ndarray,
byte_offsets: Sequence[int] | ndarray,
*,
coalesce_gap: int = 0
) -> np.ndarray
Read traces from heterogeneous files, preserving caller order.
Parameters:
-
file_ids(Sequence[int] | ndarray) –1-D array, same length as
byte_offsets. Index intoself.readers. -
byte_offsets(Sequence[int] | ndarray) –1-D array of trace-header offsets within each file.
-
coalesce_gap(int, default:0) –Forwarded to each per-file read.
Returns:
-
ndarray–(n, n_samples)float32, ordered by the caller.
sweep_io.segy.read_segy ¶
read_segy(
path: str | PathLike, *, iline: int = 189, xline: int = 193, strict: bool = False
) -> Tuple[np.ndarray, dict]
Read an entire SEG-Y volume into memory (via the optional segyio dep).
sweep_io.segy.write_segy ¶
write_segy(
path: str | PathLike, data: ndarray, *, dt: float, delay_recording_time: float = 0.0
) -> None
Write a 2-D (n_traces, n_samples) SEG-Y via segyio.
sweep_io.segy.write_segy_minimal ¶
write_segy_minimal(
path: str | PathLike, data: ndarray, *, dt: float, sample_format: int = 5
) -> None
Write a SEG-Y file with empty text/trace headers — useful for fixtures.
No segyio required. Produces a valid layout that :class:SEGYReader
can parse:
- 3200-byte zeroed text header
- 400-byte binary header with dt / n_samples / sample_format set
- one (zeroed trace header + sample data) per row
Only formats 1 (IBM) and 5 (IEEE32) are supported here; we don't write integers from this helper.
Header index¶
sweep_io.segy_index.build_segy_index ¶
build_segy_index(
paths: Sequence[str | Path],
*,
byte_map: dict | None = None,
source_depth_m_override: float | None = None,
receiver_depth_m_override: float | None = None,
coord_scalar_override: float | None = None,
num_workers: int = 4,
show_progress: bool = False
) -> SEGYIndex
Scan headers of paths in parallel and build a unified :class:SEGYIndex.
Parameters:
-
paths(Sequence[str | Path]) –SEG-Y file paths in the order you want them assigned
file_id. -
byte_map(dict | None, default:None) –Trace-header byte-offset overrides. Defaults to SEG-Y rev1 standard (:data:
SEGY_REV1_BYTES). Override per-key, e.g.{"source_depth": 44}for an OBN airgun depth. -
source_depth_m_override(float | None, default:None) –If the SEG-Y file's depth bytes are zero (or unreliable), use these instead. Common case for marine streamers.
-
receiver_depth_m_override(float | None, default:None) –If the SEG-Y file's depth bytes are zero (or unreliable), use these instead. Common case for marine streamers.
-
coord_scalar_override(float | None, default:None) –If set, ignore the per-trace
coord_scalarfield and apply this constant multiplier to (sx, sy, rx, ry). Use when you know the files have inconsistent or wrong scalar values. -
num_workers(int, default:4) –Parallel header-scan workers.
find -name "*.sgy" | wc -lis a good upper bound; 4-8 is typical for SSD / Lustre. -
show_progress(bool, default:False) –Print one line per finished file. Useful when scanning hundreds.
Returns:
-
SEGYIndex–With
len(paths)distinct file_id values.
sweep_io.segy_index.SEGYIndex ¶
SEGYIndex(file_paths: list[str], file_id: ndarray, byte_offset: ndarray, shot_id: ndarray, receiver_id: ndarray, sx_m: ndarray, sy_m: ndarray, sz_m: ndarray, rx_m: ndarray, ry_m: ndarray, rz_m: ndarray, dt_s: float, n_samples: int, sample_format: int = 5, meta: dict = <factory>, SCHEMA_VERSION: int = 1)
Per-trace catalog backed by 1-D numpy columns.
All trace-level arrays have the same length = total trace count across
all files. file_id[i] indexes :attr:file_paths. shot_id[i] is
the group key (usually FFID); the same value across all traces of one
shot. receiver_id[i] is the trace's index within its shot (0-based
after sort).
Per-shot fast lookup is built on construction via
:meth:_build_shot_lookup. lookup_shot(shot_id) returns the slice
in O(log n).
SCHEMA_VERSION
class-attribute
¶
int([x]) -> integer int(x, base=10) -> integer
Convert a number or string to an integer, or return 0 if no arguments are given. If x is a number, return x.int(). For floating point numbers, this truncates towards zero.
If x is not a number or if base is given, then x must be a string, bytes, or bytearray instance representing an integer literal in the given base. The literal can be preceded by '+' or '-' and be surrounded by whitespace. The base defaults to 10. Valid bases are 0 and 2-36. Base 0 means to interpret the base from the string as an integer literal.
int('0b100', base=0) 4
sample_format
class-attribute
¶
int([x]) -> integer int(x, base=10) -> integer
Convert a number or string to an integer, or return 0 if no arguments are given. If x is a number, return x.int(). For floating point numbers, this truncates towards zero.
If x is not a number or if base is given, then x must be a string, bytes, or bytearray instance representing an integer literal in the given base. The literal can be preceded by '+' or '-' and be surrounded by whitespace. The base defaults to 10. Valid bases are 0 and 2-36. Base 0 means to interpret the base from the string as an integer literal.
int('0b100', base=0) 4
lookup_shot ¶
Return all per-trace + source info for one shot. O(log n_shots).
to_physical_geometry ¶
Materialise a :class:sweep_io.geometry.PhysicalGeometry.
Expects every shot to have the same receiver count (typical for streamers). For OBN / variable-receiver-count layouts, slice the index by shot_id first.
to_seismic_plan ¶
Convenience: build a :class:sweep_io.seismic_plan.SeismicPlan.
All kwargs forward to :func:sweep_io.seismic_plan.build_seismic_plan.
Use grouping='csg' (default) for shot-gather organisation
(typical 2-D / 3-D streamer) or grouping='crg' (plus
receiver_quantize_m) for common-receiver gathers (OBN).
Examples:
sweep_io.segy_index.IndexedShotGatherDataset ¶
IndexedShotGatherDataset(
index: SEGYIndex,
reader: MultiFileSEGYReader | None = None,
*,
shot_ids: Sequence[int] | None = None,
coalesce_gap: int = 0
)
Per-shot dataset backed by :class:SEGYIndex + :class:MultiFileSEGYReader.
Each __getitem__ issues one coalesced byte-offset read for the
selected shot's traces (sorted ascending, runs merged). Output shape
is (n_recv, n_samples) float32.
Pair with :class:sweep_io.prefetch.Prefetcher or
:class:sweep_io.cuda_prefetch.CUDAPrefetcher for I/O / compute overlap.
Parameters:
-
index(SEGYIndex) –Pre-built SEGYIndex.
-
reader(MultiFileSEGYReader | None, default:None) –Optional :class:
MultiFileSEGYReader. Built fromindex.file_pathsif omitted. -
shot_ids(Sequence[int] | None, default:None) –Subset (and order) of shot IDs to expose. Default: all shots, sorted.
-
coalesce_gap(int, default:0) –Forwarded to
reader.read_traces; merges adjacent reads with at most this many bytes of gap.
Plans and readers¶
sweep_io.seismic_plan.build_seismic_plan ¶
build_seismic_plan(
index,
*,
grouping: str = "csg",
shot_ids: Sequence[int] | None = None,
receiver_quantize_m: float | None = None,
offset_min_m: float | None = None,
offset_max_m: float | None = None,
max_traces_per_group: int | None = None,
seed: int = 0,
build_label: str | None = None
) -> SeismicPlan
Build a :class:SeismicPlan from a :class:SEGYIndex.
Parameters:
-
index–A :class:
sweep_io.segy_index.SEGYIndex— the full per-trace catalog. Not mutated. -
grouping(str, default:'csg') –"csg"(default) → one group per uniqueshot_id(FFID)."crg"→ one group per receiver cell (usereceiver_quantize_mto define the cell size). -
shot_ids(Sequence[int] | None, default:None) –Optional whitelist of shot IDs to keep (intersected with what the index has).
None= all shots in the index. -
receiver_quantize_m(float | None, default:None) –For
grouping="crg": quantization tolerance for collapsing per-trace receiver positions to a finite set of cells. Required for CRG; ignored for CSG. -
offset_min_m(float | None, default:None) –Per-trace source-receiver horizontal-offset filter (meters). Traces falling outside
[offset_min_m, offset_max_m]are dropped pre-grouping. -
offset_max_m(float | None, default:None) –Per-trace source-receiver horizontal-offset filter (meters). Traces falling outside
[offset_min_m, offset_max_m]are dropped pre-grouping. -
max_traces_per_group(int | None, default:None) –Hard cap on rows per group after grouping. Excess rows are randomly sub-sampled (seeded via
seed). -
seed(int, default:0) –RNG seed for the per-group cap sub-sampling.
-
build_label(str | None, default:None) –Optional human-readable tag stored in
build_metafor provenance.
sweep_io.seismic_plan.SeismicPlan ¶
SeismicPlan(grouping: str, files: list[Path], trace_size_per_file: ndarray, sample_format: int, samples_per_trace: int, dt_s: float, row_file_id: ndarray, row_trace_offset: ndarray, row_source_xyz: ndarray, row_receiver_xyz: ndarray, group_id: ndarray, group_xyz: ndarray, group_offsets: ndarray, build_meta: dict = <factory>)
A filter + grouping view over a :class:SEGYIndex, persisted to npz.
Carries enough metadata to read the underlying SEG-Y traces lazily
(via :class:PlanReader) but holds NO sample data itself —
on-disk size is O(n_rows × 80 bytes), not O(n_rows × nt × 4).
Construct via :func:build_seismic_plan or load via :meth:load.
filter_rows ¶
Drop rows where row_mask is False; recompute group_offsets.
Groups that lose every row survive with a zero-length slice;
compose with :meth:drop_empty_groups to remove them.
filter_groups ¶
Drop groups where group_mask is False; recompute offsets + row arrays.
drop_empty_groups ¶
Drop groups with zero rows; preserves order of surviving groups.
sweep_io.seismic_plan.PlanReader ¶
PlanReader(
plan: SeismicPlan,
*,
mmap: bool = True,
cache_all: bool = False,
trace_cache_bytes: int = 0,
coalesce_gap: int = 0
)
Lazy SEG-Y trace reader bound to a :class:SeismicPlan.
The plan tells what to read (file_ids + byte offsets, grouped);
this reader does the actual IBM-float decoding via
:class:MultiFileSEGYReader. Trace order within a group is
preserved (matches plan.row_*[group_slice(g)]).
Parameters:
-
plan(SeismicPlan) –A :class:
SeismicPlan(typically loaded from disk). -
mmap(bool, default:True) –Whether the underlying SEG-Y reader should mmap files. Default
True— the OS page cache then shares pages across processes in distributed runs. -
cache_all(bool, default:False) –When
True, read every plan row into a single(n_rows, nt)float32 tensor at construction and serve subsequentread_group(g)calls from the cache. Useful for small datasets (e.g. 2-D Viking — ~700 MB fits in RAM trivially). For large 3-D surveys leave itFalseso memory stays bounded.
Parameters:
-
plan(SeismicPlan) –A :class:
SeismicPlan(typically loaded from disk). -
mmap(bool, default:True) –Whether the underlying SEG-Y reader should mmap files.
-
cache_all(bool, default:False) –When True, eagerly read every plan row into a
(n_rows, nt)float32 tensor at construction. -
trace_cache_bytes(int, default:0) –Optional in-RAM LRU cache for decoded traces, keyed by
(file_id, byte_offset). Used by :meth:read_rowsto absorb cold-Lustre re-reads when the same physical traces are sampled across multiple iterations (the OBN multisource supershot pattern).0disables.-1means unbounded (matches the legacy--trace-cache-bytes -1default for large surveys). Any positive value is a byte budget. Ignored whencache_all=True(the full cache is already materialised). -
coalesce_gap(int, default:0) –Bytes-within-the-same-file gap below which two miss-traces get merged into a single
preadcall.0disables. Sensible value: 4× the trace stride (~240 + nt*4bytes) so any two sequentially-indexed traces in the same file get coalesced. Critical for Lustre where per-call latency dominates small random reads. Only used by :meth:read_rows(the cached path);read_groupalways reads the full group slice which is already contiguous.
trace_cache_stats
property
¶
Cache hit/miss counters + current byte usage, or None when
trace_cache_bytes=0 was set at construction.
read_groups ¶
Return one (n_in_g, nt) array per group id in gs.
read_rows ¶
Return (n_picked, nt) float32 for an arbitrary set of plan rows.
Used by sampler-driven loaders (e.g. CRG shared-shot batches) that
pick row indices across multiple groups rather than reading whole
groups. Row order in the output matches row_idx.
Three code paths, in priority order:
cache_all=Trueat construction → fancy-index slice on the in-RAM(n_rows, nt)tensor (single numpy op, instant).trace_cache_bytesset (per-trace LRU +coalesce_gap) → :func:sweep_io.crg_dataset.read_traces_cached. This is the legacy OBN supershot IO path: hot traces (re-sampled across iters) come from RAM, cold traces are coalesced into biggerpreadcalls. Cuts wait_io about five-fold on a large survey.- Otherwise → raw :meth:
MultiFileSEGYReader.read_traces(onepreadper miss, no caching). Slow on Lustre with random access patterns; only useful for one-shot reads.
sweep_io.seismic_plan.sample_shared_shots_from_plan ¶
sample_shared_shots_from_plan(
plan: SeismicPlan,
rng: Generator,
*,
batch_size: int,
source_lines_per_group: int = 0,
max_traces_per_sourceline: int = 0,
min_coverage: int = 0,
eligible_groups: ndarray | None = None,
max_retries: int = 10,
precomputed_group_unique_keys: list[ndarray] | None = None
) -> SharedShotBatch
Pick batch_size groups whose physical-shot sets intersect.
Operates on a grouping='crg' :class:SeismicPlan: each group is a
virtual source (e.g. one OBN node, or a quantised receiver cell), and
its rows are the physical shots that group recorded. The intersection
across the chosen groups gives the shots EVERY picked group recorded —
the invariant needed for source-encoded supershot FWI.
This is the unified-plan port of
:func:sweep_io.crg_plan.sample_crg_shared_shots; behaviour is
deliberately byte-equivalent so existing runner code can swap one for
the other with no FP drift (assuming the same RNG state).
Optional sub-sampling within the intersection:
source_lines_per_group > 0— keep that many random source-linefile_ids (= source-line SEG-Y files).max_traces_per_sourceline > 0— keep that many random shots per kept source line.0for either disables that level.
Parameters:
-
plan(SeismicPlan) –A
grouping='crg':class:SeismicPlan(CSG plans cannot be sampled this way — each CSG group has a single source, so the intersection is trivially that one source). -
rng(Generator) –Caller-owned
np.random.Generator(advanced in-place). -
batch_size(int) –Number of groups (= virtual sources) per batch.
-
min_coverage(int, default:0) –Drop groups whose row count is below this.
0disables. -
eligible_groups(ndarray | None, default:None) –Optional pre-curated set of group indices to sample from. When
Noneandmin_coverage > 0, computed from :meth:SeismicPlan.per_group_row_counts. -
max_retries(int, default:10) –Re-roll group selection up to this many times when the picked groups' shot intersection is empty (rare with
min_coverageon but possible for partial-coverage edge nodes).
Raises:
-
ValueError–When
plan.grouping != 'crg', or when the shot intersection is empty aftermax_retriesattempts, or when sub-sampling empties the intersection.
sweep_io.seismic_plan.sample_percrg_independent ¶
sample_percrg_independent(
plan: SeismicPlan,
rng: Generator,
*,
batch_size: int,
source_lines_per_group: int = 0,
max_traces_per_sourceline: int = 0,
min_coverage: int = 0,
eligible_groups: ndarray | None = None
) -> PerCRGBatch
Pick batch_size CRG groups; each keeps its OWN sub-sampled rows.
Unlike :func:sample_shared_shots_from_plan, this does NOT intersect the
groups' shot sets — there is no shared receiver geometry. The sub-sampling
knobs (source_lines_per_group → keep that many random source-line
file_ids; max_traces_per_sourceline → that many random shots per
kept line) are applied PER GROUP on that group's own recorded shots. Use
for the per-shot (non-encoded) FWI path so each CRG node inverts its full
aperture; over iterations the stochastic per-node draw sweeps each node's
entire coverage.
Grid-cell dedup of the receivers is deliberately left to the caller — it needs the model frame / grid origin this plan-space sampler is agnostic to.
Data and model windows¶
sweep_io.plan.DataPlan ¶
DataPlan(
shot_start: int = 0,
shot_stop: int | None = None,
shot_stride: int = 1,
shot_indices: Sequence[int] | ndarray | None = None,
receiver_stride: int = 1,
offset_min_m: float | None = None,
offset_max_m: float | None = None,
abs_offset: bool = True,
dt_target_s: float | None = None,
time_decimate: int | None = None,
t_start_s: float | None = None,
t_end_s: float | None = None,
)
Which shots / receivers / time samples to feed into FWI.
Attributes:
-
shot_start, shot_stop, shot_stride–Standard range-style shot selection. Ignored when
shot_indicesis given. -
shot_indices–Explicit list of shot indices to keep. Overrides the range fields.
-
receiver_stride–Keep every
N-th receiver per shot (uniform across shots). -
offset_min_m, offset_max_m–Receiver-offset window in meters (signed if
abs_offset=False, else compared to|offset|). Offsets are computed in the horizontal plane only —ndim=2uses|rx - sx|,ndim=3usessqrt((rx-sx)^2 + (ry-sy)^2). -
abs_offset–Treat offsets as absolute values when comparing to the window.
-
dt_target_s–Time-resample the observed data to this dt (delegates to
sweep_io._resample.resample_time). Mutually exclusive withtime_decimate. -
time_decimate–Keep every
K-th time sample (no anti-alias filtering — that's on the caller). Mutually exclusive withdt_target_s. -
t_start_s, t_end_s–Trim the time axis to this window (inclusive of start, exclusive of end), applied after any resample / decimate.
abs_offset
class-attribute
¶
bool(x) -> bool
Returns True when the argument x is true, False otherwise. The builtins True and False are the only two instances of the class bool. The class bool is a subclass of the class int, and cannot be subclassed.
receiver_stride
class-attribute
¶
int([x]) -> integer int(x, base=10) -> integer
Convert a number or string to an integer, or return 0 if no arguments are given. If x is a number, return x.int(). For floating point numbers, this truncates towards zero.
If x is not a number or if base is given, then x must be a string, bytes, or bytearray instance representing an integer literal in the given base. The literal can be preceded by '+' or '-' and be surrounded by whitespace. The base defaults to 10. Valid bases are 0 and 2-36. Base 0 means to interpret the base from the string as an integer literal.
int('0b100', base=0) 4
shot_start
class-attribute
¶
int([x]) -> integer int(x, base=10) -> integer
Convert a number or string to an integer, or return 0 if no arguments are given. If x is a number, return x.int(). For floating point numbers, this truncates towards zero.
If x is not a number or if base is given, then x must be a string, bytes, or bytearray instance representing an integer literal in the given base. The literal can be preceded by '+' or '-' and be surrounded by whitespace. The base defaults to 10. Valid bases are 0 and 2-36. Base 0 means to interpret the base from the string as an integer literal.
int('0b100', base=0) 4
shot_stride
class-attribute
¶
int([x]) -> integer int(x, base=10) -> integer
Convert a number or string to an integer, or return 0 if no arguments are given. If x is a number, return x.int(). For floating point numbers, this truncates towards zero.
If x is not a number or if base is given, then x must be a string, bytes, or bytearray instance representing an integer literal in the given base. The literal can be preceded by '+' or '-' and be surrounded by whitespace. The base defaults to 10. Valid bases are 0 and 2-36. Base 0 means to interpret the base from the string as an integer literal.
int('0b100', base=0) 4
sweep_io.plan.apply_data_plan ¶
apply_data_plan(
plan: DataPlan,
geom: PhysicalGeometry,
obs: ndarray,
*,
time_axis: int = -2,
receiver_axis: int = -1
) -> tuple[PhysicalGeometry, np.ndarray, np.ndarray]
Apply the data plan to a (geometry, obs) pair.
Parameters:
-
plan(DataPlan) –:class:
DataPlanto apply. -
geom(PhysicalGeometry) –Source-of-truth geometry in physical units.
-
obs(ndarray) –Observed data tensor. Shape must include shot axis at
0, a time axis attime_axis, and a receiver axis atreceiver_axis. The default layout matches sweep's binding output:(nshots, nt, nrec)(or(nshots, nt, nrec, nchannel)). -
time_axis(int, default:-2) –Position of the time axis (default
-2). -
receiver_axis(int, default:-1) –Position of the receiver axis (default
-1).
Returns:
-
geom_out(PhysicalGeometry) –Geometry restricted to the selected shots (receiver dim untouched because the receiver mask is per-shot non-uniform when offset filters are active).
-
obs_out(ndarray) –Observed data after shot subset + time resample.
-
receiver_keep_mask(ndarray) –(nshots_kept, nreceivers)boolean mask. The caller applies it per-shot toobs_outalongreceiver_axisbecause the kept count may vary per shot (offset window depends on the source).Trueentries indicate "use this receiver".
sweep_io.plan.ModelPlan ¶
ModelPlan(
x_window_m: tuple[float, float] | None = None,
y_window_m: tuple[float, float] | None = None,
z_window_m: tuple[float, float] | None = None,
drop_outside_sources: bool = True,
drop_outside_receivers: bool = True,
)
Crop a velocity model to a region of interest (axes are zero-indexed in meters).
Coordinate convention matches sweep: the model array has shape
(nz, nx) for 2-D or (nz, ny, nx) for 3-D, and the physical
origin of the array (index 0) is at origin_xyz_m (defaults to
zero on every axis).
Attributes:
-
x_window_m, y_window_m, z_window_m–(low, high)inclusive windows in meters.Nonekeeps the full axis. -
drop_outside_sources, drop_outside_receivers–When
True(default) acquisition geometry entries whose physical position falls outside the kept window are dropped.
drop_outside_receivers
class-attribute
¶
bool(x) -> bool
Returns True when the argument x is true, False otherwise. The builtins True and False are the only two instances of the class bool. The class bool is a subclass of the class int, and cannot be subclassed.
drop_outside_sources
class-attribute
¶
bool(x) -> bool
Returns True when the argument x is true, False otherwise. The builtins True and False are the only two instances of the class bool. The class bool is a subclass of the class int, and cannot be subclassed.
sweep_io.plan.apply_model_plan ¶
apply_model_plan(
plan: ModelPlan,
vp: ndarray,
dh: tuple[float, ...],
geom: PhysicalGeometry | None = None,
*,
origin_xyz_m: tuple[float, ...] | None = None
) -> tuple[np.ndarray, PhysicalGeometry | None, np.ndarray | None]
Crop a velocity model and (optionally) drop out-of-window acquisition.
Parameters:
-
plan(ModelPlan) –:class:
ModelPlan. -
vp(ndarray) –Model array, shape
(nz, nx)(2-D) or(nz, ny, nx)(3-D). -
dh(tuple[float, ...]) –Grid spacing tuple matching the model axes:
(dz, dx)for 2-D or(dz, dy, dx)for 3-D. -
geom(PhysicalGeometry | None, default:None) –Optional acquisition. If given, sources / receivers outside the cropped window are dropped per the plan's flags. The returned geometry has its physical positions rebased so that the new cropped origin sits at
(0, 0). -
origin_xyz_m(tuple[float, ...] | None, default:None) –Coordinate origin of the input model. Defaults to zero on every axis. Note that axis order matches
vp.shapeso it's(z_origin, x_origin)for 2-D.
Returns:
-
vp_out(ndarray) –Cropped model.
-
geom_out(PhysicalGeometry | None) –Cropped + rebased geometry, or
Noneif no geom was given. -
keep_shots_mask(ndarray | None) –(nshots,)boolean — surviving shots.Noneif no geom given ordrop_outside_sources=False. The corresponding obs slice isobs[keep_shots_mask].
Geometry¶
sweep_io.geometry.Geometry ¶
Geometry(sources: ndarray, receivers: ndarray, dt: float | None = None, nt: int | None = None, dh: Tuple[float, ...] | None = None, meta: dict[str, Any] = <factory>)
Source / receiver acquisition geometry.
Attributes:
-
sources–(nshots, ndim)grid indices.ndim=2for 2-D(x, z),ndim=3for 3-D(x, y, z). -
receivers–(nshots, nreceivers, ndim)grid indices. -
dt–Time sampling interval (s). Optional.
-
nt–Number of time samples. Optional.
-
dh–Spatial sampling per axis (same length as
ndim). Optional. -
meta–Arbitrary user metadata that gets serialized verbatim.
sweep_io.geometry.PhysicalGeometry ¶
PhysicalGeometry(sources_xyz_m: ndarray, receivers_xyz_m: ndarray, dt: float, nt: int, meta: dict[str, Any] = <factory>)
Source/receiver positions in meters — independent of any grid.
For multi-scale FWI the same geometry is reused across many dh;
holding the physical form and calling :meth:to_grid per stage is the
canonical pattern.
Attributes:
-
sources_xyz_m–(nshots, ndim)source coordinates in meters.ndim=2is(x, z),ndim=3is(x, y, z). -
receivers_xyz_m–(nshots, nreceivers, ndim)receiver coordinates in meters. -
dt–Time sampling interval (s).
-
nt–Number of time samples.
-
meta–Optional metadata dict.
to_grid ¶
to_grid(
dh: Tuple[float, ...] | float,
*,
origin_xyz_m: Tuple[float, ...] | None = None,
dedupe: bool = True,
dedup_method: Literal["nearest", "first"] = "nearest"
) -> Tuple[Geometry, np.ndarray]
Snap to nearest grid cell for the given dh.
Parameters:
-
dh(Tuple[float, ...] | float) –Grid spacing in meters. Scalar broadcasts to all axes; otherwise a tuple of length
ndim. -
origin_xyz_m(Tuple[float, ...] | None, default:None) –Coordinate origin (the position of grid index
0). Defaults to zero on every axis. -
dedupe(bool, default:True) –When two or more receivers in the same shot snap to the same grid cell, drop the redundant ones. Sources are not deduped (each source already targets a distinct shot).
-
dedup_method(Literal['nearest', 'first'], default:'nearest') –"nearest"— keep the receiver whose true position is closest to the grid-cell center."first"— keep the first receiver in input order.
Returns:
-
geom(Geometry) –Grid-index geometry with
dhpopulated. -
keep_mask(ndarray) –Boolean array of shape
(nshots, nreceivers).Truewhere the receiver survived dedupe. Use this to slice observed-data tensors along the receiver axis so they line up with the surviving receivers per shot::gg, mask = pg.to_grid(dh=(25.0, 25.0)) obs_aligned = [obs[s][..., mask[s], :] for s in range(pg.nshots)]
apply_rotation ¶
Return a copy with the horizontal (x, y) coordinates rotated
from UTM into the model frame defined by frame.
Only valid for ndim == 3. The vertical z coordinate is
passed through unchanged. The returned geometry shares dt /
nt and carries a meta['rotation'] payload recording the
frame parameters so downstream consumers can round-trip back to
UTM via :meth:RotatedFrame.to_utm.
sweep_io.geometry.RotatedFrame ¶
RotatedFrame(
origin_xy_utm: ndarray,
rotation_matrix: ndarray,
target_axis: Literal["x", "y"] = "x",
inline_shift: float = 0.0,
crossline_shift: float = 0.0,
)
UTM ↔ model-frame 2-D rotation.
Maps model_xy = (utm_xy - origin_xy_utm) @ R^T + shifts where
R is a 2×2 rotation matrix. The model frame is what the
propagator's axis-aligned grid consumes; the UTM frame is what
field-data SEG-Y headers (and the OBN survey metadata) live in.
Sign convention matches the legacy rotation_metadata.json schema
produced by fwi_workflow-dev: rotation_matrix is the matrix
that rotates a UTM displacement vector (relative to origin_xy_utm)
onto the model-frame axes. target_axis="x" means the inline
azimuth points along model +x; "y" means inline points along
model +y.
Parameters:
-
origin_xy_utm(ndarray) –(2,)UTM origin that maps to model-frame(0, 0)before any inline / crossline shift. -
rotation_matrix(ndarray) –(2, 2)rotation matrix. If only a rotation angle is known, construct with :meth:from_rotation_deg. -
target_axis(Literal['x', 'y'], default:'x') –"x"(default) or"y"— which model axis is the inline. -
inline_shift(float, default:0.0) –Optional rigid offsets along the rotated inline / crossline axes (meters). Default
0.0. -
crossline_shift(float, default:0.0) –Optional rigid offsets along the rotated inline / crossline axes (meters). Default
0.0.
crossline_shift
class-attribute
¶
Convert a string or number to a floating point number, if possible.
inline_shift
class-attribute
¶
Convert a string or number to a floating point number, if possible.
target_axis
class-attribute
¶
str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str
Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.str() (if defined) or repr(object). encoding defaults to sys.getdefaultencoding(). errors defaults to 'strict'.
from_rotation_deg
classmethod
¶
from_rotation_deg(
*,
origin_xy_utm: ndarray | Tuple[float, float],
rotation_deg: float,
target_axis: Literal["x", "y"] = "x",
inline_shift: float = 0.0,
crossline_shift: float = 0.0
) -> RotatedFrame
Build a frame from a single rotation angle (degrees, CCW).
from_metadata
classmethod
¶
Build a frame from a rotation_metadata.json path or dict.
Compatible with the legacy fwi_workflow-dev schema:
origin_xy, rotation_matrix, optional inline_shift,
crossline_shift, target_axis.
Missing shifts default to 0.0; missing target_axis
defaults to "x".
fit
classmethod
¶
fit(
points_xy: ndarray,
*,
target_axis: Literal["x", "y"] = "x",
shift_inline_to_zero: bool = False,
shift_crossline_to_zero: bool = False,
shift_points_xy: ndarray | None = None
) -> RotatedFrame
Derive a frame from scattered UTM points via a 2-D PCA line fit.
This is the companion constructor to :meth:from_rotation_deg /
:meth:from_metadata: those consume a known rotation, while
fit derives one from the acquisition geometry. It estimates
the dominant survey-line azimuth from points_xy with
:func:principal_direction and builds the frame that rotates that
azimuth onto the model target_axis.
Ports the rotation estimation of
fwi_workflow.geometry.line2d.rotate_csg_index: identical SVD
principal direction, origin = mean(points), and
[[cos, -sin], [sin, cos]] matrix convention. The resulting frame
therefore serialises (via :meth:to_dict) to the same
rotation_metadata.json schema and reproduces the legacy
inline / crossline coordinates bit-for-bit.
Parameters:
-
points_xy(ndarray) –(npoints, 2)UTM xy used to estimate the line azimuth (e.g. one point per shot for a 2-D streamer line). -
target_axis(Literal['x', 'y'], default:'x') –Model axis the inline direction is rotated onto:
"x"(default, 2-D streamer convention) or"y". -
shift_inline_to_zero(bool, default:False) –Shift the inline coordinate so its minimum is
0(line starts at the grid edge). Matches the legacy 2-D streamer default. -
shift_crossline_to_zero(bool, default:False) –Shift the crossline coordinate so its minimum is
0. Combine withshift_inline_to_zerofor the 3-D OBN first-quadrant layout. -
shift_points_xy(ndarray | None, default:None) –Points used to compute the shift minima; defaults to
points_xy. Pass the union of sources + receivers to reproduce the legacyrotate_csg_indexshift, which always minimised over both even when the rotation was fit on shots only.
Returns:
-
RotatedFrame–Frame whose :meth:
to_modelaligns the fitted line totarget_axisand applies the requested shifts.
to_utm ¶
Inverse of :meth:to_model — model-frame xy back to UTM xy.
sweep_io.geometry.load_rotation_metadata ¶
Build a :class:RotatedFrame from a JSON metadata file.
Thin wrapper around :meth:RotatedFrame.from_metadata for the
file-path call style; matches the legacy
fwi_workflow.geometry.line2d.load_rotation_transform.
Velocity models¶
sweep_io.models.load_velocity ¶
load_velocity(
path: str | Path,
*,
shape: Tuple[int, ...] | None = None,
dtype: str | dtype = "float32",
order: str = "C"
) -> np.ndarray
Load a velocity/parameter model.
Parameters:
-
path(str | Path) –File path. Format is dispatched by extension.
-
shape(Tuple[int, ...] | None, default:None) –Used only for raw-binary loads (e.g.
.bin,.raw, or unknown extensions). Ignored for self-describing formats. -
dtype(Tuple[int, ...] | None, default:None) –Used only for raw-binary loads (e.g.
.bin,.raw, or unknown extensions). Ignored for self-describing formats. -
order(Tuple[int, ...] | None, default:None) –Used only for raw-binary loads (e.g.
.bin,.raw, or unknown extensions). Ignored for self-describing formats.
Returns:
-
ndarray–Array as stored on disk. The caller is responsible for reshape / transpose —
sweep-iodoes not impose(nz, nx)vs(nx, nz).
sweep_io.models.save_velocity ¶
Save a velocity/parameter model. Format dispatched by extension.
header is an optional metadata dict; honored only by formats that
can store it (HDF5). Otherwise silently ignored.
Prefetching and datasets¶
sweep_io.prefetch.Prefetcher ¶
Prefetcher(
source: Iterable[T],
*,
queue_depth: int = 2,
name: str | None = None,
daemon: bool = True
)
Bases: typing.Generic
Yield items from source while a background thread reads ahead.
Parameters:
-
source(Iterable[T]) –Any iterable. Must be safe to consume from a single background thread (most generators are; only iterators with hard thread-affinity requirements are not).
-
queue_depth(int, default:2) –Maximum number of items the worker may hold ahead of the consumer.
2is the canonical double-buffered default. Higher values trade peak RAM for more slack against bursty compute / I/O. -
name(str | None, default:None) –Optional thread name (helps in py-spy / gdb).
-
daemon(bool, default:True) –Whether the worker is a daemon thread. Default
Trueso Python doesn't wait for it on interpreter exit.
Lifecycle
The worker starts in __init__ and runs until the source is
exhausted, an exception is raised, or close() is called.
Reusing a consumed instance is not supported; build a new one.
Examples:
>>> def load(i):
... # imagine slow disk I/O here
... return i * 2
>>> with Prefetcher((load(i) for i in range(5))) as pf:
... out = list(pf)
>>> out
[0, 2, 4, 6, 8]
close ¶
Signal the worker to stop, drain the queue, join the thread.
sweep_io.prefetch.ThreadPoolPrefetcher ¶
ThreadPoolPrefetcher(
load_fn: Callable[[int], T],
indices: Sequence[int] | Iterable[int],
*,
num_workers: int = 2,
queue_depth: int = 2,
name: str | None = None
)
Bases: typing.Generic
Order-preserving prefetcher with num_workers parallel loaders.
Use this when each item is loaded by an independent call to a
function load(idx), the calls are I/O-bound (so threading
helps), and there's enough parallel bandwidth to benefit from
multiple in-flight reads.
Parameters:
-
load_fn(Callable[[int], T]) –Callable
(idx) -> item. Must be thread-safe; if it mutates shared state, you must serialize access yourself. -
indices(Sequence[int] | Iterable[int]) –Sequence of indices to feed
load_fn. Items are yielded in the order of this sequence (not the order workers finish). -
num_workers(int, default:2) –Worker thread count.
2is a sensible default for OS-cached spinning disk;4-8for NVMe or Lustre. -
queue_depth(int, default:2) –How far ahead of the consumer the worker pool may run. Roughly: peak RAM ≈
queue_depth × item_size.
Examples:
>>> def load(i): return i * i
>>> with ThreadPoolPrefetcher(load, range(5), num_workers=2) as pf:
... out = list(pf)
>>> out
[0, 1, 4, 9, 16]
sweep_io.prefetch.TimingPrefetcher ¶
Bases: sweep_io.prefetch.Prefetcher
Variant that also records per-item wait times.
Records, on the consumer side, how long __next__ blocked waiting
for the worker. wait_times accumulates one entry per consumed
item. Useful for benchmarking the prefetch / compute overlap:
- all-zero wait times → consumer is the bottleneck (compute-bound)
- large wait times → worker can't keep up (I/O-bound)
sweep_io.cuda_prefetch.CUDAPrefetcher ¶
CUDAPrefetcher(
source: Iterable[Any],
*,
device: str | device = "cuda",
queue_depth: int = 2,
stream: Stream | None = None
)
Wrap an iterable so each yielded item lands on GPU before compute uses it.
Parameters:
-
source(Iterable[Any]) –Any iterable. Items can be numpy arrays, torch tensors, or nested dict/list/tuple of those. Non-tensor leaves pass through.
-
device(str | device, default:'cuda') –Target CUDA device (
"cuda","cuda:0",torch.device). -
queue_depth(int, default:2) –How many items the background thread keeps ready.
2for double buffering;3+if compute is very fast / bursty. -
stream(Stream | None, default:None) –CUDA stream to issue H2D on. Defaults to a fresh stream on
device.
Notes
- The consumer must use the returned tensors on the current stream. We insert a GPU-side wait so the current stream sees the copy as completed; the host returns immediately.
- For maximum throughput, pair this with
torch.cuda.set_per_process_memory_fractionor a persistent pool to avoid pinned-allocator pressure under high queue depth.
Examples:
>>> import numpy as np
>>> def load(i): return np.random.randn(2048, 256).astype("float32")
>>> pf = CUDAPrefetcher((load(i) for i in range(N)),
... device="cuda:0", queue_depth=2)
>>> for shot in pf:
... pred = solver(shot)
sweep_io.datasets.ShotGatherDataset ¶
ShotGatherDataset(
geometry: Geometry,
obs: ndarray | Tensor | Callable[[int], Any],
*,
transform: Callable[[dict], dict] | None = None
)
Bases: torch.utils.data.dataset.Dataset
Per-shot dataset: each item is (sources_i, receivers_i, obs_i).
Parameters:
-
geometry(Geometry) –Acquisition geometry. Provides
sourcesandreceiversslices. -
obs(ndarray | Tensor | Callable[[int], Any]) –Observed data, either: - a numpy array / torch tensor of shape
(nshots, ...), or - a callableobs(shot_index) -> array_likefor out-of-core access (e.g. one SEG-Y per shot). -
transform(Callable[[dict], dict] | None, default:None) –Optional
transform(sample) -> sampleapplied lazily.
iter_prefetched ¶
iter_prefetched(
indices: Sequence[int] | Iterable[int] | None = None,
*,
queue_depth: int = 2,
num_workers: int = 0
) -> Iterator[dict[str, Any]]
Iterate samples with a background prefetch thread.
Parameters:
-
indices(Sequence[int] | Iterable[int] | None, default:None) –Sequence of shot indices to iterate. Defaults to
range(len(self)). -
queue_depth(int, default:2) –How many samples to keep ready ahead of the consumer.
-
num_workers(int, default:0) –0→ single background thread (good when the underlying__getitem__does one big I/O call).>0→ a thread pool of that size (good when__getitem__can run in parallel, e.g. distinct SEG-Y files per shot).
Notes
This is not torch.utils.data.DataLoader. There's no batch
collation here — one item per next. Wrap with your own
batching if you want it. For the typical FWI loop you want one
shot at a time anyway, so the loop is:
for sample in ds.iter_prefetched():
pred = solver(sample["sources"], sample["receivers"], ...)
...
CRG plan cache¶
sweep_io.crg_build.build_crg_plan_from_segy ¶
build_crg_plan_from_segy(
paths: Sequence[str | Path],
*,
byte_map: dict | None = None,
source_depth_m_override: float | None = None,
receiver_depth_m_override: float | None = None,
coord_scalar_override: float | None = None,
num_workers: int = 16,
use_mpi: bool | None = None,
mpi_stage_dir: str | Path | None = None,
receiver_quantize_m: float = 0.5,
receiver_stride: int = 1,
receiver_max: int | None = None,
receiver_ids: Sequence[int] | None = None,
shots_per_source_line: int | None = None,
source_line_stride: int = 1,
rng_seed: int = 20260509,
show_progress: bool = False
) -> CRGPlan | None
End-to-end: scan raw SEG-Y headers, then group into a :class:CRGPlan.
Convenience wrapper around :func:sweep_io.segy_index.build_segy_index
(parallel header scan) followed by :func:build_crg_plan_from_index
(receiver grouping + sub-sampling).
Parameters mirror both helpers. See
:func:sweep_io.segy_index.build_segy_index for the SEG-Y header
knobs (byte_map, source_depth_m_override etc.) and the
arguments of :func:build_crg_plan_from_index for the receiver-side
knobs.
Parallelism modes (mutually exclusive):
- MPI (set
use_mpi=Trueor run undermpiexecwithuse_mpi=Noneauto-detect): rank 0 partitions the file list, every rank scans a contiguous slice single-threaded, rank 0 gathers + merges. Non-root ranks returnNone. - ThreadPool (default,
use_mpi=False): one process,num_workersbackground threads for the per-file scan.
The num_workers knob is ignored under MPI (MPI ranks ARE the
parallelism; nesting threads inside MPI workers competes for
file-descriptor cache without helping).
sweep_io.crg_build.build_crg_plan_from_index ¶
build_crg_plan_from_index(
index: SEGYIndex,
*,
receiver_quantize_m: float = 0.5,
receiver_stride: int = 1,
receiver_max: int | None = None,
receiver_ids: Sequence[int] | None = None,
shots_per_source_line: int | None = None,
source_line_stride: int = 1,
rng_seed: int = 20260509
) -> CRGPlan
Group an :class:SEGYIndex by quantised receiver position, then
apply optional receiver / shot sub-sampling, into a :class:CRGPlan.
Quantisation: traces whose (rx, ry, rz) collapse to the same cell
at receiver_quantize_m tolerance are considered the same OBN node.
A regular node grid can use e.g. 0.5 m.
Parameters:
-
index(SEGYIndex) –Loaded :class:
SEGYIndexover all source-line SEG-Y files. Built by :func:sweep_io.segy_index.build_segy_index. -
receiver_quantize_m(float, default:0.5) –Cell size (m) for receiver position quantisation.
-
receiver_stride(int, default:1) –Keep every N-th unique receiver (after quantisation).
1= all. -
receiver_max(int | None, default:None) –Cap the kept-receiver count after striding.
-
receiver_ids(Sequence[int] | None, default:None) –Explicit receiver indices to keep (overrides stride / max). Indices refer to the receiver-id ordering AFTER quantisation + sort (same order
CRGPlan.receiver_xyz_usedis written in). -
shots_per_source_line(int | None, default:None) –Per-receiver cap on shots picked from each source-line file.
Nonekeeps all; e.g.4keeps at most four shots per line when building the canonical plan from CRG. -
source_line_stride(int, default:1) –Drop every Nth source line per receiver before the per-line cap.
1keeps all lines. -
rng_seed(int, default:20260509) –Deterministic RNG seed for the per-line shot sub-sampler.
Returns:
-
CRGPlan–Ready to feed
obs.crg_plan.plan_pathin a sweep-tasks FWI YAML, or tosave_crg_planfor the on-disk cache.
sweep_io.crg_plan.CRGPlan ¶
CRGPlan(
files: list[Path],
trace_size_per_file: ndarray,
sample_format: int,
samples_per_trace: int,
dt_s: float,
receiver_indices: ndarray,
receiver_xyz_used: ndarray,
plan_file_ids: ndarray,
plan_trace_offsets: ndarray,
plan_sx: ndarray,
plan_sy: ndarray,
plan_sz: ndarray,
plan_offsets: ndarray,
)
Slim per-virtual-source CRG plan loaded from a crg_fwi_plan_v1 cache.
Each i ∈ [0, n_receivers_used) is one virtual source. Its
physical-shot rows are plan_*[slot_slice(i)].
Attributes:
-
files–SEG-Y source-line files (per-file
read_tracesdispatch keys areplan_file_ids). -
trace_size_per_file–(n_files,)byte stride per trace (header + samples). -
sample_format, samples_per_trace, dt_s–SEG-Y trace parameters; identical across files in a plan.
-
n_receivers_used–Number of virtual sources (OBN nodes) in the plan.
-
receiver_xyz_used–(n_receivers_used, 3)virtual-source positions in UTM (m). These become the FWI sources. -
plan_file_ids, plan_trace_offsets–Flat per-row arrays of (file index, byte offset) tuples handed to :meth:
sweep_io.segy.MultiFileSEGYReader.read_traces. -
plan_sx, plan_sy, plan_sz–Flat per-row physical shot positions in UTM (m). These become the FWI receivers (one per recorded trace).
-
plan_offsets–(n_receivers_used + 1,)CSR-style row pointer:plan_offsets[i] : plan_offsets[i+1]are the row indices of sloti's gather.
virtual_source_xy
property
¶
(n_receivers_used, 2) virtual-source horizontal coordinates (m).
slot_row_count ¶
Number of physical shots that hit virtual source slot.
per_slot_row_counts ¶
(n_receivers_used,) shot count per virtual source.
slot_receiver_xyz ¶
Position of the slot-th OBN node (= virtual source position).
slot_shot_xyz ¶
(n_shots_for_slot, 3) physical shot positions (= virtual recv).
filter_by_row_mask ¶
Return a new plan keeping only flat rows where row_mask is True.
The per-slot plan_offsets are recomputed so each slot still
addresses a contiguous run of its surviving rows. Slots that lose
ALL their rows survive with a zero-length slice — caller should
compose with :meth:filter_by_coverage to drop them entirely.
Used by the OBN CRG path to drop physical shots whose model-frame position falls outside the (optionally cropped) inversion grid.
filter_by_coverage ¶
Return a new plan keeping only slots with >= min_shots shots.
Mirrors the legacy min_coverage filter — drops edge / partial-
coverage OBN nodes that record so few shots they would hurt
source-encoded supershot inversions (per crg_data._sample_crg_iter_shared_shots).
sweep_io.crg_plan.load_crg_shot_plan_cache ¶
Load a crg_fwi_plan_v1 npz cache into a :class:CRGPlan.
SEG-Y paths recorded in the cache are remapped via
:func:_remap_segy_path so the same cache works across hosts when
FWI_SEGY_ROOT or FWI_SEGY_REMAP is exported.
sweep_io.crg_dataset.CRGBatchDataset ¶
CRGBatchDataset(
plan: CRGPlan,
*,
slot_indices: ndarray | None = None,
max_shots_per_slot: int | None = None
)
Dataset yielding one virtual-source gather per __getitem__ call.
Designed for torch.utils.data.DataLoader with num_workers > 0
+ :func:pad_collate_fn + :func:worker_init_fn.
Parameters:
-
plan(CRGPlan) –Loaded :class:
CRGPlan. -
slot_indices(ndarray | None, default:None) –Subset of slot indices to expose (default: all). Use this to shard slots across distributed ranks before the DataLoader.
-
max_shots_per_slot(int | None, default:None) –Truncate each slot's shot list to this length (or
Noneto keep them all). Truncation is deterministic (firstmax_shots); per-iter random sub-sampling belongs upstream in the runner.
Wavelets¶
sweep_io.wavelet.load_wavelet_npz ¶
load_wavelet_npz(
path: str | Path,
*,
prefer_keys: Sequence[str] = (
"wavelet",
"optimized_siren_wavelet",
"direct_causal_wavelet",
"initial_causal_wavelet",
),
explicit_key: str | None = None
) -> WaveletNPZ
Read a wavelet from an .npz file.
The selected array key is, in order: explicit_key (if non-None);
otherwise the first key in prefer_keys present in the npz. Raises
KeyError if none match.
The sample interval is taken from dt_s (scalar) when present;
otherwise inferred from time_s via the median diff (matches the
legacy fwi_workflow-dev behaviour). If neither is present a
KeyError is raised — sweep-stack will not invent a dt.
Parameters:
-
path(str | Path) –Path to the
.npzfile. -
prefer_keys(Sequence[str], default:('wavelet', 'optimized_siren_wavelet', 'direct_causal_wavelet', 'initial_causal_wavelet')) –Ordered fallback list of npz keys for the wavelet array.
-
explicit_key(str | None, default:None) –Force-select this key; bypasses
prefer_keys.
sweep_io.wavelet.WaveletNPZ ¶
Loaded source-wavelet payload.
Attributes:
-
samples–1-D
float32array of lengthntwith the wavelet amplitudes. -
dt_s–Sample interval in seconds.
-
source_delay_s–Optional zero-prepad in front of the wavelet (s). The SIREN wavelet-fit pipeline writes this so downstream FWI can left-pad observed traces by the same amount;
Nonewhen absent. -
key–Which npz key the samples were read from (helpful for logs).
key
class-attribute
¶
str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str
Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.str() (if defined) or repr(object). encoding defaults to sys.getdefaultencoding(). errors defaults to 'strict'.