Backends¶
SWEEP supports two main user-facing backend families:
torchjax
Within the Torch family, PropTorch now selects the actual implementation:
impl=None(default — equivalent toimpl="auto"): probesweep.is_torch_binding_available(); pick"c"when PyTorch, a visible CUDA GPU and a CUDA core (the wheel's prebuilt one, a cached local build, or an nvcc to build one) are present and the equation has compiled kernels, otherwise transparently fall back to"eager". A plainimport sweep._Calways succeeds and proves nothing.impl="eager": pure-PyTorch implementation (no build step required).impl="c": the compiled CUDA implementation — the prebuiltlibsweep_core.sothe wheel ships (sweep/lib/cu12/orcu13/, picked by your torch's CUDA major) driven by the pure-Python ctypes layersweep.backend.c; nothing compiles afterpip install. If no core fits and none can be built, or the equation has no compiled kernels,PropTorchfalls back to"eager"with aUserWarningso the slowdown is visible.
The wheel ships two cores: cu12 (sm_70–sm_90 SASS + sm_90 PTX) and cu13
(sm_75–sm_120 SASS + sm_120 PTX; no V100, driver >= 580). The loader picks
lib/cu<torch CUDA major>/ and checks its core.json (ABI 2, CUDA major, archs
and PTX gated on the driver version). A local core build happens only when no
shipped core fits — a torch of another CUDA major (e.g. cu11), a GPU older than
sm_70 on cu12, an sdist/clone install — on the first impl="c" use, or ahead
of time with python -m sweep.build. It needs an nvcc of torch's CUDA major
(for CUDA 12: >= 12.4, >= 12.8 for Blackwell targets). A GPU older than
sm_75 under a cu13 torch is refused rather than built for; use a cu12 torch
there. See Building the CUDA core.
User-Facing Backend Families¶
torch: PyTorch-based propagation and differentiationjax: JAX-based propagation and differentiation
cuda is not a separate top-level backend alongside torch and jax. It is a
device choice. impl="c" runs CUDA kernels only: the models must be CUDA
tensors (a host tensor is refused with a clear error). There is no compiled CPU
path: on a CPU device, or on a machine with no GPU, impl=None resolves to
"eager" and an explicit impl="c" falls back to eager with a warning.
Typical Torch-family usage:
from sweep.propagator.torch import PropTorch
solver_eager = PropTorch(..., backend="torch", impl="eager")
solver_c = PropTorch(..., backend="torch", impl="c")
Compiled CUDA core (sweep._C)¶
sweep._C is a lazy entry point: importing it is free, and the first attribute
access loads the prebuilt CUDA core through ctypes (sweep.backend.c —
loader.py, adapt.py, entries.py, runners.py, the generated abi.py;
jit.py decides which core). The core covers the acoustic (2-D/3-D), VRZ,
LSRTM, VTI 1st-order, elastic (2-D/3-D), elastic TTI (SG 2-D/3-D, 2nd), elastic
VRR, DAS and visco-acoustic propagators — sweep list equations is the
authoritative list. CUDA tensors only.
You can inspect backend capability from Python:
import sweep
sweep.backend.torch.is_available()
sweep.backend.jax.is_available()
sweep.backend.torch.cuda.is_available()
sweep.backend.torch.binding.is_available()
sweep.backend.torch.binding.diagnostics()
sweep.backend.torch.cuda.is_available() only answers whether PyTorch can see CUDA.
sweep.backend.torch.binding.is_available() answers whether impl="c" is
usable here: PyTorch present, a CUDA GPU visible, and a CUDA core at hand (the
shipped one for your torch's CUDA major, SWEEP_CORE, a cached local build, or
an nvcc of torch's CUDA major to build one). It loads and compiles nothing.
Example diagnostics output:
{
"usable": True,
"reason": "ok",
"shim": "ctypes", # "pybind" under SWEEP_JIT_FULL=1 or with a prebuilt extension
"cuda_home": "/usr/local/cuda-12.9", # None is fine: a shipped core needs no nvcc
"already_compiled": False, # True once the core is loaded (sweep.precompile())
"prebuilt": False, # a SWEEP_BUILD_CUDA=1 ahead-of-time extension on disk
"shipped_core": {"path": ".../sweep/lib/cu12/libsweep_core.so",
"reason": "ok", "tag": "cu12", "available": ["cu12", "cu13"]},
}
Equation-Level Binding Support¶
Not every equation exposes the compiled binding path.
Use the CLI to inspect support — the listing is generated live from the installed package, so it always reflects the current environment:
See CLI · sweep list equations for the full
table. Each row tells you:
- Torch Binding — whether the equation's source declares compiled-extension support.
- Binding Ready — whether
impl="c"can run right now (sweep.is_torch_binding_available(): PyTorch, a visible CUDA GPU, and a fitting or buildable CUDA core).
Choosing Between Torch and JAX¶
Choose torch when:
- you want the Torch ecosystem and autograd workflow
- you want to use
PropTorch(..., backend="torch", impl="eager") - you want to use the compiled C++/CUDA path through
PropTorch(..., backend="torch", impl="c")
Choose jax when:
- you want JAX transformations and device placement
- you plan to use
PropJax
Extension Kernels in the Torch Family¶
Use the c implementation inside the Torch family when:
- a CUDA core is available (
sweep.backend.torch.binding.is_available()) - your equation supports the binding
- you want hand-written CUDA kernels or memory modes such as boundary
saving, disk-backed boundary storage, or
ccheckpointing
c memory modes are equation-specific. Full-wavefield storage is available
across the c-backed solvers, and boundary saving across all of them except
ViscoAcoustic and DASZhao3D (their c default is full storage);
checkpoint modes are available where the equation exposes the corresponding
backward implementation. The user-facing entry point is: