Skip to content

HPC

Memory strategies, compressed boundary storage, one shot per GPU, and one model split across GPUs.

  • Memory strategies

    Memory · strategies


    Same forward + backward step run under every memory strategy (eager full / chunk-ckpt / boundary-saving on gpu × dtype vs. c full / boundary-saving on {gpu, cpu, disk} × dtype / chunk-ckpt / recursive-ckpt) with side-by-side peak-memory and wallclock charts.

    Open notebook

  • FWI boundary compression

    FWI · boundary compression


    storage_dtype (fp16/bf16/int8) shrinks the saved boundary wavefield while compute stays FP32. Marmousi FWI across the full {gpu, cpu, disk} × dtype matrix on the compiled path, plus gpu × dtype on eager — identical convergence, plus a runtime GPU-memory breakdown.

    Open notebook

  • Multi-GPU DDP

    Multi-GPU · DDP vs 1 GPU


    torchrun --nproc_per_node=N driver that shards shots across GPUs and syncs gradients via torch.distributed, timed against a single-GPU baseline on a two-layer toy model (the saved run: 3.53× on 4 GPUs).

    Open notebook

  • Domain decomposition

    HPC · Domain decomposition


    One model, several GPUs: ModelParallel slices it into tiles and exchanges a halo every step, so a single shot is solved cooperatively rather than replicated. Gradients are bit-identical to the single-domain run — the notebook checks that, tile by tile.

    Open notebook

  • DD on Overthrust 3-D

    HPC · DD on Overthrust 3-D


    The same split on a real 3-D benchmark, 2 × 2 tiles across four GPUs. The gradient is sliced three ways straight across the cut planes, where a halo bug would show as a stripe — and compared against the single-GPU answer.

    Open notebook