S-exp GPU
All pages
Docs · The command lineMarkdown

bundle and remote runs

sexpgpu bundle writes one directory that runs the experiment on any machine the binary runs on: the binary, every source the compile read, every local file the loaders read, and a run.sh with the selection baked in. The other machine needs no toolchain and no source layout; it needs the CUDA driver for SEXPGPU_DEVICE=cuda and the environment.

sexpgpu bundle my-run.sx -o out/my-run --binary ./sexpgpu-linux-cuda
scp -r out/my-run gpu-box:
ssh gpu-box 'SEXPGPU_DEVICE=cuda SEXPGPU_METRICS=jsonl:my-run/metrics.jsonl \
    SEXPGPU_STATUS=my-run/status.json \
    SEXPGPU_CHECKPOINT_DIR=s3://my-bucket/runs/my-run my-run/run.sh --resume latest'
flagmeaning
-o <dir>where to write; a directory that exists is refused (exit 2)
--binary <path>the binary to pack instead of this one, e.g. a Linux CUDA build packed from a laptop
--variant, --set, --steps, --eval-every, --diagnosticsbaked into run.sh
$ sexpgpu bundle my-run.sx -o /tmp/my-run --steps 100
wrote /tmp/my-run: 3 source files and 2 data files, 10.8 MiB; binary /usr/local/bin/sexpgpu (20.0 MiB, built without --features cuda: SEXPGPU_DEVICE=cuda would refuse)
  run it: /tmp/my-run/run.sh [run flags], with SEXPGPU_* in the environment

Layout

/tmp/my-run/sexpgpu                          the binary
/tmp/my-run/run.sh                           exec sexpgpu run files/my-run.sx --steps 100 "$@"
/tmp/my-run/files/my-run.sx                  the run, at the top under its own name
/tmp/my-run/files/deps/mine-5745a31f.sx      every required file, flat
/tmp/my-run/files/deps/blocks-e6e29e73.sx
/tmp/my-run/files/data/train.parquet         local loader files
/tmp/my-run/files/data/val.parquet
  • A dep is <stem>-<hash>.sx, the hash the first eight hex digits of the SHA-256 of its source path relative to the sources' deepest common directory, a NUL byte, and the rewritten text. Two files with the same text that require different files never share a name, and distinct modules keep separate bindings even when their text is identical.
  • Data keeps its layout under its deepest common directory.
  • Every require and loader path, including one given with --set, is rewritten to where the file now is: (require "./deps/blocks-e6e29e73.sx" ...). run.sh lists where each dep came from. The manifest of a bundled run names the bundle's paths and hashes.
  • A string naming a directory the data lies under is rewritten too, so a run that binds (defvar data-dir "../data/") and builds (str data-dir "train-" i ".parquet") bundles. A path computed from anything else, (str "../data/train-" i ".parquet"), cannot be rewritten: the bundle compiles the rewritten run with the baked selection and refuses (exit 1, removing what it wrote) when a loader would read anything but a file under files/. Write the directory as its own string, or read the data from an s3:// URI.
  • Standard modules are in the binary and not copied. s3:// loader files are read on the other machine with its credentials. A local loader file missing here fails the bundle (exit 1).

run.sh

run.sh passes its own arguments through: ./run.sh --resume latest starts the run the first time and continues it every time after, and ./run.sh --eval-every 50 changes the cadence. Set SEXPGPU_DEVICE, SEXPGPU_METRICS, SEXPGPU_CHECKPOINT_DIR and credentials in the machine's environment, never in the bundle. See environment.

Running remotely

Whatever runs the process (ssh with nohup, a job runner, a scheduler) also stops it, restarts it and keeps its log. With SEXPGPU_STATUS set the run leaves a status file saying where it is. Put --resume latest on the command line from the first launch: the first time there is no checkpoint and the run starts fresh; every restart continues from the newest one.

With the binary and the sources already on the machine, no bundle is needed:

sexpgpu run my-run.sx --device cuda --metrics syvain:<folder-id> \
    --checkpoint-dir s3://my-bucket/runs/my-run --resume latest

Preemptible jobs

A bundle is what a job runner starts again after a preemption. With SkyPilot, the bundle is the job's workdir, the environment sets SEXPGPU_DEVICE=cuda, the metrics target and an s3:// checkpoint location, the keys are secrets, and the command is ./run.sh --resume latest. A spot A100 job:

name: my-run
workdir: bundle
resources:
  accelerators: A100:1
  use_spot: true
  job_recovery:
    recover_on_exit_codes: [130, 143]
envs:
  SEXPGPU_DEVICE: cuda
  SEXPGPU_METRICS: syvain:<folder-id>
  SEXPGPU_STATUS: status.json
  SEXPGPU_CHECKPOINT_DIR: null   # required at launch
secrets:
  AWS_ACCESS_KEY_ID: null
  AWS_SECRET_ACCESS_KEY: null
  SYVAIN_METRICS_API_KEY: null
setup: chmod +x run.sh sexpgpu
run: ./run.sh --resume latest
sky jobs launch job.yaml --workdir bundle \
    --env SEXPGPU_CHECKPOINT_DIR=s3://my-bucket/runs/my-run \
    --secret AWS_ACCESS_KEY_ID --secret AWS_SECRET_ACCESS_KEY \
    --secret SYVAIN_METRICS_API_KEY

A managed job's workdir keeps no file modes, hence the chmod in setup. On an image whose CUDA toolkit is older than 13, setup also installs CUDA 13's NVRTC and cuBLAS (and cuDNN 9.13, optionally); see devices. Run on a VM image rather than a container: a container's shell ignores the shutdown SIGTERM, and the run is killed without its checkpoint.

When the VM is taken the run gets SIGTERM, checkpoints the step and exits 143; the job runs again elsewhere and continues in the same metrics experiment, with interrupted by SIGTERM at step <n> and resumed from <dir> marking the seam. The checkpoint location must outlive the VM, so it is s3://, never local. List 130 and 143 in the job's recover_on_exit_codes, or SkyPilot reads the exit as the program failing. A checkpoint larger than the shutdown notice allows (a GCP spot VM gets thirty seconds) is cut off and passed over for the one before it; the checkpoint annotation's seconds says how long a run's checkpoints take. See checkpoints.

Reading a preempted run

Open the experiment: its annotations, in the order they arrived, tell the whole story. A job whose spot VM was stopped 80 s into training:

- experiment opened
0 lowering: cuda compute_80 NVIDIA A100-SXM4-40GB, patterns generated
34 interrupted by SIGTERM at step 34
34 checkpoint s3://.../step-00000034 (8.349 s)
- experiment opened
34 resumed from s3://.../step-00000034
34 lowering: cuda compute_80 NVIDIA A100-SXM4-40GB, patterns generated
60 checkpoint s3://.../step-00000060 (8.411 s)

Each process opens the experiment once. The lowering block names the machine a process ran on, so a job that comes back in another zone or on another GPU says so at the step it resumed. A checkpoint at the same step as interrupted says the interrupting checkpoint fit the notice, and nothing is repeated. A resumed from with no interrupted before it was a hard kill, and the steps after the last checkpoint appear twice. The annotations are in the events table; the syvain-metrics-api-client Python package reads them.

Related: run, devices, doctor.