All pages
bundle and remote runs
sexpgpu bundle writes one directory that runs the experiment on any
machine the binary runs on: the binary, every source the compile read,
every local file the loaders read, and a run.sh with the selection baked
in. The other machine needs no toolchain and no source layout; it needs the
CUDA driver for SEXPGPU_DEVICE=cuda and the environment.
sexpgpu bundle my-run.sx -o out/my-run --binary ./sexpgpu-linux-cuda
scp -r out/my-run gpu-box:
ssh gpu-box 'SEXPGPU_DEVICE=cuda SEXPGPU_METRICS=jsonl:my-run/metrics.jsonl \
SEXPGPU_STATUS=my-run/status.json \
SEXPGPU_CHECKPOINT_DIR=s3://my-bucket/runs/my-run my-run/run.sh --resume latest'
| flag | meaning |
|---|---|
-o <dir> | where to write; a directory that exists is refused (exit 2) |
--binary <path> | the binary to pack instead of this one, e.g. a Linux CUDA build packed from a laptop |
--variant, --set, --steps, --eval-every, --diagnostics | baked into run.sh |
$ sexpgpu bundle my-run.sx -o /tmp/my-run --steps 100
wrote /tmp/my-run: 3 source files and 2 data files, 10.8 MiB; binary /usr/local/bin/sexpgpu (20.0 MiB, built without --features cuda: SEXPGPU_DEVICE=cuda would refuse)
run it: /tmp/my-run/run.sh [run flags], with SEXPGPU_* in the environment
Layout
/tmp/my-run/sexpgpu the binary
/tmp/my-run/run.sh exec sexpgpu run files/my-run.sx --steps 100 "$@"
/tmp/my-run/files/my-run.sx the run, at the top under its own name
/tmp/my-run/files/deps/mine-5745a31f.sx every required file, flat
/tmp/my-run/files/deps/blocks-e6e29e73.sx
/tmp/my-run/files/data/train.parquet local loader files
/tmp/my-run/files/data/val.parquet
- A dep is
<stem>-<hash>.sx, the hash the first eight hex digits of the SHA-256 of its source path relative to the sources' deepest common directory, a NUL byte, and the rewritten text. Two files with the same text that require different files never share a name, and distinct modules keep separate bindings even when their text is identical. - Data keeps its layout under its deepest common directory.
- Every
requireand loader path, including one given with--set, is rewritten to where the file now is:(require "./deps/blocks-e6e29e73.sx" ...).run.shlists where each dep came from. The manifest of a bundled run names the bundle's paths and hashes. - A string naming a directory the data lies under is rewritten too, so a
run that binds
(defvar data-dir "../data/")and builds(str data-dir "train-" i ".parquet")bundles. A path computed from anything else,(str "../data/train-" i ".parquet"), cannot be rewritten: the bundle compiles the rewritten run with the baked selection and refuses (exit1, removing what it wrote) when a loader would read anything but a file underfiles/. Write the directory as its own string, or read the data from ans3://URI. - Standard modules are in the binary and not copied.
s3://loader files are read on the other machine with its credentials. A local loader file missing here fails the bundle (exit1).
run.sh
run.sh passes its own arguments through: ./run.sh --resume latest
starts the run the first time and continues it every time after, and
./run.sh --eval-every 50 changes the cadence. Set SEXPGPU_DEVICE,
SEXPGPU_METRICS, SEXPGPU_CHECKPOINT_DIR and credentials in the
machine's environment, never in the bundle. See
environment.
Running remotely
Whatever runs the process (ssh with nohup, a job runner, a scheduler)
also stops it, restarts it and keeps its log. With SEXPGPU_STATUS set the
run leaves a status file saying where it is. Put
--resume latest on the command line from the first launch: the first
time there is no checkpoint and the run starts fresh; every restart
continues from the newest one.
With the binary and the sources already on the machine, no bundle is needed:
sexpgpu run my-run.sx --device cuda --metrics syvain:<folder-id> \
--checkpoint-dir s3://my-bucket/runs/my-run --resume latest
Preemptible jobs
A bundle is what a job runner starts again after a preemption. With
SkyPilot, the bundle is the job's workdir, the environment sets
SEXPGPU_DEVICE=cuda, the metrics target and an s3:// checkpoint
location, the keys are secrets, and the command is
./run.sh --resume latest. A spot A100 job:
name: my-run
workdir: bundle
resources:
accelerators: A100:1
use_spot: true
job_recovery:
recover_on_exit_codes: [130, 143]
envs:
SEXPGPU_DEVICE: cuda
SEXPGPU_METRICS: syvain:<folder-id>
SEXPGPU_STATUS: status.json
SEXPGPU_CHECKPOINT_DIR: null # required at launch
secrets:
AWS_ACCESS_KEY_ID: null
AWS_SECRET_ACCESS_KEY: null
SYVAIN_METRICS_API_KEY: null
setup: chmod +x run.sh sexpgpu
run: ./run.sh --resume latest
sky jobs launch job.yaml --workdir bundle \
--env SEXPGPU_CHECKPOINT_DIR=s3://my-bucket/runs/my-run \
--secret AWS_ACCESS_KEY_ID --secret AWS_SECRET_ACCESS_KEY \
--secret SYVAIN_METRICS_API_KEY
A managed job's workdir keeps no file modes, hence the chmod in setup.
On an image whose CUDA toolkit is older than 13, setup also installs
CUDA 13's NVRTC and cuBLAS (and cuDNN 9.13, optionally); see
devices.
Run on a VM image rather than a container: a container's shell ignores the
shutdown SIGTERM, and the run is killed without its checkpoint.
When the VM is taken the run gets SIGTERM, checkpoints the step and exits
143; the job runs again elsewhere and continues in the same metrics
experiment, with interrupted by SIGTERM at step <n> and
resumed from <dir> marking the seam. The checkpoint location must outlive
the VM, so it is s3://, never local. List 130 and 143 in the job's
recover_on_exit_codes, or SkyPilot reads the exit as the program failing.
A checkpoint larger than the shutdown notice allows (a GCP spot VM gets
thirty seconds) is cut off and passed over for the one before it; the
checkpoint annotation's seconds says how long a run's checkpoints take.
See checkpoints.
Reading a preempted run
Open the experiment: its annotations, in the order they arrived, tell the whole story. A job whose spot VM was stopped 80 s into training:
- experiment opened
0 lowering: cuda compute_80 NVIDIA A100-SXM4-40GB, patterns generated
34 interrupted by SIGTERM at step 34
34 checkpoint s3://.../step-00000034 (8.349 s)
- experiment opened
34 resumed from s3://.../step-00000034
34 lowering: cuda compute_80 NVIDIA A100-SXM4-40GB, patterns generated
60 checkpoint s3://.../step-00000060 (8.411 s)
Each process opens the experiment once. The lowering block names the
machine a process ran on, so a job that comes back in another zone or on
another GPU says so at the step it resumed. A checkpoint at the same step
as interrupted says the interrupting checkpoint fit the notice, and
nothing is repeated. A resumed from with no interrupted before it was a
hard kill, and the steps after the last checkpoint appear twice. The
annotations are in the events table; the
syvain-metrics-api-client Python package reads them.