All pages
Devices
A run executes on one device kind, chosen at start with --device or
SEXPGPU_DEVICE. The file never names a device; the same IR runs on both.
| device | what it is | for |
|---|---|---|
cpu | the IR interpreter, the default | compiling, reading data, differentiating and real steps at toy sizes, on any machine |
cuda | the CUDA backend: fused generated kernels, cuBLAS, and native cuDNN product and attention graphs | training |
The CPU interpreter
The interpreter keeps every node's value alive for the whole graph and runs one thread, so it cannot hold an LM-scale shape: a full vocabulary at sequence 1024 costs an hour and several gigabytes for two steps.
A run at an LM-scale shape therefore declares a smoke variant that
shrinks it:
(defvariant smoke (layers 2) (width 128) (seq 64))
and sexpgpu run <file> --variant smoke --steps 3 --device cpu is the
local check before a GPU is paid for. A run without such a variant is
checked on the GPU instead, with --steps 3.
The CUDA device
SEXPGPU_DEVICE=cuda needs the Linux CUDA binary, sexpgpu-linux-cuda.
| requirement | why |
|---|---|
| a CUDA 13 driver, 580 series or newer | the binary loads the 13.x driver, NVRTC and cuBLAS at start |
| compute capability 8.0 or newer: A100, L4, H100 | bf16 tensor-core GEMMs need Ampere; an older card is refused when the device opens. The native cuDNN graphs are used on compute_80 only |
| cuDNN 9.13.0 for CUDA 13, optional | the native product and attention graphs; without it those regions run the ordinary lowering, and the lowering report says so |
Nothing else: no Rust, no Python, no checkout. Kernels are compiled at run
time by NVRTC from source inside the binary. sexpgpu doctor
checks the floor.
What the device did with each graph is the
lowering report. runtime/peak_bytes
reports the allocator's high-water mark every step; see metrics.
Memory
Before its first step a CUDA run works out, from the plan its device will execute, the most it will hold at once on each GPU, and chooses how many microbatches one call of the training graph runs (its stacking). Both are lines of the lowering report, and the run ends with the peak it reached against the plan:
lowering: cuda compute_80 NVIDIA A100-SXM4-80GB, patterns on
...
memory 2.7 GiB of 78.8 GiB free
stacking 1: measured at the first step, the model within 5% could not separate them (4 130.2 ms, model 130.5 ms; 2 135.2 ms, model 131.5 ms; 1 128.9 ms, model 133.6 ms)
...
memory: peak 2.8 GiB of 2.7 GiB planned (+1.8%)
- Stacking is chosen among the degrees that fit: the only one, the cost
model's prediction, a measurement at the first step when the predictions
are within 5 percent, or a measurement an earlier run of the same graph
made on the same device, remembered in
~/.cache/sexpgpu/choices.json(delete it to measure again). A measured choice can differ on another device, where the F32 products then sum in another order. - The plan keeps 1 GiB free for the libraries. It stacks F32
microbatches only when the stacked graph fits, drops cached parameter
results when only that fits, and otherwise refuses with
E-MEM-001before any initializer runs. - An allocation that fails anyway releases the device's caches and
retries, and says so in a
memory: allocation failedline and annotation. - Both lines are annotations on the metrics stream, the memory one with its parts: parameters, optimizer states, gradients, constants, the largest graph's live set, and where it peaks.
SEXPGPU_MEMORYplans against a smaller card than the one present.
Several GPUs
SEXPGPU_DEVICES=0,1 runs data parallel over those CUDA ordinals, in rank
order; unset uses every visible device, and one ordinal is the single-GPU
path. CUDA_VISIBLE_DEVICES limits what is visible.
- The global
defrun :microbatchesis split over the GPUs, so the device count must divide it. Each rank reads its own share of the loader. - The result is the same experiment: one optimizer step per step, gradients combined across ranks.
- Timing series and sampled diagnostics are rank zero's; a diagnostic's reducer folds across ranks.
- A resume needs the device count the checkpoint was written with.
Several nodes
One run can span machines: one run process on each node, each with
SEXPGPU_DEVICE=cuda and the same number of GPUs. A node's GPUs are the
global ranks after the previous nodes'.
| variable | meaning |
|---|---|
SEXPGPU_NODES | the node count; unset or 1 is one node |
SEXPGPU_NODE_RANK | this process's node, 0 to nodes - 1. Node 0 writes the metrics, the status file and the checkpoints, so a checkpoint location every node reads is an s3:// one |
SEXPGPU_RENDEZVOUS | host:port of node 0, where it listens once for the other nodes. Under SkyPilot the host is the first line of SKYPILOT_NODE_IPS |
SEXPGPU_NODE_GRADIENTS | bf16 (default) rounds each device's f32 gradients to bf16 for the sum between nodes, carrying each rounding's error into the next step, half the bytes on the wire; f32 sends them exactly, so two nodes of one GPU train bitwise as one node of two and a resume is exact. Every node sets the same |
Under SkyPilot with num_nodes, each node's run sets the three from
SkyPilot's own variables:
export SEXPGPU_NODES="$SKYPILOT_NUM_NODES" SEXPGPU_NODE_RANK="$SKYPILOT_NODE_RANK"
export SEXPGPU_RENDEZVOUS="$(echo "$SKYPILOT_NODE_IPS" | head -n1):29500"
A resume needs the node count the checkpoint was written with.
Errors
They stop run with exit 1, before or during training; doctor reports
E-DP-001 and E-DP-005.
| code | when | fix |
|---|---|---|
E-DP-001 | SEXPGPU_DEVICES does not parse, repeats an ordinal, names one that is not visible, or no GPU is visible | list visible ordinals once each, SEXPGPU_DEVICES=0,1 |
E-DP-002 | a resume's rank count, node count or loader rank differs from the checkpoint's; the message names both | resume with the checkpoint's nodes and devices |
E-DP-003 | the global microbatches do not divide over the GPUs | pick a device count that divides :microbatches |
E-DP-004 | the ranks' or the nodes' parameters or optimizer states disagree at a checkpoint | a backend bug, never the experiment's fault: that checkpoint was not written, so --resume latest continues from the one before; report it with the step and SEXPGPU_DEVICES |
E-DP-005 | SEXPGPU_NODES is not a count, SEXPGPU_NODE_RANK is not in 0..nodes, SEXPGPU_RENDEZVOUS is not host:port, SEXPGPU_NODE_GRADIENTS is not f32 or bf16, or the device is not cuda | set all three on every node, with SEXPGPU_DEVICE=cuda |
E-DP-006 | a node differs from node 0 at the rendezvous: node count, a node index twice, GPU count, experiment, gradient dtype, or the checkpoint it resumes; the message names both | the same files, flags, device count and checkpoint location on every node |
E-DP-007 | a node did not arrive at the rendezvous within 300 s, or left the run | start every node; after a loss, restart every node with --resume latest |
Related: run, doctor, defrun, environment.