S-exp GPU
All pages
Docs · The command lineMarkdown

run

sexpgpu run <file.sx> compiles the file, applies the selection and the overrides, and runs the training loop in this process. It reads the .sx, everything it requires, the loaders' files (relative to the experiment file, or s3:// URIs) and, when resuming, a checkpoint. It writes the metrics target, the status file and checkpoints, each only when given a location.

sexpgpu run my-run.sx --variant smoke --steps 3 --device cpu
SEXPGPU_DEVICE=cuda SEXPGPU_METRICS=jsonl:my-run.jsonl sexpgpu run my-run.sx

Flags

Each is an alias of an environment variable or an override of defrun:

flagmeaning
--steps Nreplaces defrun :steps for this session; recorded as knob steps, source override
--eval-every Nreplaces every pass's :every the same way; 0 leaves only the pass after the last step, so a curriculum never runs
--device cpu|cudaSEXPGPU_DEVICE; see devices
--metrics <target>SEXPGPU_METRICS; see metrics
--status <path>SEXPGPU_STATUS; without it no status file is written
--checkpoint-dir <dir>|s3://bucket/prefixSEXPGPU_CHECKPOINT_DIR; see checkpoints
--resume <dir>|s3://...|latestSEXPGPU_RESUME; see checkpoints
--jsonthe final summary as one JSON object on standard output instead of the done: block
--variant, --set, --diagnosticsthe selection

Flags win over environment variables; see environment.

Output

On standard error: one configuration line, where the numbers go, on a GPU the lowering report, one progress line per evaluation pass (a pass other than eval names itself first, probe: eval/loss ...), a timing block, on a GPU the memory: peak line (devices), and the done: block. There is no per-step line and no quiet switch; a step's numbers go to the metrics target, and the status file is the thing to poll.

$ sexpgpu run my-run.sx
run: my-run.sx slug my-run device cpu steps 20 microbatches 2 evals eval@10 diagnostics - precision f32 seed 1 checkpoints - resume - status -
step 0/20 loss - | eval/loss 3.46573 | 0 records | 0.0s
step 10/20 loss 2.66862 | eval/loss 2.62369 | 80 records | 0.1s
step 20/20 loss 2.62068 | eval/loss 2.46219 | 160 records | 0.2s
timing graph 0.012s device-open 0.000s
timing kernels 0 compiled in 0.000s, 0 cache hits
timing step-0 0.007s eval-0 0.020s
timing steps 1..20 median 0.0077s p90 0.0082s
timing eval 2 passes median 0.0195s p90 0.0195s
timing total 0.214s over 20 steps
done: 20 steps, train/loss 2.6207
  eval/loss 2.4622
  counter records 160
  counter tokens 2720

The configuration line's resume is - from the beginning, <dir> from a named checkpoint, latest=<dir> from the newest, latest=none from the beginning because there was none yet.

With --json the last block is one line on standard output instead; the fields are in the summary:

{"kind":"summary","steps":2,"train_loss":3.347761631011963,"eval":{"eval/loss":3.265638828277588},"counters":{"records":16,"tokens":272},"stage":0,"metrics":"/tmp/my-run.jsonl","status":"/tmp/my-run.status.json"}

The run slug

The slug is SEXPGPU_RUN_SLUG or the file's stem, then -<variant>, then one -<knob>=<value> per --set in knob-name order: my-run-smoke-lr=0.003. One selection always names one run, and Metrics by 041 keeps one experiment per slug, so a resume lands in the same experiment. See events.

Stopping

SIGTERM, SIGINT or SIGHUP stops the run at the next step boundary, writes a checkpoint when there is a location, and exits 128 + n; the details are in checkpoints.

Related: sweep, bundle, metrics.