All pages
run
sexpgpu run <file.sx> compiles the file, applies the selection and the
overrides, and runs the training loop in this process. It reads the .sx,
everything it requires, the loaders' files (relative to the experiment
file, or s3:// URIs) and, when resuming, a checkpoint. It writes the
metrics target, the status file and checkpoints, each only when given a
location.
sexpgpu run my-run.sx --variant smoke --steps 3 --device cpu
SEXPGPU_DEVICE=cuda SEXPGPU_METRICS=jsonl:my-run.jsonl sexpgpu run my-run.sx
Flags
Each is an alias of an environment variable or an override of defrun:
| flag | meaning |
|---|---|
--steps N | replaces defrun :steps for this session; recorded as knob steps, source override |
--eval-every N | replaces every pass's :every the same way; 0 leaves only the pass after the last step, so a curriculum never runs |
--device cpu|cuda | SEXPGPU_DEVICE; see devices |
--metrics <target> | SEXPGPU_METRICS; see metrics |
--status <path> | SEXPGPU_STATUS; without it no status file is written |
--checkpoint-dir <dir>|s3://bucket/prefix | SEXPGPU_CHECKPOINT_DIR; see checkpoints |
--resume <dir>|s3://...|latest | SEXPGPU_RESUME; see checkpoints |
--json | the final summary as one JSON object on standard output instead of the done: block |
--variant, --set, --diagnostics | the selection |
Flags win over environment variables; see environment.
Output
On standard error: one configuration line, where the numbers go, on a GPU
the lowering report, one progress line
per evaluation pass (a pass other than eval names itself
first, probe: eval/loss ...), a timing block, on a GPU the
memory: peak line (devices), and the done: block.
There is no per-step line and no quiet switch; a step's numbers go to the
metrics target, and the status file is the thing to poll.
$ sexpgpu run my-run.sx
run: my-run.sx slug my-run device cpu steps 20 microbatches 2 evals eval@10 diagnostics - precision f32 seed 1 checkpoints - resume - status -
step 0/20 loss - | eval/loss 3.46573 | 0 records | 0.0s
step 10/20 loss 2.66862 | eval/loss 2.62369 | 80 records | 0.1s
step 20/20 loss 2.62068 | eval/loss 2.46219 | 160 records | 0.2s
timing graph 0.012s device-open 0.000s
timing kernels 0 compiled in 0.000s, 0 cache hits
timing step-0 0.007s eval-0 0.020s
timing steps 1..20 median 0.0077s p90 0.0082s
timing eval 2 passes median 0.0195s p90 0.0195s
timing total 0.214s over 20 steps
done: 20 steps, train/loss 2.6207
eval/loss 2.4622
counter records 160
counter tokens 2720
The configuration line's resume is - from the beginning, <dir> from a
named checkpoint, latest=<dir> from the newest, latest=none from the
beginning because there was none yet.
With --json the last block is one line on standard output instead; the
fields are in the summary:
{"kind":"summary","steps":2,"train_loss":3.347761631011963,"eval":{"eval/loss":3.265638828277588},"counters":{"records":16,"tokens":272},"stage":0,"metrics":"/tmp/my-run.jsonl","status":"/tmp/my-run.status.json"}
The run slug
The slug is SEXPGPU_RUN_SLUG or the file's stem, then -<variant>, then
one -<knob>=<value> per --set in knob-name order: my-run-smoke-lr=0.003.
One selection always names one run, and
Metrics by 041 keeps one experiment per slug,
so a resume lands in the same experiment. See
events.
Stopping
SIGTERM, SIGINT or SIGHUP stops the run at the next step boundary,
writes a checkpoint when there is a location, and exits 128 + n; the
details are in checkpoints.