All pages
Environment
Configuration that belongs to the machine, not the experiment, is read from the environment once, at run start.
Variables
| variable | values | meaning |
|---|---|---|
SEXPGPU_DEVICE | cpu, cuda | cpu, the interpreter, is the default; cuda needs the Linux CUDA binary. See devices |
SEXPGPU_DEVICES | 0, 0,1, ...; unset is all visible | CUDA ordinals in rank order; see several GPUs |
SEXPGPU_NODES, SEXPGPU_NODE_RANK, SEXPGPU_RENDEZVOUS, SEXPGPU_NODE_GRADIENTS | see several nodes | one run across machines |
SEXPGPU_METRICS | jsonl:<path>, stdout, syvain, syvain:<folder-id> | where numbers go; unset is nowhere. See metrics |
SEXPGPU_STATUS | a path | the status file; unset is none |
SEXPGPU_DIAGNOSTICS | globs, attention/*,logits/* | replaces defrun :diagnostics for every compile; --diagnostics wins over it |
SEXPGPU_RUN_SLUG | a name | the base of the run slug; default the file's stem |
SEXPGPU_CHECKPOINT_DIR | a directory or s3://bucket/prefix | where checkpoints go; without it :checkpoint-every does nothing |
SEXPGPU_RESUME | a directory, s3://bucket/prefix/step-<n>, or latest | a checkpoint to continue from; see resume |
SEXPGPU_EVAL_LIMIT | a positive integer; default 10000000 | the compile-time work budget; E-EVAL-041 when exceeded. See limits |
SYVAIN_METRICS_API_KEY | a key | required by the syvain target |
SYVAIN_METRICS_FOLDER_ID | a folder id | the syvain target's folder; syvain:<id> wins over it |
SEXPGPU_MEMORY | GiB, 12 or 7.5 | the free device memory the memory plan assumes instead of the device's own, to try a shape against a smaller card |
SEXPGPU_EXPLAIN | kernels | adds every launch of each graph, with its shape, dtype and model path, to the lowering report of run and explain |
S3 credentials
Two things read object storage: the loader, when a source names an s3://
file or manifest, and checkpoints, when the location is an s3:// prefix.
Each is a role, and each of the four values a role needs resolves on its
own, first match wins:
SEXPGPU_DATA_AWS_<name> / SEXPGPU_CHECKPOINT_AWS_<name> the role's own
> SEXPGPU_AWS_<name> shared by both roles
> AWS_<name> what every other tool reads
for <name> in ACCESS_KEY_ID, SECRET_ACCESS_KEY, ENDPOINT_URL and
REGION. Plain AWS_* is enough on most machines; the SEXPGPU_* names
are for when AWS_* is set for another tool, or when data and
checkpoints live in different buckets. An s3://
checkpoint location with any of the four missing is refused before the
first step, and the message names it.
Precedence
For the five settings that exist as a flag and a variable, the flag wins,
and there is no defrun key for any of them:
--device / --metrics / --status / --checkpoint-dir / --resume
> SEXPGPU_DEVICE / SEXPGPU_METRICS / SEXPGPU_STATUS / SEXPGPU_CHECKPOINT_DIR / SEXPGPU_RESUME
For the two the file states, the flag wins, and there is no variable:
--steps / --eval-every > defrun :steps / defeval :every
Diagnostics exist in all three places:
--diagnostics > SEXPGPU_DIAGNOSTICS > defrun :diagnostics
An override of the file is recorded in the manifest as a knob of source
override, so the IR, the metrics metadata and the status file all say
the run was cut short or diagnosed. Knob values follow their own rule,
--set over the variant over the default; see
knobs.