S-exp GPU
All pages
Docs · The command lineMarkdown

Environment

Configuration that belongs to the machine, not the experiment, is read from the environment once, at run start.

Variables

variablevaluesmeaning
SEXPGPU_DEVICEcpu, cudacpu, the interpreter, is the default; cuda needs the Linux CUDA binary. See devices
SEXPGPU_DEVICES0, 0,1, ...; unset is all visibleCUDA ordinals in rank order; see several GPUs
SEXPGPU_NODES, SEXPGPU_NODE_RANK, SEXPGPU_RENDEZVOUS, SEXPGPU_NODE_GRADIENTSsee several nodesone run across machines
SEXPGPU_METRICSjsonl:<path>, stdout, syvain, syvain:<folder-id>where numbers go; unset is nowhere. See metrics
SEXPGPU_STATUSa paththe status file; unset is none
SEXPGPU_DIAGNOSTICSglobs, attention/*,logits/*replaces defrun :diagnostics for every compile; --diagnostics wins over it
SEXPGPU_RUN_SLUGa namethe base of the run slug; default the file's stem
SEXPGPU_CHECKPOINT_DIRa directory or s3://bucket/prefixwhere checkpoints go; without it :checkpoint-every does nothing
SEXPGPU_RESUMEa directory, s3://bucket/prefix/step-<n>, or latesta checkpoint to continue from; see resume
SEXPGPU_EVAL_LIMITa positive integer; default 10000000the compile-time work budget; E-EVAL-041 when exceeded. See limits
SYVAIN_METRICS_API_KEYa keyrequired by the syvain target
SYVAIN_METRICS_FOLDER_IDa folder idthe syvain target's folder; syvain:<id> wins over it
SEXPGPU_MEMORYGiB, 12 or 7.5the free device memory the memory plan assumes instead of the device's own, to try a shape against a smaller card
SEXPGPU_EXPLAINkernelsadds every launch of each graph, with its shape, dtype and model path, to the lowering report of run and explain

S3 credentials

Two things read object storage: the loader, when a source names an s3:// file or manifest, and checkpoints, when the location is an s3:// prefix. Each is a role, and each of the four values a role needs resolves on its own, first match wins:

SEXPGPU_DATA_AWS_<name>  /  SEXPGPU_CHECKPOINT_AWS_<name>     the role's own
  >  SEXPGPU_AWS_<name>                                        shared by both roles
  >  AWS_<name>                                                what every other tool reads

for <name> in ACCESS_KEY_ID, SECRET_ACCESS_KEY, ENDPOINT_URL and REGION. Plain AWS_* is enough on most machines; the SEXPGPU_* names are for when AWS_* is set for another tool, or when data and checkpoints live in different buckets. An s3:// checkpoint location with any of the four missing is refused before the first step, and the message names it.

Precedence

For the five settings that exist as a flag and a variable, the flag wins, and there is no defrun key for any of them:

--device / --metrics / --status / --checkpoint-dir / --resume
  >  SEXPGPU_DEVICE / SEXPGPU_METRICS / SEXPGPU_STATUS / SEXPGPU_CHECKPOINT_DIR / SEXPGPU_RESUME

For the two the file states, the flag wins, and there is no variable:

--steps / --eval-every  >  defrun :steps / defeval :every

Diagnostics exist in all three places:

--diagnostics  >  SEXPGPU_DIAGNOSTICS  >  defrun :diagnostics

An override of the file is recorded in the manifest as a knob of source override, so the IR, the metrics metadata and the status file all say the run was cut short or diagnosed. Knob values follow their own rule, --set over the variant over the default; see knobs.

Related: run, doctor.