All pages
doctor
sexpgpu doctor checks whether a run can start on this machine: one line
per check on standard output, exit 0 or 1. Run it first on a new
machine.
$ sexpgpu doctor
sexpgpu 0.1.0 /usr/local/bin/sexpgpu
cuda build this binary has no --features cuda; SEXPGPU_DEVICE=cuda would refuse
gpu none (nvidia-smi is not on PATH or found no device)
patterns none without --features cuda
env no SEXPGPU_* variable is set; the defaults are cpu and no metrics
metrics key SYVAIN_METRICS_API_KEY not set; SEXPGPU_METRICS=syvain would refuse
s3 no AWS_* or SEXPGPU_*_AWS_* variable is set; an s3:// loader or checkpoint would refuse
disk 72Gi free in the working directory
run ok: a run can start here
On a machine with a GPU a floor line follows gpu, saying whether the
driver and the compute capability meet the
CUDA requirements, and the CUDA binary adds a
devices line with the CUDA ordinals a run would use and a patterns line
saying whether the pattern kernels are on. A nodes
line shows where a run over several nodes puts
this process. The floor line is informational: a card or driver below
the floor does not make doctor exit 1, but a run refuses the card when
it opens the device, so read it before trusting run ok. Keys are
reported as set or not, never by value.
When it exits 1
When any of these holds, each of which would stop a run before its first step:
- an unreadable
SEXPGPU_*configuration; SEXPGPU_DEVICE=cudain a binary without CUDA, or with no GPU;- an
SEXPGPU_DEVICESthat does not parse or names an ordinal that is not visible,E-DP-001; - a node topology that does not parse,
E-DP-005; - an
s3://checkpoint location without a full set of credentials; - a local
SEXPGPU_RESUMEthat is not a directory.
The device refusals apply only when SEXPGPU_DEVICE=cuda. doctor reads
the same environment a run would.