S-exp GPU
All pages
Docs · The command lineMarkdown

doctor

sexpgpu doctor checks whether a run can start on this machine: one line per check on standard output, exit 0 or 1. Run it first on a new machine.

$ sexpgpu doctor
sexpgpu      0.1.0 /usr/local/bin/sexpgpu
cuda build   this binary has no --features cuda; SEXPGPU_DEVICE=cuda would refuse
gpu          none (nvidia-smi is not on PATH or found no device)
patterns     none without --features cuda
env          no SEXPGPU_* variable is set; the defaults are cpu and no metrics
metrics key  SYVAIN_METRICS_API_KEY not set; SEXPGPU_METRICS=syvain would refuse
s3           no AWS_* or SEXPGPU_*_AWS_* variable is set; an s3:// loader or checkpoint would refuse
disk         72Gi free in the working directory
run          ok: a run can start here

On a machine with a GPU a floor line follows gpu, saying whether the driver and the compute capability meet the CUDA requirements, and the CUDA binary adds a devices line with the CUDA ordinals a run would use and a patterns line saying whether the pattern kernels are on. A nodes line shows where a run over several nodes puts this process. The floor line is informational: a card or driver below the floor does not make doctor exit 1, but a run refuses the card when it opens the device, so read it before trusting run ok. Keys are reported as set or not, never by value.

When it exits 1

When any of these holds, each of which would stop a run before its first step:

  • an unreadable SEXPGPU_* configuration;
  • SEXPGPU_DEVICE=cuda in a binary without CUDA, or with no GPU;
  • an SEXPGPU_DEVICES that does not parse or names an ordinal that is not visible, E-DP-001;
  • a node topology that does not parse, E-DP-005;
  • an s3:// checkpoint location without a full set of credentials;
  • a local SEXPGPU_RESUME that is not a directory.

The device refusals apply only when SEXPGPU_DEVICE=cuda. doctor reads the same environment a run would.

Related: devices, install.