S-exp GPU
All pages
Docs · Observing a runMarkdown

Diagnostics

A diagnostic is a number that says how the model is doing rather than how well: the largest logit, the entropy of an attention head, the radius of a point on a manifold. It is written next to the thing it examines, often in the model or optimizer library, and costs nothing until a run selects it by name.

metricdiagnostic
computesevery microbatch of every step, every batch of every passonly when selected, only on sampled steps
foldsalways the meanthe reducer it states: :mean, :min, :max or :sum
nameas calleddiagnostics/<name>
offneverunless selected

Writing them is writing diagnostics.

Turning diagnostics on

explain lists what a file defines: every diagnostic, its reducer, its graphs, and whether this run selects it.

$ sexpgpu explain my-run.sx --diagnostics 'logits/*'
...
diagnostics:
  selected         logits/*, 3 of 8
  sampled          every step and the last
  attention/entropy          mean  train, eval
  attention/max-probability  max   train, eval
  residual/rms               mean  train, eval
  logits/max-abs             max   train, eval  selected
  logits/rms                 mean  train, eval  selected
  logits/grad-rms            mean  train  selected
  optimizer/grad-rms         max   update default
  optimizer/update-ratio     max   update default

Select with globs over the /-separated name, as optimizer groups select parameters: * inside one segment, ** across any number, so attention/*, **/rms, and ** for everything. The later place wins:

wherehowfor
the file(defrun ... :diagnostics ["attention/*" "logits/*"])what this experiment always looks at
the environmentSEXPGPU_DIAGNOSTICS='attention/*,logits/*'a launched run, a run.sh, a sweep
the command line--diagnostics 'attention/*,logits/*'one session; bundle bakes it into run.sh

A selection from the environment or the command line is recorded as the knob diagnostics of source override, and the selection itself is run.diagnostics in the open event. A glob that names nothing the file or its libraries define is E-CONTRACT-016, and the fix lists what there is.

When they compute

(defrun :steps 800 :diagnostics ["geometry/*"] :diagnose-every 5 :diagnose-first 20)
  • A training step is sampled when it is among the first :diagnose-first, a multiple of :diagnose-every, or the last. The defaults, 1 and 0, sample every step.
  • A sampled step computes the selected diagnostics on its first microbatch (rank zero's under data parallel); the others run without them.
  • An evaluation pass computes them on its first batch when its step is sampled; the final evaluation always is.

Every graph a selected diagnostic is in compiles twice: the plain graph for unsampled steps and the diagnosed one for sampled steps, so an unsampled step does no extra work. With nothing selected the run is bit-for-bit the same as without diagnostics. A diagnostic never changes a parameter, a state or any other metric.

Where they go

To the metrics stream only, one metric event per series per sampled step, with the usual split, pass and stage. Not to the summary, the status file, the progress line or the curriculum, since unsampled steps have no value.

What they cost

A second compile of each graph a selected diagnostic is in, plus its own computation and one download per value on sampled microbatches only.

Built-in diagnostics

namereducerfromwhat it is
attention/entropymeangpt, per :layerentropy of the attention rows over batch, heads and positions
attention/max-probabilitymaxgpt, per :layerthe largest attention probability
residual/rmsmeangpt, per :layerroot mean square of the residual stream after each block
logits/max-absmaxgptthe largest logit
logits/rmsmeangptroot mean square of the logits
logits/grad-rmsmeangpt, training onlyroot mean square of the logits' gradient
optimizer/grad-rmsmaxsgd, momentum-sgd, adamw, muonthe gradient's root mean square, the group's largest
optimizer/update-ratiomaxthe samethe step's root mean square over the parameter's

Your own libraries can carry their own set the same way, for example geometry/* diagnostics beside a hyperbolic model and its Riemannian optimizer.

Related: writing diagnostics, defrun, metrics.