All pages
Diagnostics
A diagnostic is a number that says how the model is doing rather than how well: the largest logit, the entropy of an attention head, the radius of a point on a manifold. It is written next to the thing it examines, often in the model or optimizer library, and costs nothing until a run selects it by name.
metric | diagnostic | |
|---|---|---|
| computes | every microbatch of every step, every batch of every pass | only when selected, only on sampled steps |
| folds | always the mean | the reducer it states: :mean, :min, :max or :sum |
| name | as called | diagnostics/<name> |
| off | never | unless selected |
Writing them is writing diagnostics.
Turning diagnostics on
explain lists what a file defines: every diagnostic, its reducer, its
graphs, and whether this run selects it.
$ sexpgpu explain my-run.sx --diagnostics 'logits/*'
...
diagnostics:
selected logits/*, 3 of 8
sampled every step and the last
attention/entropy mean train, eval
attention/max-probability max train, eval
residual/rms mean train, eval
logits/max-abs max train, eval selected
logits/rms mean train, eval selected
logits/grad-rms mean train selected
optimizer/grad-rms max update default
optimizer/update-ratio max update default
Select with globs over the /-separated name, as optimizer groups select
parameters: * inside one segment, ** across any number, so
attention/*, **/rms, and ** for everything. The later place wins:
| where | how | for |
|---|---|---|
| the file | (defrun ... :diagnostics ["attention/*" "logits/*"]) | what this experiment always looks at |
| the environment | SEXPGPU_DIAGNOSTICS='attention/*,logits/*' | a launched run, a run.sh, a sweep |
| the command line | --diagnostics 'attention/*,logits/*' | one session; bundle bakes it into run.sh |
A selection from the environment or the command line is recorded as the
knob diagnostics of source override, and the selection itself is
run.diagnostics in the open event. A glob that names nothing the file
or its libraries define is E-CONTRACT-016, and the fix lists what there
is.
When they compute
(defrun :steps 800 :diagnostics ["geometry/*"] :diagnose-every 5 :diagnose-first 20)
- A training step is sampled when it is among the first
:diagnose-first, a multiple of:diagnose-every, or the last. The defaults,1and0, sample every step. - A sampled step computes the selected diagnostics on its first microbatch (rank zero's under data parallel); the others run without them.
- An evaluation pass computes them on its first batch when its step is sampled; the final evaluation always is.
Every graph a selected diagnostic is in compiles twice: the plain graph for unsampled steps and the diagnosed one for sampled steps, so an unsampled step does no extra work. With nothing selected the run is bit-for-bit the same as without diagnostics. A diagnostic never changes a parameter, a state or any other metric.
Where they go
To the metrics stream only, one metric event per series per sampled step,
with the usual split, pass and stage. Not to the summary, the status
file, the progress line or the curriculum, since unsampled steps have no
value.
What they cost
A second compile of each graph a selected diagnostic is in, plus its own computation and one download per value on sampled microbatches only.
Built-in diagnostics
| name | reducer | from | what it is |
|---|---|---|---|
attention/entropy | mean | gpt, per :layer | entropy of the attention rows over batch, heads and positions |
attention/max-probability | max | gpt, per :layer | the largest attention probability |
residual/rms | mean | gpt, per :layer | root mean square of the residual stream after each block |
logits/max-abs | max | gpt | the largest logit |
logits/rms | mean | gpt | root mean square of the logits |
logits/grad-rms | mean | gpt, training only | root mean square of the logits' gradient |
optimizer/grad-rms | max | sgd, momentum-sgd, adamw, muon | the gradient's root mean square, the group's largest |
optimizer/update-ratio | max | the same | the step's root mean square over the parameter's |
Your own libraries can carry their own set the same way, for example
geometry/* diagnostics beside a hyperbolic model and its Riemannian
optimizer.
Related: writing diagnostics, defrun, metrics.