# Diagnostics

A diagnostic is a number that says how the model is doing rather than how
well: the largest logit, the entropy of an attention head, the radius of a
point on a manifold. It is written next to the thing it examines, often in
the model or optimizer library, and costs nothing until a run selects it
by name.

| | `metric` | `diagnostic` |
|---|---|---|
| computes | every microbatch of every step, every batch of every pass | only when selected, only on sampled steps |
| folds | always the mean | the reducer it states: `:mean`, `:min`, `:max` or `:sum` |
| name | as called | `diagnostics/<name>` |
| off | never | unless selected |

Writing them is [writing diagnostics](https://sexpgpu.041.io/docs/writing-diagnostics.md).

## Turning diagnostics on

`explain` lists what a file defines: every diagnostic, its reducer, its
graphs, and whether this run selects it.

```console
$ sexpgpu explain my-run.sx --diagnostics 'logits/*'
...
diagnostics:
  selected         logits/*, 3 of 8
  sampled          every step and the last
  attention/entropy          mean  train, eval
  attention/max-probability  max   train, eval
  residual/rms               mean  train, eval
  logits/max-abs             max   train, eval  selected
  logits/rms                 mean  train, eval  selected
  logits/grad-rms            mean  train  selected
  optimizer/grad-rms         max   update default
  optimizer/update-ratio     max   update default
```

Select with globs over the `/`-separated name, as optimizer groups select
parameters: `*` inside one segment, `**` across any number, so
`attention/*`, `**/rms`, and `**` for everything. The later place wins:

| where | how | for |
|---|---|---|
| the file | `(defrun ... :diagnostics ["attention/*" "logits/*"])` | what this experiment always looks at |
| the environment | `SEXPGPU_DIAGNOSTICS='attention/*,logits/*'` | a launched run, a `run.sh`, a sweep |
| the command line | `--diagnostics 'attention/*,logits/*'` | one session; `bundle` bakes it into `run.sh` |

A selection from the environment or the command line is recorded as the
knob `diagnostics` of source `override`, and the selection itself is
`run.diagnostics` in the `open` event. A glob that names nothing the file
or its libraries define is `E-CONTRACT-016`, and the fix lists what there
is.

## When they compute

```lisp
(defrun :steps 800 :diagnostics ["geometry/*"] :diagnose-every 5 :diagnose-first 20)
```

- A training step is sampled when it is among the first `:diagnose-first`,
  a multiple of `:diagnose-every`, or the last. The defaults, `1` and `0`,
  sample every step.
- A sampled step computes the selected diagnostics on its first microbatch
  (rank zero's under data parallel); the others run without them.
- An evaluation pass computes them on its first batch when its step is
  sampled; the final evaluation always is.

Every graph a selected diagnostic is in compiles twice: the plain graph
for unsampled steps and the diagnosed one for sampled steps, so an
unsampled step does no extra work. With nothing selected the run is
bit-for-bit the same as without diagnostics. A diagnostic never changes a
parameter, a state or any other metric.

## Where they go

To the metrics stream only, one `metric` event per series per sampled step,
with the usual `split`, `pass` and `stage`. Not to the summary, the status
file, the progress line or the curriculum, since unsampled steps have no
value.

## What they cost

A second compile of each graph a selected diagnostic is in, plus its own
computation and one download per value on sampled microbatches only.

## Built-in diagnostics

| name | reducer | from | what it is |
|---|---|---|---|
| `attention/entropy` | mean | `gpt`, per `:layer` | entropy of the attention rows over batch, heads and positions |
| `attention/max-probability` | max | `gpt`, per `:layer` | the largest attention probability |
| `residual/rms` | mean | `gpt`, per `:layer` | root mean square of the residual stream after each block |
| `logits/max-abs` | max | `gpt` | the largest logit |
| `logits/rms` | mean | `gpt` | root mean square of the logits |
| `logits/grad-rms` | mean | `gpt`, training only | root mean square of the logits' gradient |
| `optimizer/grad-rms` | max | `sgd`, `momentum-sgd`, `adamw`, `muon` | the gradient's root mean square, the group's largest |
| `optimizer/update-ratio` | max | the same | the step's root mean square over the parameter's |

Your own libraries can carry their own set the same way, for example
`geometry/*` diagnostics beside a hyperbolic model and its Riemannian
optimizer.

Related: [writing diagnostics](https://sexpgpu.041.io/docs/writing-diagnostics.md), [defrun](https://sexpgpu.041.io/docs/defrun.md),
[metrics](https://sexpgpu.041.io/docs/metrics.md).

---

SexpGPU documentation. Every page: https://sexpgpu.041.io/llms.txt
