# Writing diagnostics

`(diagnostic "name" scalar :reduce :max :key value ...)` records one scalar
under `diagnostics/<name>` and returns it. Put it next to the thing it
examines; it computes only when a run selects it. How selection and
sampling work is [diagnostics](https://sexpgpu.041.io/docs/diagnostics.md).

```lisp
(let ((logits (head (norm-out ...))))
  (diagnostic "logits/max-abs" (max (abs logits)) :reduce :max)
  (diagnostic "logits/rms" (rms logits) :reduce :mean)
  (tap-gradient logits
                (lambda (g) (diagnostic "logits/grad-rms" (rms g) :reduce :mean))))
```

- It works where `metric` does (`prepare`, the model body, `objective`,
  `evaluate`) and in an optimizer body, so a library model or optimizer
  carries its own and every run that uses it can turn them on.
- Name it by what it examines, `attention/entropy`, not `layer3-thing`:
  the name is what a run selects by.
- Keywords after `:reduce` are metadata, as for `metric`: `:layer 3` makes
  one series per layer.
- A diagnostic in the model body reports from the training graph and from
  each evaluation pass; `split` tells them apart.

## Reducers

`:reduce` is required. It is how one series folds wherever several values
meet: the batches of a pass, the parameters of an update group, the ranks
of a data-parallel run.

| reducer | folds to | for |
|---|---|---|
| `:mean` | the sum over the graph runs that reported, divided by their number | typical values |
| `:min` | the smallest | a lower bound, a margin |
| `:max` | the largest | a worst case: constraint error, largest logit |
| `:sum` | the total | a count |

A `:min` or `:max` that meets a NaN reports NaN, because that is the worst
case. A diagnostic without `:reduce`, with an unknown one, or one name
declared with two reducers is `E-CONTRACT-015`.

## Optimizer diagnostics

In an optimizer body a diagnostic sees everything the update does: the
parameter before, `grad`, every state, the new moments, the proposed step
and the new parameter. The built-in optimizers report theirs through
`update-diagnostics` from `sexpgpu/optim`:

```lisp
(defoptimizer sgd (param grad ctx &key (lr 0.01) (weight-decay 0.0))
  (let ((next (- param (* lr (+ grad (* weight-decay param))))))
    (update-diagnostics param grad next)
    (optimizer-update next)))
```

- The body is traced once per parameter, so each diagnostic has one value
  per parameter, and the reducer folds them across the update group: `:max`
  is the worst parameter, `:mean` the typical one.
- Each series carries `param`, the group's name (`default`,
  `model.embed.*`), beside `split` `train`. Metadata the body passes itself,
  `:part "plus"` for half of a packed parameter, keeps apart what should
  not fold together.
- On a sampled step values stay on the device while every update graph
  runs and download once the update is done.

A parameter the loss does not depend on gets a gradient of rounding noise,
and its update ratio is noise over noise, of order one and different on
every device. An attention key bias is one: softmax ignores a shift shared
by a whole row. Where a group holds one, its `:max` can be that parameter;
give such parameters a group of their own, or no bias.

## Gradient diagnostics

`(tap-gradient x (lambda (g) ...))` returns `x` unchanged and hands the
lambda `g`, the adjoint of `x`: the gradient of the objective with respect
to it, as autodiff computes it for this microbatch.

- The adjoint is of the objective as written: an objective that averages
  has the division in `g`; one that sums does not.
- Only the training graph has an objective, so a tap in a model body is
  traced there and nowhere else; an evaluation pass returns `x` and never
  calls the lambda.
- The lambda's `diagnostic` calls are its only output; its return value is
  dropped. A `metric` inside it is `E-CONTRACT-017`, because a gradient
  exists only on the sampled microbatch.
- An unselected tap is removed with its diagnostics before autodiff and
  keeps nothing alive.

Related: [diagnostics](https://sexpgpu.041.io/docs/diagnostics.md), [sexpgpu/optim](https://sexpgpu.041.io/docs/optim.md),
[writing optimizers](https://sexpgpu.041.io/docs/optimizers.md).

---

SexpGPU documentation. Every page: https://sexpgpu.041.io/llms.txt
