All pages
Writing diagnostics
(diagnostic "name" scalar :reduce :max :key value ...) records one scalar
under diagnostics/<name> and returns it. Put it next to the thing it
examines; it computes only when a run selects it. How selection and
sampling work is diagnostics.
(let ((logits (head (norm-out ...))))
(diagnostic "logits/max-abs" (max (abs logits)) :reduce :max)
(diagnostic "logits/rms" (rms logits) :reduce :mean)
(tap-gradient logits
(lambda (g) (diagnostic "logits/grad-rms" (rms g) :reduce :mean))))
- It works where
metricdoes (prepare, the model body,objective,evaluate) and in an optimizer body, so a library model or optimizer carries its own and every run that uses it can turn them on. - Name it by what it examines,
attention/entropy, notlayer3-thing: the name is what a run selects by. - Keywords after
:reduceare metadata, as formetric::layer 3makes one series per layer. - A diagnostic in the model body reports from the training graph and from
each evaluation pass;
splittells them apart.
Reducers
:reduce is required. It is how one series folds wherever several values
meet: the batches of a pass, the parameters of an update group, the ranks
of a data-parallel run.
| reducer | folds to | for |
|---|---|---|
:mean | the sum over the graph runs that reported, divided by their number | typical values |
:min | the smallest | a lower bound, a margin |
:max | the largest | a worst case: constraint error, largest logit |
:sum | the total | a count |
A :min or :max that meets a NaN reports NaN, because that is the worst
case. A diagnostic without :reduce, with an unknown one, or one name
declared with two reducers is E-CONTRACT-015.
Optimizer diagnostics
In an optimizer body a diagnostic sees everything the update does: the
parameter before, grad, every state, the new moments, the proposed step
and the new parameter. The built-in optimizers report theirs through
update-diagnostics from sexpgpu/optim:
(defoptimizer sgd (param grad ctx &key (lr 0.01) (weight-decay 0.0))
(let ((next (- param (* lr (+ grad (* weight-decay param))))))
(update-diagnostics param grad next)
(optimizer-update next)))
- The body is traced once per parameter, so each diagnostic has one value
per parameter, and the reducer folds them across the update group:
:maxis the worst parameter,:meanthe typical one. - Each series carries
param, the group's name (default,model.embed.*), besidesplittrain. Metadata the body passes itself,:part "plus"for half of a packed parameter, keeps apart what should not fold together. - On a sampled step values stay on the device while every update graph runs and download once the update is done.
A parameter the loss does not depend on gets a gradient of rounding noise,
and its update ratio is noise over noise, of order one and different on
every device. An attention key bias is one: softmax ignores a shift shared
by a whole row. Where a group holds one, its :max can be that parameter;
give such parameters a group of their own, or no bias.
Gradient diagnostics
(tap-gradient x (lambda (g) ...)) returns x unchanged and hands the
lambda g, the adjoint of x: the gradient of the objective with respect
to it, as autodiff computes it for this microbatch.
- The adjoint is of the objective as written: an objective that averages
has the division in
g; one that sums does not. - Only the training graph has an objective, so a tap in a model body is
traced there and nowhere else; an evaluation pass returns
xand never calls the lambda. - The lambda's
diagnosticcalls are its only output; its return value is dropped. Ametricinside it isE-CONTRACT-017, because a gradient exists only on the sampled microbatch. - An unselected tap is removed with its diagnostics before autodiff and keeps nothing alive.
Related: diagnostics, sexpgpu/optim, writing optimizers.