S-exp GPU
All pages
Docs · Observing a runMarkdown

Writing diagnostics

(diagnostic "name" scalar :reduce :max :key value ...) records one scalar under diagnostics/<name> and returns it. Put it next to the thing it examines; it computes only when a run selects it. How selection and sampling work is diagnostics.

(let ((logits (head (norm-out ...))))
  (diagnostic "logits/max-abs" (max (abs logits)) :reduce :max)
  (diagnostic "logits/rms" (rms logits) :reduce :mean)
  (tap-gradient logits
                (lambda (g) (diagnostic "logits/grad-rms" (rms g) :reduce :mean))))
  • It works where metric does (prepare, the model body, objective, evaluate) and in an optimizer body, so a library model or optimizer carries its own and every run that uses it can turn them on.
  • Name it by what it examines, attention/entropy, not layer3-thing: the name is what a run selects by.
  • Keywords after :reduce are metadata, as for metric: :layer 3 makes one series per layer.
  • A diagnostic in the model body reports from the training graph and from each evaluation pass; split tells them apart.

Reducers

:reduce is required. It is how one series folds wherever several values meet: the batches of a pass, the parameters of an update group, the ranks of a data-parallel run.

reducerfolds tofor
:meanthe sum over the graph runs that reported, divided by their numbertypical values
:minthe smallesta lower bound, a margin
:maxthe largesta worst case: constraint error, largest logit
:sumthe totala count

A :min or :max that meets a NaN reports NaN, because that is the worst case. A diagnostic without :reduce, with an unknown one, or one name declared with two reducers is E-CONTRACT-015.

Optimizer diagnostics

In an optimizer body a diagnostic sees everything the update does: the parameter before, grad, every state, the new moments, the proposed step and the new parameter. The built-in optimizers report theirs through update-diagnostics from sexpgpu/optim:

(defoptimizer sgd (param grad ctx &key (lr 0.01) (weight-decay 0.0))
  (let ((next (- param (* lr (+ grad (* weight-decay param))))))
    (update-diagnostics param grad next)
    (optimizer-update next)))
  • The body is traced once per parameter, so each diagnostic has one value per parameter, and the reducer folds them across the update group: :max is the worst parameter, :mean the typical one.
  • Each series carries param, the group's name (default, model.embed.*), beside split train. Metadata the body passes itself, :part "plus" for half of a packed parameter, keeps apart what should not fold together.
  • On a sampled step values stay on the device while every update graph runs and download once the update is done.

A parameter the loss does not depend on gets a gradient of rounding noise, and its update ratio is noise over noise, of order one and different on every device. An attention key bias is one: softmax ignores a shift shared by a whole row. Where a group holds one, its :max can be that parameter; give such parameters a group of their own, or no bias.

Gradient diagnostics

(tap-gradient x (lambda (g) ...)) returns x unchanged and hands the lambda g, the adjoint of x: the gradient of the objective with respect to it, as autodiff computes it for this microbatch.

  • The adjoint is of the objective as written: an objective that averages has the division in g; one that sums does not.
  • Only the training graph has an objective, so a tap in a model body is traced there and nowhere else; an evaluation pass returns x and never calls the lambda.
  • The lambda's diagnostic calls are its only output; its return value is dropped. A metric inside it is E-CONTRACT-017, because a gradient exists only on the sampled microbatch.
  • An unselected tap is removed with its diagnostics before autodiff and keeps nothing alive.

Related: diagnostics, sexpgpu/optim, writing optimizers.