# Metrics

A run reports in two independent channels. Metrics are the time series:
one event per number, sent as it happens to a file, a pipe or
[Metrics by 041](https://metrics.041.io/). The [status file](https://sexpgpu.041.io/docs/status.md) is
the single current state, rewritten in place. Consume the metrics stream;
poll the status file.

## Targets

`SEXPGPU_METRICS`, or `--metrics`, read once at run start:

| target | where numbers go |
|---|---|
| unset | nowhere |
| `jsonl:<path>` | one JSON object per line, appended; parent directories are created |
| `stdout` | the same events on standard output, unbuffered, one per line |
| `syvain` | [Metrics by 041](https://metrics.041.io/) |
| `syvain:<folder-id>` | the same, in that folder |

`syvain` needs `SYVAIN_METRICS_API_KEY`; the folder is
`SYVAIN_METRICS_FOLDER_ID`, and one named in the target wins. Anything else
is an error at start. The run prints where its numbers go before it trains
(`metrics: /tmp/my-run.jsonl`, or the experiment's URL), and the same string
is in the summary, the status file and `sweep`'s table. With `syvain` the
run waits up to five seconds for the experiment id, then trains without a
URL to print.

Events are delivered in order, from one thread, so a run that reports
faster than Metrics by 041 accepts falls behind. When a run ends or
stops it waits up to thirty seconds for the queue, so a finished run
catches up, but a preempted one whose machine goes first loses what was
still queued, its `interrupted` and `checkpoint` annotations included; the
checkpoint itself is written before the wait.

`stdout` is the plugin interface: `sexpgpu run ... | my-forwarder` sends
events anywhere without the backend knowing. Standard output then carries
nothing but events, unless `--json` is also given: then its last line is the
[summary](https://sexpgpu.041.io/docs/status.md#the-summary), of kind `summary`, which is not an event
and has no `ts_ms`. A forwarder filters on `kind`. The event format is in
[JSONL events](https://sexpgpu.041.io/docs/events.md).

## Reporting from the experiment

`(metric "name" scalar :key value ...)` records one scalar under a name,
takes the keywords after it as metadata, and returns the value, so it wraps
an expression. It works anywhere a step runs:

| where | what it becomes |
|---|---|
| `objective`, `prepare` | averaged over a step's microbatches, emitted at that step, beside `train/objective` |
| `evaluate` | averaged over the batches of a pass, emitted at the step it ran, tagged with the pass, offered to the curriculum by name |
| the model body | observed in every graph the model is applied in: a training-step average and one per evaluation pass, under one name told apart by `split` and `pass`; a library model reports the same way |
| an optimizer body | refused, `E-CONTRACT-011`: use a [diagnostic](https://sexpgpu.041.io/docs/writing-diagnostics.md#optimizer-diagnostics) |
| `curriculum` | not read; the curriculum is a decision, not a report |

```lisp
(metric "probe/logit-rms" (sqrt (mean (* logits logits))) :layer "out")
```

A number about how the model is doing rather than how well, that nobody
needs on every microbatch of every run, is a [diagnostic](https://sexpgpu.041.io/docs/diagnostics.md):
it costs nothing until selected.

Metadata values are strings, keywords, numbers or booleans, recorded as
strings; a tensor is refused, because metadata is fixed at compile time.
Metadata is capped at 32 keys, 128-byte keys, 512-byte values and 4 KiB
of canonical JSON. The compiler sizes every `metric` exactly,
with the runtime's keys and the longest stage name counted, and refuses
one over the limit as `E-CONTRACT-012` rather than letting the run stop at
its first event.

## Runtime metadata

The runtime adds its own keys to every event of a series:

| key | on | value |
|---|---|---|
| `split` | every metric | `train` from the training graph and counters, `eval` from every pass and `curriculum/stage` |
| `pass` | every metric of an evaluation pass, and its `curriculum/stage` | the `defeval` name, `eval` or `probe` |
| `stage` | every metric, when the training loader has several stages | the stage it was measured in |
| `param` | every optimizer diagnostic | the update group its values folded over |

A series is its name and its metadata together, so `train/loss` from the
`easy` and the `hard` stage are two series. An experiment that writes
`:split`, `:pass` or `:stage` itself keeps its own value.

## Metric names

| name | what it is |
|---|---|
| whatever `metric` was called with | an observation, averaged over a step's microbatches or a pass's batches |
| `diagnostics/<name>` | a selected [diagnostic](https://sexpgpu.041.io/docs/diagnostics.md), on sampled steps, folded by its reducer |
| `train/objective` | the objective's own value, whatever its reduction and name |
| `counter/<name>` | a declared [counter](https://sexpgpu.041.io/docs/data.md#counters)'s running total, every step |
| `rate/<counter>` | what the counter grew by this step, per second of `runtime/step_seconds` |
| `runtime/step_seconds` | the whole step, first batch to end of update, device synchronized, evaluation excluded |
| `runtime/train_seconds` | the sum of `runtime/step_seconds` so far in this process |
| `runtime/eval_seconds` | one evaluation pass, synchronized, with its `pass` |
| `runtime/peak_bytes` | the device allocator's high-water mark of live bytes since the start; absent on the CPU |
| `curriculum/stage` | the stage index an evaluation ran in, with every evaluation |

Every run reports the runtime series with nothing to select. The runtime
knows no domain, so `(counter :tokens ...)` gives `rate/tokens`, tokens
per second. On the CPU timings are wall time; under data parallel they are
rank zero's. `train/loss` and `eval/loss` mean something to the runtime; see
[objective](https://sexpgpu.041.io/docs/objective.md#two-names-the-runtime-reads).

Related: [JSONL events](https://sexpgpu.041.io/docs/events.md), [diagnostics](https://sexpgpu.041.io/docs/diagnostics.md),
[the status file](https://sexpgpu.041.io/docs/status.md).

---

SexpGPU documentation. Every page: https://sexpgpu.041.io/llms.txt
