All pages
JSONL events
The metrics stream, as written to a jsonl: target or to stdout, is one
complete JSON object per line, each with ts_ms, the Unix epoch in
milliseconds. The same events go to
Metrics by 041. Targets are in
metrics.
Kinds
| kind | fields | when |
|---|---|---|
open | slug, description, metadata | the first line of the stream |
metric | step, name, value, metadata | one per number |
annotation | step, annotation, metadata | the lowering report, the memory plan, a stage change, a checkpoint, a resume, an interruption, a failure |
done | slug | the last line of a run that finished, after which the file is flushed; a run stopped by a signal ends without one |
A metric carries what metric was called with, its metadata plus the
runtime's. A non-finite value is still written to a JSONL target; the
Syvain target drops it. run --json's summary line goes to standard
output and never into a JSONL file; with --metrics stdout it follows the
last event on the same stream.
{"ts_ms":1790189597015,"step":4,"kind":"metric","name":"eval/loss","value":3.2047,"metadata":{"pass":"eval","split":"eval","stage":"easy"}}
{"ts_ms":1790189597016,"step":4,"kind":"annotation","annotation":"stage easy -> hard","metadata":{"from":"easy","to":"hard"}}
{"ts_ms":1790189597025,"step":4,"kind":"metric","name":"train/loss","value":2.7658,"metadata":{"split":"train","stage":"hard"}}
{"ts_ms":1790189597025,"step":4,"kind":"metric","name":"rate/tokens","value":14530.17,"metadata":{"split":"train","stage":"hard"}}
{"ts_ms":1790189597055,"step":8,"kind":"annotation","annotation":"checkpoint /tmp/ckpt/step-00000008","metadata":{"checkpoint":"/tmp/ckpt/step-00000008","seconds":0.041}}
{"ts_ms":1790189597074,"kind":"done","slug":"my-run"}
A consumer that wants numbers filters on "kind":"metric"; one that wants
one stage filters on metadata.stage too. Metric names are in
metrics.
Annotations
Only the runtime writes annotations; an experiment cannot.
| text | metadata | when |
|---|---|---|
lowering: <device> and the rest of the block | device | a GPU run before its first step, at the step it starts from, so a resume on another machine records that machine's; the lowering report without the kernel list. The Metrics API takes 16,000 UTF-16 units: a longer block drops graph lines from the end, says ... <n> more lines and keeps the lines after the graphs; the log and explain --device cuda always have the whole block |
memory: ..., stacking ... | the plan's parts, the measured degrees | the memory plan and stacking choice before the first step; memory: allocation failed at step <n>; ... when an allocation failed and was retried; memory: peak ... of ... planned when the run ends |
stage <old> -> <new> | from, to | the curriculum moved the training loader |
checkpoint <dir> | checkpoint, seconds | a checkpoint was written; seconds runs from the first tensor leaving the device to the last byte stored, what a spot VM's notice has to hold |
resumed from <dir>, then one line per source or knob that moved | checkpoint, changed | the run continued from a checkpoint, at that step; changed holds the lines after the first, source <path> <old> -> <new>, knob <name> <old> -> <new>, or ... is new / ... is gone, and is empty for the same files under the same selection |
interrupted by <signal> at step <n> | signal | a signal stopped the run after step n; a checkpoint annotation at the same step follows when there is a location |
failed: <error> | error | the run stopped; the last line before done |
The open metadata
The open event's metadata is the whole identity of the run, the same
object in a file, a pipe and Metrics, where it is the experiment's meta.
| key | what it holds |
|---|---|
sexpgpu | the frontend version that compiled the IR |
entry | the experiment file, as the manifest records it |
variant | the selected variant, or null |
knobs | every knob with its bound value, --steps, --eval-every and --diagnostics overrides included |
run | steps, microbatches, evals (one {name, every, batches} per pass), checkpoint_every, precision, seed, diagnostics, diagnose_every, diagnose_first |
sources | one {path, sha256} per file of the manifest: the entry, what it requires, the prelude <core> and the standard modules. Paths and hashes, never contents |
device | cpu or cuda |
devices | CUDA ordinals in rank order, or [] |
The description is <slug>: <steps> steps of <entry>. If the metadata
passes 64 KiB of compact JSON, the most Metrics accepts, sources
becomes "<n> files, too large to record" and the rest survives.
A resumed run is the same experiment
The slug is the experiment. Metrics keeps one
experiment per slug, and an open on a slug it has returns that
experiment, so a resume reports into the same URL.
- A resume opens with the same description and metadata. The resume itself
is the
resumed from <dir>annotation. When the metadata does differ (another--steps), the newestopenreplaces the stored one. - Every
openis followed by a start that marks the experimentrunning, even after adone. A run stopped or killed sends nodoneand staysrunninguntil a resume of it finishes. - A run stopped by a signal with a checkpoint location repeats nothing. A run killed outright repeats the steps after its last checkpoint, and Metrics keeps both copies of those points; the annotation marks where the second process begins.
- A JSONL file a run resumes into holds one
openline per process, identical but forts_ms.
Related: metrics, checkpoints.