All pages
The status file
Where a live run is, as one JSON file, for a poller: an agent, a script over ssh, a dashboard. It is rewritten after every step and every evaluation and once more when the run ends or dies. It arrives by temporary file and rename, so a reader sees a whole document or none. A failure to write it is printed and never ends a run.
It goes where SEXPGPU_STATUS, or --status, says; without either there
is none. sweep gives each point its own, status-0.json,
status-1.json; see sweep.
$ sexpgpu run my-run.sx --status /tmp/my-run.json
$ jq -c . /tmp/my-run.json
{"slug":"my-run","entry":"my-run.sx","selection":"default","device":"cpu","state":"done","error":null,"step":20,"steps":20,"records":160,"counters":{"tokens":2720},"train_loss":2.620679,"eval":[["eval/loss",2.462187]],"wall_seconds":0.214606417,"seconds_per_step":0.007608008449999999,"remaining_seconds":0.0,"checkpoint":null,"resumed_from":null,"metrics":null,"updated_ms":1789925843464}
| key | type | meaning |
|---|---|---|
slug | string | the run slug |
entry | string | the experiment file, from the manifest |
selection | string | everything off the file's defaults, variant=small lr=0.003, or default |
device | string | cpu or cuda |
state | string | running, done, interrupted or failed |
error | string or null | what ended the run, when something did |
step | number | optimizer steps completed |
steps | number | the run's total |
records | number | loader records consumed |
counters | object | every declared counter's running total |
train_loss | number or null | the last training loss, train/loss when reported |
eval | array of [name, value] | the last observations of the pass named eval, or of the first pass, diagnostics aside |
wall_seconds | number | since the run started |
seconds_per_step | number or null | the mean of the last twenty steps |
remaining_seconds | number or null | that mean times the steps left |
checkpoint | string or null | the last checkpoint this run wrote |
resumed_from | string or null | the checkpoint this process continued from; null for a fresh start, including a --resume latest that found none |
metrics | string or null | the dashboard URL or the JSONL path |
updated_ms | number | when this document was written |
Read it with jq .state, a cat over ssh, or any JSON reader. A run
killed outright (SIGKILL, the machine gone) writes nothing more, so its
file keeps the last document and still reads running; compare
updated_ms with the clock. What a signal leaves is in
checkpoints.
Poll this file rather than parsing the log: log lines are for humans and their format is not stable.
The summary
What run prints when it ends: the done: block on standard error, or one
JSON object of kind summary on standard output with --json.
| field | meaning |
|---|---|
steps | steps completed |
train_loss | the last training loss |
eval | the last observations of the pass named eval (or the first pass), diagnostics aside, as an object |
counters | records plus every declared counter |
stage | the curriculum stage the run ended in |
metrics | the dashboard URL or the JSONL path |
status | the status file's path, or null |
Related: metrics, run, working as an agent.