# The status file

Where a live run is, as one JSON file, for a poller: an agent, a script
over ssh, a dashboard. It is rewritten after every step and every
evaluation and once more when the run ends or dies. It arrives by temporary
file and rename, so a reader sees a whole document or none. A failure to
write it is printed and never ends a run.

It goes where `SEXPGPU_STATUS`, or `--status`, says; without either there
is none. `sweep` gives each point its own, `status-0.json`,
`status-1.json`; see [sweep](https://sexpgpu.041.io/docs/sweep.md).

```console
$ sexpgpu run my-run.sx --status /tmp/my-run.json
$ jq -c . /tmp/my-run.json
{"slug":"my-run","entry":"my-run.sx","selection":"default","device":"cpu","state":"done","error":null,"step":20,"steps":20,"records":160,"counters":{"tokens":2720},"train_loss":2.620679,"eval":[["eval/loss",2.462187]],"wall_seconds":0.214606417,"seconds_per_step":0.007608008449999999,"remaining_seconds":0.0,"checkpoint":null,"resumed_from":null,"metrics":null,"updated_ms":1789925843464}
```

| key | type | meaning |
|---|---|---|
| `slug` | string | the [run slug](https://sexpgpu.041.io/docs/run.md#the-run-slug) |
| `entry` | string | the experiment file, from the manifest |
| `selection` | string | everything off the file's defaults, `variant=small lr=0.003`, or `default` |
| `device` | string | `cpu` or `cuda` |
| `state` | string | `running`, `done`, `interrupted` or `failed` |
| `error` | string or null | what ended the run, when something did |
| `step` | number | optimizer steps completed |
| `steps` | number | the run's total |
| `records` | number | loader records consumed |
| `counters` | object | every declared counter's running total |
| `train_loss` | number or null | the last training loss, `train/loss` when reported |
| `eval` | array of `[name, value]` | the last observations of the pass named `eval`, or of the first pass, diagnostics aside |
| `wall_seconds` | number | since the run started |
| `seconds_per_step` | number or null | the mean of the last twenty steps |
| `remaining_seconds` | number or null | that mean times the steps left |
| `checkpoint` | string or null | the last checkpoint this run wrote |
| `resumed_from` | string or null | the checkpoint this process continued from; `null` for a fresh start, including a `--resume latest` that found none |
| `metrics` | string or null | the dashboard URL or the JSONL path |
| `updated_ms` | number | when this document was written |

Read it with `jq .state`, a `cat` over ssh, or any JSON reader. A run
killed outright (`SIGKILL`, the machine gone) writes nothing more, so its
file keeps the last document and still reads `running`; compare
`updated_ms` with the clock. What a signal leaves is in
[checkpoints](https://sexpgpu.041.io/docs/checkpoints.md#signals).

Poll this file rather than parsing the log: log lines are for humans and
their format is not stable.

## The summary

What `run` prints when it ends: the `done:` block on standard error, or one
JSON object of kind `summary` on standard output with `--json`.

| field | meaning |
|---|---|
| `steps` | steps completed |
| `train_loss` | the last training loss |
| `eval` | the last observations of the pass named `eval` (or the first pass), diagnostics aside, as an object |
| `counters` | `records` plus every declared counter |
| `stage` | the curriculum stage the run ended in |
| `metrics` | the dashboard URL or the JSONL path |
| `status` | the status file's path, or null |

Related: [metrics](https://sexpgpu.041.io/docs/metrics.md), [run](https://sexpgpu.041.io/docs/run.md), [working as an agent](https://sexpgpu.041.io/docs/agents.md).

---

SexpGPU documentation. Every page: https://sexpgpu.041.io/llms.txt
