# run

`sexpgpu run <file.sx>` compiles the file, applies the selection and the
overrides, and runs the training loop in this process. It reads the `.sx`,
everything it requires, the loaders' files (relative to the experiment
file, or `s3://` URIs) and, when resuming, a checkpoint. It writes the
metrics target, the status file and checkpoints, each only when given a
location.

```bash
sexpgpu run my-run.sx --variant smoke --steps 3 --device cpu
SEXPGPU_DEVICE=cuda SEXPGPU_METRICS=jsonl:my-run.jsonl sexpgpu run my-run.sx
```

## Flags

Each is an alias of an environment variable or an override of `defrun`:

| flag | meaning |
|---|---|
| `--steps N` | replaces `defrun :steps` for this session; recorded as knob `steps`, source `override` |
| `--eval-every N` | replaces every pass's `:every` the same way; `0` leaves only the pass after the last step, so a curriculum never runs |
| `--device cpu\|cuda` | `SEXPGPU_DEVICE`; see [devices](https://sexpgpu.041.io/docs/devices.md) |
| `--metrics <target>` | `SEXPGPU_METRICS`; see [metrics](https://sexpgpu.041.io/docs/metrics.md#targets) |
| `--status <path>` | `SEXPGPU_STATUS`; without it no [status file](https://sexpgpu.041.io/docs/status.md) is written |
| `--checkpoint-dir <dir>\|s3://bucket/prefix` | `SEXPGPU_CHECKPOINT_DIR`; see [checkpoints](https://sexpgpu.041.io/docs/checkpoints.md) |
| `--resume <dir>\|s3://...\|latest` | `SEXPGPU_RESUME`; see [checkpoints](https://sexpgpu.041.io/docs/checkpoints.md#resume) |
| `--json` | the final summary as one JSON object on standard output instead of the `done:` block |
| `--variant`, `--set`, `--diagnostics` | the [selection](https://sexpgpu.041.io/docs/cli.md#the-selection) |

Flags win over environment variables; see
[environment](https://sexpgpu.041.io/docs/environment.md#precedence).

## Output

On standard error: one configuration line, where the numbers go, on a GPU
the [lowering report](https://sexpgpu.041.io/docs/explain.md#the-lowering-report), one progress line
per evaluation pass (a pass other than `eval` names itself
first, `probe: eval/loss ...`), a `timing` block, on a GPU the
`memory: peak` line ([devices](https://sexpgpu.041.io/docs/devices.md#memory)), and the `done:` block.
There is no per-step line and no quiet switch; a step's numbers go to the
metrics target, and the status file is the thing to poll.

```console
$ sexpgpu run my-run.sx
run: my-run.sx slug my-run device cpu steps 20 microbatches 2 evals eval@10 diagnostics - precision f32 seed 1 checkpoints - resume - status -
step 0/20 loss - | eval/loss 3.46573 | 0 records | 0.0s
step 10/20 loss 2.66862 | eval/loss 2.62369 | 80 records | 0.1s
step 20/20 loss 2.62068 | eval/loss 2.46219 | 160 records | 0.2s
timing graph 0.012s device-open 0.000s
timing kernels 0 compiled in 0.000s, 0 cache hits
timing step-0 0.007s eval-0 0.020s
timing steps 1..20 median 0.0077s p90 0.0082s
timing eval 2 passes median 0.0195s p90 0.0195s
timing total 0.214s over 20 steps
done: 20 steps, train/loss 2.6207
  eval/loss 2.4622
  counter records 160
  counter tokens 2720
```

The configuration line's `resume` is `-` from the beginning, `<dir>` from a
named checkpoint, `latest=<dir>` from the newest, `latest=none` from the
beginning because there was none yet.

With `--json` the last block is one line on standard output instead; the
fields are in [the summary](https://sexpgpu.041.io/docs/status.md#the-summary):

```json
{"kind":"summary","steps":2,"train_loss":3.347761631011963,"eval":{"eval/loss":3.265638828277588},"counters":{"records":16,"tokens":272},"stage":0,"metrics":"/tmp/my-run.jsonl","status":"/tmp/my-run.status.json"}
```

## The run slug

The slug is `SEXPGPU_RUN_SLUG` or the file's stem, then `-<variant>`, then
one `-<knob>=<value>` per `--set` in knob-name order: `my-run-smoke-lr=0.003`.
One selection always names one run, and
[Metrics by 041](https://metrics.041.io/) keeps one experiment per slug,
so a resume lands in the same experiment. See
[events](https://sexpgpu.041.io/docs/events.md#a-resumed-run-is-the-same-experiment).

## Stopping

`SIGTERM`, `SIGINT` or `SIGHUP` stops the run at the next step boundary,
writes a checkpoint when there is a location, and exits `128 + n`; the
details are in [checkpoints](https://sexpgpu.041.io/docs/checkpoints.md#signals).

Related: [sweep](https://sexpgpu.041.io/docs/sweep.md), [bundle](https://sexpgpu.041.io/docs/bundle.md), [metrics](https://sexpgpu.041.io/docs/metrics.md).

---

SexpGPU documentation. Every page: https://sexpgpu.041.io/llms.txt
