# defrun

`defrun` is the run configuration the file owns: how many steps, how a step
is split, the precision, the seed, and how often to checkpoint and sample
diagnostics. It is required, once.

```lisp
(defrun :steps 300 :microbatches 4 :checkpoint-every 100 :precision :f32 :seed 1)
```

| keyword | default | meaning |
|---|---|---|
| `:steps` | `1` | optimizer steps; at least 1 |
| `:microbatches` | `1` | microbatches per optimizer step; at least 1. Gradients accumulate over them. Under several GPUs the count must divide by the device count; see [devices](https://sexpgpu.041.io/docs/devices.md#several-gpus) |
| `:checkpoint-every` | none | steps between checkpoints; without a checkpoint location nothing is written. See [checkpoints](https://sexpgpu.041.io/docs/checkpoints.md) |
| `:precision` | `:f32` | `:f32` or `:bf16`; anything else is `E-CONTRACT-003`. See [numerics](https://sexpgpu.041.io/docs/numerics.md) |
| `:seed` | `0` | folded into every random stream key |
| `:diagnostics` | none | globs over diagnostic names, `["attention/*" "logits/*"]`; only these ever compute |
| `:diagnose-every` | `1` | a sampled step is a multiple of this |
| `:diagnose-first` | `0` | the first this many steps are sampled too, and so is the last |
| `:guard-finite` | `false` | stop, committing nothing, when a proposed parameter or state is not finite; see [the finite guard](https://sexpgpu.041.io/docs/checkpoints.md#the-finite-guard) |

## Overrides

| file key | replaced by | recorded as |
|---|---|---|
| `:steps` | `--steps N` | knob `steps`, source `override` |
| every pass's `:every` | `--eval-every N` | knob `eval-every`, source `override` |
| `:diagnostics` | `--diagnostics`, then `SEXPGPU_DIAGNOSTICS` | knob `diagnostics`, source `override` |
| `:seed` | `--set seed=N`, when the file declares no `seed` knob | knob `seed`, source `set` |

Everything else, `:microbatches`, `:checkpoint-every`, `:precision`, is the
file's alone. There is no `defrun` key for the device, the metrics target,
the checkpoint location or the resume; those are the environment's. See
[environment](https://sexpgpu.041.io/docs/environment.md#precedence).

Values can be knobs: `(defrun :steps steps :microbatches microbatches)`.

Step numbering: a training step's number is the update being computed, so
the first training step is step 0, and an evaluation or a checkpoint at
step `n` has `n` completed updates behind it. A training loop that numbers
training metrics by completed updates calls step 0 update 1; evaluation
and checkpoint steps agree.

Related: [diagnostics sampling](https://sexpgpu.041.io/docs/diagnostics.md#when-they-compute),
[run](https://sexpgpu.041.io/docs/run.md).

---

SexpGPU documentation. Every page: https://sexpgpu.041.io/llms.txt
