# The run file

An experiment file is recognized by top-level bindings with conventional
names. Anything else in the file is an ordinary definition and is ignored
by the contract. The file is read once and evaluated once, in order.

## The order of a file

```text
(require ...)          standard modules and library files the file uses
(defvariant ...)       named sets of knob overrides, before defknobs
(defknobs ...)         everything this run may vary
(defsweep ...)         named cross products over knob values
(defvar ...)           ordinary constants
train-loader           required: where the training records come from
(defeval ...)          optional, any number: named evaluation passes
prepare                required: batch -> (list :inputs ... :targets ...)
eval-prepare           optional: defaults to prepare
model                  required: a model value
objective              required: prediction, targets -> one float scalar
evaluate               required when a pass exists: emits metrics
optimizer              required: a configured optimizer
curriculum             optional: stage, ctx, observations -> new stage
(defrun ...)           required: steps, microbatches, cadence, precision, seed
```

| name | page |
|---|---|
| `require` | [modules and require](https://sexpgpu.041.io/docs/require.md) |
| `defvariant`, `defknobs`, `defsweep` | [knobs, variants and sweeps](https://sexpgpu.041.io/docs/knobs.md) |
| `train-loader`, `prepare`, `eval-prepare` | [loaders and prepare](https://sexpgpu.041.io/docs/data.md) |
| `defeval` | [evaluation passes](https://sexpgpu.041.io/docs/evaluation.md) |
| `model` | [models and parameters](https://sexpgpu.041.io/docs/models.md) |
| `objective`, `evaluate` | [objective and evaluate](https://sexpgpu.041.io/docs/objective.md) |
| `optimizer` | [optimizer groups](https://sexpgpu.041.io/docs/optimizer-groups.md), [writing optimizers](https://sexpgpu.041.io/docs/optimizers.md) |
| `curriculum` | [curriculum and ctx](https://sexpgpu.041.io/docs/curriculum.md) |
| `defrun` | [defrun](https://sexpgpu.041.io/docs/defrun.md) |

Missing a required name is `E-CONTRACT-001`; binding one to the wrong kind
of value is `E-CONTRACT-002`. The components are checked independently, so
a mistake in `objective` never hides a missing `optimizer`. The contract
names are looked up in the run file itself, so a library that defines
`model` gives the run none.

## A complete minimal run

Six names, no variation, no evaluation.

```lisp
;;; The smallest file the contract accepts: six names and nothing else.

(require "sexpgpu/nn" cross-entropy gpt)
(require "sexpgpu/optim" sgd)

(defvar train-loader
  (loader :sources [(files ["data/tiny-train.jsonl.gz"] :encoding "jsonl.gz")]
          :fields [(field :tokens :from "input_ids" :dtype :i32 :shape [17])]
          :batch-size 4
          :infinite true))

(defun prepare (batch)
  (let ((tokens (field batch :tokens)))
    (list :inputs (slice tokens 1 0 16) :targets (slice tokens 1 1 17))))

(defvar model (gpt :vocab 32 :layers 2 :dim 32 :heads 2))

(defun objective (prediction targets) (cross-entropy prediction targets))

(defvar optimizer (sgd :lr 0.5))

(defrun :steps 20 :microbatches 2 :precision :f32 :seed 1)
```

```console
$ sexpgpu check minimal.sx
ok: 32 parameters, 65 graphs, 20 steps
$ sexpgpu run minimal.sx --steps 3
...
done: 3 steps, train/loss 3.1645
  counter records 24
```

## A complete realistic run

A realistic run adds knobs and variants (including a `smoke` variant small
enough for `--device cpu`), an evaluation pass, a counter, a reported
per-token loss beside the summed objective, and optimizer groups with a
schedule each. Its selection and its optimizer:

```lisp
(require "sexpgpu/nn" cross-entropy gpt)
(require "sexpgpu/optim" adamw warmup-stable-decay)

(defvariant small (layers 2) (width 128))
(defvariant smoke (layers 2) (width 128) (seq 64))

(defknobs
  (lr 0.0015 "peak learning rate of the block matrices, the default group")
  (layers 4 "transformer blocks")
  (width 256 "model width")
  (seq 1024 "sequence length in tokens; a window holds seq + 1")
  (seed 1 "the run seed; every random stream folds it in"))

(defsweep seeds (seed [1 2 3]))

(defvar warmup 25)

(defun schedule (peak)
  (warmup-stable-decay peak :warmup-steps warmup :decay-fraction 0.7))

(defvar optimizer
  (adamw :lr (schedule lr)
         :betas [0.9 0.95]
         :eps 1e-10
         :weight-decay 0.10
         :groups [(group (select "model.embed.*")
                         :lr (schedule 0.3)
                         :betas [0.8 0.95]
                         :weight-decay 0.0)
                  (group (select "model.head.weight")
                         :lr (schedule (/ 1.0 320.0))
                         :betas [0.8 0.95]
                         :weight-decay 0.0)
                  (group (select :rank-below 2)
                         :lr (schedule 0.01)
                         :betas [0.8 0.95]
                         :weight-decay 0.0)]))
```

[`explain --variant smoke`](https://sexpgpu.041.io/docs/explain.md) reads it back in those terms: the
parameters and their element count, the four update groups (the two named
ones, `:rank-below 2`, and `default` with the remaining block matrices) and
which parameters each holds, each group's `m` and `v` states and the
constants its schedule folded to, and the microbatches of one step.

Move reusable definitions into a [library](https://sexpgpu.041.io/docs/require.md#writing-a-library)
to keep the run file short, and format it with [`sexpgpu fmt`](https://sexpgpu.041.io/docs/new.md#fmt).

---

SexpGPU documentation. Every page: https://sexpgpu.041.io/llms.txt
