# explain, diff, ir and eval

The verbs that show what the compiler made of a file. None of them trains
or reads data; only `explain --device cuda` opens a device.

## explain

The target-independent account of one file, on standard output: the
selection with each knob's source, the parameter tree with the update group
of every leaf, the optimizer groups with their states and the constants
their schedules folded to, every graph with its ops by family and its ten
largest tensors with the `file:line` that made them, one training step,
the diagnostics, and where the precision policy keeps `f32`. About eighty
lines for a small experiment. `--json` is the same report as a document.
What a device did is a hole labelled `device`.

This is how to check a run against its plan: the knob values and where each
came from, the parameter count, which parameters each optimizer group got,
tokens per step, the evaluation cadence and the memory bound.

```console
$ sexpgpu explain my-run.sx --variant smoke
my-run.sx: 33 parameters, 300 steps, precision f32  (sexpgpu-frontend 0.1.0)

selection:
  variant  smoke
  lr       0.0015      default
  layers   2           variant
  width    128         variant
  seq      64          variant
  seed     1           default

parameters: 33, 13.3M elements, 50.8 MiB as f32 master weights
  model                  13.3M
    embed.table          [50304, 128]   f32    6.44M  model.embed.*
    ...
  groups:
    model.embed.*       1 param   6.44M  states m, v  12.9M  constants 0, 1e-10, 0.05, 0.2, 0.3, ...
    default            12 params   393k  states m, v   786k  constants 1e-10, 0.0015, 0.05, 0.1, ...

graphs:
  train  439 nodes, 0 casts added by the f32 policy
    ops      input 34, const 54, unary 28, binary 119, compare 4, reduce 17, shape 156, ...
    emits    metric train/loss, counter tokens
    largest tensors:
       25.8M  [4, 128, 50304]      f32   broadcast    lib/model.sx:25
  ...

training step:
  records          4 microbatches of 4 = 16 per step, 300 steps
  counter tokens   1040 per step (static)
  evaluation       eval: every 25 steps, the whole loader in batches of 4
  checkpoints      none
  precision        f32, seed 1
  peak activation  <= 915.5 MiB for one microbatch, an upper bound: ...
```

The diagnostics section lists every diagnostic the file and its libraries
define, its reducer, its graphs, and whether this selection turns it on;
see [diagnostics](https://sexpgpu.041.io/docs/diagnostics.md#turning-diagnostics-on).

### explain --device cuda

With `--device cuda`, or `SEXPGPU_DEVICE=cuda`, on a machine with a GPU and
the CUDA binary, the report ends with the
[lowering report](#the-lowering-report) of the graphs a run would prepare
instead of the `device` hole, and `--json` puts the same numbers under
`device`. It opens the device and builds each graph's plan, which takes the
seconds a run's first step spends on it; it reads no data and writes
nothing. `SEXPGPU_DEVICES` picks the device and the rank count the
stacking line assumes, as for `run`.

## The lowering report

What the CUDA executor did with each graph a step runs, in one block. `run`
prints it on standard error before its first step and sends the same text
as a `lowering` annotation ([events](https://sexpgpu.041.io/docs/events.md#annotations));
`explain --device cuda` prints it on standard output. The interpreter
lowers nothing, so a CPU run has none.

```console
$ sexpgpu explain my-run.sx --device cuda
...
lowering: cuda compute_80 NVIDIA A100-SXM4-80GB, patterns on
  train                     1259 nodes, 302 kernels: 49 fused groups of 2 to 15 nodes (median 5), 66 GEMMs
    row programs       row 39, softmax backward 4   model.norm1, model.blocks.{0..3}.attn, model.blocks.{0..3}.norm2, 2 more
    pointwise programs pointwise                    model.blocks.0.norm1
    cross-entropy tail forward, backward            model.proj
                       as gather of log-probabilities, negated per row
  train stacked             1949 nodes, 367 kernels: 112 fused groups of 2 to 15 nodes (median 11), 66 GEMMs
  ...
  update default            24 graphs: 1536 nodes, 120 kernels: 96 fused groups of 2 to 9 nodes (median 2)
    pointwise programs pointwise 24
  stacking                  4, 2 or 1: chosen when the run opens its device
```

- **The first line** names the device and whether the pattern kernels are
  on.
- **One row per graph:** `train`, `train accumulating` (what later
  microbatches run, adding into the previous gradients), `train stacked`,
  `train diagnosed`, each evaluation pass by name, and one row per update
  group summing its graphs. A row gives the nodes, the kernels one call
  launches, the fused groups and their sizes, the GEMMs, and `captured` when
  the graph is large enough to be submitted as captured segments after its
  tenth call.
- **Under a row,** one line per pattern family that fired, with its kernels,
  how often each fired, and the model paths of the regions,
  `model.blocks.{0..11}.attn` for twelve blocks. A `fell back` line names a
  region a family matched but left to the generic lowering, and why: a
  shape, dtype or head dimension the kernel does not take, or an
  observation that reads a value inside the region. A family that
  recognizes its math however it is written adds an `as` line for each
  form it found: the vocabulary tail says whether the loss was a gather or
  a one-hot sum, of log-probabilities or of shifted ones, negated per row or
  after the reduction.
- **The last lines** are about the run as a whole: `memory` and `stacking`
  ([devices](https://sexpgpu.041.io/docs/devices.md#memory)), then whatever the pattern kernels cannot
  do on this device. `explain` has no run, so its `stacking` line lists the
  degrees a run would choose among.

`SEXPGPU_EXPLAIN=kernels` adds every launch of each row's first graph, in
execution order, with its shape, dtype and model path; the annotation never
carries that list.

## diff

`sexpgpu diff <a.sx> <b.sx>` compiles two files, or one file twice under
two selections, and prints only what differs, by section, with an `a` and a
`b` column: knobs, parameters by path with their update group, graph node
counts by op family and largest tensors, the training step, and each
optimizer group's members, states and constants. A key only one side has
reads `-`. Sections over twenty rows print twenty and a count; `--json`
prints the whole structured diff. Identical inputs print `identical`. The
selection flags apply to `b`.

```console
$ sexpgpu diff my-run.sx my-run.sx --set lr=0.25
a  my-run.sx
b  my-run.sx --set lr=0.25

knobs
  lr  0.5 (default)  0.25 (set)

optimizer groups
  default constants  0, 0.5  0, 0.25
```

Any row you did not expect is a change you did not mean to make.

## ir

`sexpgpu ir <file.sx>` writes the whole experiment document as pretty JSON
on standard output: parameters, graphs, loaders, optimizer updates, the run
configuration and the manifest. It is large, hundreds of megabytes for a
full-size model. A checkpoint does not hold it; see
[checkpoints](https://sexpgpu.041.io/docs/checkpoints.md#what-a-checkpoint-holds).

## eval

`sexpgpu eval <file.sx>` evaluates every top-level form and prints each
value, one per line; a defining form prints `name = value`. Several values
print as `#<values ...>`. It is the closest thing to a REPL: put
`(macroexpand-1 '(my-macro x))` in a file to see an expansion.

```console
$ sexpgpu eval my-run.sx
nil
nil
seq = 16
defknobs
lr-scan
train-loader = #<loader 1 fields, batch 4>
eval
prepare = #<function prepare>
model = #<model gpt as model>
objective = #<function objective>
optimizer = #<optimizer sgd>
evaluate = #<function evaluate>
defrun = #<run 20 steps>
```

Related: [check](https://sexpgpu.041.io/docs/check.md), [knobs](https://sexpgpu.041.io/docs/knobs.md), [working as an agent](https://sexpgpu.041.io/docs/agents.md).

---

SexpGPU documentation. Every page: https://sexpgpu.041.io/llms.txt
