S-exp GPU
All pages
Docs · The command lineMarkdown

explain, diff, ir and eval

The verbs that show what the compiler made of a file. None of them trains or reads data; only explain --device cuda opens a device.

explain

The target-independent account of one file, on standard output: the selection with each knob's source, the parameter tree with the update group of every leaf, the optimizer groups with their states and the constants their schedules folded to, every graph with its ops by family and its ten largest tensors with the file:line that made them, one training step, the diagnostics, and where the precision policy keeps f32. About eighty lines for a small experiment. --json is the same report as a document. What a device did is a hole labelled device.

This is how to check a run against its plan: the knob values and where each came from, the parameter count, which parameters each optimizer group got, tokens per step, the evaluation cadence and the memory bound.

$ sexpgpu explain my-run.sx --variant smoke
my-run.sx: 33 parameters, 300 steps, precision f32  (sexpgpu-frontend 0.1.0)

selection:
  variant  smoke
  lr       0.0015      default
  layers   2           variant
  width    128         variant
  seq      64          variant
  seed     1           default

parameters: 33, 13.3M elements, 50.8 MiB as f32 master weights
  model                  13.3M
    embed.table          [50304, 128]   f32    6.44M  model.embed.*
    ...
  groups:
    model.embed.*       1 param   6.44M  states m, v  12.9M  constants 0, 1e-10, 0.05, 0.2, 0.3, ...
    default            12 params   393k  states m, v   786k  constants 1e-10, 0.0015, 0.05, 0.1, ...

graphs:
  train  439 nodes, 0 casts added by the f32 policy
    ops      input 34, const 54, unary 28, binary 119, compare 4, reduce 17, shape 156, ...
    emits    metric train/loss, counter tokens
    largest tensors:
       25.8M  [4, 128, 50304]      f32   broadcast    lib/model.sx:25
  ...

training step:
  records          4 microbatches of 4 = 16 per step, 300 steps
  counter tokens   1040 per step (static)
  evaluation       eval: every 25 steps, the whole loader in batches of 4
  checkpoints      none
  precision        f32, seed 1
  peak activation  <= 915.5 MiB for one microbatch, an upper bound: ...

The diagnostics section lists every diagnostic the file and its libraries define, its reducer, its graphs, and whether this selection turns it on; see diagnostics.

explain --device cuda

With --device cuda, or SEXPGPU_DEVICE=cuda, on a machine with a GPU and the CUDA binary, the report ends with the lowering report of the graphs a run would prepare instead of the device hole, and --json puts the same numbers under device. It opens the device and builds each graph's plan, which takes the seconds a run's first step spends on it; it reads no data and writes nothing. SEXPGPU_DEVICES picks the device and the rank count the stacking line assumes, as for run.

The lowering report

What the CUDA executor did with each graph a step runs, in one block. run prints it on standard error before its first step and sends the same text as a lowering annotation (events); explain --device cuda prints it on standard output. The interpreter lowers nothing, so a CPU run has none.

$ sexpgpu explain my-run.sx --device cuda
...
lowering: cuda compute_80 NVIDIA A100-SXM4-80GB, patterns on
  train                     1259 nodes, 302 kernels: 49 fused groups of 2 to 15 nodes (median 5), 66 GEMMs
    row programs       row 39, softmax backward 4   model.norm1, model.blocks.{0..3}.attn, model.blocks.{0..3}.norm2, 2 more
    pointwise programs pointwise                    model.blocks.0.norm1
    cross-entropy tail forward, backward            model.proj
                       as gather of log-probabilities, negated per row
  train stacked             1949 nodes, 367 kernels: 112 fused groups of 2 to 15 nodes (median 11), 66 GEMMs
  ...
  update default            24 graphs: 1536 nodes, 120 kernels: 96 fused groups of 2 to 9 nodes (median 2)
    pointwise programs pointwise 24
  stacking                  4, 2 or 1: chosen when the run opens its device
  • The first line names the device and whether the pattern kernels are on.
  • One row per graph: train, train accumulating (what later microbatches run, adding into the previous gradients), train stacked, train diagnosed, each evaluation pass by name, and one row per update group summing its graphs. A row gives the nodes, the kernels one call launches, the fused groups and their sizes, the GEMMs, and captured when the graph is large enough to be submitted as captured segments after its tenth call.
  • Under a row, one line per pattern family that fired, with its kernels, how often each fired, and the model paths of the regions, model.blocks.{0..11}.attn for twelve blocks. A fell back line names a region a family matched but left to the generic lowering, and why: a shape, dtype or head dimension the kernel does not take, or an observation that reads a value inside the region. A family that recognizes its math however it is written adds an as line for each form it found: the vocabulary tail says whether the loss was a gather or a one-hot sum, of log-probabilities or of shifted ones, negated per row or after the reduction.
  • The last lines are about the run as a whole: memory and stacking (devices), then whatever the pattern kernels cannot do on this device. explain has no run, so its stacking line lists the degrees a run would choose among.

SEXPGPU_EXPLAIN=kernels adds every launch of each row's first graph, in execution order, with its shape, dtype and model path; the annotation never carries that list.

diff

sexpgpu diff <a.sx> <b.sx> compiles two files, or one file twice under two selections, and prints only what differs, by section, with an a and a b column: knobs, parameters by path with their update group, graph node counts by op family and largest tensors, the training step, and each optimizer group's members, states and constants. A key only one side has reads -. Sections over twenty rows print twenty and a count; --json prints the whole structured diff. Identical inputs print identical. The selection flags apply to b.

$ sexpgpu diff my-run.sx my-run.sx --set lr=0.25
a  my-run.sx
b  my-run.sx --set lr=0.25

knobs
  lr  0.5 (default)  0.25 (set)

optimizer groups
  default constants  0, 0.5  0, 0.25

Any row you did not expect is a change you did not mean to make.

ir

sexpgpu ir <file.sx> writes the whole experiment document as pretty JSON on standard output: parameters, graphs, loaders, optimizer updates, the run configuration and the manifest. It is large, hundreds of megabytes for a full-size model. A checkpoint does not hold it; see checkpoints.

eval

sexpgpu eval <file.sx> evaluates every top-level form and prints each value, one per line; a defining form prints name = value. Several values print as #<values ...>. It is the closest thing to a REPL: put (macroexpand-1 '(my-macro x)) in a file to see an expansion.

$ sexpgpu eval my-run.sx
nil
nil
seq = 16
defknobs
lr-scan
train-loader = #<loader 1 fields, batch 4>
eval
prepare = #<function prepare>
model = #<model gpt as model>
objective = #<function objective>
optimizer = #<optimizer sgd>
evaluate = #<function evaluate>
defrun = #<run 20 steps>

Related: check, knobs, working as an agent.