All pages
explain, diff, ir and eval
The verbs that show what the compiler made of a file. None of them trains
or reads data; only explain --device cuda opens a device.
explain
The target-independent account of one file, on standard output: the
selection with each knob's source, the parameter tree with the update group
of every leaf, the optimizer groups with their states and the constants
their schedules folded to, every graph with its ops by family and its ten
largest tensors with the file:line that made them, one training step,
the diagnostics, and where the precision policy keeps f32. About eighty
lines for a small experiment. --json is the same report as a document.
What a device did is a hole labelled device.
This is how to check a run against its plan: the knob values and where each came from, the parameter count, which parameters each optimizer group got, tokens per step, the evaluation cadence and the memory bound.
$ sexpgpu explain my-run.sx --variant smoke
my-run.sx: 33 parameters, 300 steps, precision f32 (sexpgpu-frontend 0.1.0)
selection:
variant smoke
lr 0.0015 default
layers 2 variant
width 128 variant
seq 64 variant
seed 1 default
parameters: 33, 13.3M elements, 50.8 MiB as f32 master weights
model 13.3M
embed.table [50304, 128] f32 6.44M model.embed.*
...
groups:
model.embed.* 1 param 6.44M states m, v 12.9M constants 0, 1e-10, 0.05, 0.2, 0.3, ...
default 12 params 393k states m, v 786k constants 1e-10, 0.0015, 0.05, 0.1, ...
graphs:
train 439 nodes, 0 casts added by the f32 policy
ops input 34, const 54, unary 28, binary 119, compare 4, reduce 17, shape 156, ...
emits metric train/loss, counter tokens
largest tensors:
25.8M [4, 128, 50304] f32 broadcast lib/model.sx:25
...
training step:
records 4 microbatches of 4 = 16 per step, 300 steps
counter tokens 1040 per step (static)
evaluation eval: every 25 steps, the whole loader in batches of 4
checkpoints none
precision f32, seed 1
peak activation <= 915.5 MiB for one microbatch, an upper bound: ...
The diagnostics section lists every diagnostic the file and its libraries define, its reducer, its graphs, and whether this selection turns it on; see diagnostics.
explain --device cuda
With --device cuda, or SEXPGPU_DEVICE=cuda, on a machine with a GPU and
the CUDA binary, the report ends with the
lowering report of the graphs a run would prepare
instead of the device hole, and --json puts the same numbers under
device. It opens the device and builds each graph's plan, which takes the
seconds a run's first step spends on it; it reads no data and writes
nothing. SEXPGPU_DEVICES picks the device and the rank count the
stacking line assumes, as for run.
The lowering report
What the CUDA executor did with each graph a step runs, in one block. run
prints it on standard error before its first step and sends the same text
as a lowering annotation (events);
explain --device cuda prints it on standard output. The interpreter
lowers nothing, so a CPU run has none.
$ sexpgpu explain my-run.sx --device cuda
...
lowering: cuda compute_80 NVIDIA A100-SXM4-80GB, patterns on
train 1259 nodes, 302 kernels: 49 fused groups of 2 to 15 nodes (median 5), 66 GEMMs
row programs row 39, softmax backward 4 model.norm1, model.blocks.{0..3}.attn, model.blocks.{0..3}.norm2, 2 more
pointwise programs pointwise model.blocks.0.norm1
cross-entropy tail forward, backward model.proj
as gather of log-probabilities, negated per row
train stacked 1949 nodes, 367 kernels: 112 fused groups of 2 to 15 nodes (median 11), 66 GEMMs
...
update default 24 graphs: 1536 nodes, 120 kernels: 96 fused groups of 2 to 9 nodes (median 2)
pointwise programs pointwise 24
stacking 4, 2 or 1: chosen when the run opens its device
- The first line names the device and whether the pattern kernels are on.
- One row per graph:
train,train accumulating(what later microbatches run, adding into the previous gradients),train stacked,train diagnosed, each evaluation pass by name, and one row per update group summing its graphs. A row gives the nodes, the kernels one call launches, the fused groups and their sizes, the GEMMs, andcapturedwhen the graph is large enough to be submitted as captured segments after its tenth call. - Under a row, one line per pattern family that fired, with its kernels,
how often each fired, and the model paths of the regions,
model.blocks.{0..11}.attnfor twelve blocks. Afell backline names a region a family matched but left to the generic lowering, and why: a shape, dtype or head dimension the kernel does not take, or an observation that reads a value inside the region. A family that recognizes its math however it is written adds anasline for each form it found: the vocabulary tail says whether the loss was a gather or a one-hot sum, of log-probabilities or of shifted ones, negated per row or after the reduction. - The last lines are about the run as a whole:
memoryandstacking(devices), then whatever the pattern kernels cannot do on this device.explainhas no run, so itsstackingline lists the degrees a run would choose among.
SEXPGPU_EXPLAIN=kernels adds every launch of each row's first graph, in
execution order, with its shape, dtype and model path; the annotation never
carries that list.
diff
sexpgpu diff <a.sx> <b.sx> compiles two files, or one file twice under
two selections, and prints only what differs, by section, with an a and a
b column: knobs, parameters by path with their update group, graph node
counts by op family and largest tensors, the training step, and each
optimizer group's members, states and constants. A key only one side has
reads -. Sections over twenty rows print twenty and a count; --json
prints the whole structured diff. Identical inputs print identical. The
selection flags apply to b.
$ sexpgpu diff my-run.sx my-run.sx --set lr=0.25
a my-run.sx
b my-run.sx --set lr=0.25
knobs
lr 0.5 (default) 0.25 (set)
optimizer groups
default constants 0, 0.5 0, 0.25
Any row you did not expect is a change you did not mean to make.
ir
sexpgpu ir <file.sx> writes the whole experiment document as pretty JSON
on standard output: parameters, graphs, loaders, optimizer updates, the run
configuration and the manifest. It is large, hundreds of megabytes for a
full-size model. A checkpoint does not hold it; see
checkpoints.
eval
sexpgpu eval <file.sx> evaluates every top-level form and prints each
value, one per line; a defining form prints name = value. Several values
print as #<values ...>. It is the closest thing to a REPL: put
(macroexpand-1 '(my-macro x)) in a file to see an expansion.
$ sexpgpu eval my-run.sx
nil
nil
seq = 16
defknobs
lr-scan
train-loader = #<loader 1 fields, batch 4>
eval
prepare = #<function prepare>
model = #<model gpt as model>
objective = #<function objective>
optimizer = #<optimizer sgd>
evaluate = #<function evaluate>
defrun = #<run 20 steps>
Related: check, knobs, working as an agent.