One file is the experiment
Data, model, objective, optimizer and run, in one .sx file under 120 lines. Nothing else to read, nothing else to get wrong.
041 · SEXPGPU
Agents do the research. The compiler keeps it honest.
ML research at the speed of thought.The world's first ML framework designed for agents.
An experiment is one file. The compiler owns everything else.
ML research code is the same training script written again and again, slightly wrong each time. The loop, gradient accumulation, evaluation, checkpoints and mixed precision are rewritten for every idea, and every rewrite is a place for a silent bug.
SexpGPU states the experiment and nothing else: the data, the model, the objective, the optimizer and the run. Agents iterate against a compiler that answers in milliseconds, instead of a GPU queue that answers in hours.
Everything the same in every experiment, done once.
Data, model, objective, optimizer and run, in one .sx file under 120 lines. Nothing else to read, nothing else to get wrong.
Training loop, autodiff, gradient accumulation, evaluation cadence, checkpoints, precision and kernels. None of it is in your file, so none of it is subtly wrong in your file.
Every shape is checked at compile time, on a laptop, with no GPU and no data. A diagnostic has a code, your line, expected and actual, and the fix.
Closures, backquote macros, gensyms, lambda lists, multiple values, all at compile time. A transformer block reads like the maths, and a macro can write the ablation.
Correctness means matching published loss curves within the spread of their reference trials, never the implementation agreeing with itself.
A CUDA 13 driver and nothing else. Bundle a run into one directory, copy it to any GPU box, and --resume latest survives preemption.
As fast as PyTorch.*
The CUDA device compiles the whole step from the IR into fused generated kernels, cuBLAS and cuDNN attention, with no Python anywhere in the loop. On nine of the eleven models we measure it takes less time per step than the same model under torch.compile, or the same; the other two are 4% and 5% slower.
Same model, same data, same A100. Lower is faster: a shorter bar takes less time per step. The line is torch.compile.
| model | SexpGPU s/step | torch.compile s/step | step time |
|---|---|---|---|
| Tiny hyperbolic (Lorentz) GPT, f32 | 0.0264 | 0.04425 | 0.60× |
| Tiny Mamba, f32 | 0.2227 | 0.2696 | 0.83× |
| Tiny GPT, AdamW, bf16 | 0.0274 | 0.0329 | 0.83× |
| Tiny RetNet, f32 | 0.1328 | 0.14865 | 0.89× |
| Tiny nGPT, bf16 | 0.03745 | 0.04155 | 0.90× |
| Tiny GPT, Muon, f32 | 0.1388 | 0.15345 | 0.90× |
| GPT-2 small, 124M, bf16 | 2.7883 | 3.0002 | 0.93× |
| Geometric transformer, Riemannian Adam, f32, 512 sequences a step | 62.674 | 65.239 | 0.96× |
| Tiny GPT, AdamW, f32 | 0.124 | 0.12475 | 0.99× |
| Tiny GPT, 8 experts, bf16, top-2 | 0.03675 | 0.03535 | 1.04× |
| ResNet-9, CIFAR-10, bf16 | 0.0141 | 0.0134 | 1.05× |
* Limited benchmark. Eleven training runs on one NVIDIA A100 SXM4 40 GB, each timed against the same model under torch.compile in the same job, on the same data. Step time is the mean of two run medians over steady-state steps, measured 28 September 2026; an 80 GB A100 gives the same picture within a few percent. The mixture of experts and ResNet-9 are 4% and 5% slower than torch.compile today. Every shape not listed here is unmeasured. We are actively working to be on par in every case, and treat any shape where we are slower as a bug.
Check, explain, run, bundle. Answers a program can read.
Every verb answers in a form a program reads: --json where there is structure, a stable exit code always, and a status file while a run trains. No prompts, no dashboards to scrape, no log lines to grep. These four carry an experiment from file to GPU;every verb →
Evaluates the file, traces every graph and checks every shape. No GPU, no data. --json gives agents an array of diagnostics with code, span, notes and fix; exit code 1 means fix it.
$ sexpgpu check width-grid.sx ok: 59 parameters, 119 graphs, 100 steps $ sexpgpu check bad-shape.sx error[E-DIM-001]: matmul: inner dims differ: [2, 16] x [32, 8] --> bad-shape.sx:7:15 | 7 | (lambda (x) (matmul x weight))) | ^^^^^^^^^^^^^^^^^ note: operands: `x` [2, 16] f32, `weight` [32, 8] f32 note: applying model `model` in: model -> matmul
What the compiler made of the file: every knob and where its value came from, every parameter and its optimizer group, the constants each schedule folded to, one training step and the memory bound. diff shows what changed between two runs, and nothing else.
$ sexpgpu explain width-grid.sx --variant width-512 width-grid.sx: 59 parameters, 100 steps, precision bf16 selection: variant width-512 width 512 variant layers 4 default lr 0.003 default groups: and model.blocks.** :rank-at-least 2 24 params 12.6M states momentum-buffer default 35 params 55.8k states m, v training step: records 2 microbatches of 4 = 8 per step, 100 steps precision bf16, seed 1 peak activation <= 197.2 MiB for one microbatch
One process trains one file on the CPU interpreter or the CUDA device. Numbers stream to JSONL or Metrics by 041, a status file says where the run is, and --json prints the summary as one object.
$ sexpgpu run tiny-sgd.sx --status status.json run: tiny-sgd.sx slug tiny-sgd device cpu steps 20 microbatches 2 ... step 0/20 loss - | eval/loss 3.46573 | 0 records | 0.0s step 10/20 loss 2.66862 | eval/loss 2.62369 | 80 records | 0.1s step 20/20 loss 2.62068 | eval/loss 2.46219 | 160 records | 0.2s done: 20 steps, train/loss 2.6207 eval/loss 2.4622 counter tokens 2720 $ jq -c '{state, step, train_loss}' status.json {"state":"done","step":20,"train_loss":2.620679}
One directory with the binary, every source and the local data, and a run.sh with the selection baked in. Copy it to any GPU box with a CUDA driver. Give it a checkpoint location and --resume latest, and a preempted run continues where it stopped.
$ sexpgpu bundle tiny-lorentz.sx -o out/lorentz --binary ./sexpgpu-linux-cuda wrote out/lorentz: 3 source files and 2 data files, 10.8 MiB run it: out/lorentz/run.sh [run flags], with SEXPGPU_* in the environment $ scp -r out/lorentz gpu-box: $ ssh gpu-box 'SEXPGPU_DEVICE=cuda SEXPGPU_CHECKPOINT_DIR=s3://my-bucket/runs/lorentz \ lorentz/run.sh --resume latest'
Six names. That is a whole experiment.
(require "sexpgpu/nn" cross-entropy gpt)
(require "sexpgpu/optim" sgd)
(defvar train-loader
(loader :sources [(files ["data/tiny-train.jsonl.gz"] :encoding "jsonl.gz")]
:fields [(field :tokens :from "input_ids" :dtype :i32 :shape [17])]
:batch-size 4
:infinite true))
(defun prepare (batch)
(let ((tokens (field batch :tokens)))
(list :inputs (slice tokens 1 0 16) :targets (slice tokens 1 1 17))))
(defvar model (gpt :vocab 32 :layers 2 :dim 32 :heads 2))
(defun objective (prediction targets) (cross-entropy prediction targets))
(defvar optimizer (sgd :lr 0.5))
(defrun :steps 20 :microbatches 2 :precision :f32 :seed 1)train-loaderwhere training records come frompreparea batch into inputs and targetsmodela closure over its parametersobjectiveone scalar to minimizeoptimizerconfigured with plain valuesdefrunsteps, microbatches, precision, seedA real Lisp. The experiment writes itself.
Macros and closures run at compile time and produce the experiment: one macro declares a whole ablation grid as named variants, repeat unrolls a tower of blocks, and selectors hand block matrices to Muon and everything else to AdamW. The compiler checks what comes out, so a macro can never produce a run that silently does something else.
;;; One macro writes the ablation grid; one closure writes the model.
(require "sexpgpu/nn" block cross-entropy embedding linear rmsnorm)
(require "sexpgpu/optim" adamw muon warmup-stable-decay)
;; Each value becomes a named variant: width-128, width-256, width-512.
(defmacro defgrid (knob &rest values)
`(progn
,@(map (lambda (v)
`(defvariant ,(intern (format "~a-~a" knob v)) (,knob ,v)))
values)))
(defgrid width 128 256 512)
(defknobs
(width 256 "model width")
(layers 4 "transformer blocks")
(lr 0.003 "peak learning rate"))
(defvar train-loader
(loader :sources [(files ["data/tiny-train.jsonl.gz"] :encoding "jsonl.gz")]
:fields [(field :tokens :from "input_ids" :dtype :i32 :shape [17])]
:batch-size 4
:infinite true))
(defun prepare (batch)
(let ((tokens (field batch :tokens)))
(list :inputs (slice tokens 1 0 16) :targets (slice tokens 1 1 17))))
;; A model is a closure over its parameters; `repeat` unrolls the stack.
(defmodel tower (&key vocab dim layers)
(let ((embed (embedding vocab dim))
(blocks (repeat layers (lambda (i) (block dim 4 :layer i))))
(norm (rmsnorm dim))
(head (linear dim vocab :bias false)))
(lambda (ids)
(head (norm (reduce (lambda (x b) (b x)) (embed ids) blocks))))))
(defvar model (tower :vocab 32 :dim width :layers layers))
(defun objective (prediction targets) (cross-entropy prediction targets))
;; Block matrices to Muon, everything else to AdamW, chosen by path and rank.
(defvar optimizer
(adamw :lr (warmup-stable-decay lr :warmup-steps 10)
:groups [(group (select-and (select "model.blocks.**")
(select :rank-at-least 2))
:optimizer (muon :lr 0.02))]))
(defrun :steps 100 :microbatches 2 :precision :bf16 :seed 1)$ sexpgpu variants width-grid.sx knobs: width 256 model width layers 4 transformer blocks lr 0.003 peak learning rate variants: width-128 width=128 width-256 width=256 width-512 width=512
Each variant is one run with its own slug, so --variant width-512 lands in its own metrics experiment with every knob recorded. How macros work →
How agents do accurate science with it.
Every quantity an experiment may vary is a declared knob, every configuration a named variant, every grid a sweep. Nothing is varied by editing code.
An agent iterates against the compiler, not against a GPU queue. Diagnostics are JSON with a code and a fix, so the next edit is the right one.
explain states the experiment as the compiler understood it. The agent compares it with the plan: parameter counts, optimizer groups, tokens per step, memory.
diff between two runs prints only what differs, by section. An extra row is a mistake caught before it becomes a result.
One selection is one run slug. The metrics carry every source file's SHA-256 and every knob's source, so a curve always names the text that made it.
Diagnostics written into models and optimizers cost nothing until selected, and never change a training number when they are.
This site speaks markdown.
Every page answers with markdown when asked for Accept: text/markdown or fetched with curl. The docs are short on purpose, one question per page and dense, so an agent reads what it needs without flooding its context. Working as an agent →
$ curl https://sexpgpu.041.io/ $ curl https://sexpgpu.041.io/docs/agents/ $ curl -H 'Accept: text/markdown' https://sexpgpu.041.io/docs/lisp/ $ curl https://sexpgpu.041.io/llms.txt $ curl https://sexpgpu.041.io/llms-full.txt
Want to join early? niko@041.io