S-exp GPU

041 · SEXPGPU

S-exp GPU

Agents do the research. The compiler keeps it honest.

ML research at the speed of thought.The world's first ML framework designed for agents.

45 ms
to check a full GPT-2 124M run
0.93×
the step time of torch.compile on GPT-2 124M*
0
lines of training loop in your experiment

Why · One file

An experiment is one file. The compiler owns everything else.

ML research code is the same training script written again and again, slightly wrong each time. The loop, gradient accumulation, evaluation, checkpoints and mixed precision are rewritten for every idea, and every rewrite is a place for a silent bug.

SexpGPU states the experiment and nothing else: the data, the model, the objective, the optimizer and the run. Agents iterate against a compiler that answers in milliseconds, instead of a GPU queue that answers in hours.

The good parts

Everything the same in every experiment, done once.

01

One file is the experiment

Data, model, objective, optimizer and run, in one .sx file under 120 lines. Nothing else to read, nothing else to get wrong.

02

The compiler owns the loop

Training loop, autodiff, gradient accumulation, evaluation cadence, checkpoints, precision and kernels. None of it is in your file, so none of it is subtly wrong in your file.

03

Errors at your line, before the GPU

Every shape is checked at compile time, on a laptop, with no GPU and no data. A diagnostic has a code, your line, expected and actual, and the fix.

04

Basically Common Lisp

Closures, backquote macros, gensyms, lambda lists, multiple values, all at compile time. A transformer block reads like the maths, and a macro can write the ablation.

05

Correct by published curves

Correctness means matching published loss curves within the spread of their reference trials, never the implementation agreeing with itself.

06

One binary, no Python

A CUDA 13 driver and nothing else. Bundle a run into one directory, copy it to any GPU box, and --resume latest survives preemption.

Performance

As fast as PyTorch.*

The CUDA device compiles the whole step from the IR into fused generated kernels, cuBLAS and cuDNN attention, with no Python anywhere in the loop. On nine of the eleven models we measure it takes less time per step than the same model under torch.compile, or the same; the other two are 4% and 5% slower.

Training step time against torch.compile

Same model, same data, same A100. Lower is faster: a shorter bar takes less time per step. The line is torch.compile.

Tiny hyperbolic (Lorentz) GPTf32
Tiny Mambaf32
Tiny GPT, AdamWbf16
Tiny RetNetf32
Tiny nGPTbf16
Tiny GPT, Muonf32
GPT-2 small, 124Mbf16
Geometric transformer, Riemannian Adamf32, 512 sequences a step
Tiny GPT, AdamWf32
Tiny GPT, 8 expertsbf16, top-2
ResNet-9, CIFAR-10bf16
Show the numbers
modelSexpGPU s/steptorch.compile s/stepstep time
Tiny hyperbolic (Lorentz) GPT, f320.02640.044250.60×
Tiny Mamba, f320.22270.26960.83×
Tiny GPT, AdamW, bf160.02740.03290.83×
Tiny RetNet, f320.13280.148650.89×
Tiny nGPT, bf160.037450.041550.90×
Tiny GPT, Muon, f320.13880.153450.90×
GPT-2 small, 124M, bf162.78833.00020.93×
Geometric transformer, Riemannian Adam, f32, 512 sequences a step62.67465.2390.96×
Tiny GPT, AdamW, f320.1240.124750.99×
Tiny GPT, 8 experts, bf16, top-20.036750.035351.04×
ResNet-9, CIFAR-10, bf160.01410.01341.05×

* Limited benchmark. Eleven training runs on one NVIDIA A100 SXM4 40 GB, each timed against the same model under torch.compile in the same job, on the same data. Step time is the mean of two run medians over steady-state steps, measured 28 September 2026; an 80 GB A100 gives the same picture within a few percent. The mixture of experts and ResNet-9 are 4% and 5% slower than torch.compile today. Every shape not listed here is unmeasured. We are actively working to be on par in every case, and treat any shape where we are slower as a bug.

Command line · Built for agents

Check, explain, run, bundle. Answers a program can read.

Every verb answers in a form a program reads: --json where there is structure, a stable exit code always, and a status file while a run trains. No prompts, no dashboards to scrape, no log lines to grep. These four carry an experiment from file to GPU;every verb →

sexpgpu check

Verify it compiles, in milliseconds

Evaluates the file, traces every graph and checks every shape. No GPU, no data. --json gives agents an array of diagnostics with code, span, notes and fix; exit code 1 means fix it.

zsh
$ sexpgpu check width-grid.sx
ok: 59 parameters, 119 graphs, 100 steps

$ sexpgpu check bad-shape.sx
error[E-DIM-001]: matmul: inner dims differ: [2, 16] x [32, 8]
  --> bad-shape.sx:7:15
   |
 7 |   (lambda (x) (matmul x weight)))
   |               ^^^^^^^^^^^^^^^^^
note: operands: `x` [2, 16] f32, `weight` [32, 8] f32
note: applying model `model`
in: model -> matmul
sexpgpu explain

Verify the science against the plan

What the compiler made of the file: every knob and where its value came from, every parameter and its optimizer group, the constants each schedule folded to, one training step and the memory bound. diff shows what changed between two runs, and nothing else.

zsh
$ sexpgpu explain width-grid.sx --variant width-512
width-grid.sx: 59 parameters, 100 steps, precision bf16

selection:
  variant  width-512
  width    512         variant
  layers   4           default
  lr       0.003       default

  groups:
    and model.blocks.** :rank-at-least 2  24 params  12.6M  states momentum-buffer
    default                               35 params  55.8k  states m, v

training step:
  records          2 microbatches of 4 = 8 per step, 100 steps
  precision        bf16, seed 1
  peak activation  <= 197.2 MiB for one microbatch
sexpgpu run

Run it, right here

One process trains one file on the CPU interpreter or the CUDA device. Numbers stream to JSONL or Metrics by 041, a status file says where the run is, and --json prints the summary as one object.

zsh
$ sexpgpu run tiny-sgd.sx --status status.json
run: tiny-sgd.sx slug tiny-sgd device cpu steps 20 microbatches 2 ...
step 0/20 loss - | eval/loss 3.46573 | 0 records | 0.0s
step 10/20 loss 2.66862 | eval/loss 2.62369 | 80 records | 0.1s
step 20/20 loss 2.62068 | eval/loss 2.46219 | 160 records | 0.2s
done: 20 steps, train/loss 2.6207
  eval/loss 2.4622
  counter tokens 2720

$ jq -c '{state, step, train_loss}' status.json
{"state":"done","step":20,"train_loss":2.620679}
sexpgpu bundle

Pack a remote run

One directory with the binary, every source and the local data, and a run.sh with the selection baked in. Copy it to any GPU box with a CUDA driver. Give it a checkpoint location and --resume latest, and a preempted run continues where it stopped.

zsh
$ sexpgpu bundle tiny-lorentz.sx -o out/lorentz --binary ./sexpgpu-linux-cuda
wrote out/lorentz: 3 source files and 2 data files, 10.8 MiB
  run it: out/lorentz/run.sh [run flags], with SEXPGPU_* in the environment

$ scp -r out/lorentz gpu-box:
$ ssh gpu-box 'SEXPGPU_DEVICE=cuda SEXPGPU_CHECKPOINT_DIR=s3://my-bucket/runs/lorentz \
    lorentz/run.sh --resume latest'

Code · The smallest run

Six names. That is a whole experiment.

MINIMAL.SX
(require "sexpgpu/nn" cross-entropy gpt)
(require "sexpgpu/optim" sgd)

(defvar train-loader
  (loader :sources [(files ["data/tiny-train.jsonl.gz"] :encoding "jsonl.gz")]
          :fields [(field :tokens :from "input_ids" :dtype :i32 :shape [17])]
          :batch-size 4
          :infinite true))

(defun prepare (batch)
  (let ((tokens (field batch :tokens)))
    (list :inputs (slice tokens 1 0 16) :targets (slice tokens 1 1 17))))

(defvar model (gpt :vocab 32 :layers 2 :dim 32 :heads 2))

(defun objective (prediction targets) (cross-entropy prediction targets))

(defvar optimizer (sgd :lr 0.5))

(defrun :steps 20 :microbatches 2 :precision :f32 :seed 1)
  • train-loaderwhere training records come from
  • preparea batch into inputs and targets
  • modela closure over its parameters
  • objectiveone scalar to minimize
  • optimizerconfigured with plain values
  • defrunsteps, microbatches, precision, seed

Code · Macro magic

A real Lisp. The experiment writes itself.

Macros and closures run at compile time and produce the experiment: one macro declares a whole ablation grid as named variants, repeat unrolls a tower of blocks, and selectors hand block matrices to Muon and everything else to AdamW. The compiler checks what comes out, so a macro can never produce a run that silently does something else.

WIDTH-GRID.SX
;;; One macro writes the ablation grid; one closure writes the model.

(require "sexpgpu/nn" block cross-entropy embedding linear rmsnorm)
(require "sexpgpu/optim" adamw muon warmup-stable-decay)

;; Each value becomes a named variant: width-128, width-256, width-512.
(defmacro defgrid (knob &rest values)
  `(progn
     ,@(map (lambda (v)
              `(defvariant ,(intern (format "~a-~a" knob v)) (,knob ,v)))
            values)))

(defgrid width 128 256 512)

(defknobs
  (width 256 "model width")
  (layers 4 "transformer blocks")
  (lr 0.003 "peak learning rate"))

(defvar train-loader
  (loader :sources [(files ["data/tiny-train.jsonl.gz"] :encoding "jsonl.gz")]
          :fields [(field :tokens :from "input_ids" :dtype :i32 :shape [17])]
          :batch-size 4
          :infinite true))

(defun prepare (batch)
  (let ((tokens (field batch :tokens)))
    (list :inputs (slice tokens 1 0 16) :targets (slice tokens 1 1 17))))

;; A model is a closure over its parameters; `repeat` unrolls the stack.
(defmodel tower (&key vocab dim layers)
  (let ((embed (embedding vocab dim))
        (blocks (repeat layers (lambda (i) (block dim 4 :layer i))))
        (norm (rmsnorm dim))
        (head (linear dim vocab :bias false)))
    (lambda (ids)
      (head (norm (reduce (lambda (x b) (b x)) (embed ids) blocks))))))

(defvar model (tower :vocab 32 :dim width :layers layers))

(defun objective (prediction targets) (cross-entropy prediction targets))

;; Block matrices to Muon, everything else to AdamW, chosen by path and rank.
(defvar optimizer
  (adamw :lr (warmup-stable-decay lr :warmup-steps 10)
         :groups [(group (select-and (select "model.blocks.**")
                                     (select :rank-at-least 2))
                         :optimizer (muon :lr 0.02))]))

(defrun :steps 100 :microbatches 2 :precision :bf16 :seed 1)
zsh
$ sexpgpu variants width-grid.sx
knobs:
  width   256    model width
  layers  4      transformer blocks
  lr      0.003  peak learning rate
variants:
  width-128  width=128
  width-256  width=256
  width-512  width=512

Each variant is one run with its own slug, so --variant width-512 lands in its own metrics experiment with every knob recorded. How macros work →

Science · By construction

How agents do accurate science with it.

01

The plan is the file

Every quantity an experiment may vary is a declared knob, every configuration a named variant, every grid a sweep. Nothing is varied by editing code.

02

Check in milliseconds

An agent iterates against the compiler, not against a GPU queue. Diagnostics are JSON with a code and a fix, so the next edit is the right one.

03

Read back what will run

explain states the experiment as the compiler understood it. The agent compares it with the plan: parameter counts, optimizer groups, tokens per step, memory.

04

Change one thing

diff between two runs prints only what differs, by section. An extra row is a mistake caught before it becomes a result.

05

Every number traceable

One selection is one run slug. The metrics carry every source file's SHA-256 and every knob's source, so a curve always names the text that made it.

06

Measure without disturbing

Diagnostics written into models and optimizers cost nothing until selected, and never change a training number when they are.

For agents · This site

This site speaks markdown.

Every page answers with markdown when asked for Accept: text/markdown or fetched with curl. The docs are short on purpose, one question per page and dense, so an agent reads what it needs without flooding its context. Working as an agent →

zsh
$ curl https://sexpgpu.041.io/
$ curl https://sexpgpu.041.io/docs/agents/
$ curl -H 'Accept: text/markdown' https://sexpgpu.041.io/docs/lisp/
$ curl https://sexpgpu.041.io/llms.txt
$ curl https://sexpgpu.041.io/llms-full.txt

Binaries are coming soon.

Want to join early? niko@041.io

Email us →