# S-exp GPU

> ML research at the speed of thought. The world's first ML framework designed for agents.

SexpGPU is a Common Lisp for stating machine learning experiments, a
compiler that turns one experiment file into a checked IR, and a Rust
runtime that trains it on GPUs. Documentation: https://sexpgpu.041.io/docs/index.md.
Every page: https://sexpgpu.041.io/llms.txt.

## The good parts

- **One file is the experiment.** Data, model, objective, optimizer and run, in one `.sx` file under 120 lines. Nothing else to read, nothing else to get wrong.
- **The compiler owns the loop.** Training loop, autodiff, gradient accumulation, evaluation cadence, checkpoints, precision and kernels. None of it is in your file, so none of it is subtly wrong in your file.
- **Errors at your line, before the GPU.** Every shape is checked at compile time, on a laptop, with no GPU and no data. A diagnostic has a code, your line, expected and actual, and the fix.
- **Basically Common Lisp.** Closures, backquote macros, gensyms, lambda lists, multiple values, all at compile time. A transformer block reads like the maths, and a macro can write the ablation.
- **Correct by published curves.** Correctness means matching published loss curves within the spread of their reference trials, never the implementation agreeing with itself.
- **One binary, no Python.** A CUDA 13 driver and nothing else. Bundle a run into one directory, copy it to any GPU box, and `--resume latest` survives preemption.

## As fast as PyTorch*

Step time as a multiple of the same model's under torch.compile; lower is
faster, 1.00× is par.

| model | SexpGPU s/step | torch.compile s/step | step time |
|---|---:|---:|---:|
| Tiny hyperbolic (Lorentz) GPT, f32 | 0.0264 | 0.04425 | 0.60× |
| Tiny Mamba, f32 | 0.2227 | 0.2696 | 0.83× |
| Tiny GPT, AdamW, bf16 | 0.0274 | 0.0329 | 0.83× |
| Tiny RetNet, f32 | 0.1328 | 0.14865 | 0.89× |
| Tiny nGPT, bf16 | 0.03745 | 0.04155 | 0.90× |
| Tiny GPT, Muon, f32 | 0.1388 | 0.15345 | 0.90× |
| GPT-2 small, 124M, bf16 | 2.7883 | 3.0002 | 0.93× |
| Geometric transformer, Riemannian Adam, f32, 512 sequences a step | 62.674 | 65.239 | 0.96× |
| Tiny GPT, AdamW, f32 | 0.124 | 0.12475 | 0.99× |
| Tiny GPT, 8 experts, bf16, top-2 | 0.03675 | 0.03535 | 1.04× |
| ResNet-9, CIFAR-10, bf16 | 0.0141 | 0.0134 | 1.05× |

\* Limited benchmark. Eleven training runs on one NVIDIA A100 SXM4 40 GB, each timed against the same model under torch.compile in the same job, on the same data. Step time is the mean of two run medians over steady-state steps, measured 28 September 2026; an 80 GB A100 gives the same picture within a few percent. The mixture of experts and ResNet-9 are 4% and 5% slower than torch.compile today. Every shape not listed here is unmeasured. We are actively working to be on par in every case, and treat any shape where we are slower as a bug.

## A command line built for agents

These four carry an experiment from file to GPU; every verb, from `diff` to
`doctor`, is in [the command line](https://sexpgpu.041.io/docs/cli.md).

### check: Verify it compiles, in milliseconds

Evaluates the file, traces every graph and checks every shape. No GPU, no data. `--json` gives agents an array of diagnostics with code, span, notes and fix; exit code 1 means fix it.

```console
$ sexpgpu check width-grid.sx
ok: 59 parameters, 119 graphs, 100 steps

$ sexpgpu check bad-shape.sx
error[E-DIM-001]: matmul: inner dims differ: [2, 16] x [32, 8]
  --> bad-shape.sx:7:15
   |
 7 |   (lambda (x) (matmul x weight)))
   |               ^^^^^^^^^^^^^^^^^
note: operands: `x` [2, 16] f32, `weight` [32, 8] f32
note: applying model `model`
in: model -> matmul
```

### explain: Verify the science against the plan

What the compiler made of the file: every knob and where its value came from, every parameter and its optimizer group, the constants each schedule folded to, one training step and the memory bound. `diff` shows what changed between two runs, and nothing else.

```console
$ sexpgpu explain width-grid.sx --variant width-512
width-grid.sx: 59 parameters, 100 steps, precision bf16

selection:
  variant  width-512
  width    512         variant
  layers   4           default
  lr       0.003       default

  groups:
    and model.blocks.** :rank-at-least 2  24 params  12.6M  states momentum-buffer
    default                               35 params  55.8k  states m, v

training step:
  records          2 microbatches of 4 = 8 per step, 100 steps
  precision        bf16, seed 1
  peak activation  <= 197.2 MiB for one microbatch
```

### run: Run it, right here

One process trains one file on the CPU interpreter or the CUDA device. Numbers stream to JSONL or [Metrics by 041](https://metrics.041.io/), a status file says where the run is, and `--json` prints the summary as one object.

```console
$ sexpgpu run tiny-sgd.sx --status status.json
run: tiny-sgd.sx slug tiny-sgd device cpu steps 20 microbatches 2 ...
step 0/20 loss - | eval/loss 3.46573 | 0 records | 0.0s
step 10/20 loss 2.66862 | eval/loss 2.62369 | 80 records | 0.1s
step 20/20 loss 2.62068 | eval/loss 2.46219 | 160 records | 0.2s
done: 20 steps, train/loss 2.6207
  eval/loss 2.4622
  counter tokens 2720

$ jq -c '{state, step, train_loss}' status.json
{"state":"done","step":20,"train_loss":2.620679}
```

### bundle: Pack a remote run

One directory with the binary, every source and the local data, and a `run.sh` with the selection baked in. Copy it to any GPU box with a CUDA driver. Give it a checkpoint location and `--resume latest`, and a preempted run continues where it stopped.

```console
$ sexpgpu bundle tiny-lorentz.sx -o out/lorentz --binary ./sexpgpu-linux-cuda
wrote out/lorentz: 3 source files and 2 data files, 10.8 MiB
  run it: out/lorentz/run.sh [run flags], with SEXPGPU_* in the environment

$ scp -r out/lorentz gpu-box:
$ ssh gpu-box 'SEXPGPU_DEVICE=cuda SEXPGPU_CHECKPOINT_DIR=s3://my-bucket/runs/lorentz \
    lorentz/run.sh --resume latest'
```

## The smallest run

Six names and nothing else:

- **train-loader.** where training records come from
- **prepare.** a batch into inputs and targets
- **model.** a closure over its parameters
- **objective.** one scalar to minimize
- **optimizer.** configured with plain values
- **defrun.** steps, microbatches, precision, seed

```lisp
(require "sexpgpu/nn" cross-entropy gpt)
(require "sexpgpu/optim" sgd)

(defvar train-loader
  (loader :sources [(files ["data/tiny-train.jsonl.gz"] :encoding "jsonl.gz")]
          :fields [(field :tokens :from "input_ids" :dtype :i32 :shape [17])]
          :batch-size 4
          :infinite true))

(defun prepare (batch)
  (let ((tokens (field batch :tokens)))
    (list :inputs (slice tokens 1 0 16) :targets (slice tokens 1 1 17))))

(defvar model (gpt :vocab 32 :layers 2 :dim 32 :heads 2))

(defun objective (prediction targets) (cross-entropy prediction targets))

(defvar optimizer (sgd :lr 0.5))

(defrun :steps 20 :microbatches 2 :precision :f32 :seed 1)
```

## Macro magic

One macro writes the ablation grid as named variants; a closure writes the
model; selectors give block matrices to Muon and the rest to AdamW.

```lisp
;;; One macro writes the ablation grid; one closure writes the model.

(require "sexpgpu/nn" block cross-entropy embedding linear rmsnorm)
(require "sexpgpu/optim" adamw muon warmup-stable-decay)

;; Each value becomes a named variant: width-128, width-256, width-512.
(defmacro defgrid (knob &rest values)
  `(progn
     ,@(map (lambda (v)
              `(defvariant ,(intern (format "~a-~a" knob v)) (,knob ,v)))
            values)))

(defgrid width 128 256 512)

(defknobs
  (width 256 "model width")
  (layers 4 "transformer blocks")
  (lr 0.003 "peak learning rate"))

(defvar train-loader
  (loader :sources [(files ["data/tiny-train.jsonl.gz"] :encoding "jsonl.gz")]
          :fields [(field :tokens :from "input_ids" :dtype :i32 :shape [17])]
          :batch-size 4
          :infinite true))

(defun prepare (batch)
  (let ((tokens (field batch :tokens)))
    (list :inputs (slice tokens 1 0 16) :targets (slice tokens 1 1 17))))

;; A model is a closure over its parameters; `repeat` unrolls the stack.
(defmodel tower (&key vocab dim layers)
  (let ((embed (embedding vocab dim))
        (blocks (repeat layers (lambda (i) (block dim 4 :layer i))))
        (norm (rmsnorm dim))
        (head (linear dim vocab :bias false)))
    (lambda (ids)
      (head (norm (reduce (lambda (x b) (b x)) (embed ids) blocks))))))

(defvar model (tower :vocab 32 :dim width :layers layers))

(defun objective (prediction targets) (cross-entropy prediction targets))

;; Block matrices to Muon, everything else to AdamW, chosen by path and rank.
(defvar optimizer
  (adamw :lr (warmup-stable-decay lr :warmup-steps 10)
         :groups [(group (select-and (select "model.blocks.**")
                                     (select :rank-at-least 2))
                         :optimizer (muon :lr 0.02))]))

(defrun :steps 100 :microbatches 2 :precision :bf16 :seed 1)
```

```console
$ sexpgpu variants width-grid.sx
knobs:
  width   256    model width
  layers  4      transformer blocks
  lr      0.003  peak learning rate
variants:
  width-128  width=128
  width-256  width=256
  width-512  width=512
```

## Accurate science, by construction

- **The plan is the file.** Every quantity an experiment may vary is a declared knob, every configuration a named variant, every grid a sweep. Nothing is varied by editing code.
- **Check in milliseconds.** An agent iterates against the compiler, not against a GPU queue. Diagnostics are JSON with a code and a fix, so the next edit is the right one.
- **Read back what will run.** `explain` states the experiment as the compiler understood it. The agent compares it with the plan: parameter counts, optimizer groups, tokens per step, memory.
- **Change one thing.** `diff` between two runs prints only what differs, by section. An extra row is a mistake caught before it becomes a result.
- **Every number traceable.** One selection is one run slug. The metrics carry every source file's SHA-256 and every knob's source, so a curve always names the text that made it.
- **Measure without disturbing.** Diagnostics written into models and optimizers cost nothing until selected, and never change a training number when they are.

## This site speaks markdown

```console
$ curl https://sexpgpu.041.io/
$ curl https://sexpgpu.041.io/docs/agents/
$ curl -H 'Accept: text/markdown' https://sexpgpu.041.io/docs/lisp/
$ curl https://sexpgpu.041.io/llms.txt
$ curl https://sexpgpu.041.io/llms-full.txt
```

## Install

Public binaries are coming soon. Want to join early? Email niko@041.io.
