S-exp GPU
All pages
Docs · Start hereMarkdown

Why SexpGPU

Machine learning research is mostly writing the same training script again and getting it slightly wrong. A Python experiment is a program that happens to contain an experiment: the loop, the gradient accumulation, the evaluation cadence, the checkpointing, the mixed precision and the device placement are rewritten for every idea, and each rewrite is a place for a silent bug. A shape mistake surfaces an hour into a paid GPU. A learning rate schedule off by one step changes the result and nobody sees it. Two runs differ in more than the one thing the researcher meant to change.

An agent doing research makes the same mistakes faster. It needs a feedback loop measured in seconds, errors that say exactly what to fix, and output it can read without parsing a log.

One file is the experiment

A .sx file states the experiment and nothing else: the data it reads, the model, the objective, the optimizer and the run. The compiler owns everything that is the same in every experiment: the training loop, autodiff, gradient accumulation, evaluation cadence, checkpoints, precision lowering and kernel selection. A realistic run is under 120 lines. See the run file.

Errors before the GPU

sexpgpu check compiles the whole experiment, traces every graph and checks every shape in well under a second, on a laptop, with no GPU and no data. Every diagnostic has a code, the line in your file, the expected and actual facts, and a fix when one is known. See check and diagnostic codes.

The plan and the run agree

Everything a run may vary is declared in the file as a knob, a variant or a sweep, and selected before the file is read. sexpgpu explain prints what the compiler made of it (parameters, optimizer groups, one training step, memory), so the plan can be checked against the thing that will run. sexpgpu diff shows what differs between two runs, by section, and nothing else. See knobs and explain.

A real language

The authoring language is Common Lisp, evaluated at compile time: closures, macros, gensyms, multiple values, lambda lists. A transformer block reads like the maths, a macro can write an ablation grid, and a library carries the diagnostics its author knows to look at. No Lisp runs during training. See It is Common Lisp.

Correct by published curves

Correctness is decided by matching published loss curves, never by the implementation agreeing with itself: SexpGPU reproduces published results within the spread of their reference trials. The CUDA device steps as fast as the same model under torch.compile on the same GPU.

One binary

A run is one process reading one file. The binary needs a CUDA 13 driver and nothing else: no Python, no toolchain. sexpgpu bundle packs the binary, the sources and the data into one directory that runs on any GPU box, and --resume latest makes it survive preemption. See bundle.

Next: install, then the quickstart. Agents should read working as an agent.