S-exp GPU
All pages
Docs · Observing a runMarkdown

Checkpoints and resume

A run checkpoints to a directory or an s3:// prefix, and continues from the newest complete checkpoint with --resume latest. A signal stops it at a step boundary with a checkpoint of that step, so a preempted run loses nothing.

sexpgpu run my-run.sx --checkpoint-dir s3://my-bucket/runs/my-run --resume latest

Checkpoints

  • One directory per step, step-<8 digits>, under the checkpoint location (--checkpoint-dir or SEXPGPU_CHECKPOINT_DIR).
  • Written every defrun :checkpoint-every steps and once more after the last step; without a location nothing is written.
  • Four files (below); the last written is experiment.json, and a checkpoint counts only once it exists, so one cut off by a machine going away is passed over for the one before it.
  • With s3://bucket/prefix the same files go to s3://bucket/prefix/step-<n>/: the two tensor files as concurrent multipart uploads, the others concurrently. Credentials are checked before the first step, so a run never finds out hours in that it cannot write; see S3 credentials.
  • Every checkpoint is an annotation in the metrics stream, with the seconds from the first tensor leaving the device to the last byte stored, and the status file's checkpoint.

Resume

--resume or SEXPGPU_RESUME takes:

valueresumes from
<dir> or s3://bucket/prefix/step-<n>that checkpoint; one that is not there is an error
latestthe highest complete step-<n> under the checkpoint location, or a fresh start when there is none; needs a checkpoint location

A location that cannot be listed is an error, never a fresh start. Put --resume latest on the command line from the first launch; the same command then starts, restarts and continues.

  • Refused: a parameter that is new, gone, or another shape or dtype; a different device or node count (E-DP-002).

  • Allowed, and recorded: edited sources or other knobs. The resumed from annotation names each file and knob that moved, one line each. For example, a run resumed after a comment was added to its file, with --set lr=0.25 --steps 6:

    resumed from /tmp/my-run/step-00000004
    source my-run.sx 91d77bdc76be -> b05c928e631c
    knob lr 0.5 -> 0.25
    knob steps 4 -> 6

    A bundle's sources are its own paths and rewritten files, so resuming a bundled run's checkpoint from the unbundled file lists each of them.

A resumed run reports into the same metrics experiment; see events.

What a checkpoint holds

A checkpoint promises two things: the run continues from it, and its weights load into the same model somewhere else. It holds exactly that:

params.safetensors   one tensor per parameter, keyed by its dotted path
states.safetensors   one tensor per optimizer state, "<param path>/<state>"
state.json           step, stage, records, counters, the loader's position, seed
experiment.json      the source files and their sha256, the variant, every
                     knob with its value and source, each parameter's path,
                     shape, dtype, trainable and tags, each parameter's state
                     names, the counters, precision, seed, device count and
                     the sexpgpu version

experiment.json is a few kilobytes for any model and is written last. The compiled graphs, the lowering and the kernels are not in a checkpoint: a resume compiles them again from the sources and the binary. An exact replay of a compiled artifact is what a bundle is for.

Inspecting one

sexpgpu checkpoint <dir|s3://bucket/prefix/step-<n>> prints what a checkpoint is without loading it: the identity from experiment.json, the step from state.json, and every tensor's name, dtype and shape from the two safetensors headers. An s3:// location is read with the checkpoint role's credentials, reading only the headers. A location without experiment.json is not a complete checkpoint and is an error.

$ sexpgpu checkpoint /tmp/my-run/step-00000004 | head -12
checkpoint /tmp/my-run/step-00000004
step       4
sexpgpu    0.1.0
variant    -
devices    1
precision  f32
seed       7
counters   tokens
sources
  <core>                    e4f1b2681fd7e21aa5cfe45f6f1b42618f70206f8af4becbf5adb1472edd09d9
  my-run.sx                 1a74f5546c852533510ab69ca07859527eb813001fc991db683c1ed1013f1117
  ...

Then the knobs, and each tensor file with one line per tensor; a parameter's line ends with frozen when it does not train and with its tags.

Loading one elsewhere

The safetensors files are the interchange format. A bf16 parameter is stored as f32, and experiment.json's dtype says what it means:

import json
from safetensors.numpy import load_file  # safetensors.torch.load_file for tensors

step = "/tmp/my-run/step-00000004"
identity = json.load(open(f"{step}/experiment.json"))
params = load_file(f"{step}/params.safetensors")   # {"model.net.embed.table": array, ...}
states = load_file(f"{step}/states.safetensors")   # {"model.net.embed.table/m": array, ...}
for declared in identity["params"]:
    assert list(params[declared["path"]].shape) == declared["shape"]

From a bucket, read the object's bytes and give them to safetensors.numpy.load:

import boto3
from safetensors.numpy import load

body = boto3.client("s3").get_object(
    Bucket="my-bucket", Key="runs/my-run/step-00000004/params.safetensors"
)["Body"].read()
params = load(body)

Signals

SIGTERM, SIGINT (Ctrl-C) or SIGHUP (the terminal went away) stops a run at the next step boundary:

  1. the step in flight finishes;
  2. the metrics get interrupted by <signal> at step <n>;
  3. a checkpoint of that step is written when there is a location;
  4. the stream is drained without done, and the status file reads interrupted;
  5. the process prints sexpgpu run: interrupted by <signal> at step <n> and exits 128 + n: 129, 130, 143.

A second SIGTERM or SIGINT ends the process at once, without any of it. A SIGHUP never does, because a machine shutting down sends SIGTERM and SIGHUP together. A run killed outright (SIGKILL, a second signal, the machine gone) writes nothing more: its status file keeps reading running, and its resume repeats the steps after its last checkpoint.

A spot VM's shutdown or a job runner's cancel gives the process a SIGTERM and some seconds; a checkpoint larger than that allows is passed over. See preemptible jobs.

The finite guard

(defrun ... :guard-finite true) checks every proposed parameter and optimizer state for NaN and infinity on the device before any of the step is committed: one reduction and one downloaded number per tensor. When one is not finite the run stops with the step, the parameter and the state in the error, which the status file and a last failed: <error> annotation repeat. Nothing of that step is committed, so the last checkpoint is intact and a resume from it replays the step.

sexpgpu run: step 2: the proposed parameter of model.b has 1 values that are not finite; nothing was committed

It is off by default because it adds a reduction and a download per parameter per step.

Related: run, bundle, the status file.