training infrastructure, on your cloud
Train on your cloud. Lose nothing when it fails.
GPU Zero starts GPUs when there is a job, checkpoints every step, and stops them when the work or the budget is done. The platform is free: you pay your cloud provider, not us.
train.py
run resumes from step 48,000
import ravex
@ravex.train_loop(preemption_handler=True)
def train():
... # unchangedpython train.pyplatform fee. The GPUs are billed by your provider, to your account.
A machine starts when a job needs it and is returned when the job ends.
to make a PyTorch loop resumable. The rest of the script stays as it is.
how it works
Preferences, not rentals
You do not keep machines. You say what a run may use, and GPU Zero finds it when there is something to run.
- 01
Wrap the loop
One decorator on the function that trains. Inside it, nothing changes.
- 02
Say what you accept
Which GPUs, how much memory, how much money. Not a machine to keep alive.
- 03
Run it
A GPU starts when there is a job and stops when it is done, idle, or over budget.
features
Built like infrastructure
What is marked soon is being built now; the rest runs today.
checkpoints
A lost machine costs seconds, not the run
Ravex checkpoints every step it is told to, with only what changed since the last one. Rerun the same script and it resumes: weights, optimizer, schedule, RNG and dataloader position.
fork
Start again from any step
Fork a run from a checkpoint with new hyperparameters, and the console draws the child where the parent left off.
metrics
Every curve, histogram and log line
Scalars, weight histograms as heatmaps, system metrics and what the script printed, live while it runs.
your cloud
Your accounts, one provider or several
Connect RunPod, AWS, Google Cloud or Azure. A run can span providers; the machines and the bill stay yours.
your code
From a repository, at a commit
Link a GitHub repository and run a commit. The commit is pinned to the run, so a resume or a fork restarts from the same code.
your agent
Drive it from Claude Code
Launch, follow the logs and fork from your own coding agent, with the cost said before anything starts.
open source
The parts that touch your training are open
The runtime, the checkpoint engine and the node agent are Apache-2.0. Read them, run them without us, and see exactly what runs next to your model.
PyTorch runtime
ravex
One decorator: checkpoint, resume, fork, and training across the internet.
Checkpoint engine
moonclip
Rust with Python bindings: delta checkpoints, zstd, S3 and R2, zero-copy.
Node agent
smith
Runs the scripts it was given, never commands it was sent. Standard library only.