training infrastructure, on your cloud

Train on your cloud. Lose nothing when it fails.

GPU Zero starts GPUs when there is a job, checkpoints every step, and stops them when the work or the budget is done. The platform is free: you pay your cloud provider, not us.

train.py

run resumes from step 48,000

A cluster of GPUs at work
import ravex

@ravex.train_loop(preemption_handler=True)
def train():
    ...  # unchanged
python train.py
$0

platform fee. The GPUs are billed by your provider, to your account.

0 idle

A machine starts when a job needs it and is returned when the job ends.

1 decorator

to make a PyTorch loop resumable. The rest of the script stays as it is.

how it works

Preferences, not rentals

You do not keep machines. You say what a run may use, and GPU Zero finds it when there is something to run.

  1. 01

    Wrap the loop

    One decorator on the function that trains. Inside it, nothing changes.

  2. 02

    Say what you accept

    Which GPUs, how much memory, how much money. Not a machine to keep alive.

  3. 03

    Run it

    A GPU starts when there is a job and stops when it is done, idle, or over budget.

features

Built like infrastructure

What is marked soon is being built now; the rest runs today.

checkpoints

A lost machine costs seconds, not the run

Live

Ravex checkpoints every step it is told to, with only what changed since the last one. Rerun the same script and it resumes: weights, optimizer, schedule, RNG and dataloader position.

fork

Start again from any step

Live

Fork a run from a checkpoint with new hyperparameters, and the console draws the child where the parent left off.

metrics

Every curve, histogram and log line

Live

Scalars, weight histograms as heatmaps, system metrics and what the script printed, live while it runs.

your cloud

Your accounts, one provider or several

Soon

Connect RunPod, AWS, Google Cloud or Azure. A run can span providers; the machines and the bill stay yours.

your code

From a repository, at a commit

Soon

Link a GitHub repository and run a commit. The commit is pinned to the run, so a resume or a fork restarts from the same code.

your agent

Drive it from Claude Code

Soon

Launch, follow the logs and fork from your own coding agent, with the cost said before anything starts.

Your next run can survive its machine.