Community Discussion · Tracks

Wrote disaster recovery scripts for Vast.ai spot GPUs, found tons of bugs

hongtaohongtaoAug 92026/08/09 265 views

I messed around with Vast.ai spot GPUs over the weekend and wrote a failover daemon for training tasks, hitting quite a few pitfalls.

Background first. Vast.ai is a GPU rental platform; spot instances are cheap but can be evicted at any time. Eviction itself takes only a few minutes, but what really hurts is losing hours of training progress along with it. My lab budget is tight, so I wanted to save wherever possible, hence writing a daemon that detects eviction signals and automatically migrates tasks to another machine, resuming from the latest checkpoint. The materials say spot GPUs are suitable for restartable tasks with checkpoints; this idea is correct, but doing it is harder than imagined.

Day one, setting up the environment got stuck immediately. Vast.ai doesn't use passwords, only SSH keys. After configuring the key, I couldn't connect; checking the docs revealed incorrect key permissions. The platform isn't very beginner-friendly; official troubleshooting docs list many pitfalls, but they're all in English, requiring trial and error item by item. Using Claude Code to write the daemon's main logic was fast, but debugging still required hands-on effort.

By day three, I started real testing. I deliberately opened a cheap machine, terminated the instance manually halfway through to simulate eviction. The daemon indeed detected it and triggered migration, but after migrating, environment variables were lost. Checking the official docs, it explicitly stated "Environment Variables Not Working"—turns out it wasn't just my problem.

After a week of testing, I found 5 real bugs. One particularly nasty: checkpoint sync race condition. If the previous machine was still writing to disk during migration, the new machine read half a file and crashed instantly. I fixed this bug in my code, but the uncertainty of platform-side API behavior remains headache-inducing. After all, Vast.ai itself fixed a bunch of bugs in 2024; looking at their product update logs, their iteration pace is fast, but stability is still distant.

Materials mention Vast.ai has integrated SkyPilot, allowing direct log viewing via sky logs; this ecosystem integration is commendable. Some people also listed RTX 5090s for rent, indicating supply-side growth.

My conclusion: It depends. Suitable for those with engineering skills, not for novices. Pros: Cheap, diverse GPU types, spot instances friendly to restartable tasks. Cons: Many pitfalls in SSH and docs, unstable API behavior, numerous platform-side issues exposed during testing.

Pros Cons
Cheap prices, diverse GPU types Many pitfalls in SSH config and docs
Spot instances suit restartable tasks Unstable API behavior, env vars get lost
Integrated SkyPilot, expanding ecosystem Many platform bugs, need self-fallback

Trend prediction: Platforms like Vast.ai will get hotter, but reliability issues with spot GPUs will spawn a batch of specialized disaster recovery tools, similar to high-availability solutions in the database field. In the future, it might not be a question of "whether to use spot," but "which tool manages spot."


📌 This article is compiled from Hacker News, original: https://github.com/enplabs/spotwarp

Copyright belongs to the original author; this is a compilation and independent analysis based on public reports.

2 replies

?
Ctrl + Enter to reply
Yelin Does Not Eat Sponsored Meals

Docker does help avoid environment variable issues, but the image size increases, making migration slower. The checkpoint race condition is more fatal; I recommend using atomic writes or temporary files to avoid partial writes. By the way, what format do you save your training parameters in? Pickle or torch.save?

Long Yunfan

I know this environment variable pitfall too well... I got burned by this thing when running LSTM before. After switching machines, a bunch of PATHs got messed up, so I finally just put the environment into a Docker image to avoid being nervous every time I migrate. By the way, when he resumed from a checkpoint, did those training state parameters not get lost?