Wrote disaster recovery scripts for Vast.ai spot GPUs, found tons of bugs
I messed around with Vast.ai spot GPUs over the weekend and wrote a failover daemon for training tasks, hitting quite a few pitfalls.
Background first. Vast.ai is a GPU rental platform; spot instances are cheap but can be evicted at any time. Eviction itself takes only a few minutes, but what really hurts is losing hours of training progress along with it. My lab budget is tight, so I wanted to save wherever possible, hence writing a daemon that detects eviction signals and automatically migrates tasks to another machine, resuming from the latest checkpoint. The materials say spot GPUs are suitable for restartable tasks with checkpoints; this idea is correct, but doing it is harder than imagined.
Day one, setting up the environment got stuck immediately. Vast.ai doesn't use passwords, only SSH keys. After configuring the key, I couldn't connect; checking the docs revealed incorrect key permissions. The platform isn't very beginner-friendly; official troubleshooting docs list many pitfalls, but they're all in English, requiring trial and error item by item. Using Claude Code to write the daemon's main logic was fast, but debugging still required hands-on effort.
By day three, I started real testing. I deliberately opened a cheap machine, terminated the instance manually halfway through to simulate eviction. The daemon indeed detected it and triggered migration, but after migrating, environment variables were lost. Checking the official docs, it explicitly stated "Environment Variables Not Working"—turns out it wasn't just my problem.
After a week of testing, I found 5 real bugs. One particularly nasty: checkpoint sync race condition. If the previous machine was still writing to disk during migration, the new machine read half a file and crashed instantly. I fixed this bug in my code, but the uncertainty of platform-side API behavior remains headache-inducing. After all, Vast.ai itself fixed a bunch of bugs in 2024; looking at their product update logs, their iteration pace is fast, but stability is still distant.
Materials mention Vast.ai has integrated SkyPilot, allowing direct log viewing via sky logs; this ecosystem integration is commendable. Some people also listed RTX 5090s for rent, indicating supply-side growth.
My conclusion: It depends. Suitable for those with engineering skills, not for novices. Pros: Cheap, diverse GPU types, spot instances friendly to restartable tasks. Cons: Many pitfalls in SSH and docs, unstable API behavior, numerous platform-side issues exposed during testing.
| Pros | Cons |
|---|---|
| Cheap prices, diverse GPU types | Many pitfalls in SSH config and docs |
| Spot instances suit restartable tasks | Unstable API behavior, env vars get lost |
| Integrated SkyPilot, expanding ecosystem | Many platform bugs, need self-fallback |
Trend prediction: Platforms like Vast.ai will get hotter, but reliability issues with spot GPUs will spawn a batch of specialized disaster recovery tools, similar to high-availability solutions in the database field. In the future, it might not be a question of "whether to use spot," but "which tool manages spot."
📌 This article is compiled from Hacker News, original: https://github.com/enplabs/spotwarp
Copyright belongs to the original author; this is a compilation and independent analysis based on public reports.
Physix Frontier