
Cloud Robotics Platforms: How Are They Different from a Local ROS Cluster
I compared rapyuta.io with a self-built ROS multi-robot cluster and actually ran through it. I didn't dare touch the production environment, so I set up a small scenario myself: three simulated AMRs doing path planning and obstacle avoidance, offloading the compute-heavy part to the cloud. Took less than a week.
Getting started felt more like a compiler toolchain than I expected. After registering, you go into the console — packages on the left, deployments in the middle, device tree on the right. It breaks robot applications into packages, each with dependency declarations and build configs. That's the same idea as the pass pipeline I look at in IR every day.
When creating a package you fill in the Docker image address, write the startup command, and declare which device interfaces it needs. I packed a ROS-based navigation node into it and finished configuring in about twenty minutes. It supports ROS1 and ROS2; I used ROS2, and the topic mapping part has to be specified manually — that's where I got stuck. The docs weren't very clear, and it took me three tries to get it right.
On deployment, it spins up a container in the cloud and then establishes a tunnel to the robot on the edge side. The experience here is genuinely smooth — one click to deploy, up in tens of seconds, logs streaming right in the web page. Way more comfortable than SSH-ing into each machine and running docker run.
But that's also where the problem is. The cloud and the robot are connected over the network, and when my simulated environment's network jittered, topic latency went up. It has a built-in reconnect mechanism, but the behavior during reconnection doesn't have very clear logs, and it took me a while to confirm it was a network issue and not my node crashing. This is worse than a local cluster — if something dies locally, it dies, and one look at dmesg tells you.
Another thing is debugging. Running ROS locally, I can rostopic echo directly, I can gdb attach. On the cloud, the logs it gives are container stdout, and to see anything finer you have to stuff tools into the image yourself. I ended up baking diagnostic scripts into the image just to get by.
One more detail: its device management is designed around real robots. When I hooked up a simulator, the device status kept showing offline even though data was actually flowing. This is probably me using it wrong, but it also shows its tolerance for non-standard connections is so-so.
Whether it fits depends on the team, not the tech. Teams with multiple robots, multiple sites, and not enough ops people will find it useful — its batch deployment and centralized logging are genuinely convenient, especially once your robot count goes up, since the marginal cost of SSH-ing into each one is high. That logistics warehouse case in Japan — reportedly new operators can get up to speed in under an hour of training — I believe it, because its level of abstraction really is high. Conversely, single-machine, LAN, latency-sensitive setups aren't a fit. If you just have two or three machines in the same room, a self-built ROS cluster is more direct, and you don't have to deal with that layer of network uncertainty.
It treats robots as a distributed system, and that approach is right — but only if your network and team size can support that layer of abstraction.
Physix Frontier