Dreamscale Labs Moves Robot Inference to the Cloud
Dreamscale Labs (dreamscalelabs.com) has launched a cloud inference platform for robotics, betting that the models driving the next generation of robots will run on remote servers instead of on the machines themselves. In its Y Combinator launch, the company said it serves frontier robot models such as DreamZero and MolmoAct2 40 to 50 percent faster than existing cloud providers.
The argument starts with hardware economics. Robot foundation models have grown large enough that the newest ones no longer fit comfortably on the GPUs mounted on a robot, and the cost of that onboard compute is rising. Dreamscale points to a global LPDDR memory shortage and to NVIDIA doubling Jetson prices in July as evidence that the on-robot path is getting more expensive.
Moving inference off the robot trades one problem for another, since a model running in a data center has to return an action fast enough for the robot to act on it. Dreamscale's approach is to attack latency at every layer at once: architectural changes to the models for faster inference, kernel-level work to get more out of server-grade GPUs, network optimizations for transport, and placing compute in regions close to where customers' robots are. The company says the combination lets it run DreamZero 14B at 320ms against 350ms for NVIDIA's own implementation, using half the GPUs.
The four founders are Eric Ren, who is CEO and previously worked in AI safety; Kevin Lin, chief scientist and a former quant who moved into robotics; Antoine Nguyen, CTO and previously a VLA quantization researcher at UTS; and Chris Yoo, chief product officer. The company is part of Y Combinator's Fall 2026 batch.
Dreamscale also charges by usage. Robotics research teams tend to leave GPUs idle for long stretches and then send heavy bursts of traffic, so the company offers on-demand pricing for inference, which few providers in physical AI do. It is working with early partners in deployment, research, and evaluation, and builds a custom inference stack with each one.
The company laid out its reasoning in a thesis post, acknowledging that most of the field assumes a robot's intelligence should live on the robot. Its main argument is that robot models are following the same scaling pattern as language models, with larger models generalizing better, and that the move toward world models, which predict video frame by frame, adds still more compute. By Dreamscale's figures, NVIDIA's DreamZero 14B needs two GB200 instances to reach its advertised control rate.
The post also points to the physical cost of carrying a GPU. A Jetson Thor can draw 120W, which Dreamscale estimates would use about a fifth of the battery on a humanoid like the Unitree H1, and the chip's cooling needs limit how the robot's body can be designed.
Finally, Dreamscale argues that tying model performance to GPU ownership holds back small labs and independent developers, who often wait on scarce hardware before they can test a policy. It cites Physical Intelligence, which has said cloud serving added only 10 to 15 milliseconds of network overhead in its own setup. That result came from a research environment close to the serving region, and Dreamscale is building to deliver similar latency to robots wherever they are deployed.