Robocurve Raises $10M Seed to Test AI Models on Real Robots
Robocurve (robocurve.org) has raised a $10M seed round to run independent evaluations of how well frontier AI models control physical robots. Initialized Capital led the round, with Pioneer Fund, Notable Capital, Decasonic, Y Combinator, Halcyon Futures, and others participating, according to the company's funding announcement.
The company tests frontier models on real hardware and publishes the results as a third party, the same role auditors play for text and code benchmarks but applied to machines that move. Robocurve is incorporated as a public benefit corporation with a duty to report those capabilities publicly, and it says labs do not set its research agenda, methodology, or results. Founder and CEO Jay Chooi previously worked as a researcher at the UK AI Security Institute and was a top contributor to Inspect Evals, the evaluation framework used by the UK government.
Robocurve's early findings point to general-purpose language models closing in on the specialized vision-language-action models built for robotics. The company reports that LLMs can outperform state-of-the-art VLAs on simple manipulation tasks, and that tasks models were failing a month earlier are now handled comfortably. MIT professor Phillip Isola has described the shift as the start of robot-use agents, a sequel to the computer-use agents that arrived a couple of years ago, and Anthropic has been studying Claude's robotics capabilities since 2025.
Latency remains the constraint. Controlling a robot in real time requires output token speeds that frontier models have not reached, though Robocurve tracks improvements of roughly 2.1x per month alongside inference hardware work at Cerebras, Groq, and Lamb Labs. Extrapolating those trends puts real-time control within reach somewhere between late 2026 and 2029, a projection the company is careful to hedge. Whether the results hold for embodiments with many degrees of freedom, such as hands or full humanoid bodies, is still unresolved.
In the three months since incorporating, Robocurve says its research has been viewed more than 6 million times and its open-source evaluation harness, Inspect Robots, has been downloaded over 97,000 times. Researchers from more than 200 institutions, including 19 of the top 20 universities worldwide, have signed up to build benchmarks with the company.
The funding goes toward a larger research team, a wider range of robots and tasks under test, and support for academic groups building open benchmarks. Robocurve is awarding $500,000 in total funding plus YAM arms to academic teams through its open-source benchmarking program.