Here’s A $32 Million Bet That Robots Don’t Need A Billion Dollars Of Real-World Data
by John Koetsier · ForbesFigure AI pulled the wraps off Index last week. It’s a billion-dollar bet on real-world data for robot AI training, and it’s also a distributed data engine that has already collected more than 16 million videos from 108 countries, paid out $15 million to the people who shot them, and carries a commitment of over $1 billion in data and compute over the next 12 months. The pitch is simple and expensive: if you want robots that can do anything, you need a mountain of real human beings doing everything.
But do you really?
Harry Mellsop doesn’t think that billion dollars is a waste of money. He just doesn’t believe most companies can afford it.
“If you’re a company like Figure, and you have the resources to go and do something like that, that’s fantastic,” Mellsop told me in an interview last week. “Real data is always going to be the gold standard.”
Mellsop is the co-founder of Antioch, a New York-based robot training data startup that today announced a $32 million Series A led by Greylock, with participation from A*, Category Ventures, Box Group and Icehouse Ventures. the core premise: simulated data is good enough to train tomorrow’s humanoid robots. The company also disclosed a customer most robotics startups would trade a limb for — Amazon’s Ring — alongside partnerships with cloud provider Nebius and Nvidia.
The gap Antioch sells into
The argument goes like this: digital AI has compressed software development from weeks to days, because AI can write code, compile it, run it, break it, and try again. The loop is tight and quick, and while there’s quality concerns, it is often good enough.
But physical AI has no equivalent loop.
Every change to a robot — new camera, different radar, retrained policy or skill — still needs to get re-validated on hardware, in a room, with engineers standing around it. (I just spent today in exactly that world in Munich, Germany.) This is a major bottleneck: it’s slow and expensive.
“Automating the physical world is the defining economic imperative and opportunity of our time,” Mellsop said. But it only happens fast if you hand AI the tools to test its own work.
Antioch’s platform builds high-fidelity simulations for a specific customer’s hardware, then runs it at cloud scale: thousands of parallel evaluations from one experiment. Jason Mitura, VP of software development at Amazon and Ring’s CPO, says the results have held up under adversarial conditions: “Antioch’s simulations have closely matched our physical test results, including in scenarios we deliberately held out of calibration.”
That’s important: held-out scenarios are how you find out whether a simulator has learned physics for real, or learned your test suite.
Of course, the billion-dollar approach and the sim data approach aren’t opposites
One problem for most companies with real-world data is that you can never seem to get enough.
“One of the things that we’ve learned about developing autonomous systems with end-to-end learning is no data is ever enough,” Mellsop told me. “You look at the LLM space, you look at the self-driving space, and the winners are these companies that have this multi-faceted data accumulation strategy.”
He points to Tesla as an ideal: a business where customers pay you for the privilege of generating your training set, in this case for self-driving tech. Figure’s Index is a version of that — people get their houses cleaned and Figure gets the video — but while video is valuable, it’s also incomplete. Head-mounted cameras don’t capture how hard you gripped a deformable object like a towel, or how heavy a teapot is versus a ceramic mug. Teleoperation data does more of that, but it also costs an order of magnitude or two more per hour.
Simulated data, Mellsop argues, arrives ready to go and fully labeled: “You can generate samples of data for things that you really couldn’t either feasibly or responsibly collect in the real world.”
So the pitch isn’t don’t collect real data. It’s collect less of it … and spend less on it.
“You still need to collect some real-world data, but we help you be a lot more sample efficient,” he said. “Maybe you don’t need to spend the full billion dollars on this.”
Interestingly, when you’re using simulated data, your real-world data becomes less the basic training data than the correction data. Run a robot in the world and you’ll quickly find out what simulation is bad at: motors that behave differently when they get hot, or old, or other real-world oddities that’ it’s hard to dial up in sim data.
Mellsop calls it long-tail fidelity. Feed the small, expensive, high-quality real data back into the simulator, and the simulator gets meaningfully better across everything it generates.
“The more of that data you have, the more accurate that sim is going to get — but you still get the benefit of having simulation sitting underneath it to accelerate your workflow,” he said.
It’s sort of like the flywheel Waymo, Tesla and Wayve have built for driving, but offered as a product to companies that will never have five million vehicles on the road. Or five million robots.
Not just humanoids, not just cameras
Antioch’s verticals are a useful signal about where simulation is actually getting bought right now. The company is delivering data for aerial autonomy and drone delivery, ground autonomy for AMRs and AGVs in warehouses. It’s also working with smart security systems, which is the Ring work, to better understand false positives, for instance. They’re also providing data for industrial automation and increasingly construction robotics.
“We also think a lot about how do you make things like radar, and lidar, and infrared and these types of non-visual sensors perform really well,” Mellsop said. “It’s a very multi-modal problem.”
Angels in the round include Palantir CTO Shyam Sankar, Foxglove CEO Adrian Macneil, and Nvidia executive Ian Andrews.