Robot training data collection is the practice of capturing first-person video, sensor streams, or teleoperated demonstrations that teach a machine-learning model how a physical task actually unfolds. That's the definition. Here's the part nobody had a number for until this month.
Quote two vendors for the same pick-and-place task and you'll get two numbers that don't share a decimal place. One says four dollars a demonstration. The other says a hundred and fifty an hour. Neither is lying. They're describing two different jobs wearing the same job title.
That confusion just got a lot less confusing. A new industry cost review, a factory built specifically to generate this data, and a hardware launch aimed at making capture cheaper all landed within the same few weeks. Put together, they read less like three separate stories and more like one market finding its price.
Key Takeaways
- A August 2026 Teahose cost review puts robot demonstrations at roughly $1–15 each, egocentric wearable capture at $25–60 an hour, and full teleoperation at $72–150 an hour depending on region.
- JD.com's provincial-level robot data collection center in Jiangsu, running since October 2025, was upgraded this August into a real physical-scene training ground rather than a static sample library.
- Orbbec debuted a "Robot-Free" data capture rig in Japan on August 6, 2026, letting humans generate egocentric training footage without a robot on site at all.
- RetailSMV, a 2026 dataset of 32,105 captioned clips from five supermarkets, shows vertical-specific egocentric datasets scaling into retail environments TeamPL already works in.
- The common thread: data collection is being priced, budgeted, and operationalized like a commercial service — not funded like a research grant.
The Price List Nobody Used to Publish
Ask ten people in robotics what a training demonstration costs and you'll get ten different answers. That's normal for an industry that's still arguing about how much data robots even need. NVIDIA's fine-tuning mix used 4 hours of teleoperation, Ant Lingbo's chief scientist puts the "robot GPT-1 moment" at 1,000,000 hours, and Ken Goldberg's yardstick for the gap is that vision-language models already trained on roughly 100,000 years of human-equivalent internet data while the largest reported robot teleoperation dataset is about one year (Teahose, 2026). Those estimates span five orders of magnitude. Nobody agrees on the destination.
What's new is that somebody finally priced the trip. Public sources converge on roughly $1–15 per usable simple demonstration, while per-hour quotes diverge 4–10 times: egocentric wearable capture runs about $25–60 per collected hour, real-robot teleoperation about $72–145 in China, and about $90–150 fully loaded in the US (Teahose, 2026). Teleoperation, for the uninitiated, refers to a human operator directly puppeting a robot's motions in real time so the model can later imitate the recorded demonstration. It's the most expensive way to generate training data, and now there's a public number attached to exactly how expensive.
Why does a price list matter more than a benchmark score? Because budgets need line items, not aspirations. Once a buyer can compare $30-an-hour egocentric capture against $120-an-hour teleoperation, the conversation shifts from "how much data is enough" to "which method fits this line item." That's a procurement question, not a research question — and procurement questions are exactly where field operations teams earn their keep.
Building a Building Just for the Data
If data collection is now a budgeted service, someone eventually builds a facility for it. JD.com already has. A technician collects data for a housework scenario at a robot data collection center of Chinese e-commerce giant JD.com in Suqian, east China's Jiangsu Province (China.org.cn, 2026). In operation since October 2025, JD.com's robot data collection center is a provincial-level embodied intelligent robot data collection center in east China's Jiangsu Province , and this August it went further: the facility upgraded its data collection environment to a real physical scene, transforming the center from a static sample library into a training ground for embodied intelligent robots (China.org.cn, 2026). Consider what that actually means on the ground. It's not a server room. It's kitchens, laundry rooms, and living rooms built to be lived in and recorded in, staffed by technicians whose job is literally to generate housework footage on a schedule. That's field operations with a manufacturing floor's discipline — scheduling, quality control, repeat capture — applied to something that used to happen ad hoc in a research lab. We've made this same argument before: the real cost driver in robot data isn't the sensor, it's the logistics wrapped around the sensor. JD.com just built a building around that logistics.
The Rig That Skips the Robot
So if a dedicated facility is one answer to the cost problem, what's the other? Skip the robot entirely during capture. Orbbec presented its Robot-Free Data Collection Hardware Platform for Physical AI in Japan for the first time at ROSCon JP 2026 on August 6, 2026, marking the platform's official debut in the Japanese market (PRNewswire/Yahoo Finance, 2026). Built around EGO, UMI, and WristCam, the platform supports first-person observation, detailed hand-object interaction, and wrist-level near-field data collection (PRNewswire/Yahoo Finance, 2026). Egocentric video refers to footage recorded from a first-person, head- or wrist-mounted viewpoint that approximates human sightlines and hand movement, rather than a fixed third-person camera watching from across the room. That's the whole appeal of a robot-free rig: a person walks through the task wearing the sensor, and the robot never has to leave the lab. Given the Teahose numbers above, that's the difference between paying teleoperation rates and paying wearable-capture rates for the same underlying motion data. It's also exactly the terrain we've mapped out in our look at viewpoint economics in robot data — where the camera sits changes what the footage is worth.
Retail Aisles Are Now a Dataset
Here's where this stops being a robotics story and starts being a field-operations story. RetailSMV is a 2026 dataset of 32,105 captioned retail clips from five supermarkets with synchronized staff-view egocentric and exocentric capture, predefined train/val/test splits, and a held-out protocol for video-world-model adaptation (Awesome Egocentric Video Datasets, GitHub, 2026). Five supermarkets. Staff-view footage. That's not a lab bench — that's a store floor, the exact environment TeamPL's in-store audit teams work in every week. The gap between a dataset like RetailSMV and a genuinely useful one isn't the camera. It's whether the person wearing it followed the right protocol in the right aisle at the right time. We've written about that gap before in what a robot dataset actually teaches fieldwork, and the lesson holds here: capture without a defined research question just produces expensive footage nobody can use.
What This Means If You're the One Running the Program
Zoom out and the pattern across all three stories is the same: money now flows toward whoever controls the physical logistics of capture, not just whoever owns the model. That's the market growing up. It's also visible in the funding, not just the tooling — XPeng's Dogotix robotics unit raised more than $900 million, bringing its pre-money valuation to $5 billion and post-transaction valuation to more than $6.3 billion this August (The Robot Report, 2026), and Google DeepMind unveiled Gemini Robotics 2, the latest version of its vision-language-action model in the same month (The Robot Report, 2026). Big capital and better models both still need someone to physically go collect the footage. As one recent analysis put it, "whoever can build the infrastructure to collect, standardize, and manage this data at scale will hold the key to the next era of industrial automation" (SCSP, 2026). That's true for humanoid warehouses. It's equally true for a mystery-shopping program or a GIS survey — the infrastructure question is the same one, just with a different payload. Where does that leave a field data vendor honestly? Best for: designing capture protocols around a specific research question, running consistent field logistics across sites, and adapting quickly when a client's model needs a different viewpoint or scenario mix. Where it stops: a field team isn't going to out-scale a dedicated 48-hectare data collection center or a robot-free rig built for thousands of hours of unattended capture. Different tools, different jobs. The mistake is assuming one replaces the other.
Frequently Asked Questions
What does "robot-free" data collection actually mean?
It means a person, not a robot, wears the capture hardware during recording. Orbbec's platform, launched in Japan in August 2026, uses wearable sensors like WristCam and EGO to record first-person task footage without a robot present during capture.
Why is teleoperation so much more expensive than egocentric video capture?
Teleoperation requires a human to operate an actual robot in real time, which adds hardware, calibration, and slower throughput. Industry figures from Teahose's 2026 review put teleoperation at roughly $72–150 an hour compared to $25–60 for wearable egocentric capture.
Is retail a real target for egocentric datasets, or just robotics research?
It's a real target. RetailSMV, a 2026 dataset built from five supermarkets, pairs staff-view egocentric footage with third-person capture specifically for retail-relevant model training.
- Robotics Training Data: The Companies Feeding Physical AI — Teahose, 2026
- Glimpse of JD.com's robot data collection center in Jiangsu — China.org.cn, 2026
- Orbbec's Robot-Free Data Collection Hardware Platform Makes Japan Debut — Yahoo Finance / PRNewswire, 2026
- Awesome Egocentric / First-Person Video Datasets — GitHub, 2026
- Top 10 robotics stories of August 2026 — The Robot Report, 2026
- ISF Voices 2026: The Robotics Data Gap — SCSP, 2026