Data Collection Underworld: Subcontracting, "Bang Laotai," and Cross-Border Arbitrage
By 2026, ego is becoming a more mainstream data source for the embodied intelligence industry.

Dirty Data | Embodied AI Bubble | Binocular | Real Labor
Produced by | AI Nao

Info
By 2026, ego is becoming the more mainstream data source for the embodied intelligence industry.
Ego is first-person data captured from a human point of view — the method is to strap a device onto someone's head and film whatever they're doing from that "first-person perspective." The reason ego has become so widely used is based on Physical Intelligence's π0.5+Ego, and NVIDIA's recently released EgoScale. They've proven that ego can be absorbed as pre-training signal for VLA models, and more critically, that model performance improves predictably with data volume. This makes ego something like text data for large language models: the more you use, the better your model performs.
In Silicon Valley, Figure AI and Tesla haven't fully disclosed their data compositions, but according to industry insiders, ego is now regarded as training material equally critical to code.
But unlike the internet data needed for large language models — grabbing a cup, folding a shirt, tightening a screw — these real human behaviors have never been fully digitized. So collecting ego data requires paying humans.
Supplying large quantities of ego data to Silicon Valley are two companies, Build.AI and Human Archive. They collect at $10 per hour in factories across India, Nepal, and Myanmar, then sell it at $15 per hour to OpenAI, Tesla, and Figure AI — essentially a cross-border data broker.
In China, ego data collection has already formed a complete industrial chain.
The first layer is the demand side, typically model companies and large corporations. They hold the customer relationships, control the core technology, and need massive amounts of ego data for model training.
The second layer is the wearable equipment needed for data collection, including monocular, binocular, and multi-camera ego devices. Previously most teams used monocular devices like GoPro, but because these lack stable depth and 3D geometric information, subsequent spatial reconstruction and trajectory estimation costs are high. The industry as a whole is gradually shifting to binocular devices.
The third layer is data processing: representative companies include Guanglun Intelligent and Zhiyu Jishi. They clean and annotate raw collected data, then conduct simulation learning based on the corresponding data. Guanglun's founder was previously NVIDIA's simulation lead.
This episode's practitioner, Huang Darong, sits in the middle layer — she's a data broker. She entered the business in 2026 as ego collection demand was heating up.
Data brokers like her are the critical link connecting all the pieces: on one end, dealing with local power brokers to get devices into factories and retrieve qualified data; on the other end, dealing with demand-side clients to supply data or sell equipment.
In her observation, every segment and player has their own motivations and some unspoken rules. "At the end of the day, it's still a traditional tiered subcontracting business. The core is how to distribute the money properly," Huang believes. "The reason it currently works is the massive wage differential that lets everyone profit. For example, Silicon Valley typically collects in Nepal, India, Africa — this arbitrage space has already become a business."
By estimates, embodied AI will need 100 million to 1 billion hours of ego data in the next 2-3 years, but currently the largest open-source dataset has only about 100,000 hours. The gap is enormous, though how long the business lasts also heavily depends on model development.
As for whether embodied AI will ultimately replace low-level human labor? From Huang's perspective, it probably will.
For the first time, humans are selling real labor data to AI so that robots can replace humans later. This is the most paradoxical and unsettling aspect of the entire industry. Is it morphing into a new form of exploitation?
Below is Huang Darong's oral account. As a heavy storm approaches East China, she is currently supervising work at a toy factory in Kunshan, Jiangsu.
Entering the Data Collection Factory
One day in July, Guangdong was hit by torrential rain. I received a call from Ding Quan, asking me to visit a factory in Zhongshan with him. "Come see for yourself how workers collect data wearing ego devices."
I had met Ding Quan, a boss in his forties from Wuhan, during a previous company visit. Not tall, he speaks extremely fast. He previously ran a software company, entered the humanoid robot data collection business after the 2026 Spring Festival. His mind moves faster than his mouth, and his hands faster than his mind. In less than half a year, he had already deployed a large batch of equipment across factories.
When I got his call, I was in Guangzhou. I hailed a car. The hour-long ride, the driver drove especially slowly. Outside the window was a blur of white, rain pounding the glass. The scenery gradually shifted to rows of factory buildings. Arriving in Zhongshan at noon, Ding Quan gave me a restaurant address.
Today's host was Zhao Nian, a Guangdong native in his sixties. He had worked in the system when young, then went into business running factories, now holding several manufacturing plants. The table was filled with other factory bosses from Zhongshan — seafood, sashimi, local home-style dishes spread out in abundance.
The main purpose of this meal was to persuade these bosses to let their workers wear collection devices. Their questions and concerns were simple: how much per hour, will wearing it interfere with work, do we need to change the production line? If it's not troublesome and earns extra money per hour, why not.
After eating, we went to a furniture factory workshop to see the actual data collection site.
Furniture factory work doesn't look high-tech, but it's exactly the kind of operation robots struggle most to learn. Fabric is soft, deforms with a pull. Wrapping, fitting, aligning, fine-tuning — too much force creates wrinkles, too little and it's off. The subtle adjustments human hands make unconsciously require massive amounts of real data for robots to learn.
- Toy factory
Strange Devices
In this furniture workshop, workers were already wearing ego collection helmets — modified cycling helmets with binocular lenses, 3D-printed shells, batteries and other components integrated on top. The appearance looked quite professional, supposedly the most top-of-the-line equipment currently available.
But when I chatted with the workers, the universal feedback was: wearing it on the head is pure suffering, too heavy. Fine for a short while, but as soon as you need to bend down, crouch, or work repetitively, the weight on top sways with inertia. After half a day, your neck simply can't take it.
This equipment had another fatal problem: every two hours, all workers had to stop, remove devices in unison, plug into computers, copy data, verify files. Only after confirming everything was fine could they redistribute them. This required having a dedicated person on-site full-time to supervise — such operations costs simply couldn't sustain scaling.
Ding Quan lamented to me: this kind of equipment, engineers probably think it's perfect, but in actual deployment, it's all show and no go. You can't run volume with it.
He pulled out another set of equipment — just over 100 grams, much simpler structure, all extraneous parts chopped, battery moved to the waist, using swappable SD cards for storage instead of on-site file copying. When a card fills up, just swap it and mail cards back to headquarters in batches. "This set is clearly more practical, can easily and stably collect for six or seven hours."
I chatted briefly with some workshop workers. They had no resistance to wearing ego devices, nor any expectations. Their thinking was very plain: as long as this thing doesn't interfere with work.
Because furniture factories are all piece-rate wages — more work, more pay. If wearing the device drags you down and you finish fewer pieces in a day, that directly hits your income.
After the furniture factory, Ding Quan took me to a large food factory next door, also Zhao Nian's plant, larger in scale, six buildings, the assembly lines looking neat and standardized.
I had thought this would be an ideal data collection site, but after walking through the entire building, my conclusion: data collection is basically impossible here.
Food factories require dust-free, sterile, fully enclosed protection. Workers are wrapped head-to-toe in protective suits, dust caps, gloves — external filming equipment meets neither hygiene standards nor can be disinfected, so it can't enter core production areas. Additionally, such factories are very sensitive about process leaks and workflow exposure, and won't set up filming zones for us.
After touring the six-story plant, the only areas that met collection criteria were end-stage sections like finished product boxing.
This field visit with Ding Quan completely corrected my previous misconception: it's not that the more advanced and intelligent a factory is, the more suitable it is for teaching humanoid robots. Quite the opposite. The most ordinary processing plants are the ones willing to open their doors and welcome AI data collection, because they have few patents, open processes, and relaxed environments.
- A set of binocular equipment + wrist camera
Distributing Money
At the end of 2025, I developed a visual smart toothbrush and encountered many problems collecting oral data and training models. During this period I met many friends in robot model training and data collection. They collect and gather video data in the real physical world, and have developed a methodology for dealing with "dirty data."
They also introduced some data collection projects to me, so I'd have more cash flow to support real oral data collection. Plus my previous business had always targeted the North American market, so I was also authorized to act as sales agent for some ego equipment.
In the entire data collection industry, to actually land on the ground, the first link is people like us who broker equipment — we're also called collection organizers. The responsibility is to find suitable target scenarios, rent or buy equipment, deploy to factories, train factories, set rules, manage operations, handle malfunctions, then deliver qualified data upstream to demand-side clients.
We typically find local contractors and factory bosses, break down requirements and subcontract to them. They open their factories and train workers, receiving corresponding compensation. These local contractors mostly hold a dozen or so factories; many of them previously did labor outsourcing, or have local connections controlling large factory resources.
Bypassing them to contact factories directly creates a lot of trouble.
At the very bottom are the actual data collection implementers: frontline workers. All data, actions, machine learning samples come entirely from their hands.
For example, I recently received a demand from a data company, offering 60 yuan per hour, needing tens of thousands of hours of data. I'll prepare the equipment, set the standards, then distribute at 40 yuan to the factory boss, and finally about 20 yuan reaches the worker's hands.
Recently, a friend told me something that was both laughable and exasperating.
He found a hardware assembly factory in Hebei. When he was on-site supervising, workers worked normally, collected normally, data quality was excellent. But the moment he left, within just three days, the backend suddenly showed a hundred hours of garbage data.
He investigated and found the workers had removed the ego from their heads and placed it on their workstation desktops, but kept it recording the whole time. They figured you don't have to wear the device, placing it on the table can still film.
Further investigation revealed the Baoding factory boss he partnered with hadn't distributed money to the workers — meaning workers had extra burden for nothing, unpaid labor. Who wouldn't want to cut corners?
This is how the business works. However glamorous the data, if workers don't get paid, it's all castles in the air. At its core, data collection isn't some high-end business model — it's a traditional business about distributing money.
Garbage Data
Starting in July, I purchased a batch of binocular ego equipment myself and officially deployed in factories.
From scenario selection, on-site management, profit distribution, worker cooperation, effective data duration, cutting massive raw video into standard task segments, to precise annotation of data... every day I face these granular problems.
For example, what kind of data to collect has hard metrics. Different devices also have large differences in effective collection rates. The industry has recently generally used binocular ego equipment for domestic collection, and with proper management, can achieve 80% effective collection rate. Overseas data collection in places like the Philippines only has about 10% effectiveness.
My current business mainly deploys in factories in Kunshan, Jiangsu and Yunnan. One scenario collects roughly 100 to 500 hours. Previously monocular devices like GoPro could be rented, but currently monocular data sales are already very difficult. Binocular equipment must be purchased, around 3,000–5,000 yuan per unit. For large-scale data collection, equipment costs are also substantial.
Many people assumed that with the heated funding for world models and embodied AI in the first half of the year, any company claiming to train its own model would surely purchase large amounts of ego from us. Based on my personal experience, many embodied companies simply place a few self-developed devices in factories, then deceive investors by claiming thousands of devices are collecting data, with data scale reaching millions of hours.
AI Nao reported on the world model funding chaos: What Kind of Model Is the World?
I encountered an even more shameless team, publicly claiming to be the country's largest-scale data collection company with over a million hours collected. I asked the factory how many devices they actually deployed, and the factory told me: five.
I often wonder, why don't these investors come ask us data collection people? The duck knows first when the river water warms. To be honest, looking at data procurement volume, there are currently only "1.5" domestic model companies making large-scale data purchases.
Recently, an investor from Meituan Longzhu came to me for research. I suggested that when judging whether a company is genuinely building models, they should focus on their data procurement volume.
The pre-training stage essentially uses three types of data: internet video, ego data, and some UMI data. If they haven't purchased data from data vendors or collected sufficient data themselves, what exactly are they training? Furthermore, if they don't even have tactile data for grasping cups of different shapes, how did their robot demo learn to hold a cup? It must be fake.
I can only say that in China's startup scene, there are quite a few fraudsters.
Good Data
Recently, someone in charge of ego data collection at a domestic top-tier embodied company found me. His name is Wang Nanqing, from Hebei, wears black-rimmed glasses, usually loves sports, skin tanned dark. Recently he chose to leave the company and start his own data collection business.
I asked why he left. He was quite straightforward: in big companies, projects are about delivery. Whether the data is real or can genuinely help model iteration and improvement isn't the focus of delivery.
Wang Nanqing felt this wasn't what he wanted to do. He wanted to find valuable, effective data from real labor scenarios.
He also told me that many embodied companies, to make acceptance look good, hire people to perform standardized actions according to scripts — zero mistakes, zero deviation, perfectly neat throughout. The footage looks impeccable, but such data has zero training value. Real human labor is full of errors — slipping, grabbing wrong, redoing. Robots need to learn from living humans, not a fake performance clip.
It was through chatting with him that I figured out: what exactly is real data, and what is garbage data?
First, absolutely don't look at demos and PPTs from embodied robot companies that raised hundreds of millions — much of it is bragging. The raw video they present looks beautiful, but is actually useless footage.
A truly qualified piece of data: first, it must be a complete, closed-loop task unit with beginning and end — an episode. Additionally it needs clear tasks, fixed scenarios, corresponding manipulated objects, complete success/failure outcomes, plus quality ratings and anomaly notes.
Also, what seems like simple data field annotation — adding text descriptions to video, like "cutting mango on kitchen counter, left hand steadying mango, right hand holding knife slicing mango into pieces" — seems technically low-barrier, but demands high accuracy and logical relationships. It heavily depends on data annotators' abilities. Their job is describing the physical laws of object motion for robots.
Because upstream offers low annotation prices, most domestic data annotators have technical secondary school education. Error rates are high, logical relationships poorly understood. They skip labeling simple actions, or use inconsistent semantic descriptions — the same action, different people write differently, and robot understanding becomes confused.
Regions with high data annotation quality are Nepal and Kenya. Wages are similar to domestic levels, but local top university students do the annotation. Strong logical thinking, pass rates can reach over 80%.
I heard that recently some big companies like ByteDance have started hiring master's students for annotation — definitely to improve pass rates.
Exploiting Old Ladies
Even more absurd things are happening in the industry.
A top-tier data collection company, having raised tens of billions, started "village data collection" in places like Shandong and Henan — having old ladies and aunties in villages wear PICO VR devices to do tasks like chopping vegetables and cooking, 5 yuan per hour per task.
The most outrageous part: frontline staff reported that the middlemen organizing data collection eventually just ran away, not even paying this tiny labor fee.
Setting aside money, is this data even valuable? Because humanoid robots will first land in upper-middle-class households in the early stage. Datasets from rural labor differ too much from urban family life.
And while PICO collecting embodied data comes with spatial information, saving annotation costs later and making things convenient for data companies, that doesn't mean it suits real labor scenarios. I bought PICO equipment myself — standing and wearing it to play for a while is already tiring, let alone wearing VR devices on your eyes to chop vegetables, cook, and work continuously. Many people also experience dizziness and discomfort from extended wear.
In the end, what's collected — real workflows, or "performing work" to complete collection tasks? I'm afraid only these data companies know.
More ironically, these top-tier data companies took funding, have money in their accounts, yet delay payment downstream by three to six months, pushing cash flow pressure all the way to suppliers, with the hardest-working rural aunties suffering in the end.
Scenarios not thought through, whether data has value not thought through — first turn rural aunties into cheap "data consumables."
The cutting-edge AI industry, making money by delaying 5-yuan labor payments. This isn't innovation. I think this is just the oldest scam business, wrapped in a Physical AI skin.
I believe robots will definitely land first in urban upper-middle-class households, so household-related scenarios will definitely become important data sources. For example, recently some data collection companies find time-flexible mothers, having them wear binocular devices and wrist cameras to fold clothes, organize, cook — roughly 30 yuan per hour.
On my Xiaohongshu account's backend, every day countless mothers and ordinary people message me asking if there are reliable home-based collection channels. What they want is simple:碎片 time, an extra income source.
I often think lately: large language models feed on accumulated stock data from the internet, from countless ordinary people freely contributing content, but the dividends concentrate entirely at top tech companies. Ordinary people don't get much benefit. In the end, it raises society's overall productivity, but the public gains no new income.
But embodied intelligence requires physical data that must rely on present, vivid, currently-occurring human labor data — so it must purchase labor experience, actions, and real scenarios from ordinary people.
Perhaps eventually, the data collection industry will see a platform like DiDi or Meituan emerge, because ordinary people's needs are exactly the same as food delivery riders and ride-hailing drivers: take orders, earn money. Slightly more complex is that whether a DiDi driver or rider completed an order is visible at a glance. But whether data collection is qualified, whether it has training value, requires an entirely new set of standards and review.
Perhaps this is a new type of AI job accessible to all ordinary people.
AI Nao is looking for independent contributors If you're curious about people and stories in the AI wave, and skilled at writing
Please contact: shizao1123@hotmail.com
Email should include: 1. Personal introduction; 2. Relevant articles; 3. Favorite AI Nao article Competitive compensation, professional editorial guidance
Looking forward to making noise with you
