Data companies are getting outlandishly hot
Neo Lab and street smarts, hand in hand.
Around this time last year, certain star AI application companies were under intense scrutiny. People thought they had risen too fast, too far — valued at hundreds of millions of dollars barely a year after founding. So exaggerated.
But this year, those same companies suddenly look almost demure. A wave of data companies, incorporated less than a year ago, has already produced UniPat at a $2.5 billion valuation, plus several others priced between $300 million and $800 million.
Over the past few months, whether the topic was AI applications, hardware, embodied intelligence, or world models, I've sensed a whiff of irrelevance in countless small-group discussions. But there's one domain where everyone's ears perk up the moment it's mentioned — data companies.
All kinds of urban legends are flying around. We'll get into the specifics once some of these companies produce bigger results.
These stories replay in every cycle of venture capital hype. But unlike previous VC darlings, this batch of data companies and their founders genuinely craves obscurity.
Apparently the founders once reached a tacit agreement: PR only technical capabilities, keep teams and business progress strictly confidential. Their main request of the press is simply: please don't write about us.
The money really does come fast in this business. Rags-to-riches tales circulate everywhere — second-tier data vendors pulling in $4 million a month, college students landing $100,000 monthly salaries. Many data vendors have let slip, in various settings, that they don't actually need the funding; they just want the brand-name fund on their cap table to attract talent and customers, and they want the big-name VCs to provide cover and connections. When you've discovered an unclaimed treasure, of course you wear your brocade at night.
This may also represent the fastest valuation growth ever recorded in private markets. For comparison, Moonshot AI hit $2.5 billion at its one-year mark; Melvin Chen toiled diligently for three years before recently approaching a $3 billion valuation round.
Many data companies dress up their business in grand phrases like "connecting intelligence with reality" or "accelerating the path to AGI." But it boils down to one sentence: they make practice problems for models. Think of it as Yuanfudao for the AI era.
Pre-training feeds on public, static internet data. But in real-world work scenarios, models still can't perform at the level of competent professionals. They need to learn from industry experts, or practice extensively. What data vendors do is identify problems, devise solutions, and build environments. Either they find star tutors for models — collecting real human problem-solving trajectories from domain experts — or they craft exercise sets: RL environments comprising tasks, runtime environments, and scoring criteria. Then they hand these off to model labs for specialized single-subject training.
Chinese data vendors have developed a somewhat innovative tactic: publishing benchmarks. But these aren't your ordinary benchmarks. They're essentially a marketing vehicle aimed at model labs: Your kid's failing this subject, you know. See that kid next door who ranked first? It's because they practiced with our problem sets.
For the past two years, model companies competed on compute and algorithms. This year, the marginal returns on both are declining. Once the assembly line is built, the only remaining differentiator is raw material — data. That's practically the only thing that matters for models right now. Scale AI and Surge AI in the United States surged to prominence on exactly this logic.
This is one reason data company valuations have inflated so dramatically. But if it were just data, these companies probably wouldn't be worth this much.
Some data companies are also telling a Neo Lab story, even a potential model story: with enough capital and strong enough researchers, could a data vendor rapidly acquire the algorithm and compute capabilities to become a new model lab? Or at least become a model lab's vassal lord in some vertical domain? Compared to applications currently being swallowed by models, this ceiling and floor both seem considerably higher.
Remarkable. Years of effort and tens of billions of dollars in investment — in this narrative, all of a model lab's accumulated advantages can be casually leapfrogged.
Strip away all the storytelling and technical mysticism, and data looks more like a wild-grown B2B business. Data vendors operate like craft studios, each plying its own trade, with no clear leader yet emerging. Relationships — especially with top-tier model labs — are what truly matter.
This is a trust market, and there are no widely accepted验收 standards yet. Given the high degree of specialization and massive volume, model labs can only spot-check in the early stages. The more reliable method is ex post validation: run it through the model and see if performance improves. But even this has produced anomalies: one data vendor sold the same batch of data to two model companies — one saw clear improvement, the other barely any difference.
So model labs build trust with data vendors gradually. Start with small contracts of a few hundred thousand RMB, scale up if they prove reliable. The industry-gossiped $100 million-plus deals are actually annual framework contracts, with phased procurement and验收 — substandard deliverables get sent back.
But for now, model labs care more about speed and volume than quality. The consensus: whoever secures the most data captures the most vertical domains, each corresponding to commercial value in that vertical. Better rough and abundant than precise and missed. So everyone is racing against time. One category of data often becomes obsolete in under three months. A requirements doc drops today, the data vendor must deliver the day after tomorrow. Even precious data not currently needed gets stockpiled, to prevent competitors from claiming it.
This is also an information market, where being well-connected beats being technically superior. Judgments about what data is needed when cascade down from the pyramid's apex: domestic first-tier model labs follow OpenAI and Anthropic; second-tier labs follow the first-tier. If a data vendor can access this information earlier and more accurately than rivals, it's the equivalent of trading on insider information.
For example, GPT-6 Astra was announced in September, but by June, data vendors were already sniffing around for CAD drawing sources — they had learned from top domestic or overseas model labs and begun stockpiling. When Astra later blew up, modeling data indeed became expensive.
This is, moreover, an extraordinarily lucrative market. Among model labs willing to spend on data, ByteDance sits in the first tier, reportedly with over a billion dollars budgeted; Alibaba and Tencent follow. Anxious laggards like Ant Group, Xiaohongshu, Meituan, and DiDi also have demand for industry-specific data.
I once asked a data vendor friend: what's the most critical thing for success in this business? Without hesitation: of course, maintaining good relationships with model labs. Sounds a lot like SaaS, that sales-driven industry. But he offered an analogy: AGI is God; model lab researchers are priests, because they hold interpretive authority over scripture. We mortals must pay our tithes, in exchange for more divine revelation.
Now I understood. Data vendors are essentially star salespeople draped in sacred vestments.
The circle surrounding researchers is tight. Model lab researchers overwhelmingly hail from a handful of target universities — Peking University, Tsinghua University, HKUST, Beihang University — with classmates and lab-mates lifting each other up. Between data vendors and model labs, the connections are often former colleagues, fellow alumni, classmates. This also explains why founders of leading data vendors mostly come from model labs.
For all their apparent technical sophistication, these companies' founders often possess considerable street smarts.
In industry lore, there are those who invite researchers en masse to concerts; those who accompany model lab data purchasers on month-long scenic excursions; and, like Surge AI and Mercor, those who deploy large teams of attractive sales staff to interface with researchers.
Trust, of course, runs deeper than mere wining and dining. Some data vendors lacking top-tier research capabilities can theoretically serve model labs in another capacity — as externalized free agents, helping the parent entity bear risks it cannot or would prefer not to assume directly.
Consider two hypotheticals.
Suppose you're a model company that wants to distill from OpenAI and Anthropic. But the compliance and reputational risks are substantial. Here's a safe approach: find a reliable junior partner, show them how to set up the data pipeline, have them run OpenAI and Anthropic's models, then return the generated data. Once it passes through another set of hands, the data is laundered.
Suppose you're a model company with an application product. The application appears to be yours, but using user data to train models would violate your terms of service. Here's another safe approach: find another reliable junior partner, have them scrape your product data according to your specifications, even help clean it. Per your contract, you can assume the data they deliver has resolved any copyright and compliance issues.
This partly explains the mystery and evasiveness surrounding data vendors.
Where Neo Lab meets street smarts, where star projects and underwater transactions fly together — data is the most story-rich sector I've encountered, embodied intelligence aside.
But the more compelling drama lies outside the small circles of venture capital and model companies. In an industry where data rivals oil in value, an even vaster dark web exists. More on that after the holiday.
Cover image: Edgar Degas, A Cotton Office in New Orleans, 1873, Musée des Beaux-Arts de Pau