GENE-26.5 is going viral, hailed as the most mind-blowing demo in the field this year! Is it really?

What's worth watching about GENE-26.5 is the "Embodied Artificial Intelligence version of Harness + model" behind it.

What makes GENE-26.5 worth watching is the "embodied intelligence Harness + model" behind it.

👦🏻 Author: GaKi

🥷 Editor: Koji

🧑‍🎨 Layout: NCon

On May 7, Genesis AI officially released:

GENE-26.5.

26.5 stands for May 2026. According to the official blog, this is also the first public release of the GENE series.

Just days after it came out, it sparked considerable discussion in embodied intelligence circles at home and abroad. The reason is fairly straightforward: in this round of video demos, the robots started doing things that robot demos rarely pulled off before, with a fairly visible jump in capability.

For instance, cracking an egg with one hand, slicing a tomato with both hands, using the back of a knife to transfer cut tomatoes off the cutting board:

Grabbing a pipette, inserting a tip, twisting a small tube cap, organizing cable bundles, solving a Rubik's cube, clamping several objects of different sizes in one hand simultaneously — all dropped at once:

Over the past few years, humanoid robot video demos have become plentiful enough. Walking, dancing, moving boxes, folding clothes, making coffee.

Often, the visual impression gives a kind of false sense of maturity. But ordinary people's expectations of robots tend to be grounded in something very real:

When, exactly, will it actually help me get stuff done?

Getting stuff done is unremarkable for humans, but becomes extraordinarily tricky when it comes to robots.

Because real-world labor mostly doesn't end with "getting there." Sure, robots need to walk, stay balanced, and avoid obstacles — in robotics, this falls closer to Locomotion.

But actually finishing the job usually happens in the next step. It needs to pick things up, rotate them, cut them, tighten them, insert them, fold them, and finally place them in just the right spot. This is Manipulation — operation.

GENE-26.5's official blog makes this distinction too. In Locomotion, contact mostly serves to support the body; in Manipulation, contact is the task itself.

The blog mentions many more conceptual distinctions like this, so if you treat it as just a set of promotional demos, you might miss the more important substance.


After reading and organizing the original blog content, the biggest highlight of this GENE-26.5 release may be that the competition in robot foundation models has shifted direction: from foundation models to the "embodied intelligence Harness + model."

So in this article, we want to start from this "full-stack underlying system" and share our observations.

Who is Genesis AI?

Genesis AI as a company is worth a brief introduction first.

It's still very early-stage. From public information, Genesis AI only entered outside awareness this year, but the team composition is fairly typical: Xian Zhou's background leans more toward robotics and physics simulation, while Théophile Gervet has research experience at large model companies.

This isn't a robotics company that's been in the spotlight for many years.

But for its first public release, it put out the entire "embodied intelligence Harness + model" — the model itself, the dexterous hand, training gloves, the control system, simulation evaluation, and more.

This aligns with some "unusual signals" on the funding side.

On financing, Genesis's only officially disclosed round is a $105 million seed. The company states this round was co-led by Eclipse and Khosla Ventures, with participants including Bpifrance, HSG, Eric Schmidt, and Xavier Niel, among others.

Both TechCrunch and Reuters described this round as an "unusually large seed (giant $105M seed round)."

Reuters even noted it was comparable to the outsized seed round that Mistral AI set in France. This kind of capital allocation is actually very rare for a robotics company roughly one year old with no public customer list. The investors are clearly betting on its underlying platform value.

This may be the value of full-stack robotics.

What makes GENE-26.5 worth watching is the "embodied intelligence Harness + model" behind it

One sentence in the official blog is particularly worth noting:

If the goal is human-level manipulation, the solution can't stop at model training.

This sounds ordinary, but placed in today's embodied intelligence industry, it has something of a watershed flavor.

For the past two years, it's been easy to focus attention on VLAs. Vision, language, action — it sounds like once you plug a large model into a robot, progress will naturally follow, and from there you just need data and scaling.

But in reality, there's a great deal of murky, hard-to-articulate stuff between large models and robots' real-world performance.

For example: how data is collected, whether hardware can precisely express actions with "feel," whether the control system has latency, whether model output trajectories can fully feed back to motors, whether evaluation can run at scale. Any of these layers can fail.

Facing this complex set of problems, GENE-26.5 offers answers in its blog:

GENE-26.5's answer can be roughly summarized as an "embodied intelligence Harness," which breaks down into several layers.

【1】Is synthetic data really a dead end?

The chronic shortage of high-quality data in robotics is a widely acknowledged fact.

When Google built RT-1, 13 robots collected data for 17 months, ultimately yielding around 130,000 real robot episodes. Later, the DROID multi-institutional collaboration dataset mobilized 50 collectors across 564 scenes and 86 tasks, and still only amassed roughly 350 hours of real-machine interaction data.

This data is valuable, but it also confirms the inverse: real robot data doesn't naturally scale the way text, images, and video do.

Teleoperation can provide high-quality trajectories, but it's slow, expensive, hardware-dependent, and easily becomes "collecting data for data's sake." By contrast, human first-person video has a much higher scaling ceiling.

Meta's Ego4D already reached over 3,000 hours of first-person video, from 855 wearers across 9 countries.

So when GENE-26.5 emphasizes human-centric data, it's not simply swapping data sources. It's trying to circumvent the hardest part of robot data to scale: turning human actions from real work into physical experience that robots can learn from.

According to disclosures, its data engine draws from three sources: glove data, first-person video, and third-person video. The officially disclosed data scale has exceeded 200,000 hours.

This path is fairly interesting.

Because first-person video captures natural human behavior in real tasks, third-person video expands coverage, and glove data records hand movements and tactile information in finer detail.

Put simply, what robots should really learn may not be standardized trajectories in a lab one by one. What's more valuable is the "feel" that humans accumulate through long-term interaction with the physical world.

This feel is hard to write into rules, so data is needed.

In pre-training open-loop evaluation, they verified the foundation model's Scaling Law: larger models, more data, more compute — performance keeps improving.

As for real versus synthetic data, there's another interesting detail.

In GENE-26.5's official blog, synthetic data is barely presented as a core training route. What gets the most emphasis is glove data, first-person video, third-person video, and a small amount of robot data.

Simulation is certainly important, but in this blog post, it's positioned more in closed-loop evaluation — for faster, more stable model assessment — rather than as a primary training data source.

This is somewhat close to Physical Intelligence's π series public materials.

From π0 to π0.7, the emphasis is more on real robot data, web-scale vision-language pre-training, human data, and autonomous execution data. At least from public materials, synthetic data hasn't been written up as a fully validated core scaling path.

The underlying logic is the same: for contact-rich manipulation tasks, what people want most right now is probably still real-world human actions and robot interaction data.

This to some extent reiterates that at least for contact-rich robot manipulation tasks, synthetic data isn't ready to be the "main course" yet.

【2】Hardware.

But how does human data transfer to robots?

The human hand is extraordinarily complex. Finger length, joint structure, soft contact, skin friction, palm shape — all of these affect movement. No matter how rich human hand data is, if the robot hand's morphology diverges too far, a lot of information gets lost in transfer.

This is the so-called embodiment gap. So Genesis AI emphasizes Genesis Hand 1.0. According to official introduction, this hand targets near 1:1 human hand scale, with 20 active, backdrivable degrees of freedom, palm and fingers covered in soft material to approximate the soft-contact physics of human skin.

In the official blog it's referred to as proprietary hardware:

But in communities like Reddit, some voices suggest Genesis Hand 1.0's hardware appears to use Shenzhen Wuji Technology's WUJI HAND, a high-DOF biomimetic dexterous hand developed by WUJI TECH:

Wuji Technology's official account also reposted Genesis AI's GENE-26.5 release video in the community, calling it a "partner":

Setting aside this dexterous hand mystery, GENE-26.5's general thinking is: if robots are to learn from human hand movements, the closer the hardware is to humans, the smaller the loss in data transfer.

Traditional approaches often require complex motion retargeting, remapping human hand movements into robot joint space. This process loses detail and mixes in the robot hardware's own constraints as training signals.

So the dexterous hand here isn't the model's peripheral in the traditional sense — it becomes part of the data system itself.

On this basis, GENE-26.5 conducted extensive task testing:

【3】Control.

This layer is easier to overlook, but absolutely critical in robotics.

The action intentions output by the model can't directly become real robot movements. They must pass through controllers, communication layers, low-level execution, motor drivers. Any latency or error at any point distorts what the model intended.

The official blog gives some very concrete demos. The most typical case is a piano-playing video that GENE-26.5 released:

So in dexterous manipulation, low-level control is far more than an engineering detail.

If the control layer isn't clean enough, model training gets polluted by execution errors. The model thinks it output one action, but the robot actually performs another. Over time, the model may end up learning various patches in the hardware system rather than the physical laws behind human actions.

This is why I think an important layer in future embodied intelligence may be the harness layer.

Here, harness can be understood as the entire承接 system between the model and the real robot. It includes low-latency control, motion smoothing, real-time communication, execution feedback, state estimation, and all the middleware that lets model outputs actually land.

In the language model era, many capabilities can stop at the token level. But robots must turn intention into force, position, velocity, and contact — stopping at tokens means nothing.

Of course, some room for technical attribution should be left here. GENE-26.5's demos can't be simply read as "the model won" or "control won."

So some industry insiders have cautioned that if the underlying platform itself is a mature harmonic reducer robotic arm, then the final results can't all be attributed to self-developed control.

Additionally, quite a few people on X have raised questions, noting that this round of demos contains numerous cuts, and it's hard for outsiders to judge whether some tasks were completed continuously.

Such skepticism isn't surprising.

Robot demos have always been difficult to attribute technically: model, hand, robotic arm, control stack, task data, shot selection — all affect the final presentation.

So this also reminds us that robot demos can't be treated as strict benchmarks, and there's still considerable distance from real-world deployment.

【4】Then there's the model itself.

GENE-26.5's input isn't a single modality. It receives language, vision, proprioception, and touch, and outputs action trajectories. The official mention is that it uses Flow Matching to model the joint distribution of trajectories.

Intuitively, what robot models process is much messier than text models. If text output is wrong, you can retract it. If image generation hallucinates badly, you can redo it. But once the robot hand reaches out, the object may already be dropped, the liquid already spilled, the knife already touching where it shouldn't.

So robot-native models face a harder closed loop.

They can't finish in one shot and be done. Actions change the environment, and the model must continue adjusting based on the new environmental state.

【5】Finally, evaluation.

What also got noticed this time is the Genesis World simulation platform.

This platform targets Robotics, Embodied AI, and Physical AI, capable of handling different physical phenomena including rigid bodies, liquids, gases, deformable objects, thin shells, and granular materials.

The official blog mentions that in closed-loop evaluation, a single data point corresponds to 200 evaluation settings and over 150 hours of robot execution time. If put in the real world, the full picture would require massive human-machine time.

Connecting all of the above, GENE-26.5's emphasis becomes much clearer.

It demonstrates a possible embodied intelligence scaling path:

Pre-train using human operation data, then attempt to reduce transfer loss through human-proximate hardware, with low-level control designed to let model intentions land more stably, model input receiving multimodal supervision, and simulation and real feedback used to assist iteration.

Overall, this system can be viewed as a full-stack combination of "embodied intelligence Harness + model."

If this chain can operate smoothly, the competitive logic of robot foundation models may shift somewhat.

Future competition in the same arena will focus mainly on the naturalness of data, hardware's ability to express human actions, control layer latency, and the trustworthiness of simulation evaluation.

A lead in any layer could accelerate the pace of subsequent iterations.


Overall, GENE-26.5's release this time can be seen as an observation on the industry's current state:

Models can keep getting bigger, but if the hand's degrees of freedom can't express the actions, many capabilities won't emerge. If control layer latency is too high, the model will jitter when deployed on real machines. If evaluation can't scale, there's no telling how much the next version actually improved.

Going forward, embodied intelligence may increasingly resemble a flywheel: human data, human-like hardware, control systems, simulation evaluation, real-world feedback.

This "embodied intelligence Harness + model" is the next battleground, and the underlying element of the entire system.

Crossing is looking for independent contributors to write AI product and model reviews.

If you've written similar articles: Hands-on with PixVerse C1, Hands-on with LibTV, please contact zeo0811@gmail.com. Your email should include: ① personal introduction, ② AI review articles you've written.

We offer competitive compensation. Looking forward to observing and documenting the AI era with you 🎪