MaKe | From Spring Festival Gala to NVIDIA GTC: Four MaHui Members Become "Global AI A-Listers"

On March 20, Beijing time, NVIDIA's GTC 2026 conference — dubbed the "Super Bowl of American AI" — concluded successfully in the United States, bringing together the world's top AI talent: tens of thousands of developers, researchers, and business leaders from more than 190 countries.

On March 20, Beijing time, NVIDIA GTC 2026 — dubbed the "Super Bowl of American AI" — wrapped up in the United States, bringing together the world's top AI talent: tens of thousands of developers, researchers, and business leaders from over 190 countries.

At this premier global AI industry event, four MaHui member companies in AI and robotics took the stage: Moonshot AI founder Zhilin Yang attended in person, delivering the first systematic explanation of Kimi K2.5's underlying technical reconstruction; Unitree CEO Xingxing Wang gave a keynote speech assessing the critical inflection points for embodied intelligence; Galaxy Universal's Galbot G1 became the only Chinese embodied intelligence real-world application case featured in NVIDIA founder and CEO Jensen Huang's keynote; and Li Auto unveiled its next-generation autonomous driving foundation model MindVLA-o1, achieving architectural breakthroughs through a novel 3D vision encoder.

From large model foundational research to scaled embodied intelligence deployment, from autonomous driving to robotics industry ecosystems, MaHui member companies demonstrated — each through their own technical depth — the pivotal role of China's frontier technology forces in this transformation.

Moonshot AI: Rebuilding Large Model Foundations, Unlocking New Breakthroughs Through "Old Tech"

Moonshot AI founder Zhilin Yang delivered a keynote at GTC titled How We Scaled Kimi K2.5, disclosing for the first time the technical roadmap behind Kimi K2.5.

Yang distilled Kimi's evolution into three dimensions: token efficiency, long-context capability, and agent swarms. He noted that industry-standard technical approaches currently in widespread use are essentially products from eight or nine years ago, increasingly becoming bottlenecks for scaling — breaking free from this historical baggage is key to pushing the frontier of intelligence.

Regarding the Adam optimizer, which became an industry standard in 2014, the Kimi team validated Muon's potential while identifying stability issues at hyper-scale training. They developed and open-sourced the MuonClip optimizer, solving the logits explosion problem with 2× computational efficiency over AdamW. For the full-attention mechanism born in 2017, the team proposed KDA-based hybrid linear topology architecture Kimi Linear, boosting decoding speed by 5–6× at 128K and even 1M ultra-long contexts. For decade-old residual connections, Kimi introduced Attention Residuals, enabling each layer to selectively aggregate information based on input content — prompting OpenAI co-founder Andrej Karpathy to reflect and publish publicly.

In cross-modal research, Yang shared a significant finding: Vision RL can boost pure-text benchmark performance by approximately 2.1%, confirming the positive transfer from spatial reasoning and visual logic to general cognitive capabilities. Kimi will continue its open-source path, contributing these foundational innovations to the community.

Unitree: Industry at the Tipping Point, Three Bottlenecks to Overcome

Unitree founder and CEO Xingxing Wang delivered a virtual keynote at GTC titled How to Cross the ChatGPT Moment for the Intelligent Industry, systematically assessing the industry's current stage and core challenges.

Wang's assessment: market enthusiasm continues to climb, but while general embodied intelligence models with strong generalization capabilities that can stably execute tasks in raw environments have yet to fully emerge, the industry as a whole remains "before the tipping point." The next one to three years will be the critical window for whether breakthroughs can be achieved.

He offered a definition for "embodied intelligence certainty": robots capable of completing 80% of tasks via language instruction in 80% of novel scenarios, without prior training or map collection — once this threshold is reached, the industry transitions from demonstration to scaled application.

Toward this goal, Wang identified three core bottlenecks: first, insufficient model expressiveness, making it difficult to generate and execute complex non-routine actions; second, extreme scarcity of real-world robot data, requiring dramatically improved utilization of video and simulation data; third, reinforcement learning's lack of scalable reuse mechanisms, where training outcomes fail to accumulate and transfer across tasks.

On technical approach, he is more optimistic about world models and video generation models, seeing higher ceilings and broader data sources, though virtual-to-real synchronization remains the most critical unsolved problem. Wang emphasized that embodied intelligence is a global collaborative endeavor — whichever company breaks through first, it will be a historic moment for the entire industry.

Galaxy Universal: Only Chinese Embodied Intelligence Real-World Case Enters Jensen Huang's Keynote

Galaxy Universal's Galbot G1 became the only Chinese embodied intelligence real-world application case featured in NVIDIA CEO Jensen Huang's keynote, completing its leap from "China's national stage" to "the global technology stage."

The case Huang showcased came from Galbot G1's joint deployment with NVIDIA at U.S. medical AI company PeritasAI: the robot autonomously navigated a high-density pharmacy, accurately identifying and retrieving surgical instruments to assist with complex pre-procedure preparation. This not only validated Galaxy Universal's industrial-grade reliability but also demonstrated how Chinese embodied intelligence technology can deeply adapt to the world's top industrial ecosystems.

At the exhibition booth, Galbot G1's "walnut twirling" demo became a crowd favorite. Walnut twirling is considered a "world-class challenge" in dexterous manipulation, and cracking it hinged on the dexterous hand dynamics neural cerebellar model integrated in Galaxy's AstraBrain. The training path followed a paradigm of "practicing moves in the virtual world, finding the feel in the real world": first generating training data in simulation covering hundreds of walnut sizes and friction coefficients, then using fingertip torque sensors to achieve real-world physical tactility — with task generalization rates already exceeding 98% based on this successful technical approach.

During the conference, founder and CTO He Wang systematically explained Galaxy Universal's full-stack "hardware-data-model" technical architecture, with particular emphasis on the billion-scale embodied dataset "AstraSynth" — physics-simulation-based, synthetic-data-dominant with real-robot data as supplement, achieving real-robot training efficiency 1000× higher than Tesla's, with models trained on this dataset reaching 99% task success rates.

Li Auto: MindVLA-o1 Unveiled, Comprehensive Autonomous Driving Upgrade

At GTC 2026, Li Auto foundation model lead Kun Zhan unveiled the next-generation autonomous driving foundation model MindVLA-o1, achieving a core breakthrough at the foundation level: native 3D ViT — a true three-dimensional vision encoder.

MindVLA-o1 fuses camera and LiDAR data for precise 3D perception; introduces a latent world model capable of pre-enacting future scenarios to enhance decision-making; adopts VLA-MoE architecture to generate smooth, physically-plausible driving trajectories; leverages a self-developed world simulator to dramatically accelerate iteration and model evolution; and through hardware-software co-optimization, significantly compresses deployment cycles for the full large model on in-vehicle chips.

Through five technical innovations — 3D spatial understanding, multimodal reasoning, unified behavior generation, closed-loop reinforcement learning, and hardware-software co-design — the model comprehensively upgrades autonomous driving capabilities.