Morning Star | Shengshu Technology's Luo Yihang: From Understanding Language to Understanding the World, General World Model Opens a New Chapter in AI Development | WAIC 2026

AI is moving from understanding and generating the digital world to understanding and acting in the physical world, and world models are the foundational infrastructure driving this transformation.

The "Qiming Venture Partners · Entrepreneurship and Investment Forum — From Compute Origins to Application Deployment" at WAIC 2026, hosted by Qiming Venture Partners, was successfully held on July 19 at the Shanghai World Expo Center. As a global leader in multimodal models and world model R&D, Shengshu Technology open-sourced in 2025 the first unified-architecture world action model based on video generation foundation models. Luo Yihang, co-founder and CEO of Shengshu Technology, delivered a keynote speech titled General World Model: Bridging the Virtual and Physical, Building the Foundation for Physical Intelligence, sharing his team's latest explorations in multimodal generation and world models. Luo Yihang, co-founder and CEO of Shengshu Technology

"AI is evolving from understanding and generating digital worlds to understanding and acting in the real physical world, and the world model is the core foundation of this transformation," he stated. "Shengshu Technology is the world's first company to unify digital and physical worlds through a general world model, building a comprehensive framework spanning understanding, generation, and action — providing intelligence for both digital creation and physical execution."

01/

World Models: Moving AI from Interpreting Information to Participating in the Real World

The world model is not an entirely new concept, but with the rapid advancement of video generation, multimodal models, and embodied AI, it is moving from academic research to the center of industry. In Luo's view, language models primarily address the relationship between humans and language, knowledge, and information, while world models are oriented toward the interaction between intelligent agents and the world as a whole.

He pointed out that human intelligence is not manifested solely in language processing — large areas of the brain are involved in visual perception, action planning, and decision-making. The core of this intelligence is the perception of the physical world and the ability to make action decisions based on environmental changes.

"If we rely solely on large language models to understand the world, we're using only a small fraction of human intelligence. In the era of foundation models, we need a world model that reaches the paradigm level of LLMs."

Video models need to understand how people and objects move, robots need to predict the outcomes of actions, and autonomous driving systems need to continuously perceive and make real-time decisions in complex environments. Behind these different scenarios lies a similar underlying capability: perceiving the environment, predicting changes, and taking action.

Therefore, a world model is not equivalent to any single-point technology such as video generation, 3D reconstruction, or robot control. A complete world model needs to form a continuously operating dynamic loop.

Perception, prediction, and action are not three separate modules; they should be built upon a unified world representation. The model must not only know what the world looks like, but also understand how it changes and how actions transform it.

02/

Toward Generalization: The Necessary Path for Scalable World Model Deployment

Currently, many world models are still built around specific scenarios such as household robots, industrial operations, and autonomous driving. But if every type of robot and every scenario requires training a separate model, the costs of data collection, training, and deployment will be impossible to scale.

A key lesson from large language models is to first establish general capabilities through large-scale pre-training, then transfer them to different tasks and scenarios. World models need to undergo a similar evolution.

They need to learn spatial structures, motion laws, causal relationships, and action logic from massive visual and interaction data, then transfer these capabilities to different digital content, robot embodiments, and real-world tasks.

"World models must not only form a closed loop of perception, prediction, and action, but also be sufficiently general and generalizable, like language models."

Different intelligent agents may have different forms, and digital and physical worlds have different output modalities, but the underlying intelligence driving them should be unified. The real competition in world models is not just about who achieves higher success rates on individual tasks, but about who can build a general foundation model with scale effects, transferability, and emergent intelligence potential.

03/

Video Is Becoming a Critical Entry Point for Understanding the World

To build world models, AI must first see and understand the world. Text is an abstract representation of reality; images capture a moment in time; but video continuously records how the world changes. How people move, how objects respond to forces, how spaces transform, how one action leads to subsequent results — all this information is embedded in continuous footage.

When a video model predicts the next frame, it is learning not just pixel generation but also the transition relationships between world states. Thus, the value of video models is extending beyond content production itself. Their capacity to model temporal-spatial structures, motion laws, and physical consistency is becoming an important foundation for world models.

On this foundation, world models need to further absorb simulation data, first-person perspective data, human operation data, and real robot data, translating understanding of the world into real action. Shengshu Technology has thus built a unified general world model foundation: first establishing the ability to understand and predict the world, then decoding differently for digital and physical spaces.

04/

One Body, Two Wings: The Same General Foundation Driving Digital Generation and Physical Action

Centered on a unified world model foundation, Shengshu Technology has developed a product system covering world generation, real-time interaction, and world action.

The Vidu Q series is oriented toward digital content generation, continuously enhancing the model's understanding and prediction of characters, scenes, motion, and physical laws while adding creative and imaginative capabilities; the Vidu S series pushes video from offline generation toward real-time, continuous interaction, unlocking scenarios where humans interact in real time with digital content in physical space, such as companionship, interaction, and gaming; Motubrain faces the physical world, translating the model's understanding of environments and tasks into robot actions to achieve real-time interaction in physical space.

The three product lines serve different scenarios, but share a unified modeling capability for time-space, causality, and action. This means video generation and embodied intelligence are not two independent technical paths.

"Agents in digital and physical worlds can vary enormously, but the intelligence behind them should be general, generalizable, and unified."

In the digital world, world models can push video, games, and virtual characters from static content toward real-time, continuous, interactive dynamic environments; in the physical world, they can enable robots to break free from fixed instructions and preset workflows, upgrading to intelligent agents that understand goals, plan autonomously, and adapt to environmental changes.

Whether virtual characters, game agents, content agents, or robots in real environments, all need to understand their surroundings, continuously update their state, and make decisions based on feedback. The difference is that the former translates this cognition into content generation and real-time interaction, while the latter further translates it into predictions of action outcomes and executable physical movements.

They ultimately point to the same core question: how can machines build cognition of the world and, based on this cognition, complete generation, interaction, and action.

05/

From Generating Worlds to Acting in Worlds

Large language models enabled machines to begin understanding and using language; video generation models gave machines visual creation capabilities; and what general world models aim to advance is enabling machines to further understand the world, predict it, and act within it.

In Luo's view, robots of different forms in the future will all become agents in the physical world, achieving autonomous planning and execution just as digital agents do. And the real breakthrough point for achieving this goal remains the model itself.

"The general world model is the most promising technical path for physical intelligence. In the future, agents in digital and physical worlds will be diverse, but the intelligence behind them will be general — communicating like humans, creating like humans, and acting like humans."

From processing information to understanding change; from generating content to taking action — artificial intelligence is entering a new stage of development. And the general world model is the key infrastructure connecting digital and physical intelligence, driving AI to truly enter the real world.


Past Highlights

Qiming Stars | Shengshu Technology Completes New Round of Hundreds of Millions of RMB Financing, Led by Qiming Venture Partners

Qiming Headlines | WAIC Qiming Venture Partners · Entrepreneurship and Investment Forum — From Compute Origins to Application Deployment Successfully Held

Qiming Stars | Highlight Moments! Qiming Venture Partners and Portfolio Companies Win Major Awards at WAIC 2026


Founded in 2006, Qiming Venture Partners currently manages 11 USD funds and 7 RMB funds, with total assets under management reaching $9.5 billion. Since its inception, the firm has focused on investing in outstanding early and growth-stage companies in Technology and Healthcare innovation.

To date, Qiming Venture Partners has invested in over 580 high-growth innovative companies, of which more than 210 have listed on the New York Stock Exchange, NASDAQ, Hong Kong Exchanges and Clearing Limited, Shanghai Stock Exchange, and Shenzhen Stock Exchange, or exited through M&A and other means. More than 80 portfolio companies have become recognized unicorns or super-unicorns.

Many companies in Qiming Venture Partners' portfolio have grown into the most influential players in their respective fields, including Xiaomi (01810.HK), Meituan (03690.HK), Bilibili (NASDAQ:BILI, 09626.HK), Zhihu (NYSE:ZH, 02390.HK), Roborock (688169.SH), Hesai Technology (NASDAQ:HSAI, 02525.HK), UBTECH (09880.HK), WeRide (NASDAQ:WRD, 00800.HK), HyperStrong (688411.SH), Insta360 (688775.SH), Unisound (09678.HK), Biren Technology (06082.HK), Zhipu AI (02513.HK), Gan & Lee Pharmaceuticals (603087.SH), Tigermed (300347.SZ, 03347.HK), Zai Lab (NASDAQ:ZLAB, 09688.HK), CanSino Biologics (688185.SH, 06185.HK), Schrödinger (NASDAQ:SDGR), MicroPort EP MedTech (688617.SH), Sanyou Medical (688085.SH), Amoy Diagnostics (300685.SZ), SinoCellTech (688520.SH), Insilico Medicine (03696.HK), Hope Medicine, Yuanxin Technology, MediLink Therapeutics, LaNova Medicines, StepFun, among others.