Qiming Star | Infinigence AI's Yu Wang: Exploring the Next Frontier of Edge Intelligence
The future of edge intelligence requires a two-way commitment between academia and industry. Only when hardware innovation and algorithmic breakthroughs form a closed loop can we truly unlock AI's potential to transform the physical world.

Editor's Note: Recently, Yu Wang, professor and department chair of the Department of Electronic Engineering at Tsinghua University and founder of Infinigence AI (a Qiming Venture Partners portfolio company), delivered a keynote speech at the inaugural ModelScope Developer Conference. He analyzed the core contradictions facing the development of on-device large language models and proposed a hardware-software co-design breakthrough path spanning the model, software, and hardware layers. On industry trends, he noted that edge devices will no longer rely solely on cloud-based inference, instead becoming capable of independently handling more complex tasks — thereby delivering more efficient solutions for embodied AI, autonomous driving, and other scenarios. Professor Wang also shared his vision for the future of intelligent systems and the two key pathways for building next-generation intelligent data infrastructure.
Reprinted with authorization from the Qiming Venture Partners WeChat official account.

Yu Wang, professor and department chair of the Department of Electronic Engineering at Tsinghua University, and founder of Infinigence AI
Recently, the inaugural ModelScope Developer Conference opened in Beijing. With the theme "Model Power Drives Leapfrog Progress, Open Source Fuels Innovation," the conference brought together over a thousand representatives from top global universities, research institutions, and technology companies. Yu Wang, professor and department chair of the Department of Electronic Engineering at Tsinghua University and founder of Infinigence AI (a Qiming Venture Partners portfolio company), was invited to attend and deliver a keynote speech titled "The Next Stop for Edge Intelligence: Hardware Innovation and Technical Breakthroughs and Challenges in On-Device AI." Centering on hardware-software co-design as the core methodology and embodied AI as the future anchor point, the speech focused on the challenges and solutions for deploying large models on edge devices, systematically presenting a panoramic view of on-device AI from challenges to technical breakthroughs.

Professor Wang began with the rapid development of edge intelligence in the AI 2.0 era, identifying a core contradiction in current on-device large model development: the ever-expanding size of cloud-based large models has created a massive gap with the limited compute power available on terminals. Meanwhile, traditional chip processes are hitting physical limits — hardware development simply cannot keep pace with model growth, making systematic breakthrough solutions urgently needed.

Accordingly, Professor Wang proposed a hardware-software co-design breakthrough path. At the model level, efficient small models should be built by integrating compute, data, and developer community resources. For example, Megrez-3B-Omni, an on-device full-modality large model jointly launched by Infinigence AI and the ModelScope community, achieved deep hardware-software co-optimization during R&D, delivering an excellent balance between inference speed and accuracy alongside industry-best inference performance at the time across vision, language, and audio multimodal tasks. At the software level, inference optimization software for general scenarios needs to be developed. Take Infinigence AI's Mizar intelligent terminal inference acceleration engine as an example: this independently controllable large model hardware-software adaptation platform, built for PCs, compute boxes, and other intelligent terminals, achieves substantial inference speed improvements across multiple application scenarios while significantly reducing power consumption and memory footprint — truly transforming AI capabilities into the intrinsic DNA of terminal devices. At the hardware level, customized accelerators and novel devices/computing paradigms are needed to break through traditional architecture limitations and substantially improve energy efficiency and processing speed. For instance, Infinigence AI's self-developed LPU (Large Model Processing Unit) IP, a dedicated large model inference processor, supports multimodal large models for text-to-text, text-to-image, and text-to-video generation, and can support 3D-stacked DRAM. On low-end process/low-compute FPGAs, it achieves compute power and energy efficiency surpassing high-end process/high-compute GPUs. Through "algorithm-software-architecture-process" co-optimization, it substantially leads mainstream domestic and international chips, achieving major improvements in on-device large model performance and energy efficiency.

Professor Wang noted that in the AI 2.0 era, models' knowledge density will continue to increase. Through pre-trained small models and lightweight techniques, on-device small models with 4o/o1-level capabilities can be built with model sizes reduced to 3-13B parameters, adapting to the hardware resource constraints of edge devices. For large model inference in broad on-device application scenarios, future inference demand will exceed 100 tokens/s to meet real-time application requirements. This trend indicates that edge devices will no longer rely solely on cloud-based inference, becoming capable of independently handling more complex tasks — thereby delivering more efficient solutions for embodied AI, autonomous driving, and other scenarios.

Looking ahead, Professor Wang believes the future direction of intelligent systems may lie in embodied AI and collective intelligence. Embodied AI transforms decision-making capabilities into real-world productivity through deployment in physical systems and environmental interaction; collective intelligence enhances overall system capabilities by collaboratively expanding perception, decision-making, and execution spaces. Therefore, building next-generation intelligent data infrastructure requires focusing on two pathways: first, optimizing compute infrastructure to make compute power accessible to all, supporting research and industrial development; second, establishing data infrastructure and supporting hardware to underpin the future development of embodied AI. "The future of edge intelligence requires mutual commitment from both academia and industry," Professor Wang emphasized. "Only when hardware innovation and algorithmic breakthroughs form a closed loop can we truly unleash AI's potential to transform the physical world."
Past Reviews

Founded in 2006, Qiming Venture Partners currently manages 11 USD funds and 7 RMB funds, with total assets under management reaching $9.5 billion. Since its inception, the firm has focused on investing in early- and growth-stage outstanding enterprises in Technology and Healthcare.
To date, Qiming Venture Partners has invested in over 580 high-growth innovative companies, of which more than 210 have gone public on the New York Stock Exchange, NASDAQ, Hong Kong Exchanges and Clearing Limited, Shanghai Stock Exchange, and Shenzhen Stock Exchange, or exited through M&A and other means. Over 80 portfolio companies have become recognized unicorns or super-unicorns.
Many companies in Qiming Venture Partners' portfolio have grown into the most influential players in their respective fields, including Xiaomi (01810.HK), Meituan (03690.HK), Bilibili (NASDAQ:BILI, 09626.HK), Zhihu (NYSE:ZH, 02390.HK), Roborock (688169.SH), UBTECH (09880.HK), WeRide (NASDAQ:WRD), Insta360 (688775.SH), Gan & Lee Pharmaceuticals (603087.SH), Tigermed (300347.SZ, 03347.HK), Zai Lab (NASDAQ:ZLAB, 09688.HK), CanSino Biologics (688185.SH, 06185.HK), Schrödinger (NASDAQ:SDGR), MicroPort EP MedTech (688617.SH), Sanyou Medical (688085.SH), Amoy Diagnostics (300685.SZ), Berry Genomics (000710.SZ), GenScript ProBio (688520.SH), Yuanxin Technology, Insilico Medicine, MediLink Therapeutics, LaNova Medicines, Zhipu AI, StepFun, Biren Technology, and others.