Infinigence AI Raises Nearly RMB 1 Billion in Total Funding, Aiming to Become the Go-To "Compute Operator" in the LLM Era | Oasis Vitality

Advisor on Vitality

Recently, Infinigence AI completed a nearly RMB 500 million Series A funding round, bringing its total raised to nearly RMB 1 billion within just one year and four months since its founding. The round was co-led by the National Social Security Fund's Zhongguancun Independent Innovation Special Fund (managed by Legend Capital), Qiming Venture Partners, and Hongtai APlus. Strategic investors following on include Lenovo Capital and Incubator Group, Xiaomi, and Softtone High-Tech; state-backed funds such as China Development Bank Technology Innovation, Shanghai AI Industry Investment Fund (managed by Shanghai Lingang Science and Technology Investment Management), and Xuhui Science and Technology Investment; and financial institutions including Shunwei Capital, Fortune Venture Capital, DT Capital Partners, Shangshi Capital, Senruo Yukun, Shenwan Hongyuan, and Zhengjing Capital. Oasis Capital participated in Infinigence AI's previous round. The fresh capital will go toward recruiting technical talent and advancing R&D to maintain its edge in hardware-software synergy and heterogeneous computing; accelerating commercialization of its Infini-AI heterogeneous cloud platform to keep it tightly aligned with market demand; and deepening ecosystem partnerships to activate heterogeneous cluster computing resources, building an AI compute foundation that supports "M models" on "N chips," serving as a "super amplifier" for AI model compute, with the goal of becoming the preferred "compute operator" in the large model era.

Infinigence AI's investor lineup

Lixia Xia, co-founder and CEO of Infinigence AI, said: "We're grateful for the confidence from so many investors, which gives us greater conviction on this entrepreneurial journey blessed with the right timing, favorable conditions, and strong team alignment. The new '80-20 rule' brought by the AI 2.0 wave — where the Transformer architecture has unified the technical paradigm, meaning solving just 20% of key technical problems can support 80% of vertical scenario generalization — presents a rare opportunity for standardizing and scaling hardware-software joint optimization. The supply-demand tension and uneven resource distribution in China's compute ecosystem creates a timely opening for us to pull upstream and downstream partners together to achieve efficient integration of diverse heterogeneous compute. And our 'composite' team, rooted in Tsinghua University's Department of Electronic Engineering with over a decade of technical accumulation and rich industry experience, has become a talent 'gravity well' in the AI field, forming Infinigence AI's distinctive human capital advantage."

Leveraging Hardware-Software Synergy and Heterogeneous Computing to Become a "Super Amplifier" for AI Model Compute

The actual industrial scale that large models can support depends on the practically available compute for AI models — a domain with higher barriers, scarcer players, and greater value. Based on its deep understanding and long-term practice in the AI industry, Infinigence AI made an early judgment that practically available compute for large models depends not only on a chip's theoretical compute capacity, but can also be amplified through optimization coefficients that improve utilization efficiency, and through cluster scale that expands overall compute capacity. From this, the company proposed the formula: Chip Compute × Optimization Coefficient (Hardware-Software Synergy) × Cluster Scale (Heterogeneous Computing) = AI Model Compute. Following this formula, Infinigence AI will continuously improve chip utilization in large model tasks through hardware-software joint optimization, and expand overall industry compute supply through heterogeneous compute adaptation technology that raises cluster utilization.

Infinigence AI's AI Model Compute formula

On the hardware-software optimization front, Infinigence AI's self-developed inference acceleration technology FlashDecoding++ substantially improves utilization on mainstream and heterogeneous hardware, surpassing previous state-of-the-art results. The company has completed adaptation of multiple mainstream open-source large models across more than 10 compute cards including AMD, Huawei Ascend, Biren, Cambricon, Suiyuan, Hygon, Tianshu Zhixin, MetaX, Moore Threads, and NVIDIA, achieving industry-leading inference acceleration on some cards — efficiently meeting surging large model inference demand across industries. Based on these optimization results, Infinigence AI has signed a strategic partnership with AMD to jointly advance commercial AI application performance.

In heterogeneous compute adaptation, Infinigence AI also possesses rare industry capabilities in heterogeneous adaptation and clustering. Its large-scale heterogeneous distributed hybrid training system HETHUB, released in July, marked the industry's first heterogeneous compute hybrid training at kilo-card scale across six chip types in a "4+2" configuration — Huawei Ascend, Tianshu Zhixin, MetaX, and Moore Threads plus AMD and NVIDIA — achieving peak cluster utilization of 97.6%, averaging about 30% above baseline solutions. This means Infinigence AI can compress total training time by 30% under the same multi-chip datacenter or cluster conditions.

Building the Infini-AI Heterogeneous Cloud Platform: From Heterogeneous Compute Utilization to Large Model Application Development

Internationally, the model layer and chip layer have gradually converged toward a "dual-headed consolidation" pattern, while China's model and chip layers continue to show an "M×N" structure of "M models" and "N chips." However, different hardware platforms require adaptation to different software stacks and toolchains, and heterogeneous chips have long suffered from incompatible "ecological silos." As more domestic heterogeneous compute chips are deployed in regional compute clusters nationwide, the difficulty of effectively utilizing heterogeneous compute has become increasingly severe, gradually emerging as a bottleneck for China's large model industry development.

Leveraging its hardware-software synergy and heterogeneous computing advantages, Infinigence AI has built the Infini-AI heterogeneous cloud platform on a foundation of diverse chip compute. The platform is downward-compatible with heterogeneous compute chips, effectively activating dormant heterogeneous compute across the country, with currently operational compute resources covering 15 cities. Additionally, the Infini-AI platform includes a one-stop AI platform (AIStudio) and a large model service platform (GenStudio). AIStudio provides machine learning developers with cost-effective tools for development and debugging, distributed training, and high-performance inference, covering the full lifecycle from data hosting, code development, model training, to model deployment. GenStudio provides large model application developers with high-performance, easy-to-use, secure and reliable multi-scenario large model services, comprehensively covering the full workflow from large model development to service deployment, effectively lowering development costs and barriers. Since launch, leading large model industry customers including Moonshot AI, LiblibAI, Liepin, Shengshu Technology, and Zhipu AI have stably used heterogeneous compute on the Infini-AI platform and benefited from Infinigence AI's large model development toolchain services.

Becoming the Preferred Compute Operator in the Large Model Era: Prospering the Heterogeneous Compute Ecosystem and Accelerating AGI Accessibility

The Infini-AI heterogeneous cloud platform not only helps downstream customers easily abstract away hardware differences and seamlessly access powerful underlying heterogeneous compute capabilities, but also works to break China's heterogeneous compute ecosystem deadlock, accelerating upper-layer application migration toward heterogeneous compute foundations, effectively integrating and expanding the scale of practically available compute for the domestic large model industry — truly converting heterogeneous compute into accessible, sufficient, and usable large-scale compute, helping build a localized heterogeneous compute ecosystem with Chinese characteristics. Following its compute utilization improvement approach and combining its hardware-software joint optimization capabilities, Infinigence AI has also made early moves in edge-side large models and LPU IP, aiming to build closed-loop "edge model + edge chip" capabilities. The company firmly believes in the inevitable trend of rapid growth and application explosion in edge scenarios, where AI PCs and AI phones will become important human-computer interaction interfaces, helping every terminal achieve AGI-level intelligent emergence.

With the mission of "Unleashing Boundless Compute, Making AGI Within Reach," Infinigence AI is committed to becoming the preferred "compute operator" in the large model era. It is currently aggressively pursuing strategic partnerships with the most valuable customers in the industry chain, then expanding to broader markets for standardized, batch replication to build scale advantages. Through activating diverse heterogeneous compute and hardware-software joint optimization, Infinigence AI aims to reduce large model deployment costs by 10,000×, making it as accessible as "water, electricity, and gas" — a new quality productive force within easy reach and broad benefit for the industry, accelerating the democratization of AGI.