Morning Star | Lixue Xia on Building the 'Most Efficient Token Factory'

China will become the world's "token factory" in the future. In the past, Made in China meant manufacturing. Now, it's AI Made in China.

Editor's Note: As the AI industry enters the Token Economy era, compute infrastructure has become the central hub for industrial deployment and value circulation. Infinigence AI, a portfolio company of Qiming Venture Partners and a benchmark enterprise in the AGI infrastructure track, recently completed a new funding round exceeding RMB 700 million, bringing its total raised to over RMB 2.2 billion. As the company enters a steady implementation phase, it has anchored itself to a new positioning as "the most efficient Token factory" — moving beyond the singular perspective of an AI acceleration service provider. Through a multi-heterogeneous, software-hardware coordinated technical approach, it has solved the mixed scheduling challenge of domestic and overseas compute, reshaping the evaluation standards for AI infrastructure efficiency. Standing at the inflection point of industry cash flow closure and supply-demand transformation, future Token consumption will continue to grow rapidly, a pricing cycle is approaching, and with the empowerment of various model ecosystems, domestic compute is accelerating into the commercialization mainstream. China is also positioned to leverage its energy and industrial chain advantages to become a global hub for AI Token production. Infinigence AI co-founder and CEO Lixue Xia predicts: "In the past, Made in China meant manufacturing. Now it's AI Made in China."

This article is an exclusive interview with Lixue Xia by China Entrepreneur, providing an in-depth analysis of the company's technical layout, business model, and industry trend assessments.

Compared with the difficult exploration of the previous two years, Infinigence AI co-founder and CEO Lixue Xia has entered a state of "low-drag supersonic" entrepreneurship over the past year.

"In the first two years, Token scale hadn't truly taken off yet. We had to face direction selection, rhythm layout, and other problems with no standard answers. Now the track and demand are much clearer than before. Although uncertainty still exists, what's different is that many things can start to land and be measured. The initial judgments are being validated bit by bit, and we can run at full speed toward clear goals — this is already a very ideal entrepreneurial rhythm."

On May 7, Infinigence AI, a Qiming Venture Partners portfolio company and AGI infrastructure service provider, announced that it had raised over RMB 700 million in new funding.

In an exclusive interview with China Entrepreneur, Lixue Xia stated: "The company started this funding round in the second half of 2025. At that time, we believed that model capabilities had broken through the commercialization threshold, and large models were transforming from good technology to good products, and then to good industries. We predicted at that time that we should stockpile more ammunition."

Lixue Xia judges that the AI industry has entered the cash flow closure stage. Revenue earned by enterprises can be reinvested into production, manufacturing and outputting high-value Tokens, which are then monetized through commercialization to form sustainable returns, achieving a profitable, cyclical, and expandable mature industrial chain.

And within the entire AI industry chain, the Infra layer plays a key role. It is the "Token factory" that integrates chips and energy, encompassing hardware facilities such as data centers, cooling systems, and network architecture. It is also the critical layer in the five-layer cake proposed by NVIDIA CEO Jensen Huang — energy, chips, infrastructure, models, and applications.

Lixue Xia believes that in a market where supply falls short of demand, compute may remain insufficient for a long time. "It's not the richest people who will occupy the highest industry positions, but those who best understand optimization."

Infinigence AI's previous funding round was six months ago, in November 2025, when the company completed a RMB 500 million Series A+ round. Going further back, in August 2024 it announced the completion of a nearly RMB 500 million Series A round. Combined with its angel round, Infinigence AI's publicly disclosed total funding has exceeded RMB 2.2 billion.

Infinigence AI was founded in May 2023. Its founder is Yu Wang, a professor at the Department of Electronic Engineering, Tsinghua University. Lixue Xia, co-founder and chief scientist Guohao Dai, and chief technology officer Boxun Li were all students of Wang.

In a speech in September 2025, Yu Wang mentioned that optimizing Token efficiency per unit of energy consumption will be the core proposition of infrastructure and system design in the AI 2.0 era. The core metric for evaluating infrastructure efficiency has changed — the traditional "computations per joule (TOPS/J)" is being replaced by "effective Tokens processed per joule (Tokens/J)."

Infinigence AI has set its sights on becoming "the most efficient Token factory" and a Token economy hub. This positioning is clearer and more focused than the "AI acceleration" and "shovel seller" concepts the company previously proposed.

Facing the industry reality where domestic chips coexist with overseas high-end compute, with uneven performance and ecosystems, Infinigence AI has carved out a unique path centered on multi-heterogeneous, software-hardware coordinated, and autonomous AI approaches. Currently, the Infinigence AI platform has integrated leading large models including Moonshot AI, Zhipu AI, DeepSeek, Qwen, and MiniMax.

Lixue Xia candidly states that domestic chips don't need to pursue one-step replacement of overseas solutions, but rather to improve while running, optimize while using. True efficiency breakthroughs come from putting different compute resources where they excel most.

Through heterogeneous mixed training and inference, Infinigence AI has achieved mixed usage of NVIDIA and domestic chips, reasonably splitting large model prefill and decoding, training and inference, complex operators and routine computation, maximizing the value of every unit of compute. This allows large model vendors to gradually increase the proportion of domestic chip承载 without losing 3-6 months of iteration cycles.

The Token-centered business model is exciting the entire AI industry. On this, Lixue Xia provides several key data points: First, from the end of last year to the end of April this year, Infinigence AI's MaaS platform model inference Token usage grew 20-fold, with growth primarily coming from large-scale commercialization and the most intelligent models.

Second, in the next six months, Token usage will be in short supply and will maintain this growth rate.

Third, a Token price increase wave is coming. Lixue Xia says: "When you叠加 price increases and cost reductions, you'll find this is a break-even line issue: Token prices rising while costs fall will turn certain loss-making businesses in some areas into profitable ones. So AI's ultimate break-even line can achieve positive returns in more scenarios. Once positive returns are achieved, the supply-demand flywheel will start spinning."

The release of DeepSeek-V4 has also brought a critical inflection point to this path. The Pro and Flash dual versions released with V4 balance extreme performance with accessible costs, providing the best vehicle for domestic chips to achieve scaled deployment.

Lixue Xia believes that DeepSeek's true value lies not only in hardware-friendly optimization, but in using open-source ecosystems and tiered product strategies to truly bring domestic chips into the commercialization mainstream. More domestic chips are expected to be efficiently activated, used at scale, and continuously iterated.

He predicts that with factors such as electricity and model cost-performance, China will become the world's "Token production factory." "In the past, Made in China meant manufacturing. Now it's AI Made in China."

The following are selected excerpts from the conversation, reprinted with authorization on the Qiming Venture Partners WeChat official account.

01 The Industry Is Still Growing at 10x Speed

China Entrepreneur: As the Infra layer between upstream and downstream in the industry chain, do you expect competition to be particularly fierce this year, with some players being eliminated?

Lixue Xia: I don't think so. For an industry to eliminate companies, the core reason is demand saturation leading to cutthroat competition. But currently AI industry demand is not only not saturated, it's growing substantially, driving both upstream and downstream. Since the entire industry has a bigger pie to divide, every stage and every layer in between will have a larger market to share.

Specifically for the Infra track, its value lies in extracting greater production capacity from underlying resources. Only if domestic chips were sufficient now could we talk about saturation. But now and for a long time to come, it's in a state of extreme shortage.

Jensen Huang described infrastructure in his speech, placing it within a five-layer cake system of "energy, chips, infrastructure, models, and applications." Everyone can feel this is a trillion-dollar market. Achieving tens of billions or even hundreds of billions in revenue within it would be a very good state.

The greater focus now should be on whether our technology can be further improved, whether we can provide industrial value, whether customers recognize our value, and whether we can continuously launch better product iterations. These matters are far more important than competitive relationships.

China Entrepreneur: So AI Infra is still a marathon-style competition where everyone's catching up?

Lixue Xia: It may not even be competition yet. The pie is big enough that you can圈 a piece of land anywhere and first build your own city. Everyone is still choosing which market segment to address, far from the stage of needing to fight with bayonets.

China Entrepreneur: Within the entire Infra layer, what is Infinigence AI's core value proposition compared to competitors?

Lixue Xia: At this point in time, those occupying the highest industry positions aren't the richest, but those who best understand optimization.

The underlying logic of the Token factory is optimizing for Tokens produced per unit of resource and the productivity level each Token brings. Therefore, we've always operated Infra by connecting technical value with industrial value.

In Jensen Huang's five-layer AI industry structure, infrastructure sits in the middle layer. Upward, it needs to incorporate algorithm and application know-how, business traffic, scale, and distribution into the optimization space; downward, it needs to consider chip architecture and even energy construction. So infrastructure is a layer that requires very strong full-stack technical capabilities.

We have a strong algorithm team and a strong hardware-biased team. We can both preserve the best and most important computations from algorithms and make these computations run smoothly on the hardware structures they excel at. The connection between these two is the core competitiveness of Infrastructure and also the most unique aspect of Infinigence AI in the industry.

From our founding, we've positioned ourselves on software-hardware coordination and multi-heterogeneous core technology, doing joint optimization between M models and N chips. All of this declares our stance: to squeeze every unit of compute and every second on every chip to the extreme — this is the value we bring to this industry.

China Entrepreneur: A domestic large model vendor said that if algorithm iteration needs to适配 domestic chips, it would lose at least 3 to 6 months. Based on domestic compute and heterogeneous chips, how do you minimize this time gap and achieve advanced performance or efficiency?

Lixue Xia: The most important thing is to make reasonable division and拆解 of tasks. Taking DeepSeek-V4 as an example, it has two versions, Pro and Flash, with parameter counts of 1.6T and 284B respectively, facing different application scenarios with different division of labor.

Our multi-heterogeneous approach, besides mixing Chip A and Chip B domestically, more importantly enables mixing domestic chips with NVIDIA chips. They also form a division of labor, extracting sub-tasks that domestic chips can handle from tasks that are large, heavy, and demand-maximizing for operator libraries, performance, and bandwidth; while complex tasks that domestic chips aren't yet good at and need time to address are handed to NVIDIA chips.

We've also done even harder things before: splitting training tasks to let two types of chips cooperate to complete training, with a combined degradation rate below 3%, achieving 97.6% mixed training efficiency.

Now, we can split large model inference, such as Prefill and Decode, onto two different chips for heterogeneous PD separation.

This is Infinigence AI's value: through task拆解, letting each unit of compute do what it excels at, not making users "wait." For large models, "waiting" is a terrifying opportunity cost. As long as they don't have to wait, they can improve while running.

China Entrepreneur: Won't improving while running affect customer experience?

Lixue Xia: First, customers need to physically recognize that domestic chips are usable. Only after improving while running can there be direction for improvement, because the Token factory itself has an important flywheel: the more business it runs, the more optimization space it can discover.

For us, the value of the entire Token factory is that after accumulating better optimization, it provides more cost-effective Tokens. Getting more people to use it, the flywheel starts spinning.

So, the ecological闭环 of domestic chips is very important. The core value Infinigence AI provides is that through task splitting, we打通 communication libraries between chips, enabling chip fault tolerance and SLA to stably reach usability, and finally deliver uniformly.

China Entrepreneur: How do you evaluate the release of DeepSeek-V4?

Lixue Xia: First, it's a quite usable open-source model. DeepSeek continues to promote the open-source model ecosystem, and we'll certainly see more applications爆发 in the open-source ecosystem.

Second, the V4 model has quite a lot of optimization technology and also considers hardware friendliness. For example, its optimizations for Cache are well done, enabling current hardware to support very long Token contexts. Future Token usage will grow further and faster, bringing more demand to the Infra layer.

Third, simultaneously releasing Pro and Flash models is healthy model planning. The larger Pro model pursues AGI realization; the Flash model, which is usable but not as costly, can better put domestic chips to use.

Users also vote with their feet. The reason DeepSeek spent effort releasing the Flash model is that they must have discovered this version can also meet many industry needs. This means the entire AI track is moving toward a healthier state, where it's not only the most cutting-edge models that people are willing to use — 200B-scale models also have many takers.

02/

Letting Domestic Chips Produce Tokens at Maximum Efficiency

China Entrepreneur: Infinigence AI is building the most efficient "Token factory." You previously did AI acceleration, the so-called "shovel seller" — is this an upgrade in positioning?

Lixue Xia: I don't know if "upgrade" is too strong a word, but the core of our technology hasn't changed. We've always been researching how to maximize the value of every unit of compute.

When more practitioners were in the model training period, what we provided was how to better use existing resources, more like a "shovel passer" job.

Now, the technology's own goals haven't changed, but the business has changed, and product forms and business models will naturally transform accordingly: massive demand comes from Agents and customers across industries. At this point in time, only providing an "engine," customers may not necessarily be able to assemble the best "production line," so it's better for us to build the entire "production line" ourselves.

Since Token is already a commodity form with volume, pricing, and certain standardization trends, we can fully leverage our technical advantages to provide the market with the most efficient, high-quality Token production capacity.

China Entrepreneur: Now, is your biggest guiding target Token?

Lixue Xia: It's Token production efficiency and the value Tokens generate, with the most typical target being Tokens/second. We're also trying various methods to make the Tokens/second metric better. All optimizations can ultimately come back to this.

Operator optimization directly improves Tokens generated per second on chips; stability optimization and operations work ultimately also serve to improve Tokens/second.

The reason we use various heterogeneous chips is also to make the resource coefficient of "Tokens/second" larger, letting more chips contribute to "Tokens/second." One-sentence description: letting all chips in China that can be used produce Tokens at the highest efficiency — this is our most important current goal.

We're also investing some effort to help small and medium entrepreneurs who don't use Tokens most efficiently but have good creativity and product capabilities: they can use our tools to do well the环节 from Token to productivity, letting them spend 100% of their energy on releasing Token productivity and launching their own products.

China Entrepreneur: Some time ago, you also launched a Lobster Box, building an enterprise-level Token factory. Compared to other deployment solutions on the market, what pain points does the Lobster Box solve in reducing Token costs and improving output efficiency?

Lixue Xia: The Lobster Box is a product form, currently still in early stages. What we care more about is the kernel of technical value. The most important point of this product is that it reflects our ultimate solution for Token-to-productivity conversion efficiency. This involves coordination between different models and security issues between different data domains.

The Lobster Box's core selling point focuses on the optimization target of "productivity released per Token." Because Tokens can be tiered, some tasks use the strongest models while others use more cost-effective models. The box can run small models, with the key pain point being data security during transmission.

This can be used both on terminal Lobster Boxes — the pain point it addresses is data not wanting to be uploaded to the cloud. In the future it can also be used in another scenario — running large models in the cloud while coordinating with small clusters, so it more represents our new layout and breakthrough in technical路线.

We previously mentioned "heterogeneous,异域, and异属" — one network, three异. Heterogeneous solves how to run together if there are two different chips in the same cluster. 异域 solves how two clusters跨越 a certain distance (up to 4,000 kilometers) can run together. 异属 solves how resources in two different data zones can run together. The Lobster Box is also a落地 of this technical路线.

China Entrepreneur: Alibaba, ByteDance, and Tencent have all established Token departments. Do you have such dedicated Token teams internally?

Lixue Xia: I proposed a concept very early on called the "Model Power Resources Department," following the思路 of "Human Resources Department," because in the future AI is an extension of people. Currently, using AI to write code internally is basically 100% coverage; we're also using AI for operations; and we even have tools to help everyone use AI to make PPTs.

Now many companies have departments specifically responsible for AI applications, with考核 indicators possibly being the company's and employees' daily Token usage. Although somewhat overcorrecting, and this may not ultimately be the form, in early stages it's完全可以 fine to run this way first.

China Entrepreneur: You mentioned that Token usage doubled every two weeks in the past. Will this growth trend continue for the next year or several years?

Lixue Xia: Call volume is still constrained by supply. Future Token call volume growth represents users' acceptance of Token cost-performance, or the speed of supply cost reduction.

For the next 3 to 6 months, it will likely maintain the current supply-demand state; after 6 months, there may be a new wave of Token usage爆发. This is because supply capacity is expected to expand significantly: including both new-structure domestic chips and joint optimization from models to hardware. At this point, Token cost-performance and technical optimization space will also同步 increase, both making more resources available and enabling more advanced chips to have higher cost-performance Token output rates.

Just like in the previous traffic era, users went from spending hundreds of megabytes of traffic per month to using several gigabytes, without spending 10 times more money. Token usage growth also brings prosperity to the entire industry, with costs continuing to decline significantly.

China Entrepreneur: How do you internally evaluate Token metrics? Do you look at usage volume, quantity scale, growth rate, or customer payment it brings? Which is the primary metric for AI Infra company value?

Lixue Xia: Metrics are certainly different at different stages. In rapid growth stages, the scale of high-value Token usage is most important. Meanwhile, trillion-parameter models are still likely quite expensive, representing that Tokens and infrastructure produce valuable, rewarding output for the industry.

The greater the usage, the more optimization space can be seen; if optimization technology is good, it can produce better cost-performance, usage will further increase, creating a flywheel effect.

As CEO, I care more about whether the company is running well, looking at technical depth and customer recognition: whether we can maintain the most advanced leading position in a technology-driven track, and whether customers recognize our product value. Externalized metrics are paid volume or call volume for high-value models.

China Entrepreneur: From your platform's Token usage growth rate, which industry customers and scenarios does it mainly come from? What's the approximate proportion from Agents?

Lixue Xia: Over 95% is generated by Agents. The industries are also quite diverse, with code writing being the largest, plus content creative generation, etc.

03/

China Will Become the World's Token Factory

China Entrepreneur: Everyone's talking about Token price increases now. Do you think Tokens should increase in price? Or how do you think they should be priced?

Lixue Xia: Most domestic model price levels and increases are below overseas models, but intelligence is already quite good, so there's room for price increases.

More importantly, the logic behind price increases is user willingness to pay — if people are still willing to buy after the increase, that's刚性 demand.

Price increases叠加 cost reductions move the break-even line, turning certain loss-making businesses in some areas into profitable ones, then entering the supply-demand growth flywheel, ultimately bringing benefits to users.

China Entrepreneur: Do you think in the long run Tokens will be oversupplied, or that there will be too many Tokens for the market to absorb, leading to a new round of price wars? Will there be such an inflection point?

Lixue Xia: In the future, Tokens will be tiered. One tier is higher-quality Tokens that generate greater value; another tier may be cutthroat competition pursuing extreme cost-performance Tokens. This is very much like internet advertising traffic, which eventually all bills by CPM (cost per thousand), with everyone understanding which channels' exposure is more valuable. The Token economy is even clearer in this regard, because model intelligence level is reflected in Token quality.

What we see as Infra vendors is that high-quality Tokens will remain severely in short supply in the future. This is true worldwide, and in China the scarcity is actually even higher.

China Entrepreneur: At the end of March, Yahui Zhou, founder of Kunlun Tech, told us that Mobile Internet CPM rose over ten years, with customer acquisition costs getting higher and higher, possibly increasing 10 times. In this Token era it might be the same — Token costs seem to be getting lower and lower, but prices might correspondingly rise 10 times.

Lixue Xia: CPM rose because advertising platforms launched ROI-targeted optimization models that could "guarantee conversions."

Tokens are the same. Future pricing may be tiered by model type, or by Token input/output, or even by SLA tiered pricing. But essentially it's all tiering by the conversion value Tokens produce. Since the value converted to productivity is higher, the Token itself is more valuable, and the price can be higher.

China Entrepreneur: You've mentioned that China will become the world's Token factory — in the past Made in China was manufacturing, now it's AI Made in China.

Lixue Xia: China has abundant energy structural advantages, a complete AI industry chain, and the world's largest-scale AI application market, fully capable of replicating the successful path of "Made in China."

Working backward from the end, since value exists, what needs to be solved are the methods, approaches, and chain issues.

China Entrepreneur: Some people say electricity is compute, electricity is Token. How would you evaluate this view?

Lixue Xia: In a stable state in the future, this will indeed be true. For example, in chip selection, at least several chip vendors already have considerable market share. At this point, building a "Token factory" means the main cost is raw materials, not the "house."

NVIDIA is still too expensive, equivalent to a "house" built with gold bricks, and the optimization value of electricity hasn't fully emerged. But in about two years, "house" costs will become controllable, and then evaluating Token factory production efficiency will certainly look at the conversion efficiency from "raw materials" to "finished products."

Therefore, future electricity costs and electricity-to-Token conversion rates will become more critical. China's advantages in energy will certainly demonstrate enormous industry value globally.

Source | China Entrepreneur Reporter | Yan Junwen Editor | He Yifan Associate Editor | Li Yuan


Previous Reviews

Qiming Star | Infinigence AI Receives Nearly RMB 1 Billion in Total Funding, Qiming Venture Partners Co-Leads Series A Qiming Star | Infinigence AI Releases "Infini-AI" Large Model Development and Service Platform, Reaches Multiple Strategic Partnerships Qiming Perspectives | Three AI IPOs at Start of Year, Conversation with Qiming Venture Partners' Alex Zhou: Following Power Law, Firmly Layout Technology Targets in Long Snow Slopes, Big Tracks

Qiming Venture Partners was founded in 2006. Currently, Qiming Venture Partners manages 11 USD funds and 7 RMB funds, with total assets under management reaching $9.5 billion. Since its founding, it has focused on investing in outstanding early and growth-stage enterprises in Technology and Healthcare innovation.

To date, Qiming Venture Partners has invested in over 580 high-growth innovative enterprises, of which more than 210 have listed on the New York Stock Exchange, NASDAQ, Hong Kong Exchanges and Clearing Limited, Shanghai Stock Exchange, and Shenzhen Stock Exchange, or exited through mergers and acquisitions, with more than 80 enterprises recognized as unicorns or super-unicorns in their industries.

Among Qiming Venture Partners' portfolio companies, many have grown into the most influential companies in their respective fields, including Xiaomi (01810.HK), Meituan (03690.HK), Bilibili (NASDAQ:BILI, 09626.HK), Zhihu (NYSE:ZH, 02390.HK), Roborock (688169.SH), Hesai Technology (NASDAQ:HSAI, 02525.HK), UBTECH (09880.HK), WeRide (NASDAQ:WRD, 0800.HK), HyperStrong (688411.SH), Insta360 (688775.SH), Unisound (09678.HK), Biren Technology (06082.HK), Zhipu AI (02513.HK), Gan & Lee Pharmaceuticals (603087.SH), Tigermed (300347.SZ, 03347.HK), Zai Lab (NASDAQ:ZLAB, 09688.HK), CanSino Biologics (688185.SH, 06185.HK), Schrödinger (NASDAQ:SDGR), MicroPort EP MedTech (688617.SH), Sanyou Medical (688085.SH), Amoy Diagnostics (300685.SZ), GenScript ProBio (688520.SH), Insilico Medicine (03696.HK), Hope Medicine, Yuanxin Technology, MediLink Therapeutics, LaNova Medicines, StepFun, and others.

Qiming Honors | Qiming Venture Partners Wins Zero2IPO 2025 China Venture Capital Institution Third Place and 7 Other Major Awards