AI inference deployment solution provider **Qingmao Intelligence** closes tens of millions of RMB in Pre-A+ funding | Oasis Vitality

Counselor on Vitality

Qingmao Intelligence, an AI inference deployment solutions provider, recently closed a Pre-A+ round of tens of millions of RMB, led by Qifu Capital and Fortune Venture Capital, with existing investor MiraclePlus following on. In 2023, Oasis Capital led Qingmao Intelligence's angel round as the sole investor. According to sources, the new funding will primarily go toward building out the team, product R&D, and commercial deployment.

Long before the wave of large language models took hold, the "three highs" — high inference latency, high inference cost, and high resource consumption — along with hardware compatibility at the compute layer, had been the last-mile problems plaguing model deployment. This is especially pressing now that AI-powered consumer hardware has become an industry trend: getting large models to run on terminal devices with limited compute has become an urgent challenge for many device manufacturers.

Yet these pain points correspond to a gap upstream in the solutions market. On one hand, the mainstream players in inference deployment toolchains are mostly concentrated in North America. On the other, most middleware vendors provide adaptation services for overseas hardware like NVIDIA. As domestic alternatives gradually become the primary compute solution in China, the pain point of adapting large models to domestic chips has remained largely unaddressed.

As one of the earliest domestic players in inference deployment toolchains, Qingmao Intelligence was founded in October 2022. It reduces the cost and barrier of deploying and using AI models for downstream customers by providing an inference and deployment optimization toolchain.

As early as June 2022, riding the wave of AIGC models like Stable Diffusion, the Qingmao Intelligence team began developing its model deployment and inference optimization toolchain. Targeting smart scenarios such as AIoT and autonomous driving, the company launched its first-generation AI model inference optimization toolchain, MLGuider. Beyond NVIDIA, MLGuider also supports deployment on domestic and international chips including AMD, Qualcomm, and Ascend.

Based on market demand, MLGuider's functionality and framework have undergone continuous iteration. Qingmao Intelligence CEO Chaoyu Guan told 36Kr that the early MLGuider was primarily designed for edge chips and traditional small models, employing a series of optimization methods including quantization, distillation, and sparsification. As market demand for large models exploded, Qingmao Intelligence combined optimization technology stacks spanning model optimization, distributed optimization, and compiler optimization to build a full-link toolchain oriented toward foundation models and underlying compute hardware, with particular emphasis on iterating functionality for large model and domestic AI chip adaptation and optimization.

Taking domestic leading hardware Ascend as an example, at the 2024 Ascend Developer Conference this year, Qingmao Intelligence debuted as an Ascend partner representative with MLGuider-Ascend, a toolchain based on Ascend's native development environment. It addresses problems including model-compute mismatch, complex technology stacks, and high migration and optimization costs that AIGC models face when deploying on domestic Ascend hardware.

Beyond the model inference deployment optimization toolchain, Qingmao Intelligence has also launched a matrix of solutions including an enterprise-grade foundation model development and deployment platform (LLMOps), large model integrated hardware solutions, and large model localization and edge deployment solutions.

Guan believes that the dilemma for middleware vendors often lies in how to achieve commercialization at scale. To address this, while directly providing solutions to enterprise customers, Qingmao Intelligence is also focused on establishing ecosystem partnerships with chip vendors and local compute centers. "We can provide customers with end-to-end integrated solutions by connecting ecosystem partners across chips, servers, and model solution providers," Guan explained.

Model inference deployment toolchains handle the hardware-software adaptation between the compute layer and the model layer, which is why they're also called middleware. Guan believes the middleware layer's task is to push model runtime performance infinitely closer to the hardware's peak performance, fully unlocking the potential of both models and hardware.

On whether middleware risks being absorbed by upstream or downstream players, Guan told 36Kr that from the perspective of the model layer and chip layer, each has its own focus — improving the performance of the model or chip itself. Meanwhile, the proliferation of model choices and fragmented hardware environments are making the model-middleware-chip ecosystem collaboration increasingly clear.

On the talent and organizational front, Qingmao Intelligence's core team hails from institutions and companies including Tsinghua University, Huawei, and Alibaba. Founder and CEO Chaoyu Guan graduated from Tsinghua University's Department of Computer Science, was a 2021 Siebel Scholar (fewer than 100 worldwide), and led the development of AutoGL, the world's first automated graph learning project. Scientific advisor Wenwu Zhu is a professor in Tsinghua University's Department of Computer Science and Technology, and previously served as a principal researcher at Microsoft Research Asia and chief scientist at Intel China Labs.

Source: Intelligent Emergence, by Xinyu Zhou