StepFun's Major "Step": Over 5 Billion Yuan in B+ Round Funding, Qi Yin Formally Appointed Chairman

StepFun's answer is AI + endpoints.

StepFun's answer is AI + endpoints.

👦🏻 Author: Jingshan

🥷 Editor: Koji

🧑‍🎨 Designer: NCon

The "exhibition games" are over; the "elimination rounds" have begun. This may be one of the best opening lines to describe "AI 2026."

If you look back at the AI industry over the past two years with a relatively sober eye, one fact becomes unmistakably clear: the vast majority of players have been crowded onto the same track, doing remarkably similar things: chasing leaderboards, chasing benchmarks, fighting for SOTA in every standardized test.

Model capability improvements certainly matter, but as leading foundation model companies have gradually entered commercialization, by 2026 this mode of competition has started to look rather thin.

On one hand, the returns on model capability gains have hit diminishing marginal returns. On the other, the market is no longer willing to pay a steep premium for "bigger and stronger" as an end in itself.

Attention has gradually shifted to more pragmatic questions:

Who can actually land these capabilities? How far along is AI on hardware endpoints?

It is at precisely this juncture that StepFun has suddenly sent two very clear signals: a B+ round of unusual scale has closed, and Qi Yin has moved into a pivotal position, formally becoming StepFun's chairman — giving StepFun a more direct connecting channel between "model capability" and "endpoint scenarios."

This B+ round, exceeding 5 billion RMB, not only set a new record for single-round financing in the large model space over the past year — it even surpassed the IPO fundraising totals of some peers.

Participants in this B+ round include multiple strategic investors. Huaqin Technology, for instance, is a leading giant in the mobile terminal space, well-aligned with StepFun's direction in "endpoint scenarios." Investors such as China Life Private Equity Investment Company Limited, Xuhui Capital, Wuxi Liangxi Fund, Xiamen ITG, Huaqin Technology, Shanghai SDIC Pioneer Fund, China Life Private Equity Investment Company Limited, and Shanghai Pudong Venture Capital Co., Ltd. are all considered to have high investment thresholds and rarely make moves casually. Among these, we also see major players from the Yangtze River Delta hardware supply chain; Tencent and Qiming Venture Partners, as existing shareholders, also chose to double down and follow on — demonstrating consistent "endorsement" from both internet giants and professional financial investors.

Putting these two signals together, it is not hard to see that StepFun has already switched to a different logic: using large models as core capability, extending toward endpoints and hardware.

🚥

The question the industry must now answer is: Who can put AI into the real world and make it work sustainably.

This article takes this perspective to examine a question that has become increasingly unavoidable: What does "AI + hardware" actually mean for foundation model companies?

AI Is Entering "All Kinds of Endpoints"

In an interview this year, Sam Altman shared his view on AI's trajectory. He put it plainly: AI won't stay in massive data centers forever, nor will it run only through the cloud.

Going forward, AI will gradually make its way into our phones, computers, and other devices — landing on specific hardware and dedicated equipment.

Look at past technological development and you'll spot a pattern:

Technologies that truly matter eventually find their hardware载体.

Look back at the problems AI applications encountered over the past year, and one issue stands out: models themselves have gotten increasingly powerful, but using them hasn't become much more convenient.

AI now has strong comprehension and reasoning capabilities, but it remains constrained by how it's invoked, how users interact with it, and in what scenarios it can be used — so the range of things it can actually do remains limited.

Most AI applications today still take the form of API interfaces or chat boxes, with very constrained applicable scenarios.

Once you enter the complex and variable conditions of the real world, no matter how smart the model is, it can't directly translate into practical action. The process gets stuck in the middle, and the business can't close the loop.

This is why Agents that can enter phone operating interfaces and Agents that can work on computer desktops are starting to gain serious attention.

Because they don't require rebuilding entire systems from scratch to be useful.

For example, GUI Agents can work directly on existing phone interfaces by understanding screens and controlling operations.

The direction of AI development is shifting. And to get there, hardware is basically unavoidable.

This is one of the underlying logics of "AI + endpoints."

Following this line of thought, if we use StepFun as an example, we can see they've laid out a fairly clear path:

StepFun believes the endgame of AI + endpoints will be a software-hardware integrated solution driven by a super assistant + cross-device OS. In concrete terms, they use their self-developed large model as the brain, responsible for complex thinking and judgment, then partner with terminal device and industry players like Qianli Technology to put these capabilities into actual hardware and real-world scenarios for testing.

This approach, put simply, is "strong model + strong endpoint."

The strong model figures out what to do — understanding problems, breaking down tasks, making decisions, all handled at the model level.

The strong endpoint actually gets things done. Devices like phones and cars that we use every day already operate in real-world scenarios, constantly generating feedback and sending results back to the model.

When these two parts work together, AI starts to become a system that can run continuously and keep learning.

Of course, StepFun isn't the only one taking this road.

Many foundation model companies will end up heading in the same direction: getting AI into the real world to start doing concrete work.

In observing this "AI + endpoints" race, a team like StepFun makes for a suitable observation sample: its path is relatively complete, and it's moving relatively fast.

The real question is: how do you get AI to run stably in real scenarios, in real hardware endpoints?

In this race, StepFun has gradually built out a "1+2" technical system.

Unlike the common approach of simply piecing together vision models and language models, StepFun has emphasized native multimodal models and unified understanding-generation from its founding. The advantage of this design is that transitions between understanding problems, planning, and executing actions are smoother, making real-world interaction easier — in other words, progressing toward VLA scenarios.

On the foundation model front, StepFun is one of the few startup teams in China to have trained trillion-parameter models, and the only one to invest heavily in building its own AI Infra. The Step-3 reasoning model achieves a relatively balanced tradeoff between efficiency and performance, with stable showings across multiple public benchmarks. Thanks to innovations at the system architecture level, Step-3 achieves inference efficiency on domestic chips up to 300% that of DeepSeek-R1, while remaining friendly to all chips — giving it extremely high practical utility.

Meanwhile, in multimodal and VLA directions, StepFun has successively launched 30 models with fairly comprehensive capability coverage.

In areas like speech, multimodal understanding and generation, and GUI operation, model forms that can be directly validated are already visible. Speech is the most universal human-machine interaction mode, and StepFun has been exploring the frontiers of speech technology. Recently, its native speech reasoning model Step-Audio-R1.1 took first place globally on the Artificial Analysis Speech Reasoning leaderboard. On the commercialization front, StepFun partnered with Geely to achieve the industry's first end-to-end speech model mass production deployment in vehicles.

Viewed holistically, this "1+2" structure looks more like laying groundwork in advance for long-term operation and real-world scenarios.

Under this logic, when AI truly moves toward endpoints, it tends to land first on devices with high usage frequency, complex operations, and continuous data generation.

The most classic, and most presence-rich in daily life, are:

Phones and cars.

When AI Starts Operating Screens "Like a Person"

If the "AI + endpoints" strategy still sounds somewhat abstract and complex, the changes on the phone side have already become quite intuitive in actual experience.

First, the phone is the endpoint closest to users, with the highest usage frequency.

Nearly all valuable behavior happens on phones. Chatting, payments, content consumption, travel arrangements, work collaboration — none of it happens without phones. If AI can't get into phones, it's hard for it to truly participate in users' daily judgments and operations.

More importantly, phones themselves are already mature.

They have complete app ecosystems, system-level permissions, notification mechanisms, and clear input and output methods. GUIs are already stable, interaction paths are well-defined, and there's massive installed base — for AI, this is a ready-made environment.

In Crossing's hands-on test of StepFun's Step-GUI earlier, we saw a usage pattern somewhat different from previous AI. It no longer requires users to click buttons step by step, walking through processes — it can directly understand goals and complete operations autonomously.

For example, in our test, a single sentence was enough to have Step-GUI autonomously control a phone, opening real apps according to task requirements:

"Help me earn coins in Kuaishou Lite and withdraw them to WeChat."

To actually get this solution running, StepFun took a fairly pragmatic architectural route, with edge-cloud division of labor at its core.

[1] Edge side

Small models run locally on the phone, mainly handling high-frequency operations like tapping and swiping, while processing privacy-related data to ensure response speed and reduce users' "privacy anxiety."

[2] Cloud side

More complex judgment, task decomposition, and overall planning are handed off to large models in the cloud.

The whole design, put simply, is: simple, frequent, latency-sensitive tasks stay local; complex tasks requiring global judgment go to the cloud.

This route isn't for show.

According to public information, StepFun's endpoint Agent call volume has maintained nearly 170% growth for three consecutive quarters. In actual deployment, they've partnered with 60% of mainstream domestic phone makers, including OPPO, HONOR, and ZTE, with relevant models already installed on over 42 million flagship devices, serving nearly 20 million daily active users. Users employ these capabilities for intelligent search, photo Q&A, writing Moments posts, generating personalized themes.

This counts as a decent report card.

At the same time, this means many users have already started encountering this new AI interaction pattern in daily use, perhaps without realizing how much it has changed. For a long time, the core battleground for phones was hardware specs, but going forward, user choice may well tilt toward which phone offers the stronger AI experience.

Cars: AI's Next Super Endpoint?

If we pull back from phones, cars are almost certainly among AI's most important next landing points.

The reasons are easy to understand.

Car scenarios can be divided into intelligent cockpit and intelligent driving. Compared with phones, cars have less than one-tenth the commonly used applications, so the intelligent cockpit's service ecosystem and experience may be restructured by AI even earlier than phones, while intelligent driving faces more complex real-time environments. Road conditions change, surrounding pedestrians and vehicles move, demands on judgment, planning, and execution are higher, and the cost of mistakes is greater.

Precisely because of this, cars naturally become a testing ground.

Here, AI must truly judge, decide, and execute actions. This makes cars a direct test of whether AI genuinely possesses agency.

Previously, the Crossing team attended Qianli Technology's launch event in Chongqing, where we saw a marked change: in complex urban environments with many variables, an intelligent driving system connected to StepFun's large model capabilities performed more steadily overall, with decision rhythms more like someone who'd been driving for many years.

One concept was repeatedly mentioned at the time: "veteran driver feel."

But for StepFun, cars aren't merely vehicles for autonomous driving.

Through deep partnerships with Qianli Technology and Geely Auto, they're pushing a complete intelligent cockpit solution centered on Agent. A key factor in why this route can work is that Qi Yin has personally stepped in.

Qi Yin simultaneously stands at the core of both a model company and a complete vehicle system — this looks significant from the outside.

It means questions of how to design model capabilities, how to architect systems, and at which layer to embed capabilities can all be approached from the start around a unified co-creation logic.

A unified large model, a unified system architecture, allowing AI to participate in vehicle operation logic from the bottom up, truly achieving software-hardware integration rather than, under traditional tech supply logic, simply layering AI on top. For a terminal as complex as a car, this difference is genuinely substantial.

From results so far, this system has already entered scaled deployment.

Relevant mass-production models continue to hit the market, and the number of vehicles equipped with StepFun large model capabilities is growing rapidly. More vehicles means real driving and usage data continuously flowing back.

This data comes directly from real road conditions and real users; model performance in actual scenarios gets repeatedly calibrated and strengthened. Over time, a clear positive cycle forms: more usage, more data, models increasingly aligned with real-world demands.

For example, the Geely Galaxy M9 has sold nearly 40,000 units in about three months since launch, and has already begun entering overseas markets. At the current pace, the scale of StepFun large models "getting into cars" this year will likely exceed one million vehicles.

Precisely because of this, for many pure software companies, cars are not an easy battlefield to enter — what matters here isn't just how good your model is.

Whether model, system, and complete vehicle can run on the same line from day one matters enormously. Because this competition is itself a very long process.

And it is under this premise that StepFun's ability to get where it is today owes much to team structure as a critical factor.

If summarized simply, this team can be understood as a "1 + 3" combination.

We've also specifically梳理ed relevant backgrounds.

"1" industry operator: Qi Yin.

He has fully experienced AI's journey from technology wave to industrial landing.

As chairman of Qianli Technology, he has been working to打通 scenarios where AI has been difficult to land — namely, cars. As the key figure, Qi Yin has connected auto industry resources including Geely to StepFun. This provides StepFun with a realistic anchor point for advancing "AI + endpoints."

"3" technical cores:

StepFun CEO Daxin Jiang: As StepFun CEO, he was formerly a Microsoft corporate vice president, working long-term in search and NLP, with deep familiarity across the complete technical chain from underlying data to upper-layer applications.

StepFun Chief Scientist Xiangyu Zhang: As one of the authors of the ResNet paper, his research background provides the team with a relatively stable technical ceiling at the model architecture level.

StepFun CTO Yibo Zhu: Responsible for systems direction, previously led AI Infra-related work at ByteDance, one of the very few people in China with 10,000-card training experience.

Viewed together, this "1 + 3" combination happens to cover several critical links: industry judgment, model capability, systems engineering, and endpoint landing.

Precisely because of this, StepFun has both the ability to work on long-term foundation models and the capacity to try pushing these capabilities step by step into real products and real endpoints.


For a long time, people have been saying "AI + endpoints" is an attempt to change how AI participates in life.

In StepFun's vision, this future intelligent experience can be supported by a complete system.

Put simply:

A multimodal large model responsible for understanding and decision-making, an OS that can run across different devices, and an Agent continuously working across various scenarios.

Under this structure, AI is no longer confined within some app, but will naturally distribute across phones, cars, and more endpoint devices.

But let me be clear here.

Putting large models, OS, and Agent together doesn't automatically become a profitable, self-sustaining commercial闭环, but it does have the potential to form new traffic entry points.

For all foundation model companies, this remains a path that needs validation through time and real-world landing — there are no ready-made answers yet.