Dialogue | StepFun: Charting the Uncharted Territory of AI Entrepreneurship

The technical development of AI models is still on a very steep upward trajectory, and multimodal models that integrate understanding and generation are critically important.

As artificial intelligence technology moves from technical exploration into the deep waters of industry, Chinese companies' technical strength and development potential are becoming increasingly apparent. Qiming Venture Partners portfolio company StepFun is a leading foundation model startup, actively exploring practical pathways for large models across multiple domains. On the industry application front, StepFun focuses on intelligent terminal agents and has established deep partnerships with sector-leading companies in key application scenarios including automotive, mobile phones, embodied artificial intelligence, and IoT.

At Qiming Venture Partners' 16th RMB Fund Annual Meeting and Investor Summit, Alex Zhou, Managing Partner at Qiming Venture Partners, and Daxin Jiang, Founder and CEO of StepFun, held a dialogue themed "StepFun: Exploring the 'Uncharted Territory' of AI Entrepreneurship." The two discussed topics including the definition of AGI (Artificial General Intelligence) and its development roadmap, current large model technical progress, why AI agents exploded in popularity in 2025, and StepFun's layout in the AI agent domain.

Alex Zhou, Managing Partner at Qiming Venture Partners (left), and Daxin Jiang, Founder and CEO of StepFun (right)

Jiang noted that AI model technology development remains in a very steep upward phase, and that unified multimodal models integrating understanding and generation are critically important. In terms of agent strategy, StepFun focuses on the intelligent terminal direction. He emphasized that the key capability of agents lies in understanding the user's environment and task context through multimodal interaction, then proactively and autonomously helping users complete tasks. One of StepFun's important goals is to build an intelligent terminal platform that enables more people to access its comprehensive model matrix.

In response, Zhou observed that today's AI field's "model-as-product" philosophy differs significantly from internet-era product-building concepts: in AI scenarios, one or a set of excellent models or agents directly determines 70-80% of a product's performance; whereas in the internet era, with mature technical infrastructure, companies focused on product-level innovation.

The following is an edited and abridged transcript of the dialogue.

01/ Three Stages to Achieving AGI

Zhou: Good afternoon. I'm particularly delighted you could join our summit. After DeepSeek released its two models in January, I received countless inquiries asking whether, with DeepSeek's emergence, our portfolio companies like StepFun and Zhipu AI face enormous challenges? Today I'd like you to help address these challenging questions.

In April, Xi Jinping visited the "Model Speed Space" large model innovation ecosystem community in Shanghai's Xuhui District, where four companies gave one-on-one briefings. StepFun was the only foundation model company among them.

Jiang: It was a rare opportunity. As a Shanghai-based AI foundation model company, StepFun was the first to present. We introduced the latest progress in foundation model technology and applications, demonstrating how multimodal large models combined with intelligent terminal scenarios can bring convenience and create value for everyone's daily lives.

Zhou: The industry previously often mentioned the concept of China's "Six Little Tigers" of large models, plus several major tech companies, as the main forces developing foundation models in China. Recently some media have proposed a "New Top Five" competing for dominance, with three being companies of substantial scale: ByteDance, Alibaba, DeepSeek, Zhipu AI, and StepFun, suggesting these five enterprises will continue striving toward AGI. What's your view? How do you define AGI? What is StepFun's vision? How should we approach AGI?

Jiang: What is AGI? Actually there's no industry consensus. If we were discussing this ten years ago, talking about when AGI might be achieved would have seemed like fantasy. Even five years ago, before large models emerged, people felt this wasn't within the realm of discussion. Now with more large models, more people believe AGI will arrive within five years, with timelines varying from 2026 to 2030.

What is the criterion for testing whether AGI has arrived? In April, an American university conducted a test using the traditional Turing test on OpenAI's GPT-4.5, finding that 30% of people couldn't distinguish whether it was AI or human, and in 73% of cases it successfully convinced people it was human. By the original definition of the Turing test, this means GPT-4.5 passed. We feel that this standard alone doesn't mean AGI has arrived. So when I communicate with friends in Silicon Valley, they offer a new AGI definition benchmarked against human intelligence: the percentage of existing human work that models can complete. What this percentage should be varies by person. If I were to set it, a conservative number would be 50%.

If models in 2030 can complete 50% of existing human work, I believe AGI will have arrived. When StepFun was founded, we set ourselves the goal of achieving AGI. Our founders drew a roadmap for realizing AGI, dividing it into three stages: simulating the world, exploring the world, and inducing the world.

So-called world simulation involves imitation learning as the learning method. We feed all internet data into large models and have them comprehend the internal structure and characteristics of data through very simple tasks. In this stage, the primary focus is learning representations of various modalities — from speech to sound to images to video to 4D physical spacetime. The core question here is how we use deep neural networks to achieve representations of modalities from simple to complex. This is the work to be completed in the first stage.

After learning to represent the world, in the second stage we want models to help us solve problems, particularly complex ones. For example, writing code or doing math problems often requires strong chain-of-thought reasoning. When humans solve such complex problems, we employ a capability called slow thinking. For instance, when we do a math problem, we typically don't blurt out the answer immediately but break it down into different steps. If the initial approach seems wrong, we reflect and consider new solutions. So this is a process of trial and error. How to enable machines to possess such slow-thinking capability? The underlying algorithm is reinforcement learning. The concept of reinforcement learning isn't particularly new. Coincidentally this year, the Turing Award went to two reinforcement learning experts: Andrew G. Barto and Richard S. Sutton. The latter wrote a famous article called The Bitter Lesson, which OpenAI staff reportedly read daily like scripture.

AlphaGo, which defeated human Go master Lee Sedol in 2016, was a typical representative of reinforcement learning. This year, the well-known DeepSeek also employed reinforcement learning algorithms, greatly enhancing the model's reasoning capabilities.

But reinforcement learning isn't the endpoint of intelligence. Going further, intelligence can evolve to autonomous learning, where models can discover new patterns alongside scientists in cutting-edge fields like biology, nuclear energy, and quantum computing — discovering physical laws humans haven't found. We call this stage inducing the world.

Last August, OpenAI published five levels of intelligence evolution: earliest Chatbot, then Reasoner, Agent, Innovator, and Organization. If we examine their definitions carefully, we find these five levels are logically consistent with our three stages, just described differently.

We see that although OpenAI and foreign large companies have released many models, if we look along this roadmap, their models are progressively covering key nodes along this path. Today, from simulating the world to exploring the world, we see this trend becoming increasingly clear, so our confidence is growing stronger.

Zhou: Speaking of large models, returning to the opening topic of DeepSeek — we're all large model companies. What exactly are StepFun's differentiated characteristics compared to companies like DeepSeek?

Jiang: Let me introduce our work over the past two years. We've released some large models. Although all are called foundation models, their functions and directions differ. We can categorize them as: language models and multimodal models. Within language, there are foundation models and reasoning models. In the multimodal domain, we can classify by different modalities: speech models, music models, image models, and video models.

By function, they can also be divided into understanding models and generation models. One of our major characteristics is our strong emphasis on multimodal capabilities, continuously strengthening them. StepFun adheres to full-modality coverage and native multimodalism, which is a non-consensus in the industry, but we consistently believe multimodality is the inevitable path to AGI.

Actually, AGI is defined by analogy to human intelligence. Beyond linguistic symbolic intelligence, humans naturally possess visual intelligence, spatial intelligence, and motor intelligence. These intelligences cannot be achieved through language alone and must be expressed through multimodality. Beyond the AGI concept, once we reach application domains — whether C-end or vertical B-end applications — we want models that can hear, see, and speak like humans, so they can better understand users' physical environments and communicate with users naturally. From these two perspectives, we believe lacking any modality would delay AGI's realization. So compared to other companies, being able to persist in self-developing comprehensive foundation models and building a complete model matrix is rare even among major tech companies, let alone startups. This is our characteristic and advantage.

02/ Technical Progress and Development Trends

Zhou: Some of the previously mentioned Six Little Tigers have publicly announced completely abandoning pre-training to focus only on post-training and other R&D. It seems choices are diverging increasingly. From your perspective, what major technical progress does StepFun see? Where will we go next?

Jiang: On one hand, model capabilities are indeed continuously improving. Whether reasoning models or multimodal models, they are continuously advancing, catalyzing application deployment. After DeepSeek emerged, people felt that many previously poorly executed application scenarios could now be achieved through stronger models. Model capabilities have unlocked many application scenarios. Additionally, we feel model development isn't slowing down.

After the Spring Festival, possibly influenced by DeepSeek, five leading American companies released many models. OpenAI first released o3 and GPT-4o solutions, and OpenAI's release timing generally aligns with Google's, which simultaneously released the Gemini series. There was also Claude 3.7 Sonnet. In just two months, five leading foreign model companies released models in rapid succession. So model progress isn't slow, and through these model releases, we can still discern the overall development trend.

First, models are moving from original world-simulation, imitation-learning models toward reinforcement learning models.

The earliest reinforcement learning model was OpenAI's o1 model released last September, followed by the full version in December, and then DeepSeek's R1 model during the Spring Festival. This basically declared reasoning models a paradigm rather than just a trend. Looking at the models released by these leading foreign companies, they basically integrate reasoning capabilities. StepFun has also done some work in reasoning. In January, we released a small Step R-Mini model that at the time surpassed OpenAI's o1 preview model. We will also release a full-version reasoning model in the future. In reasoning models, we see much work still advancing. For example, how to further improve reasoning efficiency. People now feel the chain of thought is very long, but some thinking is ineffective.

Second, a critical question: how does reinforcement learning generalize reward functions in domains with clear right and wrong like mathematics and code, versus many domains where correctness and values cannot be clearly judged? And how can chain-of-thought data be synthetically generated and placed into pre-training? These are currently very hot topics in industry and research.

Reasoning models will continue developing over the next one to two years. Simultaneously, we see another trend: reasoning models are not only applicable in text domains but now also achieve reasoning in multimodal domains. Take OpenAI's o3 model as an example: netizens gave it an image and asked it to guess the location, and it truly acted like Sherlock Holmes, inferring from details what place was in the picture. Here I'd like to demonstrate our recently released image reasoning model. Given an image, it judges which Chinese Super League team's home stadium and match this is.

If you've played with image recognition before, you'd find that previous-generation visual models simply searched training data for similar content — still a fast-thinking process of seeing an image and judging where you've seen it before. This isn't reasoning.

The current model can find the two competing teams' logos from the scoreboard on the field. It also looks at fan clothing colors in the stands to determine whose home field it is, at which point it can already infer which stadium. Additionally, through the stadium's architectural style, such as the roof structure, it confirms exactly which stadium.

It doesn't blurt out an answer at first glance but reasons through details and sensory recognition combined with internal knowledge bases. So reasoning capabilities are becoming increasingly powerful.

We also see an interesting trend: multimodal fusion moving toward unified understanding and generation. First, let me explain what unified understanding and generation means.

In language models, such as DeepSeek, if we give it an article and ask it to answer questions or generate summaries, these are typical understanding tasks. Conversely, if given a title and asked to create, this is a generation task. People typically don't distinguish these two tasks but use the same model to complete them. However, in multimodal domains, these two are separate. For judging image content as just mentioned, one would use models like GPT-4V or GPT-4o; for generation, one would use models like Sora. So unified understanding and generation hasn't been achieved in the visual domain.

Why is this problem so important? For example, when a teacher writes on a blackboard with chalk, the hand movement including the trace of chalk contacting the blackboard — Sora can simulate this. If the teacher stops halfway and we ask what will be written next, this requires an understanding model to predict, while generation model Sora lacks this capability. This is what we mean by understanding and generation not being unified.

From the generation perspective, current generation models aren't controlled by understanding. From the understanding perspective, what counts as true understanding? If I cannot create, then I haven't understood. Only when I can truly create autonomously does it show I've achieved true understanding. As Richard Feynman said — "What I cannot create, I do not understand."

In text domains, generation tasks are Predict Next Token, while models can also understand knowledge across the entire internet, understanding this vast world. If we translate this to visual domains, Predict Next Frame still can't be achieved. Computer vision research has proceeded for decades without achieving this. This prevents many subsequent things — for example, generating a relatively long video that conforms to physical laws and logic is currently impossible. Similarly, making a general-purpose robot that can complete diverse tasks given a single instruction is currently impossible, because the visual domain still can't achieve true generalization.

So unified understanding and generation is extremely important. Currently we see a promising trend: models represented by GPT-4o, where users give instructions and it generates an image, and users can continuously input instructions for continuous editing. The required capability here is unified understanding and generation. First it must understand instructions; second, it must achieve editing according to instructions. When the model generates images, it must understand both text and images — this is extremely difficult. Although OpenAI hasn't disclosed details, we can see it has made great strides in unified understanding and generation. StepFun also has some progress in this area — we recently open-sourced a model capable of multi-round image editing.

We now feel that model technology development remains in a very steep upward phase. Every six months we discover the emergence of extremely disruptive technologies. On one hand, we see technology has matured enough for application deployment; simultaneously, we cannot ignore that this technology is still advancing rapidly.

03/

Building an Intelligent Terminal Platform

Enabling More People to Access StepFun's Comprehensive Model Matrix

Zhou: Large models remain very hot, but one direction is even hotter this year — AI Agent. How is StepFun positioned in this area?

Jiang: Agents are indeed very hot. Many say 2025 is the year of the Agent. I actually think the term Agent emerged in 2023, when there was an Agent architecture diagram. Why didn't it take off then, but became very hot in 2025? Its success relates closely to two factors:

First, AI Agents can handle very complex problems, which requires models with very powerful reasoning capabilities. Reasoning models emerged in the second half of last year, and by early this year, Agents gradually matured.

Second, they require multimodal capabilities, because Agents need to understand users' environments and task contexts, requiring models' multimodal capabilities.

This is the technical driving force behind AI Agents' explosive popularity.

As for what is an Agent? I think everyone has their own view, with some very lengthy descriptions of what an Agent is from various aspects. In my view, a very condensed definition is: a system capable of autonomously helping humans complete complex tasks is called an Agent. Looking further at what autonomy means: it includes two layers — automatic and proactive. Automatic means when completing a complex task, it does so as independently as possible, reducing or eliminating need for human intervention. Given a task, it can run on its own and deliver a result at the end — this is the automation process.

Proactivity is harder to achieve. People are accustomed to thinking about who can help them complete something and manipulating interfaces to achieve it, with the user typically being the task initiator. Imagine if there were a meeting software that automatically starts recording when a meeting begins and automatically generates summaries when it ends. During the meeting, if a superior suddenly raises a question you haven't prepared for, it automatically aggregates relevant materials and presents them — what a wonderful Agent this would be. So an Agent must possess both automaticity and proactivity.

Zhou: How is StepFun positioned in this domain?

Jiang: Currently we are focusing on intelligent terminal Agents.

Intelligent terminals are often extensions of human perception and experience. Now there's a very popular hardware device called Plaud, with tens of millions of dollars in revenue. It's a voice recorder, very cleverly designed to stick on the back of an iPhone and go everywhere with you. It can record anytime, such as during phone calls — this is an extension of human ears, allowing you to collect and organize information you hear anytime, anywhere. This shows hardware as an Agent can proactively understand users' environments and grasp task contexts — this capability is critical. So many smart devices possess this attribute as extensions of eyes or ears. For example, Qiming Venture Partners portfolio company Insta360 (688775.SH) is an extension of eyes. We also hope it can further become an Agent, where taking photos doesn't require pressing a button but simply saying "take a photo," or it understands when to shoot and when not to.

Additionally, smart devices often can help people complete tasks. For example, microwave ovens now have hundreds of functions that are difficult to operate without reading the manual. Suppose a chip is embedded in a microwave oven — it could be very user-friendly. The user says "help me steam an egg," and it completes the task itself. Its characteristic is being able to interact with users through natural language, understand users' environments and intentions, and automatically help users complete tasks. We ultimately hope to build an intelligent terminal platform that enables more people to access StepFun's comprehensive model matrix.

Zhou: As introduced earlier, I believe models are still rapidly evolving and iterating, with technical infrastructure rapidly changing and becoming more intelligent. Some investors I respect who experienced the internet era, perhaps for various reasons, believe one shouldn't invest in model companies but only in application companies with real revenue and commercialization capabilities. I feel China entered the internet era in its second half, when any internet startup hardly needed to worry about technical infrastructure issues and could focus on product-level innovation. The internet industry chain was very short: traffic above, advertising and other commercial monetization methods below. Today's AI era is still in its first half. There's still substantial optimization space at the model level or technical infrastructure. In a sense, as "model-as-product" demonstrates, a good Agent or model determines 70-80% of a product. In this era, will super-application companies emerge from enterprises like StepFun that master underlying model capabilities?

Jiang: I very much agree with what you said. I've also talked with many product managers who feel that product managers successful in the internet era may need to relearn everything in the AI era. In the internet era, technology was relatively certain while products were uncertain. Now both directions are uncertain — for example, how far technology can develop, and even harder, judging what intelligence level technology will reach in six months. Product development requires some forward-thinking; if building products based on existing technology, next-generation technology may disrupt existing products when it emerges.

So product managers' greatest dilemma is how to build new products on a highly uncertain technical platform. This is a question everyone must consider, and precisely because of this, this era is the best era.

Zhou: Thank you for your wonderful sharing.

Past Reviews

Qiming Star | StepFun and Yuanli Lingji Reach Strategic Cooperation

Qiming Viewpoint | Alex Zhou, Qiming Venture Partners: 2025 Will Be a Big Year for Comprehensive AI Application Deployment

Qiming Headlines | Qiming Venture Partners and Portfolio Companies Win Multiple Awards Including ChinaVenture 2024 Top 5 Best VC Firms

Qiming Venture Partners was founded in 2006. Currently, Qiming Venture Partners manages 11 USD funds and 7 RMB funds, with total assets under management reaching $9.5 billion. Since its founding, it has focused on investing in early and growth-stage outstanding enterprises in Technology and Consumer (T&C) and Healthcare sectors.

To date, Qiming Venture Partners has invested in over 580 high-growth innovative companies, of which more than 210 have listed on the New York Stock Exchange, NASDAQ, Hong Kong Exchanges and Clearing Limited, Shanghai Stock Exchange, and Shenzhen Stock Exchange, or exited through mergers and acquisitions. Over 80 companies have become recognized unicorns or super-unicorns in their industries.

Many Qiming Venture Partners portfolio companies have grown into the most influential companies in their respective fields, including Xiaomi (01810.HK), Meituan (03690.HK), Bilibili (NASDAQ:BILI, 09626.HK), Zhihu (NYSE:ZH, 02390.HK), Roborock (688169.SH), UBTECH (09880.HK), WeRide (NASDAQ:WRD), Insta360 (688775.SH), Gan & Lee Pharmaceuticals (603087.SH), Tigermed (300347.SZ, 03347.HK), Zai Lab (NASDAQ:ZLAB, 09688.HK), CanSino Biologics (688185.SH, 06185.HK), Schrödinger (NASDAQ:SDGR), MicroPort EP MedTech (688617.SH), Sanyou Medical (688085.SH), Amoy Diagnostics (300685.SZ), Berry Genomics (000710.SZ), Sinocelltech (688520.SH), Yuanxin Technology, Insilico Medicine, MediLink Therapeutics, LaNova Medicines, Zhipu AI, StepFun, Biren Technology, among others.