AIGC 'Peak Series' | ModelBest's Li Dahai: We Are Entering a New Era of Collective Intelligence
A single fish doesn't have much intelligence, but a school of fish can demonstrate collective intelligence as a group.


The ChatGPT boom earlier this year rapidly kicked off an AIGC entrepreneurship wave brimming with imagination. A wave of large model startups emerged in China. Now, half a year later, the conversation around large models and AI has evolved into a different phase: from "discussing concepts and talking about the future" to "showing products and looking at real-world deployment." With the underlying technical logic validated and widely accepted, AI entrepreneurs are repeatedly testing application products built on large models, while those outside the industry remain on the sidelines, wondering just how much these deployed products can actually change specific business operations across industries and enterprises.
Recently, at a Code Brain · Private Chat session, Source Code Capital and Amazon Web Services jointly invited Li Dahai, co-founder and CEO of ModelBest and Zhihu partner and CTO, to share his views on AGI. He proposed a development trend from individual intelligence toward collective intelligence, outlined six typical characteristics of AI agents, and exchanged ideas with numerous entrepreneurs on large model-related startups in China.
Li Dahai graduated from Peking University's mathematics department and was among the founding members of Google China. He later served as engineering lead at Yunyun and search technology lead at Wandoujia, with over a decade of continuous entrepreneurial experience. In addition to his role as CEO of ModelBest, he currently serves as Zhihu partner and CTO, leading a team of more than 1,500 people, with extensive experience in building top-tier technical systems, strategic planning, technical management, and commercialization.
Below are the main takeaways from the private chat session:
01 AGI Is Humanity's Fourth Technological Revolution
Large models represent a concentrated release of years of accumulated progress across three elements: AI technology, data, and computing power.
From the initial proposal of the Transformer architecture, to Google's introduction of the BERT pre-training paradigm and its effective deployment, to OpenAI's unwavering commitment to the "compression is intelligence" approach that brought this dawn to the world — it may look like a series of coincidences, but it's actually the result of gradual accumulation, with each generation building on the shoulders of those before them.
This year's accumulation of large model capabilities marks what I'd call year one of AGI.
AGI's emergence has been described as the fourth wave of human technological revolution because, since the Industrial Revolution, the core of technology has always been about improving the efficiency of humanity's tools for shaping the world. We once used hoes to till fields, then switched to oil-driven machinery. Though the energy sources and tools changed, the underlying logic remained constant: using tools and energy as leverage to boost efficiency.
The information revolution achieved a kind of virtual reality mapping, creating a digital world that serves as our digital twin of the physical world. The AGI revolution, meanwhile, sits at the frontier of seamlessly connecting these two worlds. At present, the physical and digital worlds remain disconnected in many respects, but the rise of AGI signals this gap steadily narrowing. We currently rely on complex business intelligence systems and reporting systems to extract insights from digital abstractions for decision-making; advances in AGI will directly enhance this decision-making process, making interactions between the two worlds more direct and natural.

Many decisions and tasks still require humans to serve as the bridge between these worlds, which remain somewhat estranged. The AGI revolution can bring about complete integration of these two worlds — a remarkable development that will substantially boost both humanity's efficiency in transforming the world and in transforming itself.
02 Six Typical Characteristics of AI Agents
Speaking of large language models, especially ChatGPT's development and OpenAI's latest releases, we can see this technology moving toward real-world application. Agents are certainly an excellent way to deploy large language models, because they make application deployment more convenient and efficient.
Here are what I consider the more important characteristics of agents:
First, external memory. AGI essentially combines information with large models through retrieval; abstractly speaking, this is also a form of memory, similar to a computer's hard drive, while the large model itself is more like RAM.
Then, external tool use. Using tools can dramatically expand a large model's capabilities. Using tools to check the latest stock information and help users execute corresponding stock operations based on their instructions clearly expands its capabilities in financial work. It can also check flight schedules, book tickets, or order takeout. With the ability to call external tools, large models become capable of virtually anything.

Agents have six typical characteristics, or six dimensions we need to consider: IQ, EQ, persona, growth potential, perception, and values. Perception is particularly important. For instance, large language models currently only perceive text, but to communicate well with people, they need to at least continuously see people's expressions through video to better understand emotional changes.
Persona relates to stability, because sometimes EQ and IQ don't necessarily go hand in hand. You may have recently seen an interview with Character.AI where they explicitly stated they don't need particularly high IQ in their models and haven't specifically optimized for it, but they've invested considerable effort in EQ optimization — EQ absolutely must be high, because they're focused on companionship. Growth potential is also a crucial capability; future embodied AI will place heavy emphasis on this.
03 From Individual Intelligence to Collective Intelligence
AI agents can do many things. What we can already see includes digital human livestreaming, virtual characters, ACG content, and emotional companionship, among others. Additionally, using AI agents as industry experts holds considerable promise. Some banks and government service halls, for example, are already using digital humans and intelligent agents as virtual staff for customer service. Agents can assist people with tedious, repetitive labor.
AI agents can also serve as offline assistants. Once embodied AI combines with robots that have locomotive capabilities, this will certainly give rise to massive new endpoints — an opportunity with tremendous potential. But when exactly will this explode, when will the technology fully mature? Keep watching this space.
What I've discussed so far is individual intelligence, which naturally leads to a question: Can agents evolve from individuals to groups?
Not long ago, I met a professor from The Chinese University of Hong Kong who studies oceanography. He researches swarm intelligence in fish and discovered that a single fish doesn't have particularly high intelligence, but a school of fish as a collective can generate high intelligence.
In nature, whether it's schools of fish, bee swarms, or ant colonies, the amount of pheromone exchanged between individuals seems negligible. Yet it's precisely these trace amounts of pheromone exchange that enable the entire group to find food sources via the most efficient routes. Faced with obstacles, ants can cooperate to build bridges to overcome them, demonstrating advanced intelligence at the group level — intelligence that naturally emerges from collective behavior in complex environments.
Whether individual or collective intelligence, agents based on large models will become the dominant form of internet applications in the future. The future internet will be an era where AI agents connect everything. This is a concept we've recently proposed, called Internet of Agents, or IoA.

I believe that as time goes on, every phone manufacturer will eventually build their own integrated agent to accompany users — this could happen as soon as next year, but it will come sooner or later. Similarly, home refrigerators, televisions — all will become agents in the future, capable of conversing with people like humans and handling complex tasks and demands.
There's a question on Zhihu: "What would the world be like if all animals could talk?" It's quite unimaginable. I think the future will really become like this — all appliances in your home can become agents. We can interact with them in completely natural ways, have them work for us, and even chat with them for comfort when we're feeling down.
ModelBest attaches great importance to AI agent development and has been making significant moves in this direction for quite some time. We currently have three flagship innovations in large model-driven AI agents.
First, AgentVerse, an agent collaboration platform. It constructs a rich virtual space populated with numerous AI agents playing various expert roles — marketing expert, programming expert, design expert, bug-fixing expert, and so on. When users enter with a task, agents immediately initiate a team formation process, a strategic recruitment phase to determine which experts should be assigned to a particular task. Once these experts form a team, they begin negotiating task details and clarifying division of labor. After negotiation, they move to execution, with each agent completing work according to its role, then integrating results.
Throughout this process, there's also a strategic planner ensuring all agents' work remains coordinated and converges toward a final deliverable. This deliverable is compared against user requirements, and if significant deviations exist, it may enter iterative improvement. This framework's generality allows us to conduct extensive work on its foundation. In one AgentVerse application, we designed a virtual classroom scenario including a professor lecturing and students participating through raising hands and discussion, simulating the interactive process of a traditional classroom.
Second, XAgent, a super-agent capable of handling relatively complex work. Based on planning with dynamic instruction, it distributes and executes tasks, playing the role of an agent expert. XAgent can plan according to human needs, determining how many steps are required to fulfill a user's request. On this planning foundation, if user input is insufficient, it can interact with the user to obtain necessary information. After completing the plan, at each execution step, it can assess whether additional work is needed after completing that step — it's a dynamic structure of this sort.
For example, when you give XAgent the instruction "I have some friends coming over this weekend, can you recommend some restaurants?" this advanced language model won't immediately spit out a long list of restaurants. Instead, it will first probe your preferences, asking whether you prefer a quiet environment or specific types of cuisine, to understand your needs. Its first step is to interact with you, not immediately execute. Then, based on your response, it searches for restaurants; next, it organizes search results and presents several options with pros-and-cons analysis. Once the options are ready, it displays them in table form for your selection. After you choose, it connects to APIs to directly book the restaurant for you. This differs from the single-step Q&A mode we're typically familiar with; it demonstrates a superior experience that agents can provide.
Third, ChatDev, a collective intelligence software development platform based on the AgentVerse multi-agent framework. Building on the underlying agent technology logic, ChatDev helps us construct a virtual AI software company, setting up agents in roles such as CEO, CTO, product manager, programmer, and designer, connected through a communication network called the "conversation chain." These roles' interaction flow aligns with the waterfall model in software development, including software design, system testing, and documentation. We have these AI agents interact according to clear division of labor to meet users' specific needs. When we narrowed this multi-agent system's application scope from broad to specific, the product quickly gained market traction. Within just two months, ChatDev's GitHub stars surged past 18,000, topping the Trending rankings for multiple consecutive days.
Our three-pillar AI agent lineup is gradually being deployed, though there's still more room for improvement. We're continuously deepening R&D and regularly releasing iterative features and versions. When a product's functionality is so broad that users struggle to choose, they often feel overwhelmed; a product with clear positioning and specific functionality is more likely to gain enthusiastic market response. This echoes our product development philosophy from early entrepreneurship: a universal product isn't necessarily truly valuable, while solutions focused on specific needs tend to hit the target more effectively.
ModelBest
ModelBest spun out from Tsinghua University's Department of Computer Science NLP Lab. The founding team comprises experienced entrepreneurs, industry-leading AI technologists, and renowned NLP scientists and scholars. Core technical team members come from China's top natural language processing research labs, all holding doctoral and master's degrees from prestigious universities, with over a hundred papers published in authoritative domestic and international journals and conferences, and multiple patents granted — their research and technical capabilities rank among China's leading tier. Additionally, senior talent with extensive industry product and commercialization experience has joined to spearhead large model applications across scenarios and domains.




