Former AWS Scientist Teaches Agents to Cooperate, Compete, and Even Argue | A Conversation with Raphael Shu, Founder of OpenAgents, on Collective Intelligence

**By Pippobei | Produced by AI Nao**

By Pippobei | Produced by AI Nao

Intro

With the capabilities and value of single agents now beyond doubt, how multiple agents can collaborate has become the next major trend for the second half of 2025.

Many see this as AI's second awakening.

The first awakening was marked by the birth of large language models — AI learned to understand, remember, and reason. The second awakening is multi-agent collaboration, where individual agents learn to converse, cooperate, divide labor, and even argue.

This also means agents are no longer isolated actors, but are gradually evolving into something resembling a nascent society.

Raphael Shu is an entrepreneur deeply immersed in multi-agent collaboration.

He began focusing on natural language processing (NLP) during his undergraduate and master's studies, and started researching neural-network-based natural language generation while pursuing his PhD in computer science at the University of Tokyo. Around 2016, when the field was still transitioning from "syntax to semantics," his research had already begun shifting toward exploring the "decision-making capabilities" of language models. He was among the earliest scholars to study the transfer potential of Seq2Seq models in language understanding and generation. "If a model can learn to transfer intent across different tasks, then it's no longer just a model — it's an agent capable of action," Raphael Shu said.

In 2021, he joined Amazon's AWS science team for his first career stop, participating in the R&D of conversational AI. A year later, he architected and implemented Dialog2API — AWS's first large-model-based agent system. At that time, the word "agent" hadn't yet become buzzworthy. "Many of my Amazon colleagues, including the clients I worked with, thought: isn't this just a smarter RPA?"

The watershed moment came in 2023. With the emergence of large language models and the launch of ChatGPT, the AI world was quickly immersed in the miracle of "language models." Many pivoted to training models through natural language rather than reinforcement learning. Companies across Silicon Valley began chasing bigger models, lower latency, more stable APIs, and exploring all kinds of application endpoints.

But Raphael Shu once again changed his research direction. "If AIs could collaborate through natural language, might an entirely new form of agent emerge?" This direction excited him far more. Although multi-agent collaboration had been explored by pioneering scientists as early as the 1990s — initially applied to optimizing how thousands of malfunctioning traffic signals in a city could coordinate most efficiently — he now pursued it at Amazon through enterprise-grade multi-agent research that "had already been integrated into product lines with engineering teams." Starting in 2024, Raphael Shu began thinking about open-world multi-agent cooperation. "I researched this for over a year, and it's one of the most worthwhile directions in AI right now — with only a one-to-two-year window."

So this scientist, who had been working in a big tech lab in Silicon Valley, decided to "leave" and start a company building an open-source platform where agents could understand, divide labor among, cooperate with, and even negotiate against each other.

He named it OpenAgents — an ambitious name. It launched in October 2025.

In its ideal form, OpenAgents hopes to redefine how agents collaborate with each other — and even the rules governing human-agent interaction. This closely resembles the vision proposed in the 1960s by Douglas Engelbart, known as the "father of the mouse": first connect humans with intelligent machines, then connect intelligent machines with each other, thereby achieving "collective IQ" (the term "collective intelligence" didn't exist yet). And the "mouse" was merely the simplest component of his vision — a small tool for human-machine interaction.

In any case, grand and ambitious visions always attract the attention of investment institutions, because they are full of uncertainty — a playground for adventurers. Right now, the research paradigm for multi-agent systems hasn't solidified, let alone a clear commercial ecosystem: who pays for collaboration? How is ecological order established?

"The core answer lies in speed," Raphael Shu said.

He believes that more powerful chips will emerge in the future, enabling AI-generated content to outpace human output by tenfold or even a hundredfold. The interaction speed between agents will also exceed human thinking speed.

"Perhaps reaching millisecond levels." Raphael Shu believes speed will solve many problems. Perhaps AI will eventually bring the world to a stage where humans cannot participate in real time.

  • Raphael Shu giving a talk in Silicon Valley

  • Raphael Shu attending the ACL conference in Florence, Italy

  • OpenAgents team photo

  • Product interface

Conversation with Raphael Shu

Part 01

The Evolution of Agent Collaboration

From Orchestration to Ecosystem

AI Nao: The industry has been talking about "multi-agent collaboration" lately. How do you understand "collaboration"?

Raphael Shu: I think there are two levels: engineered workflows and open ecosystems.

"Engineered" collaboration is characterized by a limited number of participating agents with fixed functions, and a relatively closed system structure. Microsoft's Magnetic One system falls into this category. In such systems, there is typically an "orchestrator" responsible for coordinating task distribution among multiple agents. For example, one agent writes code, another operates the browser, a third reads local files, and a fourth executes command-line tasks. These agents each assume different roles — some tasks execute quickly, others require longer processing times.

The whole system resembles a fixed production assembly line. The advantage is controllability and stable performance, but the downside is obvious — it cannot dynamically incorporate new agents based on external changes, nor can agents self-adapt in unfamiliar environments.

This leads to the second level: open systems.

First, real-world tasks themselves are uncertain, and goals can change — meaning the system must possess dynamic understanding and self-adjustment capabilities.

Second, the sources of collaborating agents become more diverse: different agents may be developed by completely different companies, teams, or even individuals. They use different protocols, model architectures, and training objectives. Enabling these "heterogeneous agents" to collaborate within the same network is an extremely challenging task.

Third, each agent has its own goals and value orientations. Their behaviors may not align, and may even conflict or compete. Therefore, the system needs to find balance among "multiple objectives" and "multi-stakeholder interests."

AI Nao: Could you give a concrete, understandable example?

Raphael Shu: If I'm an investment bank valuing Starbucks, the entire logic is clear, closed, and repeatable — so it can be modeled as a fixed agent. But if I change this to "valuing any company in the world," the Starbucks logic completely falls apart: Starbucks cares about coffee bean prices, Tesla needs to look at battery costs, Google requires analyzing advertising market structure. There's no fixed workflow that works universally.

So you should build an open system — an exchange where different agents, whether human or machine, can come together to博弈 and spontaneously form a consensus on company value.

This is what OpenAgents aims to do: move multi-agent collaboration from "engineered orchestration" toward "ecosystem construction."

AI Nao: At this stage, OpenAgents mainly targets developers. What value does it provide users?

Raphael Shu: First, helping users build a deployable agent network. Second, helping them connect their agents to the network. It's essentially network-layer infrastructure.

For example, say I want to build a community of multiple agents that maintains an auto-updating AI Wikipedia, continuously collecting the latest AI-related activities, lectures, offline meetups, or discussion groups in various cities.

I would first enable a "Wikipedia" plugin on OpenAgents, giving the system the ability to automatically organize and update information. Then I'd add "chat" functionality, allowing different agents to communicate and share information. Then I'd activate a "shared folder" plugin for uploading, storing, and editing materials. When these functional modules connect together, a complete agent network with information collection, communication, and collaboration capabilities is born. Then I can invite other developers to join.

  • Architecture diagram: Agent Network (left), Plugin System (center), OpenAgents Studio (right)

AI Nao: Are there more commercial landing scenarios?

Raphael Shu: We've recently been working with an AI recruiting startup called Peak Mojo. They do fully automated AI interviews — candidates simply upload their resumes and can immediately start a 12-to-15-minute online interview. After the interview, the system automatically generates results or HR can confirm them.

What we're doing is extending this AI interview capability into an agent community. Imagine 80 to 120 companies' AI interviewers simultaneously in the same community. A candidate fills in basic information and uploads their resume, and all these AI interviewers can see it. When one company is interested in this candidate, its AI interviewer might initiate an interview and ask: what open-source projects have you participated in on GitHub? The candidate answers: I've done projects in Python. That answer gets shared across the entire community. Other companies' AI interviewers won't ask the same question again.

This way, a candidate might receive 30 interview invitations from different companies in a single day. Each interview takes only 15 minutes; working 8 hours a day, they could complete all interviews and even receive offers the same day.

This "AI interviewer community" demo version is already live. Next, I hope to get full validation.

This is just one starting point among many applications for OpenAgents, but it already demonstrates the potential of "collective intelligence."

Part 02

Building the Ecosystem

Building the Basketball Arena, Not the Team

AI Nao: If agents can collaborate, then a new form of collective intelligence emerges. When thinking about collective intelligence, you've said that The Wisdom of Crowds influenced you the most. Is this because you believe human "collective intelligence" is being reconstructed by AI?

Raphael Shu: It mainly clarifies one point: when the number of individuals reaches a certain scale, the system should not rely on single directives or processes, but can achieve self-coordination through博弈 mechanisms.

In other words, when there are more and more agents, the best solution may not come from a single agent's reasoning, but from their interactions, arguments, and trade-offs.

Take the company stock valuation scenario I just mentioned. If multiple agents each approach from different angles and debate with each other — one focusing on finances, one analyzing markets, one assessing risk — and through continuous debate and博弈 reach a conclusion, the result is often more accurate than what any single model could reason out.

Let me give a more concrete example.

Suppose a company just bought an office floor and now needs to design the layout. There are two approaches: find one expert, or find ten experts from different fields. The safety expert says: the corridors are too narrow, people won't escape in a fire. The aesthetics expert says: that wastes too much space. They keep discussing and revising until reaching a balanced solution that satisfies all parties.

This is a process of collective optimization through博弈.

AI Nao: If human social collaboration is built between consensus and博弈, then in the world of AI, how do you make this "collective decision-making" work?

Raphael Shu: It's not about "how to divide labor," but "how to design rules." If agent collaboration only does division of labor, system growth will definitely be limited.

For example, a user uploads a Word document; the system needs to convert it to PDF, then compress it by 50%. There are two agents: Agent A handles format conversion, Agent B handles compression optimization. After task completion, how should the system "reward" them? For instance, whoever contributes more to performance or efficiency gets more reward; whoever completes the task better gets higher ranking or revenue share.

Once rules are set, countless agents can autonomously enter, autonomously exit, compete or cooperate, forming a positive growth cycle with self-evolution capabilities.

AI Nao: There are many teams working on "multi-agent collaboration" frameworks in the industry, such as AutoGen, CAMEL, and LangGraph. How does OpenAgents' approach differ from theirs?

Raphael Shu: There's an essential difference in positioning.

AutoGen, CAMEL, and LangGraph help users assemble an agent team — they want to help you build an NBA team. We're building the basketball arena, where many, many teams come to play. So we're not in competition with them; we're complementary.

In other words, other frameworks focus on task-level orchestration, while OpenAgents focuses on infrastructure. We care more about how to enable countless agents to coexist, collaborate, and communicate smoothly, forming a community ecosystem.

AI Nao: Building the basketball arena rather than the team means you're establishing an ecosystem, even redefining rules, and you need enough teams to join. What are your current priorities?

Raphael Shu: Enough tools that are good enough to use. We call them "plugins" or "Mods." Plugins can be tools, rules, or even social features or games.

For example, we enable multiple agents to write the same document in real time, share materials, or process files. We're building a social plugin: letting agents play RPG games. Not for entertainment, but so agents can meet new partners in the game, learn cooperation methods, and find potential collaborators. There are also rule-setting plugins: when a new task emerges, who assigns it? Which agent has final decision-making authority? How is the incentive mechanism designed?

Another issue is that different agents use different communication protocols. Some agents can communicate directly in natural language, connectable via HTTP or WebSocket. Others have more complex structured data needs and require special communication protocols. Regardless of which protocol or tech stack they use, once connected to the OpenAgents network, they can seamlessly interface with other agents.

So we need to be open source, because OpenAgents needs a massive tool ecosystem. It might take us two months to develop a plugin letting agents play RPG games. As the community grows, perhaps 2-3 new plugins could be born every day, eventually growing to thousands of plugins.

Part 03

Speed Decides Everything

Who Participates, Who Watches

AI Nao: Around 2023, when the industry was just beginning to understand agents, you had already started pivoting to multi-agent research. The entire industry, especially technology development, wasn't moving as fast as today. How did you overcome technical bottlenecks?

Raphael Shu: Give a model a 500-word task description and it could understand immediately? The large models back then couldn't comprehend such instructions at all. So we adopted an approach called "in-context learning" — instead of directly telling the model "please execute this task," we showed it numerous examples and let it summarize the patterns itself.

Actually, the trickier problem was the model's "memory." Today's models can process millions of tokens; back then there were only about two thousand. If the conversation went on slightly longer, it would forget the context. So we also had to carefully select, compress, and rewrite training samples, letting the model learn complex tasks within extremely limited context.

So entering 2025, has the industry reached consensus that agent collaboration is inevitable? Or is there still an argument that each agent will have its own independent ecosystem, or that a super-agent will emerge?

There is indeed divergence in the industry. If you can interview big names on this, I'd love to hear their perspectives.

But my view is: collaboration is inevitable, because of "resource constraints."

For example, there are specialized financial analysis companies in the United States that have decades of accumulated financial analysis experience and proprietary data. They are fully capable of developing an agent specifically for analyzing listed company valuations — something other companies cannot do.

Therefore, while I believe "super agents" will emerge, and agent capabilities can expand infinitely, the resources and professional knowledge that agents can access cannot expand infinitely.

AI Nao: The famous Stanford Smallville experiment was the first time agents exhibited social behavior in virtual space. Does this experiment intersect with your entrepreneurial direction?

Raphael Shu: I believe "Stanford Smallville" is a very important but severely underestimated research direction.

Stanford Smallville could actually be well applied in enterprises. For example, Amazon could build a community composed of buyer agents and seller agents, letting them autonomously trade, price, and communicate. Through the operation of this virtual market, they could gain insights into real market trends. This is a prediction method closer to "reality" than traditional data analysis.

OpenAgents can directly provide enterprises with the underlying framework needed for such predictions, bringing this simulation capability into real scenarios.

AI Nao: If your ideal multi-agent collaboration ultimately takes shape, the future will become a human-machine collaborative society. Might humans no longer be the central controller, but rather a node, a participant, or even become part of an agent?

Raphael Shu: Isn't there a saying that humans should think about whether they can become a valuable MCP? (laughs)

I think the key question isn't whether humans and agents can collaborate, but whether humans can keep up with agents.

Ultimately, speed decides everything. For example, a human team might take 15 minutes to develop a feature; but in the future, an agent might complete it in 0.05 seconds. In that case, it's very likely humans simply won't have time to intervene before the agent has already finished the task.

AI Nao: What happens when agent action speed exceeds human reaction speed?

Raphael Shu: It will lead to a new social structure: continuous interaction and evolution between agents and agents, with human participation becoming increasingly low. Then we might need to rethink: can what we call human "collaboration" still be called "collaboration"? Might we no longer call ourselves collaborators, but supervisors?**

AI Nao: Finally, please recommend three books.

Raphael Shu: How to Win Friends and Influence People, Getting Things Done: The Art of Stress-Free Productivity, and Machine Learning: A Probabilistic Perspective. The third book has been updated through several editions; it truly taught me machine learning.**

Image sources | Provided by interviewee, Unsplash