Between a world of temptation and going all-in, Nano AI's choice is...

Ditching the "One-Size-Fits-All" for Specialized Expertise

Ditching "Universal" for "Specialized"

👦🏻 Authors: Jingshan, Xiaoju

🥷 Editor: Koji

🧑‍🎨 Layout: NCon

General-purpose AI agents can do everything, but that creates its own problem: users often can't remember when to use them.

Leading products are each finding their own solutions:

Manus launched Playbook, turning common scenarios into templates and leveraging community power to continuously output new ones.

Genspark launched Slides, Sheets, Docs, Pods... as individual standalone products, reinforcing distinct mental models for each scenario.

Nami AI also made a clear choice: focus on one specific scenario and take "using AI to make videos" to the extreme.

On August 6, Zhou Hongyi launched a major upgrade to Nami AI in Beijing: "Nami AI Multi-Agent Swarm." The entire event emphasized just one thing: generate complete videos from a single sentence.

Nami AI is going all-in on this single point. No more pushing the idea that Nami's agent "can do everything" — instead, it's laser-focused on one goal: letting ordinary users who can't edit or write scripts create high-quality videos with one click.

This is a deliberate strategic convergence:

  • Ditching "universal" for "specialized"
  • Using the actual product to reshape user mental models: "For AI video, use Nami"

Koji (far left) was invited to this 150-minute launch event

🚥

Throughout the event, Zhou repeatedly emphasized: "One-sentence input, one-sentence generation, one-sentence usability."

To kick things off, Nami played a 10-minute long video generated entirely by its AI in one go, demonstrating its agent capabilities:

Then, for the first half hour of the launch, the company kept playing various Nami AI-generated videos, letting the output speak for itself.

This time, Nami AI Multi-Agent Swarm introduced a new concept and chose a more foundational approach — orchestrating Agent Infra — while also doing quite a bit of innovative work.

Next, we'll share our in-depth review of this product.

How Do You Get an Agent to Generate The Godfather 4?

The traditional AI agent answer: manually integrate various tools.

The new answer: one sentence.

Every new AI tool chooses a specific scenario to demo its capabilities. This time, Nami AI's "main battlefield" is "generate a complete video from one sentence." Under the Swarm framework, Nami AI uses the A2A protocol, a swarm-style scheduling center, swarm-style collaboration mechanisms, and so on. The end result: the multi-agent swarm achieves a pretty impressive success rate.

When I first tried this tool, I kept wondering: how "complete" is its "complete"? After all, video generation models like Veo3 and Runway Gen-3 already produce pretty impressive visuals.

I briefly summarized Nami AI's differentiation: it integrates all of these functions together — research, scriptwriting, style selection, aspect ratio adjustment, voice selection, copywriting, voice synthesis, image generation, background music, lip-syncing, camera angles, video generation, and more.

Enough talk — first, check out the results of our 4-minute 13-second black-and-white The Godfather 4.

My prompt was simple:

Make a black-and-white film video starring Marlon Brando as the Godfather, write a "Godfather 4" movie

Nami AI link: https://bot.n.cn/share/mcp?id=og6lsz&from=pc&src=360_llq

The consistency across shots and continuous scenes in this cinematic AI video is remarkably high — especially the storyboard pairings with BGM, voiceover, and designed story scenarios.

Honestly, was there a moment when you didn't realize this was AI-generated?

I captured a few clips from the video so you can see for yourself what this "completeness" looks like. In the funeral scene, Nami AI chose a pull-back shot; for the segment about the Godfather's final life decision, it paired a detailed contemplation scene.

The direct-to-camera monologue was handled quite naturally:

There's also a plot point about a distant relative sending a letter to Vito. Look closely at this transition effect, and the camera push-in as the frame moves to the envelope.

From character close-up to envelope push-in, the transition is fluid and natural, very much in the language of cinema:

After this complete video was generated, I discovered that Nami AI had already carried out hundreds of steps of thinking, execution, and integration work.

I roughly counted: behind a 4-minute video lies approximately 200 steps, consuming 11,332,835 tokens. That over-10-million-token information integration volume is genuinely staggering.

Next, let's do a deep dive into how Nami AI works under the hood.

Traditional AI video generation models can typically only produce "single shots, single style" — still a long way from "complete." To create a "complete video" with narrative arc, rhythm, and plot development, users usually still need to stitch shots together themselves, conceive scripts, composite voiceovers, add music, and edit and color-grade. The whole workflow is complex, with a high barrier to entry.

So first off, Nami AI provides many L3-level agents. Some of these are available off-the-shelf online, but the ones that actually work well are basically all self-trained within Nami AI's own framework.

This time, Nami AI Multi-Agent Swarm (my understanding: it can call upon large numbers of agents to work simultaneously) refers to L4-level agents — for example, "generate a blockbuster from one sentence":

Generally speaking, Nami AI's task execution starts with DeepResearch, which was also the foundational feature when we tested Nami AI previously. Because it supplements so many information sources, I don't need "good prompts" at all — a single short sentence suffices.

After receiving the prompt, Nami AI activates its first agent — "Master Screenwriter" — then uses this agent's search capabilities for deep research, and based on this information rapidly simulates a video script and BGM:

After dozens of steps of thinking and refinement, large numbers of other agents from various vertical task domains begin working.

Such as:

【1】Storyboard agent creates video storyboard scripts (47 shots);

【2】Voiceover agent uses MiniMax MCP to voice dozens of shots;

【3】Nami AI image generation agent generates images for each scene based on prompts;

【4】Camera agent generates dozens of videos from images;

【5】Editing agent is called in to add voiceover to video.

Among these, the script is the core.

I have to say, this video storyboard script: The Godfather 4: Echoes of Destiny is genuinely well-written.

For example, Nami AI chooses to open directly with "the secret revealed after the Godfather's funeral" to spark audience curiosity. And the structure is tight, with progressive information delivery. Every page advances the plot; there are no wasted frames.

What I personally found most impressive and most worth noting: many steps throughout this process run in parallel rather than rigidly proceeding step by step; during each agent's task execution in the intelligent assistant swarm, there are intervals of dozens of "large model thinking steps."

Words alone are somewhat pale here, so I recorded a GIF that's more intuitive:

Throughout the workflow, hundreds of steps execute continuously, yet the final success rate remains high, with no large-scale "AI hallucination stacking" — genuinely surprising.

This "Agent Swarm" feels very much like hiring a professional team that understands efficient collaboration; they work together seamlessly.

Because Nami AI's "complete video task" capability is genuinely intriguing, other members of the Crossing team also tried it out extensively on their own.

Writing a New Episode for Doraemon

We again selected the "generate a blockbuster from one sentence" intelligent agent swarm, and input this prompt — just this:

Write a new episode for Doraemon

The script structure here is somewhat special: Nami AI divided it into video dialogue and image generation prompts. The former collaborates with a "super voice actor agent" for subsequent voiceover, while the latter feeds instructions to an "intelligent drawing expert agent" to ensure consistency in image generation.

The image generation for each line of video dialogue is already confirmed:

Check out this storyboard below. The overall effect is quite good — strong character consistency throughout, and the storyline fits the original characters well.

The full video is 41 seconds, using 15.3 million tokens.

Nami AI link: https://bot.n.cn/share/mcp?id=x1dpfr&src=nm_share_agent_chat

Donald Trump Falls in Love with Me, the White House Cleaner

Let's look at a recent short drama that's reportedly extremely profitable and has blown up globally: "Donald Trump Falls in Love with Me, the White House Cleaner."

We gave it a very detailed prompt:

Donald Trump falls in love with me, a cleaner. In the resplendent White House hall, a young female cleaner is diligently wiping the marble floor. Suddenly, Trump strides in with exaggerated, confident steps, preparing to deliver a grand speech, but freezes the moment he sees the cleaner. He slowly approaches and, clumsily yet gentlemanly, offers a silk handkerchief. Their eyes meet under the lights, time seeming to slow to a crawl. Then, Trump takes the cleaner on a tour of the White House's corridors and gardens; roses and golden hair intertwine in the sunset into a strangely romantic tableau. At night, they gaze at fireworks together; Trump tentatively reaches out, and the cleaner smiles shyly, gently taking his hand. Finally, against the backdrop of night and interweaving fireworks, their silhouettes gradually disappear down a garden path. Visual style: American comic style. Background: resplendent White House hall, majestic corridors, warm dreamy sunset garden, brilliant fireworks in the night sky. Character design: Trump with golden hair, in a suit, exaggerated expressions; young, plain cleaner in neat uniform, with a shy smile. Cinematography: wide shots, close-ups, extreme close-ups, low angles, high angles, POV shots interwoven, soft lighting, dramatic compositions. Voiceover suggestion: Narrator: humorous yet warm adult male voice, with a tone of gentle teasing and unexpected romance. Trump's voice: confident, slightly exaggerated American male voice, occasionally with comically deliberate pauses. Cleaner's voice: natural, sincere, slightly surprised young female tone. Background music: light strings and piano interweaving, shifting from light humor to warm romance as the plot develops, ending with gentle harp overlay to create a fairy-tale conclusion.

This video is truly dramatic:

Nami AI link: https://www.n.cn/share/mcp?id=hxv86i&from=pc&src=360_llq

Especially these two shots — peak artistry achieved:

The AI not only grasps the narrative rhythm of this short drama genre but also makes some fairly creative choices in the details. The design of several key shots is genuinely interesting.

Agent Swarm Factory

Beyond these already "packaged" agent swarms, Nami AI — to handle complex tasks — also excels at assembling teams for collaborative work. In the "Create Agent" interface below, we can build agent swarm workflows by dragging and connecting nodes. The overall flow should be quite familiar to everyone.

Entry point is "Create Agent" in the upper right corner

Select "Create Multi-Agent Swarm"

However, there are some differences.

Traditional AI workflow platforms still require low-code for some nodes, especially critical ones. But in Nami AI, all nodes are designed as agents (i.e., large model nodes).

For example, we designed a "Crossing AI Digital Human Presenter" agent swarm.

All nodes inside are L3-level agents available within Nami AI's internal options. Each L3 agent basically comes with pre-configured prompts already, so I essentially just did "workflow conception" and "connecting lines" — simple operations.

For instance, "Douyin Trending Tracker" is an agent that Nami AI has designated as L3-level. Its system prompt is already quite complete, no need to modify; MCP tool configuration is also already done:

The workflow logic roughly works like this:

【1】User inputs a prompt and image about wanting to research the AI field; agent extracts keywords;

【2】Six agents query information from various platforms;

【3】All information is integrated, scripted, voiced, and digital human created;

【4】Video is generated.

The most critical intermediate node is: Self-Media Content Quality Analyst, whose role is to integrate all information.

As demonstrated below, if you connect the "workflow lines" between nodes, it automatically senses inputs — unlike traditional workflow platforms where you'd need to manually select. Not a single node requires code input; you only need to feed prompts to the AI:

The full workflow overview looks like below — a relatively simple type. After trying it out, we found it contains many possibilities; everyone should explore more:

After setting up the workflow, I randomly captured a photo of Koji and a brief prompt, telling it to help me research "AI hardware":

Then this agent swarm began using various information-querying agents to gather content from various platforms:

After integration through a series of Nami AI video agents, a quite commercially polished digital human promotional video came out — the AI hardware it researched was: AI acceleration cards.

Video result:

Nami AI link: https://bot.n.cn/share/mcp?id=3j116n&from=pc&src=360_llq

Clearly, the overall copywriting, digital human (facial expressions, body movements, lip-syncing) are all quite well-executed.

Nami AI Multi-Agent Swarm is quite well-suited for this kind of "short video self-media workflow."

In our testing, we found that Nami AI maintains a high success rate even after hundreds of steps of execution. Behind this is actually a simple math probability problem: if a single agent's success rate is 90%, and a single task requires over 100 steps, the success rate drops to just 59% after five steps, and to two in a million after 100 steps.

In multi-step workflows, error rates compound exponentially.

So to achieve a 95% task success rate, each step needs at least 99% completion — this is the confidence behind Nami AI's multi-agent swarm.

Now, let's briefly summarize Nami AI's technical capabilities:

Under the swarm framework, through the A2A communication protocol, swarm-style scheduling center, swarm-style collaboration mechanisms, and so on, Nami AI's L4 multi-agent swarm achieves very impressive performance metrics:

【1】Continuous execution of 1,000-step tasks

【2】Token consumption: 5-30 million

【3】Task success rate: 95.4%


At the live event, we frequently heard Zhou mention a key phrase: user imagination.

From "The Godfather 4" to "Doraemon," from short dramas to AI presenters — each case isn't technical showboating, but rather repeated attempts to answer one question:

If a group of agents knows how to collaborate, can they help an ordinary person turn the inspiration in their head directly into reality?

We believe: the value of AI isn't in replacing humans, but rather in giving everyone their own "right to create" as the barriers to expression keep lowering:

"I have a story" — AI can help me tell it.

This time, Nami AI is betting on video, betting on creative scenarios, betting on users' genuine desire to express. Not all technology needs to be the most comprehensive, but someone always needs to go the deepest, the most stable, the most thorough in a particular direction.