A Conversation with Zeng Xi, Former Senior Director at Doubao and Founder of Chance AI: The New Battleground for Visual AI Isn't Image Recognition — It's Mind Reading

**Biology | Visual Nerves | Human Consensus | Interpretation | Entry Point**

Biology | Visual Neural Networks | Human Consensus | Interpretation | Entry Point

Produced by | AI Nao

01.

Picture this: you're sitting in a Bangkok restaurant, phone pointed at a Thai menu. In 2026, AI won't just translate Thai into Chinese — that already feels like last decade's trick — it'll remember you're vegetarian, pick out the three meatless dishes from twenty options, place your order in standard Thai, and slip the highest-rated dessert shop nearby into your afternoon itinerary.

It's like hiring a personal assistant who gets your taste, speaks the local language fluently, and happens to be a food blogger. Except it lives in your phone, accessible through a single snapshot.

This is the bet many entrepreneurs are making for 2026: not just getting AI to accurately recognize things, but getting it to understand why you're photographing this thing and what you want to do next.

Chance AI is切入 this direction. Founder Xi Zeng has an interesting background: he pursued a PhD in cognitive science and contemporary art in Barcelona, researching a question that sounds artsy but is fundamentally hardcore — why do humans feel melancholy when looking at Picasso's Blue Period paintings?

It touches on the essence of human visual systems: our eyes are cameras, but our brains convert visual signals into emotion, memory, meaning.

Now Zeng wants AI to learn this same skill: "building a visual reasoning system for AI that transforms visual signals into valuable, judgment-based interpretations."

Chance AI's product logic is straightforward: user takes a photo, app auto-identifies, then provides interpretation. As AI learns more about you, interpretations become increasingly personalized.

Take the same Picasso photo at an exhibition:

  • A child might get: Who was Picasso? Let's review what yesterday's art class covered, and here's something you could try painting tonight.
  • An art enthusiast might get: Similarities and differences between Picasso and Munch, which other exhibitions in the city suit your taste? Want me to book tickets now?

Zeng has a formula for Chance AI's core logic: Visual signal recognition + Personalized context + Social consensus = Meaningful value

Sounds abstract? Let's break it down.

Say you casually photograph a concert poster. For AI, this isn't just "a piece of paper with text and graphics" — it's an engineering problem waiting to be solved:

What concert is this? (Visual signal recognition)

Are you a fan of this artist? (Personalized context)

Are tickets easy to get? Is it worth going? (Social consensus)

Then, derive action:

When do tickets go on sale?

Add to calendar?

Set a reminder for sale day?

"We want AI to grow eyes that can think," Zeng says. "See the unseen — perceive what's beneath the surface."

Zeng carries a fascinating hybrid energy. He can explain visual cortex mechanics in neuroscience terminology, riff on the aesthetic philosophy of British versus Chinese royalty with dark humor, discuss supply chains and PMF in hardware jargon, and profess his love for Orange Ocean — an English-only band from Qingdao, Shandong.

After graduating, he worked at OnePlus, then OPPO, with his last role as Senior Director on ByteDance's Flow team.

In 2024, when GPT-4o's multimodal model launched, Zeng received a clear signal — this technical direction was converging on the problem he'd studied during his PhD: how human visual systems generate meaning.

That's the origin story of Chance AI.

  • Xi Zeng

  • Exhibitions are a common use case

02.

Chance AI has accumulated 200,000 users so far, 40% in North America. The usage barrier is minimal: photograph, identify, interpret.

At the technical foundation, Zeng made a contrarian choice. "The biggest misconception in the industry right now is trying to solve complex visual reasoning with a single model. That's impossible."

He modeled the engineering side after biological visual mechanisms, breaking reasoning into four steps — just as the human brain processes visual information through primary visual cortex, shape recognition, semantic understanding, decision planning, and other stages.

How effective is this approach? According to Zeng, on MMMU-Pro — currently the most rigorous professional-grade multimodal reasoning benchmark — Chance AI scored 86.07%, the highest known score to date. For comparison:

  • Gemini 3 Pro: 81.00%
  • GPT-5.4: 78.00%
  • Claude Opus 4.6: 75.00%

To dispel skepticism about "internal testing," the team recently open-sourced their underlying API, packaging it as a CLI tool callable by other agents, inviting academia and developers to verify scores themselves.

Chance AI is still at a very early stage. Zeng acknowledges that large-scale explosion of VLM (vision language model) applications still requires waiting for (or proving) three things:

First, "seeing" is not low-frequency behavior. Vision will become the next generation's interaction entry point — just as touchscreens replaced keyboards a decade ago.

Second, converting "seeing" into genuine "action." Recognition is step one, understanding is step two, but ultimate value lies in — can AI help you get things done?

Third, establishing irreplaceability outside giants' system capabilities.

"Conversation with Xi Zeng"

AI Nao First question — an AI vision popular science question: AI can already help us write papers, solve Olympiad math, but still struggles to judge "a cup of steaming water shouldn't be touched." Why is something even babies understand high-difficulty for AI?

Xi Zeng I need to introduce a concept first: anything humans see has more than just its surface layer.

Why does a Bugatti Veyron cost more than its weight in gold? Or a hypebeast T-shirt — maybe just an extra logo — costs way more than an ordinary tee?

So getting AI to truly understand requires three layers:

Layer one is perception, visual recognition.

Layer two is context — where did this come from? What have you experienced? Why does it matter?

Layer three is social consensus. Driving a Bugatti Veyron, for instance, signals wealth.

Perception + context + consensus determines a thing's value.

But most industry products today stay at layer one, because the common approach uses one model to solve complex visual reasoning — completely impossible.

What we're doing is getting AI into layers two and three, like understanding why in human society a Bugatti Veyron costs more than its weight in gold.

AI Nao

You previously made an interesting judgment — breakthroughs in visual understanding can't rely solely on bigger models and more compute, but need to reference human biological mechanisms?

Xi Zeng

I believe the next wave of AI technical breakthroughs will come from referencing solutions in other disciplines and translating them over — the hardest and most challenging part.

For AI visual understanding, we reference biological visual systems. Biology moves from "seeing" to "understanding" in four steps.

Step one is acquisition — mapping visual signals from the real world onto the retina;

Step two is transformation — converting visual signals into neural signals;

Step three is transmission — sending neural signals to the thinking brain;

Step four is actual visual reasoning.

Our current approach breaks these four steps into engineering modules using different technologies, referencing Unix philosophy — decompose complex problems into many small modules, each doing one thing well.

  1. Acquisition: Capture visual signals
  2. Transformation: Convert visual signals into model-understandable formats
  3. Transmission: Establish unified communication protocols
  4. Reasoning: Deep understanding based on transformed signals

These four steps need serial communication protocols, somewhat like today's popular MCP or Skills — essentially we're also doing some infrastructure building in the visual domain.

AI Nao The photo feature in Doubao was led by you — aren't the big companies doing this?

Xi Zeng

Big companies tend toward one-model-solves-all, not from lack of capability but because they have bigger models, more data, stronger compute, and want more unified entry points. But visual understanding isn't a problem solvable by a larger parameter table. It's more like a neural pathway. Human eyes don't think — they just acquire signals. Real understanding happens after signal transformation, transmission, and brain reasoning.

Big companies want to build stronger eyes. We want to build the nervous system behind the eyes.

AI Nao

The product just launched a year ago and already has **200,000 users — how?

Xi Zeng Seed users came from a small project I explored with friends in 2024 — an AI guide for an Andy Warhol exhibition in Shenzhen. After the exhibition ended, several thousand users kept using it daily to photograph things — historic sites, plants, products, food.

After official product launch, the first batch came from community programs targeting universities in North America and India. We deliberately sought design and art students — extremely visually-oriented young demographics. Strong word-of-mouth spread on campus. Plus we won Product Hunt's Product of the Day twice in a row.

These 200,000 users came with virtually zero paid acquisition.

AI Nao What specifically do users do with it?

Xi Zeng Counterintuitively, we have almost no users aged 30 to 45.

Core users fall into two groups: young people around 15 to 25, and people 45 to 55 or older, nearing retirement. They have time and curiosity. The 30-to-45 demographic struggles to use something "unrelated to productivity."

Three main scenarios:

First, travel, especially international.

Second, daily life — AI checks your outfit, what to wear for interviews, what to wear to meet your boyfriend; afternoon tea with friends checking food calories, photographing books at bookstores to understand core concepts.

Third, hobbies. One user photographed 300+ rock photos in a day — turned out he was a mineral enthusiast with a large collection. Hobby scenarios show highest stickiness.

AI Nao The "outfit" scenario is interesting — other scenarios provide information, knowledge, relatively objective. Outfit advice involves taste, relatively subjective. How does AI handle this?

Xi Zeng Somewhat true. But my observation is more complex.

Good question. Here's how I understand it: good taste is composed of a series of high-quality decisions.

We can't tell users "looks good" or "looks bad." What we do is provide more high-quality options before they decide.

Say a user asks about a floral dress — we immediately pull trending images related to floral dresses, social media discussion hotspots, high-quality perspectives formed publicly online. What to wear ultimately remains the user's choice.

We're essentially compressing the efficiency of taste formation. Over time, users naturally develop their own taste.

AI Nao I have a concern: do we really need an AI constantly explaining life around us? How to avoid "over-explaining"?

Xi Zeng For example, I love watching soccer — Chance AI is somewhat like match commentary, different commentators create completely different experiences.

Say you're shopping, it might tell you: this blue skirt doesn't suit you well, you already have something similar at home. Why not consider that light green one? It's trending lately, makes your skin look fairer. And that green skirt happens to pair well with shoes from that shop you browsed yesterday — like having a real-time "bestie" seeing the world with you.

It helps interpret life before you, providing more options. Making your lived reality more meaningful.

AI Nao These past two years, companies building on the VLM application layer have been few — first because LLM opportunities are abundant, second because VLM application certainty remains low. Why choose this direction for entrepreneurship?

Xi Zeng Visual applications are the biggest entry point of the next era, bar none. I'm absolutely certain.

First, Google Lens added AI mode in 2025 and saw 70% user growth, all new users — objective data;

Second, Gen Z and even Alpha generations grew up immersed in Instagram, TikTok, Douyin, Xiaohongshu — they're "visually native." For them, text is supplementary;

Third, everyone will need a terminal connecting themselves to the virtual world. Today's terminals are laptops, phones, or smartwatches, smart earbuds. What will future devices evolve into? Uncertain. But I'm certain they'll be products synchronized with your senses — hearing what you hear, seeing what you see, with their own computing and communication capabilities to supplement your information.

What we want is to polish the "visual brain" before next-generation AI terminals fully take shape, then plug in seamlessly when they arrive.

AI Nao That happens to answer another question — you have hardware background, but didn't start from hardware?

Xi Zeng I know exactly how that works. iPhone sold well not because of hardware itself, but its rich ecosystem. Hardware is the end result — most people just misunderstand this.

AI Nao Not doing vertical scenarios, not doing hardware, how do you think about monetization?

Xi Zeng Our current monetization judgment is: establish the "entry point" first, then deepen monetization.

Chance AI isn't a one-time tool. We want users to first form a new habit — seeing something, instinctively using AI to understand it. If this works, business models will grow naturally.

Three clearest paths. First, premium subscriptions for high-frequency users with stronger capabilities — deeper Live mode, longer-chain visual memory, more personalized understanding and judgment, and professional Visual Agents for different scenarios. Second, B2B/licensing partnerships — art exhibitions, museums, education scenarios, even AI hardware manufacturers, all essentially need a "visual understanding system"; third, in-scenario recommendations and transactions, but done very restrained. We won't channel user curiosity toward ads or shopping links, but help with next actions after they've already achieved understanding — booking tickets, making reservations, in-store experiences, ordering food, purchasing, etc.

But everything hinges on the visual entry point working. Subscriptions, licensing, and transactions all have room; if it doesn't work, any monetization is just short-term extraction.

AI Nao Many giants are also exploring visual OS directions — they have hardware, models, distribution. What's Chance AI's moat?

Xi Zeng Giants provide VLM capabilities, hardware handles world acquisition, specific applications serve particular scenarios. Chance AI's value lies in understanding, organizing, interpreting, and triggering next actions.

What I want to build is the middle layer, the "nervous system." More like the Mi Home ecosystem — whatever Xiaomi appliances users buy, everything ultimately returns to the Mi Home ecosystem.

AI Nao What other startups are working the same direction?

Xi Zeng Haven't seen any yet — we're the only one globally. Which is slightly panic-inducing, but we haven't necessarily arrived too early, like ten years ahead. Maybe DeepMind is working on it, just unknown to outsiders (laughs).

Image sources | Interviewee, Youmind

Join Community