From "No Competitors" to "Crashing Several Times a Day" | A Conversation with Zilliz Founder/CEO Charles Xie

"I've always believed that the world of technology needs idealism.

"I have always believed that the world of technology needs idealism."

👦🏻 Podcast host: Koji

🥷 Editor: Bella

🧑‍🎨 Layout: NCon

In AI circles, if you follow infrastructure — especially vector databases — you've probably heard of Zilliz. In 2023, a shout-out from Jensen Huang at the GTC conference brought the company into mainstream view. But what really caught my attention was an interview article by Zilliz founder Charles Xie earlier this year, titled "We Had No Competitors." Such blunt confidence is rare in business, and it convinced me that he has both deep conviction and genuine technical leadership in what he's building. Digging further, I found that Zilliz is not only technically hardcore but also has a rich story behind it.

For this episode, we recorded in person at the AI Hacker House in Shanghai — also Crossing's first attempt at video content. The full video will be released simultaneously on Xiaohongshu, Bilibili, WeChat Channels, and other platforms. Search for us and tune in.

Part 1 Rapid Fire: Getting to Know Charles Xie

👦🏻 Koji

Where did you graduate from?

👨🏻 Charles Xie

Huazhong University of Science and Technology.

👦🏻 Koji

How many years has Zilliz been around?

👨🏻 Charles Xie

Eight years so far.

👦🏻 Koji

What did you do before starting Zilliz?

👨🏻 Charles Xie

Database engineer.

👦🏻 Koji

Your MBTI and zodiac sign?

👨🏻 Charles Xie

ENTP, Scorpio.

👦🏻 Koji

One sentence to describe your company and product?

👨🏻 Charles Xie

We're an AI-era data infrastructure company, focused on building unstructured data platforms.

👦🏻 Koji

Revenue and profit situation?

👨🏻 Charles Xie

Can't disclose specific numbers, but revenue grew 3.3x over the past 12 months.

👦🏻 Koji

Current team size?

👨🏻 Charles Xie

About 130 people globally.

👦🏻 Koji

When Jensen Huang mentioned you at GTC 2023, how did you feel?

👨🏻 Charles Xie

I saw it as a highlight moment for the vector database category itself. When we started in 2018, almost nobody knew this space existed — we were even questioning whether the market was real. But by 2023, the industry finally recognized that AI, especially GenAI, cannot function without vector databases.

👦🏻 Koji

Was that a turning point for you?

👨🏻 Charles Xie

Not really. Building infrastructure is a grueling path. Unlike algorithmic breakthroughs, we rarely surpass competitors or win customers through a single flash of insight. Databases demand heavy investment with slow returns — it's about long-term refinement and product compounding.

👦🏻 Koji

From GTC 2023 to now, GenAI has changed dramatically. What trends have stayed the same, and what has shifted significantly?

👨🏻 Charles Xie

What's unchanged is that AI innovation keeps accelerating, and demand for data platforms continues to rise. But there have been bumps along the way. In 2023, many companies jumped on the bandwagon and raised funding. But by October or November 2024, many hadn't found real product-market fit, and their products were fairly homogenized. Lots of companies couldn't raise their next round and went under in clusters.

👦🏻 Koji

Many of our listeners may not have technical backgrounds. Could you start with a primer on what vector databases are, then introduce Zilliz and Milvus?

👨🏻 Charles Xie

At its core, a database is a systematic way to store and manage large amounts of data. Thousands of years ago, when humans began recording information in writing, libraries were data management tools.

In the IT era, data became digitized, and relational databases emerged, widely used in finance, e-commerce, ERP, and other domains.

In the AI era, computers began processing naturally generated human information — language, images, video, and other unstructured data. The advent of deep learning enabled this content to be converted into "feature vectors" (embeddings). AI's explosive growth has brought explosive growth to this data structure as well.

With so many feature vectors, AI developers need new databases to store and manage them. This is where vector databases come in. They enable efficient retrieval of unstructured data — text, images, video — using natural language and semantic search.

Part 2 "We Once Had No Competitors": The Lonely Road Ahead of the AI Era

👦🏻 Koji

So vector databases predate generative AI, and they're not limited to it. What was your thesis when you started the company eight years ago?

👨🏻 Charles Xie

Feature vectors aren't actually a new concept. Though they've become red-hot in this latest wave of deep learning-driven AI revolution, catapulting us into the spotlight, vector databases were already widely used in AI scenarios like image recognition and natural language processing seven or eight years ago. As the fundamental "language" of neural networks, embeddings are the core data structure for communication within networks, between networks, and with external systems. So starting in 2018, we were already serving many AI companies from the previous generation. Back then, they were mainly working with convolutional neural networks (CNNs) and recurrent neural networks (RNNs).

👦🏻 Koji

Over the past three years, from the rise of generative AI to now, what changes have occurred in the vector database space?

👨🏻 Charles Xie

Significant changes, mainly in three areas.

First, data volume has grown massively. Five or six years ago, tens of millions or hundreds of millions of records was considered large. Now we're talking tens of billions or even hundreds of billions.

Second, application scenarios keep expanding. Beyond LLM knowledge base retrieval, vector databases are now used for data cleaning during model training, multimodal data processing in autonomous driving, recommendation systems in e-commerce, risk control and fraud detection, even analyzing protein structures and gene sequences in biomedicine. More and more algorithms are converting various data types into feature vectors to advance drug discovery and screening. With this explosion in data volume and use cases,

The third clear trend is that users are increasingly focused on reducing the cost of using vector databases.

👦🏻 Koji

What's your priority going forward? Scaling up, cost reduction, or something else?

👨🏻 Charles Xie

We want to help users handle data at larger scales.

Traditionally, vector databases were mainly used for real-time queries with extremely high demands on latency and accuracy. But now, more and more scenarios require offline analysis of massive data volumes.

So we're expanding from a traditional vector database to a hybrid architecture incorporating a vector lake. This means supporting online queries while building a "data lake" for unstructured data, specifically designed for offline tasks.

When data reaches hundreds of billions or even over a trillion records, real-time query-by-query processing becomes not only costly but technically challenging. In these cases, it's more suitable to process the full dataset periodically through offline jobs — daily, weekly, or even monthly cycles.

👦🏻 Koji

You mentioned data scales are getting larger. I'm curious — what's the largest-scale application you've seen so far? Which company, what product, and why do they need that much data?

👨🏻 Charles Xie

We have a client that's one of the largest IT companies globally. Their goal is to use vector databases for semantic search across the entire internet — converting every single web page into a vector.

👦🏻 Koji

What's their end purpose? What service are they ultimately providing?

👨🏻 Charles Xie

Ultimately, AI search.

👦🏻 Koji

Something like BoCha AI or the early Bing Search API?

👨🏻 Charles Xie

Yes, many large language model queries now incorporate real-time search. If you want the most precise results, ideally you can retrieve information from across the entire web.

👦🏻 Koji

So this type of customer has virtually unlimited data needs.

👨🏻 Charles Xie

Yes, and the data keeps growing. AI search and RAG essentially use similar technology. RAG is typically a private knowledge base, while AI search turns the entire internet into a public knowledge base.

RAG data volumes may not be large per database, but the sheer number of customers is also a challenge. For example, if one enterprise serves 100,000 customers, each with 10,000 knowledge items, that's a billion records.

The difficulty in these scenarios isn't total data volume but data management. The system must support large-scale multi-tenant isolation and security, ensuring each customer's data doesn't interfere with others.

👦🏻 Koji

I saw a Silicon Star article titled "We Had No Competitors." You said that. What was the context? Because I actually think competition in this space is quite fierce.

👨🏻 Charles Xie

The full quote is "we once had no competitors." When we started building vector databases in 2018, there were virtually no peers globally. It really felt like walking through a desert. If you go too long without competitors, it might actually mean you're heading in the wrong direction.

But in recent years, more and more companies have entered this track, and we've watched vector databases become a hot direction. Honestly, we feel quite happy about it.

Part 3 Zilliz's Competitive Strategy: Open Source vs. Closed Source

👦🏻 Koji

I feel like competition in vector databases is actually quite intense — take Pinecone, for example. You chose the open source route, while they're closed source, currently valued at $750 million. How do you view this difference?

👨🏻 Charles Xie

Our two companies do compete intensely. They're valued at $750 million, we're at $600 million. The biggest difference is open versus closed source. If I had to choose again, I'd still firmly choose open source. Open source drives knowledge sharing and technical exchange, and accelerates product iteration.

👦🏻 Koji

So you think your biggest competitive advantage over Pinecone is open source?

👨🏻 Charles Xie

Open source is definitely our core long-term advantage. If we just talk product and technology, we're 3 to 5x faster than Pinecone in performance. But I don't want to make that the main differentiator, because technical advantages tend to converge over time.

Our current lead in technology and product isn't accidental. Precisely because we're open source, we can attract more developers globally to use and participate. They continuously feed back real needs, help us iterate quickly, and keep us from going down wrong paths.

Future competition doesn't depend on today's starting point, but on whether we can sustain this open ecosystem and let the product evolve continuously in real-world scenarios. This is our greatest confidence when facing closed-source companies.

👦🏻 Koji

You mentioned Pinecone and Zilliz are neck-and-neck. But in the open source track, you also have other competitors like Qdrant, Faiss, and Weaviate. From your perspective, have they posed any real threat?

👨🏻 Charles Xie

Faiss is a very important vector search algorithm library. Zilliz is the largest external contributor besides Facebook itself. We extensively use Faiss as the underlying algorithmic foundation in the Milvus project, so it's more part of our ecosystem than a direct competitor.

Compared to other open source projects, Milvus's advantage is better TCO. Milvus's strong performance and high scalability mean users need less hardware to support large-scale scenarios, using fewer resources.

Second is lower development costs. Over the past seven or eight years, we've done deep integrations with mainstream global AI frameworks and large language models, supporting rich data types and query methods. Milvus is no longer just vector search — it also supports scalar filtering, hybrid queries, clustering, classification, ranking, and other complex tasks, significantly lowering the barrier for developers.

Finally, lower operational costs. We provide a complete operational toolchain, including visual interfaces, permission system integration, and connectivity with enterprise access control systems, helping users reduce maintenance overhead.

👦🏻 Koji

I understand there are two other potential competitive directions. One is traditional databases like MongoDB and Postgres adding vector search capabilities. The other is frameworks like LangChain and LlamaIndex potentially integrating vector databases into their own systems. Are you worried that standalone database companies might get swallowed up?

👨🏻 Charles Xie

First, traditional databases adding vector modules work fine for small data volumes and simple scenarios. But as scale grows and scenarios become complex, users will still migrate to dedicated vector databases. It's like "range-extended electric vehicles" versus true EVs — two fundamentally different approaches. The former is ultimately a transitional solution that can't compare with native architecture.

Second, LangChain and LlamaIndex are development frameworks, while databases are underlying infrastructure. They were never the same category of product from day one, so there's no real competition or substitution.

I believe AI-era development frameworks will only become more diverse, but databases remain the core foundational layer and won't be "wrapped" or "swallowed" by upper-layer frameworks. Just like Web-era architecture with clear separation between application layer, middleware, and database layer. We're actually strategic partners with LangChain and LlamaIndex, coordinating on ecosystem level.

👦🏻 Koji

Among all the competitors you just mentioned, is there any one you worry about most?

👨🏻 Charles Xie

I care more about ourselves. What I truly worry about isn't what competitors do, but whether we can keep innovating at a faster pace.

👦🏻 Koji

I recall Databricks co-founder Reynold Xin saying that if he could do it over, he'd choose closed source. If you could start over, would you still choose open source?

👨🏻 Charles Xie

I'd still choose open source.

Without open source, there would be no Databricks today. They built early influence through the open source community, secured funding, and got their first users. Looking at it now, I actually think in the competition between Databricks and Snowflake, Databricks's massive developer base may give it greater ecosystem potential.

👦🏻 Koji

So do you think open source for Databricks and Zilliz is a shortcut, or an unavoidable necessity?

👨🏻 Charles Xie

It's definitely not a shortcut. Open source actually requires more patience. But it can become your moat, helping you win developers' mindshare. You need to let them plug your tools into their workflows at low cost, be willing to learn your product, and keep using it. Open source products can be downloaded directly from GitHub, implementation details visible. This openness naturally earns more user goodwill.

👦🏻 Koji

But Reynold Xin felt that open source made them go through a "second startup": first do open source, then find closed-source PMF, like crossing two mountains. What do you think of this?

👨🏻 Charles Xie

The "two mountains" Reynold describes is actually an important barrier to Databricks's success today. Though this path is hard, they got through it, and competitors would find it equally difficult to replicate.

The traditional open core model open sources a core and adds enterprise services on top for the commercial version. The advantage is you only need to build once, but the problem is it's hard to convince users to pay: if the open source version works, why buy the commercial one?

Databricks adopted a dual-core model: one open source core, one closed source core. Both have nearly identical interfaces and user experience, enabling seamless migration, but the underlying implementation is completely rewritten — open source in Java, closed source in C++, with the commercial engine independently designed. This approach balances user friendliness and commercial viability, a very clever architectural design.

👦🏻 Koji

So Databricks essentially has two separate teams building two separate systems?

👨🏻 Charles Xie

Yes. They must ensure the commercialized core surpasses the open source version in functionality, performance, and design. Only then can they convince users to pay for the closed-source product — same experience, near-zero migration cost, but better performance and stronger results, so naturally they'll pay.

But this path is also hard. Fundamentally you're building two products simultaneously: one for the open source community, one closed-source product for commercial customers. The commercial version must not only be compatible with the open source version but stay consistently ahead.

And this lead is dynamic. Because the open source product keeps iterating too, you need to ensure the closed-source version stays 12–18 months ahead. This is an enormous challenge for engineering capability, product design, and organizational execution.

👦🏻 Koji

Are you on the dual-core path now too?

👨🏻 Charles Xie

Yes. We made this decision in 2018–2019. This path isn't easy. It demands extremely high execution from engineering and product teams, with strong iteration speed.

👦🏻 Koji

What do you think of DeepSeek's open source strategy, and what value has it brought them?

👨🏻 Charles Xie

DeepSeek is somewhat different from database companies like us. As a latecomer, they're more focused on how to quickly capture user mindshare. Open source helped them achieve user acquisition and attention capture. For developers, once you've installed DeepSeek, you probably won't install other models. "Open source" here is essentially a "land grab" strategy.

👦🏻 Koji

It seems the original intention behind choosing open source is no longer just about attracting developers to participate and improve the product together. Do you think open source is losing its original purity? Becoming a competitive tactic, a brand strategy, even a way to gain goodwill and attention?

👨🏻 Charles Xie

I think the nature of open source collaboration itself is changing.

Attracting external developer participation is certainly good, but too many people can actually create difficulties in project management and direction steering. So many mature open source projects today are actually led by a company behind the scenes, guiding the community and project.

The real value of open source isn't necessarily getting everyone to contribute, but making technology transparent.

Many engineers choose open source not to contribute code, but to read the source code, understand the architecture, learn design details, and grow from it.

Additionally, in overseas markets, choosing open source is often not about avoiding payment, but avoiding vendor lock-in. With a closed-source product, you're committed to one path going forward, unable to judge its future direction. Open source at least provides an exit — if the partnership ends, you can build your own team to maintain and upgrade based on the community version.

👦🏻 Koji

In your open source project, how much key code comes from external developers?

👨🏻 Charles Xie

We currently have over 300 developers in our community. Only 20% are from our company, but they contribute 80% to 90% of the code.

External developers participate more in bug fixes, tool enhancements, and ecosystem integrations. This aligns with our expectations. Database systems are complex; becoming a core contributor typically requires very long-term accumulation.

Part 4 "Breaking Down" Is the Norm in Entrepreneurship

👦🏻 Koji

How do you define success in this entrepreneurial journey? You've already invested seven or eight years, and likely many more to come. What are your expectations?

👨🏻 Charles Xie

I want to be the first person globally to explore unstructured data processing and vector databases.

When I retire, I hope we'll not only be industry pioneers, but also the ones who brought it all together — true success.

👦🏻 Koji

Are you worried? Being the industry pioneer but not making it to the end, with someone else picking the fruit?

👨🏻 Charles Xie

That fear exists. Walking in uncharted territory inherently means facing massive uncertainty and pressure from technological change — and in the AI era, this pressure is amplified and accelerated.

Innovators may try 1000 approaches, with only one succeeding; followers only need to copy that one result.

Only the ability to continuously innovate and iterate rapidly is true long-term advantage.

👦🏻 Koji

How has your company maintained this innovative capacity over the years — in management, culture, or other dimensions?

👨🏻 Charles Xie

I don't think innovation can be managed.

If you want to become an innovative company, the key is finding people who want to innovate and enjoy rapid iteration.

👦🏻 Koji

You've been an entrepreneur for eight years now. Is there anything you particularly want to say to your younger self just starting out?

👨🏻 Charles Xie

Maybe I'd talk myself out of starting a company. It's far harder than imagined, and you simply can't stop. You solve one problem, a new challenge appears the next day. Every stage brings different levels of difficulty.

If you choose this path, better treat it as a lifestyle — something you're willing to do for life. Otherwise, you might break down.

👦🏻 Koji

When were you closest to breaking down?

👨🏻 Charles Xie

At worst, several times a day. On better days, roughly every one or two weeks.

👦🏻 Koji

But your company is valued at $600 million now — many would consider that very successful. Yet you still break down frequently. Can you share a recent moment that made you feel that way?

👨🏻 Charles Xie

The past two years have been the most difficult period since I started the company.

Before that, we mainly focused on product and open source. The team was all engineers, still in our "comfort zone." But two years ago, the company began pushing commercialization and set very aggressive growth targets — immense pressure.

By 2024, the market also became turbulent. Some GenAI companies failed. This wasn't our fault, but we were still affected.

👦🏻 Koji

You had some clients that suddenly disappeared?

👨🏻 Charles Xie

Yes, we had a client that was a top-tier US AI company. They suddenly fell into difficulties, projects were suspended — a major blow to us.

After losing clients, we had to quickly find new ones, not just filling the gap but maintaining overall growth momentum.

What made it harder was that the team was doing commercialization for the first time. Much organizational structure and processes weren't in place yet, but we had to start running. Like flying a plane while swapping engines and continuing to assemble the fuselage — that state was excruciating.

👦🏻 Koji

As a founder with an engineering background, you're now frequently facing clients and driving sales. Any insights or methods to share?

👨🏻 Charles Xie

First, find the right people. No amount of investment in hiring is too much.

Second, commercialization isn't that scary. If I had to do it again, I'd probably go through the same painful stages, just with different mistakes in the details. The key is recovering quickly after errors — adjusting your own mindset, stabilizing team rhythm, and most importantly, not letting morale drop.

👦🏻 Koji

When facing setbacks, how do you maintain team morale?

👨🏻 Charles Xie

Ultimately, you raise morale by winning.

Mistakes are inevitable, but don't repeat the same ones. The key is quickly learning from failures and moving on fast.

Beyond that, strategic judgment matters too — avoid directional errors.

Also, accept your own imperfection. If you can't make peace with your flaws, often what truly defeats you isn't others, but yourself.

In many competitions, what ultimately matters isn't who does more, but who makes fewer mistakes under pressure.

Part 5 From Idealism to Reality: Charles Xie's Transformation and Persistence

👦🏻 Koji

Is there anything you firmly believed eight years ago that you completely disbelieve now?

👨🏻 Charles Xie

Before starting the company, I was 100% an idealist. But after eight years, that colorful outer layer has faded, leaving mostly gray underwear.

👦🏻 Koji

Is there any particular moment where you clearly felt yourself shifting from idealism to realism?

👨🏻 Charles Xie

In team management, I used to think the ideal state was absolute transparency, nothing hidden. I equated management with bureaucracy. But now I believe management is actually indispensable for organizational growth.

Technically too. As an engineer, idealism makes you obsess over innovation and perfection. But in business, "good enough" suffices.

Winning by just a little is enough. Like Intel and NVIDIA, making steady, incremental "toothpaste-squeezing" progress can still achieve massive success — because they grasped the rhythm between technology and business.

I have always believed that the world of technology needs idealism. Though idealism has been worn down by reality to nearly nothing, it was precisely that pursuit of perfection back then that established our product and technical advantages.

Even as we become more commercialized now, I still hope that while realism gradually dominates decision-making, we can preserve a colorful, sentimental patch of sky deep in our hearts.

👦🏻 Koji

Across the entire AI infrastructure track, from large models to databases, if you were to invest, which companies would you bet on?

👨🏻 Charles Xie

Across the entire AI track, what I'm most bullish on is actually cloud platforms, especially giants like Amazon. Because AI has entered the "energy and infrastructure" phase, the ultimate competition is about who can build and operate large-scale data centers — precisely the strength of public clouds. I believe public cloud importance will continue rising.

Large models as the foundation layer are of course unignorable, especially the top few companies.

There are also some good AI application companies. The tools I personally use most are ChatGPT, DeepSeek, and Cursor.

👦🏻 Koji

Thank you very much to Charles Xie for joining us for this very hardcore podcast episode. We wish Zilliz continued rapid growth, and look forward to having Charles back on Crossing next time.

👨🏻 Charles Xie

Thank you.