BosonAI's Mu Li: If There's Something You're Meant to Try in Life, Start as Early as Possible | Oasis Vitality

Counselor Vitality

Mu Li — co-founder of BosonAI, former Chief Scientist at Amazon, and an old friend of Oasis Capital. After leaving Amazon, he threw himself into a large language model (LLM) startup. After weathering funding turbulence, technical hurdles, and market exploration, he finally saw light at the end of the tunnel, achieving break-even in his first year of entrepreneurship. He distilled his experience into an essay titled One Year of Entrepreneurship, Three Years of Life, sharing his progress, challenges, and reflections from year one. Below is Li's account in his own words, adapted from his Zhihu and Bilibili columns. Enjoy.

Actually, by my fifth year at Amazon, I already wanted to start a company. But the pandemic derailed those plans. By year seven and a half, the itch had grown unbearable — so I handed in my notice and dove in.

Looking back, I believe that if there's something in life you're destined to try, the sooner the better. Because once you're actually on the entrepreneurial path, you discover there's an endless amount to learn. I often find myself wondering why I didn't start even earlier.

The Origin of BosonAI

Before founding the company, I worked on a series of projects codenamed "Gluon." In quantum physics, a gluon is a boson that binds quarks together tightly. The name symbolized how the project began as a joint effort between Amazon and Microsoft.

"Gluon" was the project manager's spontaneous idea, but naming things is notoriously hard for programmers. In our early days, we agonized daily over filenames and variable names. Eventually we just named the company after the boson itself, hoping people would smile at the "bosons and fermions make up the world" reference. Unexpectedly, many people misread it as "Boston." (Laughs)

"I'm in Boston, let's grab coffee?" "Huh? But I'm in the Bay Area 😅"

Fundraising: Lead Investor Bails the Day Before Signing

In late 2022, two ideas for LLM-powered productivity tools flashed through my mind. By chance, I ran into Yiming Zhang and asked his advice. After some discussion, he turned the question back on me: "Why not build the LLM itself?"

I instinctively recoiled. My team at Amazon had spent years on LLMs, wrestling with tens of thousands of GPUs and a cascade of complex challenges. But Zhang responded: "Those are short-term difficulties. You need to think longer-term."

One of my strengths is taking good advice. So I pivoted to building LLMs. We quickly assembled a founding team with leads covering data, pre-training, post-training, and architecture, then started raising money. Fortunately, we secured seed funding quickly. But it wasn't enough to buy the GPUs we needed. A Series A was urgent.

For that round, we secured a lead commitment from a major institution. We spent months preparing documents and negotiating terms. Then the day before we were set to sign, the lead suddenly pulled out — which directly caused several co-investors to withdraw. Fortunately, remaining investors stepped up and helped us close the round, giving us our ticket into the LLM race.

Reflecting now, if we'd struck while the capital markets were hot and kept raising, maybe today we'd be sitting on a billion in cash like some peers. At the time, I worried that too much funding would make exits harder and leave us trapped. But looking back, entrepreneurship is about defying fate. You need to burn your boats. Why think about retreat?

Hardware: First to Eat the Crab

With funding in hand, we went to buy GPUs. Every supplier gave the same answer: H100 delivery in a year. Facing this dead end, I had a flash of inspiration — I'd email Jensen Huang, NVIDIA's CEO, directly. Unexpectedly, I got a quick reply. Within an hour, Supermicro's CEO called me directly. We ended up paying a premium, and 20 days later we had our machines. We ate the crab early.

But that crab nearly broke us. We hit bewildering bugs. GPU power instability, fixed only when Supermicro engineers patched BIOS code. Fiber optic cleave angles slightly off, causing communication failures. NVIDIA's recommended network topology wasn't optimal, so we designed our own — which NVIDIA later adopted. To this day I don't understand: we bought fewer than 1,000 cards, a small buyer. Had bigger buyers not hit these issues? Why did we have to debug them?

We also rented an equal number of H100s and hit endless bugs there too — GPUs failing daily. We half-wondered if we were the only company eating this crab on that cloud. Then we saw Llama 3's technical report: they claimed that after switching to H100s, a single training run was interrupted hundreds of times. We felt deeply seen.

Comparing build versus rent over three years, costs are roughly comparable. Renting saves headaches. Building has two advantages: first, if NVIDIA remains dominant in three years, they can control pricing to keep GPUs valuable; second, self-built data storage is cheaper. Storage needs to be close to GPUs, and whether from major clouds or smaller GPU clouds, storage pricing is high. A single training run might use several TB for checkpoints, with training data storage starting at 10PB. AWS S3 for 10PB runs $2 million yearly — money that could self-build 100PB of capacity.

Business: Grateful to Customers, Break-Even in Year One

Thankfully, we achieved break-even in year one. Our spending went to people and compute; our revenue came from building custom models for major clients. The CEOs who adopted LLMs early showed tremendous decisiveness. They weren't deterred by high compute and labor costs — they chose us and pushed their internal teams to experiment with us. We're deeply grateful these customers gave us the chance, sparing me from running between investors again.

More companies should start experimenting with LLMs soon, whether to upgrade their products or cut costs. Technology costs are falling on one hand, and industry leaders (like our customers) are rolling out LLM-based products that are forcing the rest of the industry to move.

We're also watching LLM consumer applications. The last wave of darlings like Character.AI and Perplexity AI are still finding business models, but a dozen or so LLM-native apps are generating solid revenue. We provided models to a role-play startup targeting hardcore players — they've balanced revenue and expenses, which is impressive. Model capabilities keep evolving, with more modalities (voice, music, images, video) converging. I believe even more imaginative applications are coming.

Looking at the broader market, industry and capital remain impatient. Several companies that raised over a billion within a year have already exited this year. From technology to product is inherently a long process — two to three years is normal. Factor in user demand emerging, and it may take even longer. Our stance: focus on the present, feel our way through the fog, stay optimistic about the future.

Technology: Four Stages of LLM Understanding

My understanding of LLMs has evolved through four stages:

Stage one: BERT to GPT-3. My feeling was: new architecture, big data, it can work. At Amazon, we were among the first to do large-scale training and product deployment.

Stage two: When I first started the company. GPT-4 appeared and blew my mind — largely because the technology went closed. Based on rumors, one training run cost roughly 100 million RMB, with data labeling running tens of millions. Investors asked me the cost to replicate GPT-4; I estimated 300–400 million. One of them actually put several hundred million in at once.

Stage three: My first half-year. We couldn't match GPT-4, so we started from concrete problems. We found customers across gaming, education, sales, finance, insurance — trained models for their specific needs. At first there were no good open-source models, so we trained from scratch. Later, excellent models emerged and lowered our costs. We designed evaluation methods for business scenarios, labeled data, identified model weaknesses, and improved them surgically.

By end of 2023, we were surprised to find our Photon series (a type of Boson) outperforming GPT-4 on customer applications. Custom models' inference cost was 1/10 of API calls. Though APIs have gotten much cheaper, our technology has advanced too, so we still maintain that 1/10 ratio. Plus we control QPS, latency, and other factors better. The insight this stage: for specific applications, we could beat the best available models.

Stage four: My second half-year. Though we delivered models meeting contract requirements, they weren't customers' ideal — because GPT-4 itself is far from enough. Early this year we found that training for single applications hit a ceiling. Then it clicked: if AGI reaches average human level, customers want professional-level performance. Gaming wants expert designers and actors, education wants top teachers, sales wants star salespeople, finance wants senior analysts — this requires AGI plus domain expertise. Though we held deep respect for AGI, this path seemed unavoidable.

We designed the Higgs series (the God particle, another type of Boson) — strong general capabilities tracking the best models, but excelling in one area. We chose role-play: playing fictional characters, teachers, salespeople, analysts, etc. By mid-2024 we'd iterated to v2. On Arena-hard and AlpacaEval 2.0 (general capability tests), v2 traded blows with the best models. On MMLU-pro (knowledge testing), results were comparable.

Higgs v2 is based on Llama 3 base with full post-training. We can't spend Meta-level money on labeling, so v2's improvement over Llama 3 instruct likely comes mainly from algorithmic innovation.

We built a role-play evaluation benchmark covering character-based and scenario-based acting. Our model ranked first on our own benchmark, but never touched the evaluation data during training. The benchmark was originally for internal use — we wanted true capability reflection, avoiding dataset overfitting. But our eval team wanted to write a technical report, so we released it publicly. Interestingly, the character-play test samples came from Character.AI, yet their model ranked last.

My stage four insight: good vertical models can't be weak on general capabilities either — reasoning, instruction following, these matter vertically too. Long-term, both general and vertical models must advance toward AGI. Vertical models can just be slightly lopsided: ace the specialized courses, decent on the general ones. So R&D costs are somewhat lower, approaches differ.

Stage five is ongoing — hope to share more soon.

Vision: Human Companionship

I confess we didn't start with a vision. We focused on improving technology, building custom models for clients, and gradually found our purpose through that process — seeing what customers wanted, what we wanted, what the future might need.

Years ago, I dreamed of a robot nanny to help raise and play with my kids, because parenting was hard for me and I didn't fully grasp my child's thinking. I wanted a brilliant virtual assistant to invent new things with me at work. When I'm old, I want interesting robots keeping me company. So my prediction: as productivity tools advance, individuals will accomplish what once required teams, making people more independent, everyone pursuing their own goals — and more lonely.

Putting this together, our vision became "intelligent agents for human companionship" — agents with high emotional intelligence and solid IQ. Translated to real people, that would be a professional team. Want it to play with you? It's an expert designer plus actor. Want it for fitness? It's a motivator plus professional coach. Want to learn? It explains what you don't understand. The beauty of models: they can accompany you long-term, truly know you, and be "genuinely for you."

Current technology remains far from this vision. Today it can only chat, and often poorly — thin content, IQ and EQ sometimes both offline. These are problems to solve step by step. If you're building overseas applications in this space, reach out.

Team: Hard Things Require Teams

Only after starting a company did I truly grasp team importance. At big companies, I felt like a cog, teammates were cogs, even the team was a cog. In entrepreneurship, the team is a car — small, but it runs, carries load, turns nimbly, goes anywhere. When miHoYo's Cai visited early on, seeing our whole team in one room, he sighed: small teams are great.

Of course, small cars have inconveniences — constantly watching for "fuel," taking bad "roads" carefully so the "car" doesn't shake apart. Every member matters in a small team; there's no redundancy. One underperformer and a "tire" goes flat. People are precious — losing one might mean losing a tire.

Before, I'd choose projects I could lead development on, but that meant the problems weren't very challenging. Entrepreneurship means tackling huge problems, and that requires the team entirely. Don't let all the "I"s in this essay fool you — the work is all team. Without them, I'd probably have switched to selling courses. (Laughs)

Personal Pursuit: Fame or Fortune?

So far, I've followed my inner voice in decisions — returning for a PhD after working, making videos, starting a company. Entrepreneurship needs strong motivation to overcome endless difficulties. That requires deeper analysis of your motives, which come from desire or fear.

Ten years ago I might have chased fame and fortune more eagerly. At my age now, money's marginal utility has declined, and fame's emotional returns have shrunk.

Leaving aside the universe's vastness — in human history alone, one person's existence is a grain of sand, arriving by accident, vanishing quickly. A hundred billion humans have lived on Earth; most leave no trace. Of the thousand-plus names on my family genealogy, I barely recognize any.

So what's the meaning of a person's existence? I once fell into depression as a child, unable to solve this. In my subconscious, I want to create value, to give existence meaning. I chose "self-improvement" to build value-creating capacity. I chose long videos and textbooks to create educational value. I chose writing summaries of PhD, work, and entrepreneurship experiences — the struggles and dilemmas — to create value through example. I chose entrepreneurship to unite many people's strength to create greater value.

Postscript

Last year Hua Su and I walked at Stanford. He clapped my shoulder: "Tell me honestly, why did you want to start a company?" I was nonchalant then: "Just wanted to do something different." Su smiled without speaking.

Now I understand his meaning — he's tasted entrepreneurship's full flavor. If asked today, I'd say: "I was out of my mind."

But I'm glad I didn't realize how hard it would be, or I might never have jumped in. Otherwise you'd be reading "Ten Years of Work Reflection" instead of "One Year of Entrepreneurship Review." I still think the entrepreneurship story is more interesting.

Finally, salute to all entrepreneurs.

If you're interested in joining BosonAI, click below for openings (Bay Area + Vancouver): https://jobs.lever.co/bosonai. Also welcome to overseas application builders — contact: api@boson.ai

Click "Read Original" for the source article