Oasis Capital was invited to participate in Yicai's WAIC live broadcast "AI Geeks Talk"

Counselor Vitality

During the 2024 World Artificial Intelligence Conference (WAIC), Oasis Capital was invited to join Yicai's "AI Geeks Talk" roundtable on the theme "The Future Belongs to 'Large Models + ?'." Fellow panelists included Junhao Zhong, Secretary-General of the Shanghai Artificial Intelligence Industry Association and Secretary-General of the Shanghai Artificial Intelligence Standardization Technical Committee, and Jifan Yu, Assistant Researcher at Tsinghua University's Institute of Education and a Shuimu Scholar at Tsinghua. Notably, the interview outline itself was generated by a large language model based on the discussion theme. The guests shared their diverse perspectives on the AI wave, and we've selected highlights to share with you. We hope you find them thought-provoking. Enjoy.


Yicai: In which domains have large model technologies achieved notably significant applications so far? And what concrete value have these advances created?

Junhao Zhong: I'll frame this in two phases — from AlphaGo in 2016 to ChatGPT at the end of 2022.

In the first phase, I think something interesting happened: once you started using AI, you basically forgot it was AI. The technology had already penetrated every aspect of daily life. Unlock your phone with your fingerprint? That's AI — biometric recognition. Withdraw money at a bank counter where they scan your face? Also AI. Go to the hospital for diagnostic imaging — chest X-rays, lung scans? AI reads them first. So in that first phase, you used it and forgot about it; you didn't think of it as artificial intelligence. It had simply become embedded in our lives. That was one stage.

Then around late 2022, we got this new wave driven by LLaMA and large language models. Here's the thing: the higher the precision you demand, the harder it is to commercialize. Why do I say that? Take something I use constantly now — and I should mention, you were essentially a microphone just now, because today we're really having a Q&A with AI. I'll be upfront: I've been using an Agent developed for this year's WAIC. It's an intelligent agent, and I'm using "magic" to answer your "magic," or using a model to answer your model. Basically, for every question you gave me, I ran it through the WAIC Agent. But it functions as an assistant, as prompts — giving me outlines. I then extract what I think is correct and personally endorse, and set aside what falls outside my own knowledge base. So from a textual standpoint, I'm also using a model to answer your model. That's one thing.

From my personal perspective, in the near term I'm most optimistic about applications with high fault tolerance — they'll have the biggest market. But taking a long-term view, I don't think the issue is whether large models are 100% accurate; it's whether our own willingness to accept this thing has grown. You know it's not 100% perfect — the question is whether you're willing to try it. Yesterday I posed the same question to industry professionals who don't know much about AI: would they be willing to ride in a driverless car? Several said yes — they were willing to accept that still-imperfect driving experience.

Jinjian Zhang: I think if we offer a perspective on these two phases, they represent two fundamentally different worldviews in humanity's exploration of AI. We all know doctors treat illness — there's knowledge, experience, systems, and methods involved. How do we solve this problem?

Before the Transformer architecture emerged in 2018, AI was primarily about building input-output relationships. Like a doctor's work: symptoms on a medical record are the input, treatment solutions are the output. The mapping happens between input and output for specific tasks.

A second characteristic: it operated within constrained scopes, because expanding the range disperses the knowledge system and makes such mapping difficult. So you'd find it highly accurate but applicable to very limited scenarios. That was our earliest approach to exploration.

When the Transformer architecture came out in 2018, it completely upended this approach.

If I want to replicate a doctor, one method is to explore the relationship between what information a doctor sees and what conclusions they draw. The second method is to simply have the system learn every book the doctor has read, every patient they've seen, every case they've encountered. Then ask: isn't this person a doctor? So the second worldview isn't based on exploring the problem; it's based on exploring the growth process of a professional like a doctor. What's being replicated isn't question-and-answer pairs — it's all the information input and basic knowledge output throughout the doctor's development. It explores this process.

Like taking on an apprentice for ten years — are they you? You'd say they're very like you, but you can't say they are you. But if they've read all your books, experienced all your experiences, they'll become increasingly like you. What's the advantage? Their worldview will be broad; their generality strong, because they were trained as a doctor. What's the downside? Like any apprentice, even after ten years, they'll make mistakes, say wrong things, be inaccurate. These two frameworks represent different approaches to solving problems — different inputs and outputs.

So at this point, we've found this framework that provides the world with vast numbers of apprentices for professional talent. Today this "apprentice" has relatively weak foundational abilities, and we haven't invested sufficient energy, patience, or time in cultivating them — so today they're exhibiting all the problems of someone in the early stages of learning. But we can see these problems are iterable and developable. In the long run, we believe these apprentices will grow from junior to advanced apprentices.

You'll find this wave of large models demonstrating different capabilities — they're developing aesthetic sense, more abstract thinking and linguistic abilities. In areas like image generation, audio-video processing, emotional expression in language, and linguistic resonance, they're showing capabilities that may even surpass humans. The reason is they're simulating how a person speaks, how a person learns to be human — not merely learning speech itself. So I think past methods and current methods suit different scenarios, and these scenarios represent different worldviews. Given time, there's much to anticipate in opportunities and development.

Jifan Yu: Within my field of view, the main application is the integration of large models with education. Since the question was about what achievements large model technology can produce in these domains, I'll briefly report on Tsinghua University's approach to integrating AI to empower education.

First, in building AI applications, we're operating at something like hierarchical levels — as Zhang mentioned, it can penetrate to different depths. The most basic level is essentially subject-knowledge Q&A functionality. It can help answer questions in relatively niche or highly specialized subjects — something nearly every large model can do in general domains but hasn't yet mastered in professional domains. That's the most fundamental function.

The next level, as the two teachers just mentioned, involves simulating human skills — having large models attempt things like playing a teacher who understands a student's knowledge state, or acting as an excellent teacher to help generate problems and perform very basic educational functions. The deepest level involves environmental transformation. Recently we've been conducting some very cutting-edge explorations — some call it "driverless" in education; others see it as a fully AI-guarded classroom environment. In such an environment, we have some large models playing teachers, some playing classmates, some playing teaching assistants, constructing an online virtual environment where students can still interact purely with AI, achieving effects roughly equivalent to real classroom attendance. So at different levels, as the two teachers discussed, the integration of large models with a field can produce different effects — with fault tolerance certainly decreasing as you go deeper.

Yicai: How do you view future development trends for large model technology, and in which domains will it continue to show enormous potential?

Junhao Zhong: *Regarding the question just raised, I should speak from an industry practitioner's perspective — though industry and technology are now highly overlapping, as are industry and academia. Nearly 100% of entrepreneurs in this space were previously scientific or technical workers; the overlap is complete. So we've been researching and closely following trends in large models for quite some time. From our perspective, there are several relatively clear future directions.

First, it's definitely moving toward multimodality, especially integration with embodied intelligence. At this WAIC we saw massive numbers of robots — much coverage focused on whether they can walk, whether they have dexterous hands. But more important is to look at perception and interaction systems. Without computer vision, they can't see objects on a table. Without a good human-interaction system, without a large language model as support — whatever you say to it, the machine won't understand. This time represented a major leap forward, what I just described: from multimodality toward embodied applications. That's one trajectory.

Second, we're no longer just talking about large models — now we're discussing on-device small models, plus sparse activation networks and on-device small models. The core issue this addresses is our massive computational investment. Everyone knows large models align massive data on massive compute — training runs routinely use tens of thousands of GPUs, and inference in practical applications requires similarly enormous compute. Energy experts have calculated that by 2035, all energy on Earth going to power chips and GPUs wouldn't be enough. But personally, I say — that's why they're energy experts. They took one data point and calculated rigidly. Energy itself will continue to develop technologically; we'll have nuclear fusion and better energy systems ahead.

Moreover, they didn't consider that our energy approach must change too. We won't necessarily continue consuming massive power to support model applications. So the second trend is how to use on-device small models and sparse networks to address the overall compute problem. This leads to a further point: our human brain weighs about 1.5 kilograms. Its energy consumption? A 15-watt lightbulb. With such small compute, not particularly large volume, it can drive a body like mine — nearly 200 jin. That's the second direction.

Third, setting aside embodied intelligence — we call this biointelligence: brain-computer interfaces, brain-inspired models, etc. I see these as future directions from the technology side.

If we think about where to apply this well, based on my earlier framework and the second framework I just mentioned, I think text generation and language dialogue applications are probably the easiest to commercialize. After all, we in industry really hope to see science and technology truly empower our daily lives — not just empower, but eventually make money, have people willing to pay. That's the best stage of technology transforming into industry or product.

I believe business models may have a chance to emerge in the second half of this year or next. Because I'm already a heavy user of large models — I no longer use [certain] search engines; I open various boxes, various Agents, or what we might call interface-side small models. So I think business models will emerge. When usage reaches a certain scale, that internet business model of "wool from the pig, paid by the dog" may come into play too. That's my understanding.*

Jinjian Zhang: Everyone today has discussed many AI trends and parameters. We might consider a further question: today's AI is an industrial revolution. No one today cares what Watt named his steam engine. We can't recall the parameters of Watt's steam engine. We don't think about how Watt's steam engine also released versions 1.0, 2.0, 3.0 — from receiving the order in 1763 to completing it in 1790, the improvement process was lengthy.

In the great waves of history, what we need to care about is what this wave means, what transformations it may bring — not merely what might happen in this wave in one year, two years, three years. Because the trend has formed; the wave has already risen. But when the wave surges, much in our present remains unprepared.

We've communicated with many AI entrepreneurs; many are thinking about how to build AI businesses; many still treat AI as a tool, thinking at the tool level — how do I use this tool, how do I do things with it. If you treat it as an entire industrial revolution, it's another dimension of thinking. Just as JP Morgan rose because it was Edison's angel investor.

When we discuss future trends, we might consider three questions:

First, what will the relationship between people be in the future? For example, today AI knows you well — listens to you speak, watches you act, chats with you daily — and you increasingly trust it because it increasingly understands you. Why do we have best friends? Because they know what happened to you over the past decade; it's like your private data is backed up with them. Once you lose a best friend, you need a year to synchronize all your past experiences before they can give advice. But AI holds private data and can follow you continuously. As AI knows you better and becomes more familiar, is your best friend a person or AI? Today we see some social apps experimenting: an AI as your friend you can chat with — and within three months, this AI becomes one of your top five contacts, meaning AI can become your most trusted companion.

Second, what is the relationship between people and industry? Take lawyers — an industry with hundreds of thousands of them. If AI can train a lawyer, or be trained as a lawyer, most people would choose to train AI to be the best lawyer in the industry. Does an industry still need so many people? Could it be that instead of an industry with hundreds of thousands of practitioners, there will be thousands of top-tier professionals leading hundreds of thousands of AIs serving the world?

Third, what is the relationship between people and society? If future medical services come from a robot, your best friend is a robot, your life advice comes from a robot — what then is your relationship with society? Are you in society? Is there a "you" in society?

When we talk about potential, it may be about answering these questions. These questions contain many opportunities and potential — not merely the development of AI as a tool itself.

Jifan Yu: If the formation of human civilization initially solved the problem of humans relying solely on physiological limits for survival, and inventions like the steam engine replaced physical labor, this replacement is indeed more comprehensive — it replaces some of humans' more basic intellect-related labor.

At the same time, it may indeed create new opportunities. In my view, AI for Science — AI helping discover new knowledge — may help us quickly find new growth points. In academic research, many researchers joke that engineering research is just a checkerboard approach (laughs): find a problem, find an unused method, and you've created a new research point. AI can do this very efficiently. If applied to industry, it might create new value.

AI still can't do paradigm-level discovery of new knowledge, but it has considerable ability to quickly identify, based on existing data, things humans should have seen through rapid iteration but haven't due to time and energy constraints. This is something AI can do in the near term, and can do with rapid returns.