Oasis Capital Conversation with Professor Fisher Yu: The Anti-Moore's-Law Nature of AI Development

**Oasis Capital: ChatGPT's "Mind-Blowing" Results — Did AI Insiders See It Coming?**

Within days, OpenAI released GPT-4, Baidu unveiled ERNIE Bot, and Microsoft launched Copilot — a flurry of major announcements, one after another.

This interview with European professor Fisher Yu is a selection of highlights. Enjoy. Fisher Yu, Assistant Professor at ETH Zurich, core faculty member at the ETH AI Center and ETH Center for Robotics.

Oasis Capital: ChatGPT's results are so "mind-blowing" — did AI practitioners see this coming?

Professor Fisher Yu: This question has two sides. On one hand, people were hoping for something extraordinary; on the other, they're surprised that the results could suddenly become this good. Language processing technology has advanced very rapidly in recent years, especially after the emergence of Transformer, which made it possible to study large language models (LLMs). Although ChatGPT's performance is outstanding, researchers and practitioners in the AI field weren't entirely without expectations. Over the past few years, GPT itself has gone through several versions, and Facebook and Google have been continuously iterating their own LLMs. Throughout this process, people discovered some very interesting properties. For example, last year Google released PaLM (Pathways Language Model), which could support a wide range of tasks and even explain jokes. When the parameter count is large enough, LLMs exhibit emerging properties — making it feel as if the language model truly understands language itself or logic itself, even if strictly speaking it doesn't actually understand. Every year brings an endless stream of new technologies, enabling rapid LLM iteration. But at the same time, ChatGPT's emergence does represent a leap forward built upon years of technological jumps, because with ChatGPT, we can not only obtain large amounts of useful information, but humans have achieved seamless communication with language models. People joke that programmers won't be needed anymore — just Prompt Engineers to extract information and generate results from LLMs. ChatGPT made everyone realize that even through natural language conversation, you can get meaningful information from the model. This helps not only professionally trained engineers but also ordinary people who can feel its capabilities and usefulness. ChatGPT achieved a product-level breakthrough.

Oasis Capital: Can ChatGPT actually understand language logic itself?

Professor Fisher Yu: This is a highly controversial topic in the industry. While many believe it has developed some functional understanding of language, no one is 100% convinced that ChatGPT truly understands language itself. Because to truly learn language and logic, you need deep understanding of linguistic meaning and reasoning rules. Recently in linguistics, there's been extensive discussion about this. The renowned linguist Noam Chomsky recently stated clearly in the New York Times that strictly speaking, ChatGPT cannot understand language itself. Noam's position has been opposed by NLP (Natural Language Processing) scholars — it's not entirely correct to say it has zero understanding. However, everyone agrees on one point: while we can't guarantee that ChatGPT understands language logic the way humans do, it can execute the functions of language logic understanding.

Oasis Capital: How do the open-source community and cloud computing giants view ChatGPT?

Professor Fisher Yu: Let me talk about ChatGPT's impact on the entire tech sector. One very influential aspect is that OpenAI, through close collaboration with Microsoft, has integrated ChatGPT into Microsoft's products. At the same time, OpenAI is partnering with many small companies and service-oriented firms, trying to apply their technology across different fields. Because of this, although OpenAI hasn't open-sourced the model itself, it has maintained a very open attitude in the product environment so far. ChatGPT's API has been made available at very affordable prices, giving any developer the ability to integrate it into their own products and give them ChatGPT-like capabilities. This is precisely what has had such enormous impact on the AI business environment. It's as if every developer can reference or use such technology. Other major cloud computing providers are also catching up very quickly — Google published an LLM API based on Google Cloud just last week, ensuring that not only large companies can monopolize these advanced LLM technologies, but small companies can also embrace them as part of their products. This is one of the key reasons why related products are iterating so rapidly.

Oasis Capital: Which job positions is ChatGPT affecting?

Professor Fisher Yu: GPT-4 has impacted many jobs, and everyone can use it to some degree in their work — especially positions that require producing large amounts of content, such as customer service. Previous AI chatbots for customer service could only handle the most basic user guidance or task dispatching. It's foreseeable that as ChatGPT matures across vertical domains, demand for human customer service could decrease. Similarly, advertising copy and social media content creation will inevitably accelerate when empowered by ChatGPT. Data scientists, who previously needed large teams to mine massive datasets, can now use AI to directly extract data information and present it. Other positions like HR already have AI screening resumes and scheduling interviews. However, it must be pointed out that while ChatGPT or GPT-4 will have very profound effects on these jobs, for now it's not replacing humans but augmenting them — enhancing the efficiency of professionals across different fields, helping everyone solve basic problems. But when it comes to truly professional issues, humans are still needed. The most straightforward example: GPT-4 can score above 90% of test-takers on the bar exam. While the score is high, the biggest problem is that without deep professional knowledge, it's very difficult to know which parts of GPT-4's generated answers are wrong. This is one of the biggest obstacles to ChatGPT or GPT-4 truly replacing practitioners. It's like a lawyer can have ChatGPT draft some documents, but ChatGPT can't actually represent someone in court. It lacks strong self-awareness capabilities and cannot guarantee 100% factual accuracy.

Oasis Capital: You mentioned ChatGPT's augmentation of human efficiency, but in our experience using ChatGPT, we find that because we can't verify the truthfulness of generated content, we need more time to check and verify, which actually reduces efficiency?

Professor Fisher Yu: This isn't just ChatGPT's problem — it's a major obstacle for AI as a whole. Whether language or vision models, current AI can achieve fairly high accuracy rates, but the biggest problem is that the erroneous 10% still requires human review. The most typical example is autonomous driving: current self-driving can solve 99% of problems, but what autonomous driving companies struggle with most are the 99.9%, 99.99% cases. Although uncommon, they create huge obstacles to job replacement, which is why many intelligent driving companies are transitioning to assisted driving. A very serious problem currently facing the AI field is: how can we know whether this learning and statistics-based model can be combined with traditional language logic principles, so that the model can know whether its output is correct or incorrect, while also understanding what it doesn't know and making that explicit?

Oasis Capital: With widespread ChatGPT use, will human cognitive abilities decline?

Professor Fisher Yu: This is a very interesting and controversial question. When ChatGPT first came out, high schoolers and even undergraduates used it to write assignments, which was very discouraging for many teachers and professors. They found that although students weren't plagiarizing others' work, machine writing wasn't exercising the students' own abilities, and teachers were wasting time grading machine-generated content. So schools had to implement policies banning such technology. I believe for society as a whole, the impact of ChatGPT's emergence is a process that needs to be gradually absorbed. Like calculators — in different educational fields, opinions still differ today. Some schools allow calculators in exams, seeing no need for students to solve calculations manually, which allows exam questions to go deeper; when calculation itself isn't the bottleneck, the problem itself becomes the focus — in physics, chemistry, or applied math exams, whether you truly understand the principles matters most, while calculation becomes secondary. In our society today, calculators are available everywhere, but in the most basic elementary education, developing students' fundamental arithmetic abilities is still necessary. Similarly, even with ChatGPT's existence, in early education students still need to master basic writing and content creation abilities.

ChatGPT itself can serve as an auxiliary tool to improve people's work and learning efficiency. I believe it may actually enhance human cognitive abilities to some degree. Because previously, the bottleneck in human cognition or learning lay in basic tasks and expression. If this portion can be handled by ChatGPT, people are no longer constrained by basic problems and have time for deep thinking and further research on problems themselves. So I think there may even be courses specifically teaching people how to use ChatGPT to improve their own learning and work efficiency, enabling people to further enhance their cognitive abilities and improve their professional capabilities.

Oasis Capital: What industry changes will ChatGPT bring about?

Professor Fisher Yu: This goes beyond my professional scope (laughs), but what's foreseeable or already known is that AI applications in content creation will become increasingly widespread. AI has already permeated every aspect of daily workflows. Take Adobe, creator of Photoshop — they invest heavily every year in researching how AI can help creators better express their creativity and operate software more conveniently. This process has been ongoing all along; it's just that before ChatGPT, Adobe's research and the changes it demonstrated weren't recognized by the general public. ChatGPT and GPT-4 will accelerate this process, even producing step-function leaps — at some point, a tool will suddenly appear that makes us rethink entire workflows.

Another example is translation services, which is particularly evident in Europe with its many languages and large linguistic differences. Portable language services like Google Translate are very helpful in daily life, and I think language models' industry impact may be most apparent in this area.

Previously, when doing content creation — say, a speech — top industry managers might have their own professional teams to draft copy. But now with AI's help, every ordinary person can have their own "team" to complete expression and creation. This will certainly improve personal work efficiency and well-being, and will also give rise to new industries. Take self-media: people can not only express themselves freely, but also have the opportunity to express at high quality like major influencers.

Oasis Capital: What impact do LLMs have on your professional field?

Professor Fisher Yu: It should be said that they affect the entire AI research field enormously. Especially for natural language processing — people even feel a sense of existential crisis. Students in this specialty are very nervous, unsure whether their research still has value under LLMs. The computer vision field has the same sense of crisis, because GPT-4 can solve visual problems very well, freely generating rich descriptions based on different images, and engaging in effective dialogue through recognition — something that my visual field greatly admires.

From my professional perspective, we're also constantly thinking about how to use strong language capabilities to enhance image recognition. After all, images are the problem that computer vision studies. It's not just about image-text dialogue — we need in-depth analysis of visual information itself. For example, not just the entire image, but even at the object level, and even the dynamic information of objects throughout video sequences. Another point is that visual understanding requires not just semantic analysis, but also understanding of shape, geometry, and interactability. This scenario we call affordance — that is, when you see a chair, you need to know it's for sitting, and you can sit on it. This intersects with language understanding; there are many other aspects of visual information understanding that currently aren't well solved.

Additionally, LLMs will greatly promote the combination of language and vision. One main direction our lab focuses on is how to give robots visual recognition capabilities, enabling them to automatically generate control signals for the entire robot by observing scenes of human interaction. This is hugely helpful for vision itself and downstream applications. However, currently these dimensions cannot be fully controlled semantically.

Oasis Capital: In terms of LLMs, what are the characteristics of changes in European academia and industry compared to other countries?

Professor Fisher Yu: Europe and the United States are roughly similar in their degree of technological awareness. The academic circles don't differ much either — the resources accessible are the same, and everyone is thinking about the same question: after GPT-4 or ChatGPT's technology, what direction should our own research take?

But in industry, influenced by the overall industrial atmosphere, Silicon Valley companies — especially smaller ones — raise funding faster. Many American companies have quickly integrated APIs and iterated their products. Europe is indeed slower, because European industry is generally more traditional, and iteration on new technologies — especially at the software level — tends to lag Silicon Valley by a step.

Oasis Capital: Europe's overall data privacy protection, GDPR, is among the strictest in the world. Could this hinder LLM promotion?

Professor Fisher Yu: It's more about protecting people's private information. If personal information in LLMs isn't clearly protected, it's very possible that each of our personal information could appear in LLMs — which is terrifying. With clear data protection, companies developing LLM technology will be very cautious, rather than taking chances with personal privacy and individual interests.

Oasis Capital: How far away do you think AGI is in the broad sense?

Professor Fisher Yu: This is very difficult to predict. For strict AGI, it's hard to achieve within 10 years. Of course, if we had said 10 years ago that we'd achieve AGI in 10 years, everyone would have thought it was fantasy; now when we discuss it, there are already possible ideas about where AGI might emerge. It's hard to say whether it's technology based on pattern recognition for learning, but it has indeed produced emerging properties under massive data and parameters, making people feel it has initially acquired some intelligence.

This gives us an entry point for AGI. But as for when it will be achieved — like predictions for other AI technologies — it will always keep getting closer, but there will always be a sense of being just out of reach.

Take autonomous driving: Ford said in the 1950s that we'd have full self-driving in 20 years, but looking back, that goal certainly wasn't achieved. But at least in recent years, although we still haven't achieved full autonomous driving, expectations for it have been shortening — from 20 years to 10, from 10 to 5, from 5 to 2, with many teams even saying next year. I think Elon Musk's autonomous or assisted driving solutions will ultimately be feasible — it's just that the technology development path is hard to predict.

From the general shortening of expected timelines, we can see our technology is making great strides. But we can also see that people's expectations for AI are completely different from expectations for computers. Computer expectations have always been based on Moore's Law, predictable according to fixed patterns — even a formula could be written for what might happen. But for AI development, it's an inverse Moore's Law: each time progress is made, solving that final 10%, 5%, or even 1% may require more effort and cost than before.

Oasis Capital: Where do you think the capability boundaries of LLMs lie?

Professor Fisher Yu: Looking back at deep learning's development in recent years, although it has many supporters, there are also dissenting voices. For example, a scholar at the same school as Yann LeCun, who won the Turing Award for his deep learning contributions — Gary Marcus — wrote an article in March 2022 called Deep Learning is Hitting a Wall. His skepticism included whether language models can truly reason and have common sense. Although these doubts are reasonable, deep learning has demonstrated astonishing capabilities time and again, so it's hard to say where the capability limits are. Many years ago, I discussed with an NLP student: what would happen if we downloaded all web pages and learned basic facts? At the time, we found a very interesting phenomenon in language processing — people generally don't write "common sense" content online; it's hard to obtain "common sense" from the web. For example, bananas are yellow — when people write articles online, they won't straightforwardly write that bananas are yellow, because it lacks newsworthiness; they'll only write about discovering red bananas or other strange things. But now we find that when your data boundary expands to a certain extent, much "common sense" can also be learned.

In our conversations with ChatGPT, for very basic or obvious matters, it can speak very coherently and logically. For challenging and in-depth topics, it's at a loss. The theoretical boundaries of large language models are constantly being challenged and broken, but there will be resource and commercial boundaries. For example, if we only scale up using current technology, our data and compute are already at the limit. If we continue scaling up with current technological accumulation to develop large models, we'll encounter resource or human bottlenecks. But it's also hard to say — as attention to this problem grows and more resources can be invested, new technologies may emerge to compensate. For example, GPT's own technology, Google's Transformer, and other underlying technologies. Large companies will pay more attention to LLM capabilities and invest more. Once technological bottlenecks break, boundaries become hard to predict.

Vitality

What do you think is technological vitality?

Tireless, fearless innovation and challenge. —— Fisher Yu, Assistant Professor at ETH Zurich

Oasis Capital is a new-generation venture capital firm in China, dedicated to discovering the most vital entrepreneurs of the next decade and growing alongside them to create long-term value. "Vitality" is Oasis's vision and mission. This vitality represents both the direction of structural transformation in the era and the resilience and evolutionary power of entrepreneurs.

Oasis Capital focuses on early and growth-stage investments, with individual investments ranging from $3 million to $30 million, targeting robotics, artificial intelligence, technology services, and other fields, supporting China's technology-driven new service upgrade.