Oasis Capital Dialogue with Professor Wu Yu: Pushing Every Millimeter
***Oasis Capital: Could you share your recent and future research priorities and directions?*** ***Professor Wu:*** My recent work has focused primarily on generation. Earlier, I concentrated more on recognition tasks like detection and segmentation. As research has progressed, the field has gradually come to realize that generation seems more promising — and indeed, we've seen major breakthroughs in the past year or two.
The pace of the AI revolution is dizzying. How do you choose your direction amid the torrent of change, learn quickly, and adapt quickly?
Today we've invited Professor Yu Wu from the School of Computer Science at Wuhan University. Professor Wu went from switching PhD programs to winning a Google PhD Fellowship in just two years. From mechanical engineering to computer vision, his path of choice may be a textbook example of embracing change. Enjoy.
Oasis: Could you introduce your recent and future research focus areas and directions?
Professor Wu: My recent research has concentrated on generative models. Earlier work focused more on discriminative tasks like detection and segmentation. As research has progressed, people have gradually realized that generative directions seem more promising, and the past couple of years have indeed seen major breakthroughs. I mainly work on generation, and also on multimodal learning — things like vision-language and audio-visual cross-modal associations.
Oasis: Facebook recently open-sourced SAM and DINO. Are these models discriminative? Has your research direction shifted from discriminative to generative?
Professor Wu: SAM and DINO are discriminative. The AI boom that started around 2012–2013 was primarily about "discriminative" tasks — using AI models to solve real-world problems like human detection, tracking, face recognition, and so on. After years of development, both applications and research have become relatively mature.
SAM is indeed the most visible work in computer vision, sparking considerable excitement. But compared to ChatGPT and other LLMs, SAM's influence is clearly more limited. Initially many people thought SAM would be CV's "ChatGPT moment," but in practice, after solving segmentation and edge detection, how to truly address complex real-world visual tasks remains an open question. From our research experience, SAM doesn't solve all problems in CV — it's just a reasonably good auxiliary knowledge model.
As for why I've moved to generative tasks: discriminative approaches are already relatively mature at the application layer, while generative approaches need more breakthroughs. Generative demos are always stunning, but there's always something missing when it comes to actual deployment. For example, AI-generated albums or portraits look realistic at first glance, but careful examination reveals many errors. Moreover, we can't directly use generative tools for productive planning. So our recent research focuses on pushing the frontier of generative tasks in academia toward actual productivity tools.
Oasis: What explorations and breakthroughs do you think academia can help industry achieve? At the current stage, what kinds of research should be combined with industry and deployed quickly? What needs to be accelerated by academia?
Professor Wu: This is something people have been paying close attention to recently. For LLM-based work, it's hard to say whether academia can lead on large models — currently academia clearly doesn't have sufficient resources, so both domestically and overseas, companies and industry are the ones releasing large models. Industry is undeniably the main force behind this wave of large models, but academia still plays an important role. Going forward, the two need to integrate and develop together.
Recent industry work has had relatively little innovation at the underlying algorithm level. If you set aside engineering tricks in training, the model innovation is minimal. If AI remains stuck in the mindset of stacking parameters, it will quickly saturate and become an AI bubble, with no new directions left to pursue.
Academia has an important mission — how to combine with large models and produce papers that can truly scale and land in practice. This is a challenge for academia, and more importantly, a responsibility.
Take the generative field: Diffusion was proposed by academia, but it was really industry, with Stable Diffusion, that brought it to prominence. Industry models aren't particularly complex in architecture — through massive data and training strategies, they achieve good results. But where to go next? If we only keep stacking parameters and data, will there be another "Diffusion"-type revolution? This "next step" requires academia's participation.
For example, a recent project of ours converts fully free-form Stable Diffusion generation into customizable generation. Existing models can generate images from text descriptions — say, "a person sunbathing on a beach." But in reality, humans usually have specific goals for image generation results; not just any person lying on any beach will do. So customized image generation has broad potential. Our recent work introduces an innovative approach based on industry's Stable Diffusion — no additional training, no fine-tuning required for customized generation. Given any user-provided image (a person, object, logo), with language guidance, it generates images that satisfy both the language and visual constraints. This is essentially academia doing secondary development on top of industry's large model compute.
Oasis: Do you think there are worthwhile sub-directions in the current AIGC field?
Professor Wu: Yes. The first is controllable (customized) generation — moving from random generation in AIGC to refined generation. Some peers are doing similar work, but it requires fine-tuning. There was a well-known work at CVPR 2023 Best Paper Candidate called "Dreambooth." Compared to that, our work preserves visual features better, requires no fine-tuning or training, and is faster. The overall research trend in the field is moving from holistic generation to more specific, more controllable generation.
The second sub-direction is image editing. AIGC is pure generation. Real applications involve modifying existing images — making someone smile brighter, removing irrelevant people, removing shadows, replacing objects, and so on. Many people are exploring this direction.
The third is generated image detection. Generated images can involve infringement and misinformation. How to detect generated or tampered images is a valuable sub-direction.
None of these can be solved by brute-force large models alone — they all require thinking within each sub-domain.
Oasis: Have there been any noteworthy papers in these directions recently?
Professor Wu: For customized generation, there's Dreambooth from 2022, which already has 300 citations. On the detection side, there's recently Microsoft's DIRE for Diffusion-Generated Image Detection, for determining whether an image is real or fake — though the research approach is still debatable. Image editing work is more complex.
Oasis: You also care about the detection field. Do you think original content protection is quite a headache?
Professor Wu: Awareness in this direction is gradually increasing. I previously worked on music generation, where copyright was a major issue. Even though it was just academic research, using song data without rights meant the work couldn't be published. Now there's a category of research specifically examining whether AIGC models infringe on rights.
Oasis: Was the shift in your research direction a smooth process, or were there bottlenecks?
Professor Wu: Generally speaking, major directional shifts are somewhat difficult. Relatively speaking, transitions happen through slight modifications within one field, gradually migrating to directions with smaller changes. I've been working on multimodal learning since 2017, which is somewhat connected to AIGC. Because initially multimodal work also involved image caption generation — it's just that what was generated wasn't images, but text. Stable Diffusion is also multimodal generation: input language description, get visual expression as image. So the directional change wasn't large; what changed more was the underlying technology.
A few years ago, everyone was using VAE, LSTM, GAN; recently everyone's using Diffusion and Transformer. We started working on Diffusion in 2021. This transition wasn't a desperate pivot, but gradually discovering more interesting directions and shifting accordingly, until new technology emerged, spending sufficient time to learn it, then continuing forward.
Oasis: Where do you see multimodal landing in 3–5 years? Film, for example?
Professor Wu: For generation, one-click film generation should land relatively quickly. Some researchers have already done similar work. Rather than saying a full film, it's probably more accurate to say short videos with plots. Use LLM to write the script, use AIGC to generate each frame and scene, then string them together. The quality gap with real-world films costing hundreds of millions will certainly exist; generating a demo video is more realistic.
Second, comprehensive generation across audio, video, and image — there's no relevant work in this direction yet. We're trying to do more unified, comprehensive multimodal generation. This is necessarily more challenging than pure image generation, because correlations between modalities need to be considered. For film generation, current quality can't be guaranteed — things like whether scene transitions are smooth, and so on, need gradual improvement. Eventually it should be possible to make films from guidance. The initial quality won't be ideal, but it will eventually reach an approximate level. Colleagues in academia are all working toward this direction; I estimate it should land within three years.
(Oasis: It seems you're quite confident about generative landing directions.)
Oasis: Do you think future multimodal content generation will come straight from one model, or require several models working together?
Professor Wu: My subjective feeling is that several models working together is needed. Each modality model has its strengths, and economically this is more reasonable. With equivalent parameter counts, strictly assigning each model to different modalities, with a central controller coordinating the smaller models, should yield better results. Unless there's a major breakthrough in compute, it's unlikely that one sufficiently large model could complete the work on its own.
Oasis: There are many rumors that GPT-4 is an MoE. From an engineering perspective, unlike GPT-3 and GPT-3.5 which are just one model, it's not actually "one" — behind it there may be 16 smaller models?
Professor Wu: Achieving a qualitative improvement over GPT-3 purely by stacking model parameters would be too costly for OpenAI. If GPT-3's cost multiplied several hundredfold while the improvement was marginal, there would be no commercial value.
Oasis: From the last wave of AI to this wave, the boundary between academia and industry has become increasingly blurred. You've also experienced crossing boundaries — how will you choose this time?
Professor Wu: Both sides have advantages. Industry has massive resources to support your work, which is attractive from a pure research perspective. But compared to academia, industry is less free, and it's harder to persist in one direction. For example, with the recent large model frenzy, in industry, if my previous field was video understanding, I might now be forced to work on NLP large models. I ultimately chose academia because I want to explore more freely, to push the boundaries of scientific exploration millimeter by millimeter. What industry ultimately delivers is product applicability, practicality, and business value — not breakthroughs in model algorithms and technology.
Oasis: Fei-Fei Li proposed the direction of Embodied AI. What do you think of it?
Professor Wu: I worked on Embodied AI for a period in 2018. Recently it feels like development has become more concrete. Initially Embodied AI was somewhat conceptual, quite far from actual application — it could only achieve a little intelligence. However, Embodied AI leveraging this recent wave of large models as decision-making brains is a natural progression. So Embodied AI has recently moved up a step.
My senior fellow student Wenguan Wang at Zhejiang University focuses on Embodied AI and has had many impressive new works recently. But Embodied AI also faces the problem that translating to business scenarios still needs time. With the premise of physical AI becoming widespread, Embodied AI will have greater development. If it becomes real-world robots, people will see it as a direction with landing value, quite intelligent.
Oasis: After Elon Musk made humanoid robots, there's been high attention and much controversy. How much does LLM improve robots? Or is it just hype?
Professor Wu: LLMs help, but not decisively. Robots need to solve their own problems first; AI algorithms come second. Because robots exist in the real world, which is less idealized than the software world. Pure software scenarios are relatively simple and controllable; hardware has many errors. For example, with sensors — when training simulated robots, we generally provide precise mechanical data. But in real scenarios, mechanical sensors have errors, leading to errors in the feedback system. The problem isn't the software algorithm; the real world is too complex, everything has noise, noise accumulates, and robots struggle to walk freely.
Oasis: When recruiting students and collaborators, what traits do you mainly value?
Professor Wu: For recruiting students, I mainly look at their ideas. Especially when interviewing undergraduates, I don't judge based on what they've done or what experience they have. I usually present挫折难题 I've encountered in research — say, this method doesn't work, how would you try to think and act? That's what's valuable. If I analogize to AI algorithms, it's meta-learning, learning to learn — how to react to new problems and quickly learn new knowledge, which is precisely the core capability for doing research. We've also found that some students may have excellent past papers, but deeper collaboration reveals that their research habits and patterns aren't very scientific. Observing how someone thinks when facing difficult problems may be more important than existing achievements.
Oasis: From your personal experience, what is worth learning about choosing development directions?
Professor Wu: Change is a good thing; people should embrace change. I previously dropped out of mechanical engineering to switch fields, mainly because I felt mechanical technology was too mature, using things from hundreds of years ago, not suitable for research. Comparatively, going to faster-developing directions was also my interest. The dizzying pace of the large model era means the field is developing rapidly — a challenge for academia, but even more so an opportunity. Looking back, the past transitions went relatively smoothly; I could quickly adapt to new directions, feel interest and passion in AI, be closer to my hobbies and work, and feel deeply satisfied.
Celebrating Vitality
What do you think is technological vitality?
Embracing change, and staying curious.
——Professor Yu Wu
School of Computer Science, Wuhan University





Oasis Capital is a new-generation venture capital firm in China, dedicated to discovering the most vital entrepreneurs of the next decade and growing alongside them to create long-term value. "Celebrating Vitality" is Oasis's vision and mission. This vitality is both the direction of structural change in the era and the resilience and evolutionary power of entrepreneurs.
Oasis Capital focuses on early and growth-stage investments, with individual investments ranging from $3 million to $30 million, concentrating on robotics, artificial intelligence, technology services, and other fields, empowering China's technology-driven new service upgrade.
