Hongjiang Zhang: My Eight Observations and Views on AI and Large Models

This article is republished from Deep Web Tencent News (ID: qqshengwang)
Silicon Valley lies along a narrow, roughly 50-kilometer strip in California stretching from San Francisco through Santa Clara to San Jose. It's a critical American electronics industrial base and the world's most famous concentration of the industry. This land has always been a birthplace of legends. The earliest can be traced back to 1838, when an adventurer named Sutter discovered this fishing village, conquered the indigenous people, and became a wealthy man thanks to the land's abundance.
Ten years later, a chance discovery revealed that one area of this land was covered in gold. In an instant, Sutter became the richest man in the world. But this vast fortune also brought him catastrophic disaster. Adventurers from across the globe, contemptuous of law, plundered the place. The result: Sutter died poor, lonely, and forgotten. "Gold Rush" is what people call the craziest memory of that era and this land.
Perhaps as some kind of historical compensation, years later the gold dust no longer visible on this ground, yet a gathering of brilliant, gifted minds continues to assemble here, creating myths about technology. Since the mid-1960s, Silicon Valley has gradually formed around the rapid development of microelectronics technology. Its characteristics include reliance on nearby top-tier American research universities — Stanford, UC Berkeley, Caltech, and other world-renowned institutions — a base of high-tech small and medium-sized companies, and major corporations like Cisco, Intel, HP, Lucent, and Apple.
After the 1980s, research institutions for emerging technologies in biotech, aerospace, marine science, communications, and energy materials emerged one after another, making the region objectively the cradle of American high technology. Technological innovation in Silicon Valley has never stopped. The garage where HP's founders repaired oscilloscopes. The cozy little house where Microsoft started. The inspiration at Fairchild where a group of fiery young people created the CPU. The bar at Intel where genius flowed freely... Silicon Fever and The Silicon Valley Way also inspired a generation of Chinese internet leaders to write their own legends with the times.
Of course, technological innovation thrives endlessly on this hot soil of Silicon Valley. OpenAI has become the new legend here, opening a Stargate to a new era. Silicon Valley's architecture is unpretentious — few high-rises, mostly three- or four-story buildings hidden among green trees. And yet from this place, countless technological legends have been born.
In late July, during Qingteng's global private visit study trip to Silicon Valley, Hongjiang Zhang, founder and founding chairman of BAAI, gave a sharing session for Qingteng students titled "The Development and Outlook of AI Technology." Below is the transcript:
"Large models are the new generation of operating systems, and will bring a new ecosystem. Today all software companies, especially B2B software companies in the United States, were the earliest to take action rewriting software with AI — that's what I saw early last year. This year, American companies are all using large models to redefine their software," Zhang elaborated.
Dr. Zhang discussed eight observations:
First, the essence of large models — scaling laws;
Second, the center of computing has shifted from CPU-centric to GPU-centric;
Third, large models are operating systems and will establish new ecosystems;
Fourth, applications of large models;
Fifth, build large models or small models?
Sixth, investment in large models;
Seventh, multimodality is the ultimate model for AGI;
Eighth, multimodal large models will empower robots.

First observation: The essence of large models.
Since the Transformer appeared in 2017, there's been a series of models in this lineage. The models were all good, but they were specialized models and small models, mostly done by Google. But OpenAI, starting from GPT-1, quickly found the breakthrough. Large models have three major characteristics: first, large scale; second, emergence; third, generality.
Today when we talk about large models, the core is scaling law. What Ilya Sutskever (OpenAI co-founder and chief scientist) saw further than others was scaling law.
Second observation: As scale increases, the past CPU-centric approach has become GPU-centric.
When everyone is buying GPUs and building clusters of ten thousand cards, the impact on data centers is also significant — the architecture, operations, and design of data centers all face new challenges. If you buy 10,000 cards for a data center, efficiently using those 10,000 cards is extremely difficult. For data centers with a few thousand cards in a single cluster, very few achieve utilization rates above 50%. On one hand we're short of cards, on the other hand effective utilization is not high.
Another aspect of scaling law: when clusters become so large that people can't handle them, they'll naturally look for another solution — putting more computing power on edge computing. Going forward, we'll definitely see an intelligent architecture connecting cloud, edge, and end devices.
Third observation: Large models are operating systems and will establish new ecosystems.
This operating system has a natural language user interface, better than previous operating systems. This will also bring a new Moore's Law.
Large models are the new generation of operating systems and will bring a new ecosystem. Today all software companies, especially B2B software companies in the United States, were the earliest to take action rewriting software with AI — that's what I observed early last year. This year, looking at American companies, whether successful or not, they're all investing heavily, using large models to redefine their software.
Fourth observation: Applications of large models.
If we divide applications or models into five layers, today we're at the foundation layer, L1 or L2. At the pace of five years, we can reach L3. Whether these companies are GitHub or chat-based, customer feedback has generally been quite positive. When will these applications truly land? It can be divided into several stages:
The first stage is selling shovels — GPUs, data centers, cloud. The stocks of these companies, whether Microsoft, NVIDIA, or data center companies, that's the first wave. After that come applications, many personalized and B2B applications will start from here. Finally there's Physical AI — various robots, such as those for scientific research. That's five years out. Actually, between the first and second stages in the US, there's an inserted stage: B2B software companies. All SaaS companies are working hard to use large models to enhance software capabilities. So in B2B SaaS, in productivity — this is a market where the US surpasses China, and there's already significant investment here.
Will there be a super app? I believe there definitely will be. If OpenAI continues to make progress on products, I think its potential to become a super app is very high.

Fifth observation: As an entrepreneur, build large models or small models.
Today people say: "I can't build large models, so I'll build small models." What I want to say is that the scenarios where a small model does one thing well are very limited. I believe only by making large model performance good can true emergence appear. If you need to go vertical, you also need to put it on the edge, make it small through distillation methods and continual learning methods — rather than building a small model from the start.
To use an analogy: you send your child to vocational school, they quickly learn one skill, but that's all they know. Maybe the child does that one thing well, but as soon as new technology appears, they can't keep up. You should send your child to Harvard for a good undergraduate education. After graduation they can pursue vertical development — three years of medical school to become a good doctor, or three years of law school to become an excellent lawyer. But a good undergraduate foundation is crucial.
Sixth observation: Investment in large models.
From an investment perspective, the opportunity is enormous. Today's AI application areas are all existing markets. If AI can improve efficiency in these areas by 10% to 15% annually, after several consecutive years, the market will double. Just by improving efficiency of existing businesses, AI has a huge market. Today 60% of money goes into foundation large models; not many people are truly building applications. But look at the previous cloud technology wave — most cloud revenue was in applications and SaaS. So AI's future opportunity is also in applications.
Seventh observation: Multimodality is the ultimate model for AGI.
Recently Sora was very stunning, but multimodal large models are actually far more complex than Sora. Sora only did one thing: text-to-video, and also image-to-video. It's very realistic, making people feel amazed. For me, what's stunning is that it generated a three-dimensional world, and behind it there's no 3D model. Completely using data to train a world model — of course, a rudimentary world model.
GPT-4o is an end-to-end trained model. First, it doesn't translate speech to text, send to LLM, output text, then translate to speech — there's no such process. Input is speech, output is speech, and this achieves 200 milliseconds. And it has no translation process, no process from speech to text and text to speech, no information loss.
GPT-4o has reached this moment. Going forward, doing multimodality will definitely be end-to-end, definitely a unified model. There are many unified models today, right? Today there are audio models, video models, text models (ChatGPT 3.5), multimodal generation models, multimodal reasoning models. Only when we reach world models will there be Embodied Artificial Intelligence and general-purpose robots, and only then AGI. GPT-4o is moving in this direction.
Eighth observation: Multimodality will empower robots.
The difference between specialized robots and general-purpose robots is that specialized robots have a program, while general-purpose robots have an embodied brain. The development of multimodal large models toward world models will give robots true general thinking and action capabilities, and is also the necessary foundation for general-purpose robots.
The future is a world of autonomous intelligence — general Embodied Artificial Intelligence. Machines will not only possess human thinking and action capabilities, but also human autonomy.
What is the singularity arriving? As long as machine learning capability exceeds human learning capability, intelligence will exceed human intelligence. The so-called singularity arriving means exactly this.
Geoffrey E. Hinton (Turing Award winner) believes that digital computation will definitely surpass biological computation, meaning digital intelligence will definitely surpass biological intelligence. This means the singularity will arrive sooner or later — whether in 5 years, 25 years, or longer.




