Code Brain | Eight Takeaways on Google Gemini

A Small Step for Gemini, a Giant Leap for General Agents/Robots

In 1948, inspired by psychiatric patients, British psychiatrist Ross Ashby invented an eccentric machine — the "homeostat" — and declared that this device, costing roughly £50, was "the closest thing to a synthetic brain that the human race has yet designed."

The homeostat used four bomb-control switch gear mechanisms from Britain's Royal Air Force as its base, topped with four cubic aluminum boxes. The four small magnetic needles on top of these boxes were the machine's only visible moving parts, oscillating in small water troughs like compasses.

When activated, the needles would move in response to electrical currents from the aluminum boxes, the four needles perpetually poised in a delicate, sensitive balance. The homeostat's sole purpose was to keep these four needles centered — to maintain a state of "comfort" for the machine.

Ashby tried various methods to make the machine "uncomfortable": reversing the polarity of wire connections, flipping the needle directions, and so on. Yet the machine always found ways to adapt to its new conditions, swinging the needles back to center. In Ashby's words: the machine "actively" resisted any attempts to disrupt its equilibrium through synaptic action, executing "coordinated activity" to regain balance.

Ashby believed that one day, such a "crude device" would evolve into an artificial brain "more powerful than any human," capable of solving all the world's complex and intractable problems.

Though Ashby knew nothing of today's AGI evolution, and though four small magnetic needles as sensors are laughably inadequate for the conditions required by intelligence, his machine challenged everyone's understanding of "intelligence" at the meta-logical level — isn't intelligence simply the ability to absorb multi-modal information from the environment, correct behavior based on feedback, and process tasks?

From the eccentric "homeostat" to today, 75 years later, Gemini — claimed to surpass human capability in multi-modal task processing for the first time — accelerates its iterative advance toward billions of years of carbon-based intelligence evolution through the injection of native multi-modal big data.

The speed of machine intelligence evolution now far exceeds our imagination.

One year ago, OpenAI toppled the AI banner Google had spent years erecting, building a tower of Babel for human language through "brute-force aesthetics."

One year later, Google unveiled Gemini, "fighting fire with fire" to construct humanity's first unified cross-modal model, becoming another node accelerating AGI evolution.

Though Gemini was mired in controversy over "exaggerated video demos" on its launch day, the dawn of unified multi-modal capability has undeniably begun to glimmer. What capabilities does Gemini — this "twin star" symbolizing perceptiveness and keen curiosity — validate? How will Google's gears of fate turn? Is time on OpenAI's side or Google's? What does multi-modality mean for Agents and embodied intelligence? Are the foundations for the emergence of AGI with autonomous consciousness already in place? What revelations does Gemini hold for the future?

01

Cross-modal knowledge transfer in large models proven once again

For humans, more important than learning skills is the ability to transfer knowledge across domains and through time. If machines master cross-modal knowledge transfer, they more easily approach "generality."

In July this year, Google released RT-2, a large model-based robotic system that gave people hope for general-purpose robots. The robotic arm, drawing on the language model's "common sense," could "pick up extinct animals" from a table — demonstrating cross-modal knowledge transfer from commonsense reasoning to robotic execution.

In December, Gemini — this heavyweight move by the tech giant — once again validated large models' cross-modal knowledge transfer capability: the "common sense" of language models can transfer to the training of other non-language modalities added downstream.

Language models are the foundation of cognitive intelligence, and the most basic cognitive intelligence is "common sense."

Without common sense empowerment, many practical deployments of multi-modal large models would be difficult to achieve. Gemini smoothly transfers this "common sense" learned from the internet to downstream multi-modal tasks. Like RT-2, it achieves cross-modal integration through internet text knowledge transfer — Gemini can connect abstract linguistic concepts to understanding auditory and visual objects, even linking them with Action to become an intelligence deployment system.

From a model training perspective, compared to language models trained on massive internet data, downstream models (such as robotic models) can be trained with smaller datasets through knowledge transfer. This progressive training approach solves the downstream data scarcity problem that has plagued academia for years.

For example, to achieve the effects shown in the video (which sparked skepticism about Gemini's video understanding, though this doesn't affect the discussion of cross-modal knowledge transfer), Gemini first needed certain ontological knowledge — it had to know the concept of "duck," know what color ducks typically are, know what blue is. Only then, upon seeing a "blue duck," could it react similarly to humans and express the "common sense" that "blue ducks are uncommon."

Gemini perceives through sound and vision that the blue duck's material is rubber, and knows that rubber's density is less than water's. Based on this common sense and reasoning, when it hears a squeak, it can predict that "the blue duck can float on water."

From RT-2 to Gemini, from single-modal capability to the "fusion" of multi-modal perceptual and cognitive intelligence, from separated "five senses" modules of eyes, ears, mouth, nose, and body to an integrated, complete digital "person." Doesn't this mean that on the path to simulating human intelligent behavior, the "unified" model is the true way forward?

02

The unified multi-modal model,

finally surpassing定向优化的单模态模型

Humans perceive, cognize, and generate emotions and consciousness through multi-sensory integration. Gemini is also practicing multiple modal inputs, comprehensive brain processing, and multi-modal outputs — this comprehensive "simulation" of human intelligence is accelerating its evolution.

Previous multi-modal model training resembled a combined system with separate eyes, ears, arms, and brain — their unified coordination was not particularly strong.

The direction represented by Gemini clearly feels like the large model becoming a complete digital person — a silicon-based whole with coordinated hands, eyes, brain, and mouth.

Gemini is the first truly end-to-end multi-modal system.

Previously, models with定向优化 for single modalities typically outperformed those processing multiple modalities simultaneously, so the conventional approach was single-modal model training. Including GPT-4, which "stitched" different modalities into the whole rather than being a unified multi-modal model.

Gemini's particularly exciting aspect is that it was designed from the start as a native multi-modal architecture, with various modalities interleaved throughout the training process from the very beginning. If previous large models were brains with externally connected senses or robotic arms, now the eyes, ears, and arms grow directly from within the body, capable of effortless, natural movement.

Whether in model architecture, training process, or final presentation, Gemini achieves truly seamless multi-modal fusion.

For the first time, Gemini shows us a unified model that can handle all modalities — and perform better than models focused on a single modality! For instance, compared to Whisper, a model specifically optimized for speech recognition, Gemini shows clear accuracy improvements.

This signals the dawn of the unified multi-modal era.

In fact, Gemini wasn't the first model to validate that different modalities can help improve each other's performance. This was also demonstrated in PaLM-E: "PaLM-E trained across diverse domains, including internet-scale general visual-language tasks, shows significantly improved performance compared to single-task robot models."

Another example of modalities enhancing each other is large language models' multi-lingual processing capabilities.

If we treat different international languages as distinct细分 "modalities," the practice of language large models has proven that unified processing of native data across all languages (tokenization and its embedding) collectively built the tower of Babel for human language.

The overwhelming volume of English data in language large model training similarly benefits the model's understanding and generation of other, less-represented languages — language knowledge transfer has been repeatedly validated.

Just as a person skilled at tennis can触类旁通地 improve their squash or golf abilities.

Since the explosion of large models in February this year, many people gradually developed a faith that "unified multi-modal models will surpass single-modal models" — but this faith had never been validated by large-scale practice. This time, Google's Gemini demonstrates the prospect of this faith's realization, allowing more people to reshape and solidify this conviction.

In the future, building dedicated recognition models for speech recognition, machine translation, and similar tasks may no longer make much sense. Many generation tasks such as TTS, image generation, and others will also be unified under large models. Some may complain that large models are too expensive and slow, not suitable for all applications — but cost and speed are primarily engineering problems. In practice, we can distill unified multi-modal models down to specific modalities or scenarios.

We firmly believe that unified cross-modal large models will become the mainstream pathway to achieving AGI.

Extending further, "modality" encompasses not only sound, images, and video. Olfactory, gustatory, tactile, temperature, and humidity sensors are also different modal means of acquiring environmental information — all objects to be encompassed within the unified model.

At its core, various modalities are merely carriers of "information," a rendering, a presentation format, a means for intelligent agents to interact with the physical world. In the eyes of the unified model, all modalities can ultimately be represented by unified multi-dimensional vectors internally, enabling cross-modal knowledge transfer and information intersection, alignment, fusion, and reasoning.

When the barriers between modalities are broken through, and we dissect the core beneath all these renderings, we see the starting point of cognition — language.

03

Language is the core and主线 of the unified model

In the AGI system we imagine, is its core and主线 visual or linguistic? Some believe it's visual, but we believe more strongly that language is the core.

Stalin once said in his linguistic writings: "Any lower organism has its own language."

But no matter how many layers of variation they possess, none constitute true language. True language is unique to humans — including invented writing, symbols, and subjectively assigned meanings, then combined to form countless expressions, carrying human cognitive evolution and knowledge accumulation over millions of years.

Language is the starting point and source of cognition. Human linguistic information contains human highly abstract cognitive capabilities, while audio, images, and video are more感性, representing human emotions and concrete capabilities, more inclined toward capturing human perceptual abilities.

When humans learned cognition, combined with more感性 expressive capabilities like audio, images, and video — from perception to cognition, from emotion to logic — this is the state of our human brain. The same applies to unified multi-modality: when the鸿沟 is bridged in information processing and reasoning,融会贯通 is the natural result.

In both RT-2 and Gemini, language occupies the主线.

For example, in RT-2, the parameter scale and data volume representing the language modality far exceed those of the downstream image and action modalities.

We predict that in any future AI system, regardless of whether it's a language task, the language model will serve as a foundation model and training starting point, with other modalities or task data added for continued training — all inheriting to some degree the language model's powerful cognitive capabilities.

If this is truly achieved, perhaps this will be the language model's greatest contribution to AI, because it truly realizes researchers' original vision and positioning for it — the Foundation Model.

04

The "brute-force aesthetics" methodology has become consensus

Looking back at OpenAI's initial victory, it was primarily not an algorithmic innovation but a triumph of "brute-force aesthetics."

Today, "brute-force aesthetics" has become an industry methodology for building AI. Specifically, it manifests in two aspects: technical and organizational.

Technically, the fundamental methodology of GPT-style large models is: keep the model architecture simple, then focus intensely on data and compute.

It seems simple, but before OpenAI successfully built GPT-3, many found it hard to believe that a simple Decoder-only architecture, plus an optimization objective for next-token prediction, trained on massive unsupervised internet data through self-learning, could handle various AI tasks and thus move toward general artificial intelligence. Only OpenAI persisted in this faith and successfully implemented it in engineering.

Organizationally, OpenAI's approach was: everyone围绕 a general model, rather than letting a hundred flowers bloom.

Before large models emerged, much AI research was small workshop-style: a few researchers with a few interns building a system for a specific task. Research topics were also highly concrete — TTS, ASR, machine translation, vision, and so on — rather than general models like large models.

Previously, this small workshop organizational style was typical in Google and Microsoft's research labs, where hundreds of researchers simultaneously pursued dozens of different research topics. OpenAI, on one hand, truly believed in "brute-force aesthetics"; on the other hand, precisely because of resource constraints, it反常识地 chose to have hundreds of people all-in on a single GPT model.

The essence of "brute-force aesthetics" is minimalism and focus, then repeating and amplifying through scale.

Scale encompasses model parameters, data, compute, personnel, and other dimensions. As model parameter count and training data scale increase, performance exhibits the "emergence" phenomenon we all know today.

Although Google invented most of the underlying key technologies that today's large models depend upon — such as the Transformer architecture, Instruction Tuning, CoT, Mixture of Experts, and so on — OpenAI used these key technologies to practice the "brute-force aesthetics" methodology of the large model era, beating Google into helpless submission.

And this Gemini release has made everyone realize that perhaps Google internally has also reached consensus on the "brute-force aesthetics" methodology.

When Google, with far greater resources, awakens from slumber,认同并掌握 the "brute-force aesthetics" methodology, and focuses its energies in one direction — might even greater resources birth even greater miracles?

05

Google's sleeping lion has awakened,

the gears of the brute-force machine begin to turn

Gemini's emergence makes it clear that in this peak对决, Google has caught up. With a clear consensus on "brute-force aesthetics," when this浓眉大眼 engineer machine decides to get "brute-force," it is absolutely not a competitor to be underestimated.

First, Google has finally learned to "大力出奇迹" organizationally. Gemini's technical report spans nine full pages of author lists, with over ninety names per page — more than eight hundred people, already exceeding OpenAI's total company headcount.

For Google, with ten times OpenAI's researcher count, moving from its consistent bottom-up approach to top-down, the execution difficulty is imaginable. The organization must trigger高度统一的使命感, then rapidly adjust strategy and structure — including merging Google Brain and DeepMind, the two major AI labs, into the new Google DeepMind department, beginning to上演复仇者联盟.

"Brute-force aesthetics" organizational engineering resembles the Manhattan Project, requiring soulful leadership figures. Facing the organization's focal issue — coordination between multiple teams, where to focus, whether two teams tackle separately or融合协作 together? Even for a large enterprise like Google, facing massive resource demands, careful selection of investment directions is essential.

How to effectively allocate resources, concentrate efforts to achieve established goals, and implement at scale — this is every leader's challenge. Hassabis, as a formidable leader, has demonstrated not only his leadership abilities but also the deep organizational strength of a company like Google.

Beyond strong organization and high-density talent, Google also holds unique advantages in data scale and user scale — it is the absolute king of distributed computing.

This time, Google also simultaneously released its most efficient and scalable TPU system to date, Cloud TPU v5p, supporting the training of frontier AI models. The new-generation TPU will accelerate Gemini's development, helping developers and enterprise customers train large-scale generative AI models faster, thus bringing new products and features to market more quickly.

Google's years of full-link ecosystem cultivation and various product lines with hundreds of millions of users also provide fertile ground for unified model deployment. This gives Google the most confidence in应对 the complementary alliance of Microsoft and OpenAI.

This time, Gemini launched three versions: (1) Gemini Ultra for highly complex tasks; (2) Gemini Pro, the best model for diverse tasks; (3) Gemini Nano for on-device deployment (such as phones).

So, given Google's strength in talent, data, compute, users, and other "brute-force aesthetics" essential elements, as long as it keeps pace, when the brute-force machine's gears of fate begin to turn, it may well carry the AI arena's script toward an entirely new situation. OpenAI's局面 of riding alone, lonely and undefeated, is beginning to change.

06

Time will ultimately be AGI's friend

In the competition to come, whose friend is time more — OpenAI's or Google's? So far, OpenAI has enjoyed the enormous momentum of first-mover advantage. But undeniably, while pursuing AGI, OpenAI must also face growth bottlenecks, commercialization pressures, and investor questioning (rumors say Microsoft demands OpenAI maintain a six-month lead over Google forever). Under immense pressure, it's inevitable that actions become distorted.

The recent OpenAI palace intrigue has deeply wounded the company. Though Sam said this only delayed OpenAI's AGI dream by five days, in the AI battle where不进则退, this has cost at least several months in the race against Google.

Now that Google's lion has awakened, OpenAI will face even greater competitive pressure. More importantly, OpenAI's non-profit mission and its massive fundraising pressure remain fundamentally irreconcilable — like a ticking time bomb — and its competitive-cooperative relationship with Microsoft is also微妙异常.

Under pressure-induced distortion, what's more likely to intensify is OpenAI's internal路线之争 (effective accelerationism vs. superalignment). Other black swan events may also emerge, not uncommon in capital-intensive technology entrepreneurship — as seen in many autonomous driving company stories.

In contrast, Google as a mature, stable giant has none of OpenAI's fragile board architecture and the non-profit vs. capital contradiction behind it, nor the entanglement of微妙 investor relations. With its雄厚家底, possessing碾压级 advantages over OpenAI in researchers, data, compute, user scale, and other dimensions, once it认同并掌握 the "brute-force aesthetics" methodology, it's like a massive machine whose后发优势 may grow increasingly apparent over time. So, from a competitive perspective, might time be more Google's friend?

Of course, Google's risk lies in big company organizational disease, and the potential过分 top-down and excessive resource concentration on developing a single model after fully转向 "brute-force aesthetics," which could冲垮 the bottom-up and百花齐放 innovative culture that previously underpinned Google's success.

OpenAI will certainly respond with full force, striving to maintain its AGI leadership position. Gemini will逼仄出 an even more astonishing GPT-5, while Google under the gears of fate will continue to unleash Gemini 2.0... Under this arms race, AGI's推进步伐 will grow ever more rapid. Whether Google or OpenAI, each in its own way is螺旋式推动 AGI forward through fierce competition.

AGI's historical wheel has already rolled forward. Time will ultimately be AGI's friend.

07

Multi-modality is the foundation of Agent and embodied intelligence

Norbert Wiener, father of cybernetics, looking to the future in Cybernetics, wrote: "Human capabilities are now greatly extended by machines — radar extends the human eye, jet engines or tires extend human limbs, and the autopilot is the nervous system connecting them."

Today's large language models can encode the world's rich semantic knowledge. Their significant weakness is the lack of grounding, making "hallucinations" inevitable.

Multi-modality itself provides the foundation for grounding. With this foundation, Agents can interact with a multi-modal environment and obtain necessary feedback, making autonomous planning more reliable.

Robots and other embodied intelligent agents are also Agents — except they are not virtual but physical entities with bodies, with "hands and eyes," capable of concrete tasks in the physical world. Thus, multi-modality is the foundation of Agent and embodied intelligence, and a necessary condition for reducing hallucinations.

Hassabis revealed that Google DeepMind is already researching how to combine Gemini with robotics for physical interaction with the world. After all, to become truly multi-modal still requires touch and tactile feedback.

This path never before trodden may bring major breakthroughs in robotics. Unified multi-modal models like Gemini can become the foundation for rapid AGI innovation, promoting intelligent agents and their planning and reasoning, as well as physical robots' interaction with the environment.

Agent = brain cognition + perception + action. Agents and embodied intelligence need both perception and cognition; both brain and external support.

Today we clearly see: large language models solve high-level cognitive problems, multi-modality provides grounding foundation, Agents solve autonomous planning problems, and embodied intelligence completes final physical world actions and interaction — this组合拳 makes all elements of general Agent/robotics seem to be in place.

And the unified cross-modal model appears to be the必经之路. Gemini's small step may be a giant leap for general Agent/robotics.

08

Are the foundations for AGI with autonomous consciousness

already in place?

Before and after the explosion of large models, AGI transformed from an abstract concept that most professional researchers disdained or dared not associate with, to suddenly coalescing into mainstream consensus. Discussions about how AGI will arrive abound.

When large models exploded globally in February this year, many believed that following the "brute-force" path, simply scaling up language models would bring AGI. Now it appears this won't work.

Language models are indeed the foundation of cognition and the core of intelligence, but they are only the cornerstone of AGI.

To achieve AGI still requires coordination with many surrounding modules.

Since April, many began patching around language models, sparking a wave of Agent enthusiasm — but now this too appears to be空中楼阁. Without multi-modal加持 grounding, Agents' reasoning and planning are extremely unreliable, mere噱头 in many scenarios.

Gemini's emergence lets us see the next essential cornerstone for AGI emergence: multi-modality.

Without multi-modality, language models are "brains in vats." Moreover, AGI emergence necessarily requires native multi-modality, not stitching together multiple independent models — because the拼接 approach likely cannot achieve deep complex reasoning and seamless knowledge transfer in a unified multi-modal space. And Gemini's excellent performance on multi-modal tasks this time provides powerful endorsement for unified multi-modality.

With language-model-centric multi-modality in place, virtual and physical Agent deployment is no longer空中楼阁. Various modules added to Agents, such as memory, tool use, environment feedback, are also necessary conditions for AGI emergence.

In an interview with Lex Fridman, Hassabis expressed: "Consciousness is the feeling that comes when information is being processed." When large models' multi-modality fuses as丝滑 as human perception, when Agent modules together自如 adapt to various environments, can we deduce that machine autonomous consciousness already has the foundation for "emergence"?

If we take a longer view, perhaps the trend is already clear — a trilogy on the path to AGI: large language models laying the cognitive foundation, multi-modality/Agent/embodied intelligence solving grounding, and AGI with some autonomous consciousness will ultimately "emerge."

Conclusion

British writer Samuel Butler wrote a novel called Erewhon, containing a section "The Book of the Machines," in which a fictional thinker expresses evolutionary concerns about machine autonomous consciousness: "In the ultimate development of machine consciousness, we have no security. Who can say that the steam engine is not a conscious species?"

Clearly, machines have clear relationships of inheritance, development, and evolution among themselves — like the evolution from music box cylinders to punched paper tape, like the evolution from GPT-1 to GPT-4V.

So can machines be regarded as a "species"?

It's just that their evolutionary process requires human participation — but who can say that human creation and participation is not the unique evolutionary strategy of this machine "species"?

In Darwin's theory of evolution, we默认 that the essence of "evolution" is gene evolution at the protein-coding level, functioning to optimize organism survival. But if machines can be created by humans, extending or altering various multi-modal natural organs, can we say that machines are a new form of human evolution — replacing traditional gene evolution as a more efficient way to change human "traits"?

And when machine autonomous consciousness evolves to the day of breaking free from human dependence, when humans complete their mission of evolving AGI as a new species, can humans then exit the historical stage like ancient apes?

As the last generation of pure carbon-based beings, if we can walk at the forefront of this mission's path in our remaining years — how tragic, how fortunate!

When human-built high-rises become ruins, when ancient stone tablet inscriptions are weathered and eroded, with no one able to decipher their meaning — they are merely traces left by some species. Millions of years of history are but this species' continuous reproduction, survival, and continuation, essentially no different from today's evolution of GPT, RT-2, and Gemini, until continuously creating new species.