A 6,000-Word Postmortem: How Google Got Its AI Mojo Back — From Nano Banna, Genie 3, and Veo 3 to Gemini 2.5's Counterattack

"AI Google Time" has begun

"AI Google Time" Has Begun

👦🏻 Author: Jingshan

🥷 Editor: Zeo

🧑‍🎨 Layout: NCon

A year ago, Google was still seen as a "follower" in the AI race. When ChatGPT swept through Silicon Valley, it looked sluggish.

But just a few months later, everything changed.

Gemini 2.5 Pro dominated the leaderboards. The "Banana" model, Nano Banana, made image generation and editing effortless. The video model Veo 3 demonstrated an understanding of physical reality. Genie 3 could generate an entire virtual world from a single sentence.

With a string of blockbuster products, Google returned to center stage.

This naturally raises the question:

How did Google suddenly get so good?

This wasn't actually a sudden explosion. It was more like "the elephant turned around and monetized its technology." Google is now, with unprecedented determination and efficiency, converting decades of accumulated AI research into product power.

Put more bluntly: Google didn't suddenly get stronger. The giant that pioneered the Transformer era is jogging back into view.


Next, this article will take a deep dive into Google's AI progress and analyze why the company has "suddenly gotten so good" on the AI track.

The full piece unfolds across four core sections:

[1] Topping the Charts, Winning Gold, and Reclaiming the Throne: Gemini 2.5 Pro

[2] Banana in One Hand for Image Generation, Veo 3 Director in the Other

[3] World Model: Genie 3

[4] The Elephant Turns: Technology Becomes Product


Topping the Charts, Winning Gold, and Reclaiming the Throne: Gemini 2.5 Pro

Let's start with the foundational large language model. For most people, the "felt beginning" of Google's sudden surge was the launch of the Gemini 2.5 Pro series.

In the winter of 2022, OpenAI's experimental chatbot sparked a storm, adding a million users per day. Despite frequent factual errors and simple calculation mistakes, its potential began to "shock" all of Silicon Valley — and for the first time, Google felt the pressure of a "fire in its own backyard."

For the next year or so, Google's posture was more that of a slightly clumsy "chaser." From the hastily assembled Bard to the initial attempt at Gemini 1.0, it kept building, but the skepticism never stopped.

The outside narrative at the time went something like this:

The question shifted from "How will Google respond?" to "Is Google still in the game?"

Then a critical inflection point arrived — the official launch of Gemini 2.5 Pro. Though Gemini 2.0 before it was already powerful, it hadn't yet flipped user perception.

Only at this point could Google truly say it had "found the position that once belonged to the tech giant that defined the internet era."

1) Topping the Charts

Six months ago, in March 2025, on the third-party benchmark platform LMSys Chatbot Arena, the codenamed "nebula" Gemini 2.5 Pro burst onto the scene and seized the top spot. Its Elo rating surpassed all rivals, including GPT-4o and Claude 3 Opus.

It achieved true leaderboard dominance.

This performance was widely interpreted by media outlets as Google catching up to or even surpassing its competitors in overall model capability.

According to the LMSys team, this was "the first time in history a model has simultaneously dominated three leaderboards: text, vision, and web development" — a genuine "triple crown."

Notably, the LMSys team's "Web Development" benchmark simulates real-world development tasks, going well beyond mere coding ability. It involves building interactive web applications, covering frontend (UI), functional interactions, dependency management, and complete application architecture.

On programming capability specifically, the claim of "comprehensive crushing" remains debatable from a practical standpoint. But multiple benchmarks and developer feedback show that Gemini 2.5 Pro is on par with the industry-leading Claude 3.7 in code generation, comprehension, and debugging — and even outperforms it on certain specific tasks (like LeetCode-style problems).

Moreover, at every subsequent launch event, large or small, the Gemini series received full upgrade after full upgrade.

At this point, all doubts had evaporated.

Google had returned to the top tier in the most critical core capability of large language models.

2) Winning Gold

Beyond benchmark bragging rights, the AI community cares deeply about how a foundation model performs in arenas that capture broad public attention — in other words, how does Gemini fare in more challenging professional domains?

Here, the International Mathematical Olympiad (IMO) is worth discussing — a competition that AI companies love to trot out for "shock value" headlines.

A specially trained Gemini model with "Deep Think" capabilities reached gold-medal level at the 2025 International Mathematical Olympiad. Gemini 2.5 Deep Think scored 35 out of a perfect 42 at IMO 2025, winning gold by solving 5 of 6 problems and directly surpassing Grok 4 and OpenAI o3.

Meanwhile, before officially releasing GPT-5, OpenAI also used its latest experimental internal reasoning model to win gold at the IMO — and Gemini achieved the same score.

This result demonstrates the potential of Google AI on complex tasks requiring deep logical reasoning.

Thirty days ago, this IMO gold-medal model went live in Gemini ChatBot, with its actual performance considered to exceed even its competition-day level.

Belgian mathematician Michel van Garrel even used it in a live demonstration to show how deep thinking capabilities could be used to prove conjectures.

Overall, high scores on foundation model benchmarks vividly demonstrate model strength to developers and the technical community. Success in competitions like the IMO represents major progress for AI — and Google AI in particular — in frontier reasoning.

Therefore, the launch of the Gemini 2.5 Pro series can be seen as a clear turning point for Google in this AI race.

Gemini is now not only one of the best consumer products, but also a project capable of pushing the technological frontier.

Google began signaling to the community and the market: they were no longer followers, and their foundation models were now officially leading the industry.


Banana in One Hand for Image Generation, Veo 3 Director in the Other

If in pure-text large models Google was "catching up," then in the multimodality domain it leveraged its deep technical accumulation to demonstrate a position of "near-absolute leadership."

While Gemini models were designed from the start as natively multimodal, capable of seamlessly understanding and processing text, code, images, audio, and video, Google also possesses a range of powerful dedicated multimodal models beyond Gemini.

Let's look first at progress in the image domain.

1) The Story of a Banana

On visual reasoning, Google never paused its research from Gemini 1.5 Pro onward. By Gemini 2.5 Pro, its visual reasoning capabilities had already reached excellent levels.

And this became clearly integrated into Google's image models.

Also six months ago, in March 2025, after open-sourcing the well-received Gemma 3 model, Google quickly rolled out Gemini 2.0 "edit images with your voice" — Gemini 2.0 Flash Experimental — which exploded across the internet.

This was mainly because people discovered it could understand natural language input and offered powerful, controllable editing capabilities.

The Crossing team also immediately conducted a comprehensive evaluation of this feature: "Google Is Back! Gemini 2.0 Image Editing Tested: Can Talking Naturally Replace Meitu?"

This feature was wildly popular at the time. A flood of domestic and international reviewers dug into its potential. Precisely because it attracted all kinds of creative uses from netizens across every field, Gemini 2.0 had become one of the coolest toys around.

Interestingly, just before this feature launched, Google's proprietary image generation model Imagen 4 had landed in the industry with something of a whimper rather than the expected bang. Many people assumed that the "edit images with your voice" update was merely a clever but relatively minor product optimization.

But Google's aggressive push in the image domain didn't let up — if anything, it accelerated.

Gemini 2.5 Flash Image (Nano Banana)

The anticipated charge didn't keep people waiting long.

Just one week ago, a mysterious image model codenamed "Nano Banana" appeared on major global AI leaderboards. Its performance on generation and editing tasks quickly sparked heated discussion and speculation across communities.

At the time, one prevailing theory was:

Could this be another Google model?

The guess stemmed from the fact that it practically "crushed" nearly every comparable product on the market. The community broadly believed that only Google, with its deep multimodal expertise, could produce such a "monster" of a model.

Eventually, the mystery was solved: "Nano Banana" was Gemini 2.5 Flash Image.

It demonstrated precise understanding of "object replacement" — not merely "drawing something," but comprehending relationships within images and making edits while maintaining logical consistency. This represented a massive leap in image quality and editing capability.

Of course, the Crossing team immediately published a comprehensive review: "Nano Banana Blows Up! We Stayed Up All Night to Compile 14 Unconventional Use Cases."

One particularly popular example: using the Nano Banana model to fuse 13 input images into a single, cohesive, stylistically consistent image:

Beyond image editing, it showed exceptional location reasoning ability:

In short, the arrival of Gemini 2.5 Flash Image means this: while other manufacturers are still figuring out how to generate a pretty picture, Google has already moved on to making AI understand and reconstruct the real visual world.

This isn't hyperbole, because the capabilities of Google's video generation model Veo 3 confirm exactly this point.

2) Veo 3

Beyond text and image, Google has been equally "impressive" in the video modality.

In dynamic AI video generation, Google used Veo 3 to complete the final — and most crucial — piece of its multimodal puzzle.

Before Veo 3 emerged, all video generation models on the market (including OpenAI's Sora, Google's earlier Veo versions, Runway's Gen-3, and LUMA's Dream Machine) were stunning in effect but universally constrained by three bottlenecks: too short in duration, poor logical consistency, and weak controllability.

What they produced were closer to high-quality "moving image" clips than true "cinematic narrative."

However, in May 2025, Google officially unveiled Veo 3 at its I/O conference, changing the game entirely.

Its biggest technical innovation was achieving high-fidelity synchronized video and audio generation, including dialogue, sound effects, and ambient sound — widely seen as marking the moment AI video generation officially "stepped out of the silent film era."

A remarkably realistic Veo 3-generated talk show clip went viral across the internet at the time, leaving a deep impression:

To this day, despite months having passed since its release, Veo 3 remains virtually unmatched in the industry for long-form video generation, logical coherence, and audio-visual synchronization.

The Hollywood Reporter even published an article stating:

The emergence of Veo 3 marks the evolution of AI video generation technology from an expensive "toy" to a tool that can be integrated into professional production workflows.

Now, advertising agencies are using it to rapidly generate visual prototypes for creative scripts, while independent filmmakers employ it to create fantastical visual effects impossible to achieve through traditional shooting.

A single year of catching up has placed Google shoulder-to-shoulder with — or even ahead of — OpenAI and other top-tier AI foundation model companies in the multimodal direction.

Just last week, the prominent venture firm a16z released its latest report on the top 100 generative AI consumer apps. In this ranking, Gemini's user activity has risen to second place on both web and mobile, trailing only ChatGPT:

Google's "leadership" in multimodal AI isn't reflected merely in a single metric of an individual model, but rather in its comprehensive ability to rapidly productize cutting-edge technology and create disruptive user experiences.

Looking back over the past six months, users have experienced "Aha Moments" through Google's AI foundation models again and again — and that itself is the best possible amplifier.

World Model: Genie 3

If Gemini represents Google's deep cultivation of language and multimodal understanding, then Genie 3 demonstrates its "investment in the future" in generative AI and reality simulation. This is a pure, forward-looking investment.

This is also what all who follow AI and technology should expect from a major tech company:

This is what tech giants should be doing, and Google, more than anyone, should be doing it.

Google DeepMind's "General Purpose World Model" Genie 3 is precisely the product of this expectation.

It can generate explorable, controllable 3D virtual worlds from a single text prompt, supporting 720p resolution, 24 FPS real-time rendering, and maintaining consistency and interactive experience for several minutes.

It has even been called by Chinese and international media: the most advanced world simulator ever created.

Users can move and interact in real time within this dynamically generated world, experiencing virtual environments that remain consistent for minutes on end.

The revolutionary aspect of this technology is that it opens infinite possibilities for training more general-purpose AI agents.

Traditional AI training requires vast amounts of pre-built environments, while Genie 3 can "conjure from thin air" an endless supply of training grounds in wildly varied styles.

This capability will fundamentally transform game development and film production workflows. More importantly, it lays the groundwork for achieving general AI that can understand and adapt to complex physical worlds.

For instance, as we analyzed in our article "Can AI Turn the Tide for China's EV Bloodbath?", general world models will also play an enormous role in autonomous driving training for the automotive industry.

From the debut of Genie 1 in early 2024 with the paper "Genie: Generative Interactive Environments" to today's Genie 3, the outside world has repeatedly marveled at Google's performance in the "multithreaded AI competition."

Many people probably want to shout:

How does Google still have the bandwidth for world models? And how are they doing it this well?

It's fair to say that in the world model domain, Google has once again seized another "flag" on the path to AGI ahead of the competition.

As DeepMind CEO Demis Hassabis described it:

This simulated environment will allow agents to "learn in a virtual mental world, accelerating the path to AGI."

At this point, one can say that a "peak-era AI Google" is approaching.

The Elephant Turns, Technology Monetizes

Behind Google's AI push, of course, lie organizational restructuring and shifts in talent strategy.

Google has actually housed two top-tier technical teams over the past decade:

[1] Google Brain, founded by Jeff Dean, Stanford professor Andrew Ng, and Greg Corrado;

[2] Google DeepMind, a British AI startup acquired by Google in 2014.

The two teams were far from the "Harmony" outsiders might have imagined.

At the end of 2022, OpenAI's release of ChatGPT made it impossible for Google to keep ignoring this dissonance.

Just a few months later, in April 2023, Google announced the merger of the original Google Brain team and the DeepMind team into a new Google DeepMind division, with DeepMind co-founder Demis Hassabis as CEO. Meanwhile, Google Brain's "big shot" leader Jeff Dean was promoted to Google Chief Scientist, focusing on long-term AI research.

At the time, this merger was seen as Google's response to the OpenAI shock, aimed at concentrating its strengths, avoiding internal redundancy, and accelerating the productization of AI research.

1) Google Labs

Promoting technically-minded executives to management roles is hardly new. What's more noteworthy is Google Labs, whose status inside Google has long since surpassed that of a mere internal lab — it's now viewed as the "AI innovation gene pool" driving Google's future.

Google Labs dates back to 2002, once the symbol of Google's engineering culture and "20% time" policy, birthplace of classics like Google Maps and iGoogle. After years of dormancy, it was revived at the 2023 Google I/O conference and quickly began incubating all sorts of "weird and wonderful" AI projects.

Today's Google Labs is no longer just a creative incubator. It's a complete "big-tech native" methodology:

[1] It provides a proving ground for any Google team with a wild idea, encouraging them to build AI projects that seem "outlandish."

[2] It clears the shortest path from prototype concept to product available for public testing, ensuring innovation doesn't stall at the demo stage.

[3] It serves as a "free experimental field" for Google employees.

As we covered in our earlier article What! Google Quietly Launched a Bunch of Practical AI Tools?, which rounded up 8 particularly interesting products, the platform has birthed a series of "small but beautiful" yet highly promising products like NotebookLM and Whisk.

These successes prove that when innovators are given sufficient freedom and resources, their imagination can create tremendous value. And Google is willing to provide such a platform.

So why mention Google Labs?

Because Google has once again placed "innovation" at the top of its agenda.

In April 2025, Sissie Hsiao, the executive previously overseeing Bard and Gemini app integration, stepped down. Her replacement was Josh Woodward, VP of Google Labs.

Woodward's résumé is tightly woven into the spirit of Google Labs. He was one of the driving forces behind NotebookLM, the project that "made waves across tech communities and media platforms from day one."

Putting such a "product geek" and innovation practitioner in charge of Gemini makes Google's intent unmistakable:

Google can no longer be content with merely showcasing its models' technical capabilities. It urgently needs to convert those capabilities into user-perceivable, market-winning super-apps.

In short, even at the highest levels, Google is constantly shuffling its talent through internal competition, placing those who can best execute an "innovation strategy" in critical positions.

2) Technology No Longer Exists Just for Research

DeepMind was once known for academic research, publishing era-defining papers (AlphaGo, Transformer, etc.), while Brain also contributed abundant open-source work (TensorFlow, etc.).

But now Google prioritizes commercial competitiveness. Reports indicate that Google DeepMind has imposed stricter review on research publications to avoid leaking valuable innovations or exposing weaknesses to competitors.

The Transformer architecture, dubbed the "foundation of ChatGPT," saw all eight of its famous authors — the so-called "Transformer Eight" — leave Google in 2023 to start their own companies, shortly after its release.

This might once have been seen as Google "making wedding dresses for others" — a regrettable loss.

But the perspective has shifted: on one hand, it proves Google serves as the "Whampoa Military Academy of AI," nurturing core talent for the entire industry, its technical influence extending far beyond company walls; on the other hand, it has prompted Google to learn from this painful lesson and value "not losing a single key person" more than ever before.

In its talent war with Meta, Google has changed its attitude, going all-out to "prevent attrition."

Reports mention, for instance, that Google DeepMind offered core researchers compensation packages of up to $20 million annually, and shortened equity vesting schedules to three years.

3) An AI-First Company

Organizationally, Google has elevated AI to an unprecedented strategic height.

CEO Sundar Pichai has repeatedly emphasized that Google is an "AI-first" company, and now treats AI as the core of the company's entire future. Google has established various internal AI task forces, tilting resources from search, ads, cloud, and other divisions toward AI.

Its best engineers and largest TPU compute clusters are prioritized for core AI projects like Gemini. Every core product line — from search, ads, and cloud to Android, YouTube, and hardware (Pixel) — must answer one question:

What is your AI strategy?

Then, the old silos began to crumble.

Google Search engineers sat alongside DeepMind engineers to jointly develop Search Generative Experience (SGE). Google Cloud consolidated all its AI capabilities, from AutoML to algorithmic trading, into the unified Vertex AI (Google Cloud AI Platform), offering enterprise customers end-to-end AI solutions. This cross-departmental deep collaboration dramatically improved coordination efficiency, avoiding the previous fragmentation.

As one Bloomberg headline put it, Google DeepMind is transforming from a "research lab" into an "AI product factory."

This shift, in terms of helping Google respond to external competition and integrate internal forces, appears to be working well. After all, even for Google, launching so many AI models and product updates in such a short time would be impossible without strong coordination and execution.

In short, after consolidating every force it could, Google's AI organizational culture has undergone some changes — precisely what we mentioned at the start:

Google has begun monetizing all of its technical accumulation.

What we're seeing now is a new Google that has shed its excess, has clear goals, and executes with startling efficiency.

It's foreseeable that in the next six months to a year, we'll see a Google that is "louder, faster, stronger."

🚥

In our earlier article 2025 Silicon Valley AI War: Mid-Year Review, Lu Zhang, Founding Partner of Fusion Fund, noted a telling detail:

On the surface, OpenAI seized first-mover advantage. But many overlooked that Google is actually the deepest player among large companies — with both vertical research depth and horizontal technical breadth.

So when Google truly converts that depth and breadth into product momentum, its comeback is hardly surprising.

From the criticism that "Google hasn't produced a revolutionary product in five years," to the fear that "in the AI era, Google would be left in the dust by OpenAI, becoming the traditional enterprise everyone talks about," to now advancing on four fronts simultaneously — foundation models, multimodality, world models, and applications.

In less than a year, Google has proven to the world once again: Google is still Google, and it's pouring its long-accumulated strength, without reservation, into products.

This time, it hasn't just returned to the table. It's brought back that long-missed composure of letting technology speak for itself.

If you have firsthand observations or inside information about Google, we sincerely welcome you to share in the comments — we would love to hear more valuable insights.