Google's Back! Gemini 2.0 Image Editing Tested: Can Plain English Take Down Meitu?
When the foundation model gets updated, do you feel anxious? Or excited — as an entrepreneur?
When a foundation model updates, do you, as an entrepreneur, feel anxious? Or excited?
In the Crossing podcast's "2025 New Year Dialogue: The Critical Year for AI, the Dawn of the Agent Era", we asked Yusen, managing partner at ZhenFund, to give one piece of advice to AI entrepreneurs in 2025. He posed this "soul-searching question" above.
After Gemini 2.0 Flash Experimental went live, I'd imagine a large number of AI entrepreneurs felt not just "anxious," but something closer to "despair"...

Question: Why has Gemini 2.0 Flash EXP blown up recently? Why is everyone playing with it? Isn't it just a minor text-to-image model update?
My answer: It really deserves the hype. It understands human speech and removes every cognitive barrier for beginners.
The launch of Gemini 2.0 Flash Experimental has basically brought the Chinese AI internet back to life over the past two weeks.
Its instruction-following capability is simply too powerful.
For example:
- Through natural language input (to put it simply: plain human speech), directly removing watermarks.

Source: X user @abdiisan

Source: X user @tanayj
Thanks to Gemini 2.0 Flash EXP's powerful image parsing capabilities, using simple prompts, you can restore old photos to a certain degree.

Source: X user @literallydenis
First, a quick knowledge drop for all beginners.
What is Gemini 2.0 Flash Experimental?
It's a multimodal AI model (processing text, images, video, and other data types simultaneously) launched by Google.

Just like OpenAI's ChatGPT, Meta's Llama, and xAI's Grok series, it's Google's flagship large AI model.
Why did Gemini 2.0 Flash Experimental go viral?
The main reason: it understands natural language input and offers powerful, controllable editing.
Before this, if I wanted to play around with image editing, it was basically Photoshop, or Stable Diffusion and a bunch of other AI image editing platforms. ByteDance's SeedEdit from last year was also worth trying.
Honestly, none of these platforms truly understood natural language (to put it simply: plain human speech). What this actually reflected was that large models hadn't yet achieved high-quality prompt following. So most likely, I'd still have to go back to Photoshop. But for beginners, even this household-name application still requires a certain amount of time to master.
So the buzz around Google's Gemini 2.0 Flash Experimental this time was, in a sense, predictable.
Late last year, it started rolling out to select testers; a few days ago, it officially opened to developers. But don't worry — ordinary people can use it and have fun with it too.
Precisely because netizens from all walks of life have gathered to innovate across various use cases, Gemini 2.0 has become one of the coolest toys around right now.
Here's the link to try it out (definitely remember to use a VPN — Google's IP detection is still pretty strict):
When I Used Gemini 2.0 Flash Experimental to Raise an Iron Man Cat
After entering the main page, you need to select Gemini 2.0 Flash (Image Generation) Experimental in the model selection bar at the top.

Click the settings button in the top right corner, and remember to set Output format to Image and text.
Don't worry about Temperature for now — this parameter controls the randomness and creativity of generated content. The higher the temperature, the more the model lets loose, increasing randomness and making responses more creative, but reducing prompt adherence.

I happened to have a photo of my cat, so I uploaded it through the + button in the input bar.
Then I simply entered a natural language prompt: "Make this cat face me directly."

Gemini 2.0's performance is shown below. While it excels at instruction following, the image generation quality isn't quite breathtaking.
Take this cat for example — the neck rotation appears stiff and lacks natural fluidity, the fur transitions aren't smooth, the head-to-body proportions are off, and the overall result is visually jarring. Gemini 2.0's text-to-image model still has significant limitations in detail rendering, structural consistency, and photorealism.
If compared against platforms like Kling, Flux, Stable Diffusion, and Midjourney, Gemini 2.0 Flash EXP's generation quality can hardly be called top-tier.

As a Marvel comics fan, I often fantasize: if a cat traveled to the Avengers universe, could it replace Iron Man and beat up the Hulk? So I designed six scenes and used Gemini 2.0 Flash EXP's image generation and memory-following capabilities to create a simple comic strip.
Scene 1:
Prompt: A cat travels to the Avengers universe, becomes Iron Man, and beats up the Hulk.

Scene 2:
Prompt: "Iron Man cat" reveals its mask.

Scene 3:
Prompt: A cute little "Iron Man kitty" takes off its mask, revealing an adorable fluffy little head, and fires lasers at a chibi Hulk.

Scene 4:
Prompt: "Iron Man kitty" takes the Hulk back to the cat planet.

Scene 5:
Prompt: On this mysterious planet, the Hulk has also turned into a cat.

Scene 6:
A cat with Iron Man armor and Hulk cat live together in harmony.

For the final scene, I simply entered the prompt: "Change the background to a Chinese city, classical and elegant."
As an edited image, the core narrative was preserved while incorporating classical Chinese architectural styles.

Beyond continuous scene continuation and expansion, Gemini 2.0 Flash EXP's performance across diverse scenarios demonstrates its significantly enhanced image capabilities. This model has not only improved in image generation and natural language instruction following, but also shows potential across various industries.
I've curated several examples made by X users for showcase.
First up is the image editing domain. Take the user instruction "Turn my selfie into an influencer's Instagram profile pic" — Gemini 2.0 Flash EXP's output shows remarkable results, with the generated image undergoing a distinct stylistic transformation. To a certain extent, it aligns with Instagram's aesthetic. If Google continues rapid iteration on Gemini 2.0 Flash EXP, investing more technical resources into image fidelity and consistency.
Someday in the future, I believe AI text-to-image models might truly wipe out Meitu at the pixel level.
However, it's also worth noting that some elements from the original selfie were correspondingly removed or added.

Source: X user @wongmjane
Next is the e-commerce domain. It's worth mentioning that beyond image content removal operations, Gemini 2.0 Flash EXP has another standout feature: image subject fusion.
In traditional e-commerce, to attract customers and boost purchase intent, merchants have to stock their product showcases with large quantities of real human photos. As a cost, merchants often spend considerable time and effort on model-product matching. If using traditional AI technology, beyond the operational barrier, technical issues like subject distortion mean the artificiality of AI models can easily erode consumer trust in products.
Put simply, faced with a fake person wearing merchandise, traditional AI alone won't make consumers open their wallets.
This release of Gemini 2.0 Flash EXP has, to some degree, eased this merchant-consumer tension.
I found several real-world examples from netizens, offering a glimpse of how this model's revolutionary potential in e-commerce is already emerging.
For instance, upload two images — one of a standing male model, another of the down jacket you want him to try on.
Gemini 2.0 Flash EXP can effortlessly fuse the two while maintaining a certain degree of realism. It must be noted that the current model still has some hallucination issues — even when the prompt explicitly requests "maintain the same pose," the generated model's stance may still deviate.

Source: X user @sardo_adam
Beyond simply wearing items, Gemini 2.0 Flash EXP's performance on "Change the person's gesture, have the model hold a perfume bottle facing the camera" is also commendable.

Source: X user @KurawaDono
Beyond single-item dressing, even with multiple objects undergoing image subject fusion simultaneously, Gemini 2.0 Flash EXP delivers relatively good results.
Some netizens are already sighing: "The number of coffins in the AI startup graveyard is growing exponentially."
Indeed, as mentioned at the beginning of this article. I imagine AI-driven tech startups focused on verticals like e-commerce imagery, text-to-image, and background removal — Photoleap, Pixelcut, Photoroom, and the like — must be panicking from head to toe right now?
The capabilities Gemini 2.0 Flash EXP has demonstrated easily "drown out" and "surpass" the product value they've built.

Source: X user @HalimAlrasihi
Gemini 2.0 Flash EXP's performance in AI image editing also brings new innovation to the comics and anime industries.
Take "coloring black-and-white line art" as an example. The model demonstrates color fill capabilities, accurately identifying the structure and layers of line drafts. Even when the background lacks dramatic cloud imagery, Gemini 2.0 will interpret the prompt and render it reasonably well.

Source: X user @AEAE_94
Without harming the main subject, Gemini 2.0 Flash EXP's performance in style conversion is also noteworthy.
For example, inputting the instruction "Using the same material, convert this image to ukiyo-e style" produces the effect shown below.
When applying this same operation in Midjourney, if you input "Try to change the style of the attached image," the final result may differ significantly from the original material.

Source: X user @kaiju_ya
Because Gemini 2.0 Flash EXP can maintain high consistency across frames during image generation, using it to create local GIF animations from images has also become a viable option.
Source: X user @pandeyparul
While writing this article, there was a moment when a familiar feeling welled up inside me:
From our university days, Google was a "beacon-like" presence — whether through its "Don't be evil" philosophy or its constant stream of new products with peak user experience, it made us feel: this might just be the greatest company in the world!
But in recent years, to be honest, Google has gone somewhat quiet in the AI wave.
It inevitably faces competition from formidable rivals like OpenAI (ChatGPT, GPT-4o), xAI (Grok), and Meta (Llama).
The Gemini large model series, while never fading from public view, can hardly be said to occupy the very center of the AI arena.
This may relate to Google directing more resources toward core products (search, cloud services) or DeepMind's foundational research (like AlphaFold), rather than fully accelerating Gemini's iteration. This has also resulted in GPT-4o, Grok 3, Claude 3.7 Sonnet, and other models consistently dominating product discussions in users' minds.
But from last year's NotebookLM to the viral breakout of Gemini 2.0 Flash EXP's natural language image editing this time, it reveals that while this tech giant may be somewhat conservative in product iteration, it still maintains the technical depth of an industry titan and sharp sensitivity to user needs.
Looking forward to the launch of features like Deep Research and Canvas (supporting intelligent document and code editing), may Google Gemini this "twin star" shine with even more brilliant radiance.
What an electrifying era — giants not yet faded, new stars competing to shine, the glittering of human constellations!

