Why Is AI Creativity Trapped in Text?
Since its launch, ChatGPT has become the fastest product in history to reach 100 million users. More importantly, it has had a profound impact on the world: AI has finally entered everyday life in a way that everyone can understand.

Produced by | AI Nao
Text Is Not the Only Way
Since its launch, ChatGPT has become the fastest product in history to surpass 100 million users. More importantly, it has delivered a profound impact on the world: AI has finally arrived in everyday life in a way everyone can understand.
Yet this phenomenon has also created a narrative trap. Many people have mistakenly come to believe that text is the best output format for AI.
You input text, you get text back. This interaction seems simple and natural, but it compresses all rich contextual information into a linear structure: it cannot display structure, state, or enable multi-step collaboration.

Take a birdwatching enthusiast who wants to understand annual bird migration patterns — but only receives a text summary. Or a user planning a road trip through Yunnan who gets a text-based list of recommendations, when the ideal presentation would be a visual route plan with maps, images, and weather overlays overlaid.
The chatbot interaction model isn't wrong. The problem is that a chat box can only output text.
The human brain is inherently multimodal, capable of receiving visual, auditory, gestural, eye-tracking, tactile, and spatial information. Yet even with advanced AI, all we get is a block of text?
Is there more possibility for what AI can deliver?
Two months ago, Anthropic CEO Dario Amodei said in a public conversation that AI, for all its sophistication, still interacts like it's the 1970s. He didn't mince words: the text-centric interaction paradigm is essentially "industrial-era media inertia."
Kai-Fu Lee expressed similar concerns in an interview. He said humans have mistakenly come to believe AI's only capability is "explaining the world." The breakthrough for next-generation AI products won't happen at the model layer, but rather at the interface layer that "makes models take action."
In short: deliver results, not just textual explanations.
So what are more natural, efficient, and immersive ways to deliver value to users?
The industry has three major trend predictions:
1. AI entry points moving from "text" to "multimodal"
At this year's CHI, NeurIPS, and CVPR, multimodal papers outnumbered unimodal ones for the first time. Companies across Silicon Valley are exploring how to make vision, voice, and environmental understanding the primary — not supplementary — interaction modes for AI. Outputs could be video, audio, or more structured images.
2. From "passive Q&A" to "intent-driven"
Intent-driven interaction has been one of the most discussed topics in the industry during the second half of this year. What does intent-driven mean? Simply put: the user states what they want to do, and AI makes it happen. Throughout the process, the user doesn't need to think about which software to use, which tools to call, or how to complete the task — they only need to express their intent.
In other words, treat AI as a real assistant that gets things done for you.
3. "Interaction" must be actionable
Many engineers AI Nao interviewed this year have raised the same point: as vibe coding costs decline, why should AI only generate text answers for users instead of generating usable tools?
In our October interview with former Amazon scientist Raphael Shu, he said text is the least efficient form of expression: "You wouldn't ask a programming-savvy employee to explain their work in long paragraphs of text — you'd have them write programs, build modules, and run tasks." Former AWS Scientist Teaches Agents to Cooperate, Compete, and Even Argue: On Swarm Intelligence with OpenAgents Founder Raphael Shu
AI should directly generate small applications for us.
From Text to Multimodal
Taken together, these three trends point the industry toward the same direction: AI interaction is evolving from a "text-based Q&A system" toward an "actionable" interface generator.
Ant Group's recently launched Lingguang App represents some innovation in this interaction model.
On the surface, the product still maintains a chatbot form — the most accessible format for the general public — but the output is no longer limited to text alone.
The first type of delivery is "structured content": not just images and text, but card panels, 3D models, multi-step flowcharts, dynamic information structures, and visual analytics.

The second type is interactive, editable, shareable mini-application tools. Specific features include "Flash Apps" — mini-programs generated from a single sentence — and an built-in AGI camera with an "Open Eye" feature that can describe what it sees.
In other words: the essence of the chatbot remains unchanged, but the "delivery method" is being redefined. Where it once provided only text information, it now provides an "executable interface" and "reusable tools."
We believe this represents the industry's latest engineering thinking: rather than treating the product interface as a static container, Lingguang recognizes the interface as a "generable space" for the model.
This interaction directly expands the capability boundaries of the chatbot: from language model to results model, and further to tool generator.
To put it more plainly, Lingguang breaks the strong "chatting sensation" that chatbots give users. This small innovation in interaction makes users realize that AI products can serve as their workbench.
Individual creativity immediately overflows the boundaries of technology.
According to Lingguang's statistics, since launch users have created 3.3 million "Flash Apps" — mostly lifestyle tools: English vocabulary tools made for children, plant-watering reminders, casual decompression mini-games, cyberpunk-style violin metronomes, random food-picking tools... Some users even let their imaginations run wild, creating their own versions of Alipay, WeChat, and DiDi.

These long-tail, fragmented, highly personalized needs are being created by users for the first time — something completely impossible in the mobile internet era.
Of course, the new Lingguang product remains an evolving,阶段性 sample. But it has completed a more crucial step: through interaction innovation, it has made the general public realize that text is not all there is to AI, that AI has richer ways to play, more elegant information texture, and more possibilities.
"Humans" Are Creators of Themselves
There's an interesting phenomenon in technology development: how it lands is never determined by the inventor, but by how users interact with it.
150 years ago, when Edison invented the phonograph, he envisioned it as an "office recording tool" and "academic documentation tool" — somewhat like today's DingTalk. It wasn't until he was sixty that he finally admitted popular music was the phonograph's true purpose. Most young people bought phonographs to listen to music, and the phonograph simultaneously drove the prosperity of the record industry.
The same goes for mobile phones. Originally a communication device, it was only when Steve Jobs "obsessively" packed a camera, television, and music player into a single terminal, using touchscreen as the interaction mode, that the phone became our interface for thinking, decision-making, and receiving information.
The latest industry discussion today is that neither Edison's phonograph nor Jobs's phone is suitable for carrying AI anymore. When AI's capabilities have far surpassed what came before, we should not continue using interaction paradigms inherited from the industrial era: screens, keyboards, notification bars, input boxes.

To put it more extremely, all the interactions we're accustomed to now weren't born for AI — they were born for the internet.
Don't cage AI in.
One entrepreneur told us he believes future AI interaction should be ubiquitous: "In today's era where attention is everything, users shouldn't need to care about how tool calling works in the background. They should directly express intent, and the product assembles a perfectly adapted interface with the right presentation format — nothing extraneous. Whatever's most convenient for the user."
AI could be a phone, glasses, a camera, a webpage, or any new medium. The user expresses intent, AI automatically calls resources and tools, and independently judges how to deliver to you:
We see a flower by the roadside — taking the photo itself represents intent, with results presented as an identification card.
We complain about difficulty losing weight — it should directly generate actionable tools, not a block of text.
Even for something as simple and common as checking travel guides, it shouldn't just be a string of text. Our trigger of interest for a place is often a stunning landscape photo or a brilliant travel video.
Future possibilities also include:
Walking through an unfamiliar city neighborhood, without opening a map, the moment you pause, AI has already pointed you in the right direction;
Scanning a piece of clothing at a mall, still hesitating, AI immediately presents a "3D try-on effect" and what items at home could match it;
In a meeting or while studying, a slight furrow of the brow, and AI immediately generates an easy-to-understand mind map with case explanations.
These predictions and imaginations point to the same logic: AI's value lies not in technical showmanship, but in responding with the most suitable interaction method when a user expresses even a faint intent.

This is precisely the product philosophy Lingguang demonstrates — not rushing to pile on more capabilities, but exploring with restraint, hoping each feature point can deliver greater user value over time.
From this perspective, Lingguang in 2025 is more like a small, beautiful new experiment. The exploration it completes has clear significance: since AI can already understand images, sound, and text, the way we express intent can also be taking a photo or speaking a sentence, and what the product delivers to users goes beyond a single text response to become a form of delivery.
Compress the delivery chain as much as possible, make the delivery result as rich as possible. Expand from single text to interface, structure, and tools.
As interaction methods are repeatedly broadened, human creativity will also emerge in new forms: humans are no longer just questioners — humans can be creators of their own lives.
Design | Youmind


