Agents No Longer Rely on Language Alone to Understand the World | Vital Young
Language is not the world; it is merely humanity's abstract encoding of it.

In 1976, psychologist Harry McGurk published a paper in Nature titled Hearing Lips and Seeing Voices, sharing a counterintuitive experimental finding: when the sound heard and the lip movements seen don't match, the brain doesn't simply choose to trust either the ears or the eyes — it automatically synthesizes a third perceptual result that exists in neither input.
Among test subjects, 98% couldn't resist this behavioral inertia even when they knew the answer. Psychology and neuroscience have since termed this phenomenon the McGurk Effect.

In the fifty years since, neuroscience research has pointed toward a more fundamental explanation: multimodal fusion is the foundational architecture of human cognition. In 2002, Ernst and Banks used a Bayesian inference framework to further reveal the mathematical structure behind the phenomenon: the human brain is essentially performing probabilistic calculations. In the real world, information from any single channel is always incomplete and noisy; when multiple signals most likely originate from the same event, fusing them yields a judgment that is always more accurate than relying on any one alone.
"The more senses involved, the closer the overall solution is to optimal." This path, validated by millions of years of human evolution, is one that AI is now retracing.
In 2024, mainstream technical directions took a simultaneous step forward, teaching Agents to "see" screens — to directly observe interfaces, understand layouts, identify buttons, and determine next actions, connecting perception with action. When an Agent's input expands from text to real visual environments, its relationship with the world shifts: from "humans translating information for AI" to the Agent seeing, judging, and acting on its own.
But vision is only one dimension. When Agents gain more "senses," will the way they understand the world and the actions they take change as a result?

With this question in mind, on May 14 we co-hosted the eighth Vital Young event with MiniMax, Volcano Engine V-START Accelerator, Research AI+, and Elsewhere, inviting peers exploring multimodal Agent directions to share their thinking.

MiniMax has focused on multimodal R&D since its founding. Arthur, head of ecosystem, shared during the open mic that the ultimate goal of multimodality isn't "being able to handle everything" — it's understanding the relationships between modalities during processing, enabling text, images, video, and audio to truly cross-understand within a single model, simultaneously hearing and seeing like a human, grasping the whole. Returning to the Bayesian conclusion from the opening: fusion always outperforms single channels, and what MiniMax is doing is encoding this evolutionary law into its models.

Tian'ge Ling, founder of Gezi Interactive, offered deep insights on the sound modality. A "veteran" Gen Z entrepreneur of nearly four years, he leads his team in developing a new voice Agent engine that transforms sound into a programmable digital asset, covering generation, transformation, and real-time voice changing capabilities. As Agents move from text to voice interaction, sound ceases to be merely an accessory feature — it becomes the primary interface through which Agents establish trust with humans.

Ziteng Wang, founder of Shunji Technology, approached from the visual direction, sharing his research and observations. Traditional cameras capture at fixed frame rates, constantly exposing regardless of whether the world changes; "event cameras" operate on entirely different logic, outputting signals only when light changes occur, achieving microsecond-level response speeds and dynamic ranges thousands of times beyond conventional cameras. When this bio-inspired vision is mounted on Agents, it means they can continuously perceive real environments over extended periods like the human eye, integrating more seamlessly into daily life.
On other perceptual dimensions, Qiguang Wang, founder of Digitalmemos, brought integrated hardware-software devices enabling Agents to simultaneously absorb visual and audio inputs across scenarios, breaking them down into actionable knowledge solutions in real time; ORRIS Scent Player hopes to read physiological signals through wearable devices, using scent to help people regulate their state in real time; the marketing lead from Yimu Technology shared latest progress in visuotactile sensor R&D; and Sheng Yang, founder of Pole Interactive, drew from years of experience in entertainment and culture to explore how to build Agents' understanding of multimodal, human-machine interaction to create new worlds...

Vision, hearing, touch, smell... eight speakers pieced together the current state of multimodal Agents from their respective directions. The 1976 paper revealed that the human brain never relies on a single sense to understand the world; fifty years later, the same rule is being written into Agent engineering architectures, becoming one of the optimal choices for machines to comprehend complex worlds.

Additionally, Xinran, initiator of event partner Research AI+, shared a broader data observation: according to the Leonis AI 100 report from Leonis Capital, among the fastest-growing 100 companies out of over 10,000 AI Native firms, 82 have technical-background CEOs; of 241 founders, 86% come from research or engineering backgrounds.
In the internet era, business models drove change; in the AI era, increasingly more technical people are stepping out of papers and labs, coming to events like this, driving new consensus through collision, disagreement, and collaboration — pushing the frontier forward.
Vital Young doesn't aim for every discussion to yield answers, nor does it shy away from disagreement itself. Come anytime with questions and curiosity — keep coming to the market!






