Beyond Language: What If Agents Had Five Senses | Vital Young

May 14 — come to the fair.

You probably can't live without Agents anymore.

You use them to draft emails, crunch data, and automate entire workflows — yet every interaction still happens inside a text box. They can read, write, and call APIs, but for the longest time they couldn't see the buttons on your screen, couldn't pick up the tone in a phone call, and certainly couldn't twist open the bottle cap sitting on your desk.

This is an intelligent agent without senses. Brilliant, but separated from the real world by a pane of glass.

That glass is starting to crack.

On May 14, the Vital Young Market opens its doors again. Come imagine with us: what happens when Agents gain their "five senses"?

From late 2024 through 2025, Anthropic has steadily iterated on Claude's Computer Use capabilities — taking screenshots, identifying on-screen buttons, moving the cursor, clicking — operating a computer like a human would. OpenAI's Operator and Google's Project Mariner have pursued similar directions. Academic validation has followed in waves: GUI Agent papers have surged over the past year, with task success rates climbing from under 20% to nearly 50%.

Agents can now "see."

And it's not just vision. In 2026, Voice Agents have exploded onto the scene. End-to-end voice latency has been compressed below 300 milliseconds — meaning Agents can now react at near-human speeds. Models like Gemini are processing raw audio streams directly, understanding without first converting speech to text.

After sight and hearing, what comes next? Cross-modal understanding, digital taste modeling, spatial awareness...

These directions share one thing in common: they're all about giving Agents sensory organs.

Humans never worked through text alone. We read blueprints, listen for tone, read expressions, feel space. The information we draw on to make decisions far exceeds what any prompt or command line can carry. An Agent that lives only in text can never align with physical reality. If we want Agents to truly enter real-world scenarios — to collaborate with humans, make decisions on our behalf, even act independently — they must become multimodal, like us.

That's the question this edition of Vital Young wants to explore with you.

When an Agent gains vision, what can it comprehend? When it gains hearing, what can it perceive? When it can manipulate interfaces and enter physical environments, what can it accomplish? With all five senses in place, what scenarios get redefined?

We're partnering with MiniMax, Volcano Engine V-START Accelerator, Elsewhere, Research AI+, and other organizations. No predetermined answers — we just want to find people curious about Agent × multimodal: whether you're a researcher with a paper, a founder looking to demo, or an indie developer exploring the space... Bring your prototype, your hypothesis, or even just your questions, and come to the market!

Scan to register 👇