What startup ideas does his AI experiment give you? | Chatting with Yan Wang: Giving AI ears and eyes, using AI to buy groceries and ship packages

The future belongs to those who know how to use AI best.

The future belongs to those who use AI best.

👦🏻 Podcast interview: Koji, Ronghui

🥷 Edited by: Bella

🧑‍🎨 Layout: NCon

In the tech world, there are curious, hands-on geeks who jump on new products the moment they appear. They don't just use them — sometimes they assemble, hack, and remix them into something entirely new.

Our guest this episode, Yan Wang, is exactly that kind of geek. Recently, we learned he's been running some AI experiments. For instance, he built his own voice input method, recording everything he says 24/7 through his Apple Watch — giving AI a rich stream of voice data to better understand him.

Beyond that, he also straps an Insta360 camera to his chest, capturing video and photos around the clock — during our interview, we joked that Yan's experiments are like giving AI ears and eyes, brilliantly solving the extreme information asymmetry in human-AI communication. And it doesn't stop there: he's also gotten AI to buy groceries and ship packages for him, among other things.

In this episode, Yan shares these experiments and what he's learned from them. And after pushing efficiency to such extremes, his reflections on how humans should actually use AI.

We invited Yan to Crossing because we believe his experience as an early adopter, and all the projects he's hacked together, can offer AI entrepreneurs real inspiration for future products.

Yan was also a guest on our episode Behind the Manus Hype: 20 Questions to Actually Understand AI Agents. His personal blog, Computing Life[1], is worth a read for anyone interested in AI.

Recommended: Yan's course From User to Builder[2], which helps you evolve from an AI tool user into a capable builder — creating practical projects, boosting work efficiency, and strengthening your career competitiveness.

Listen on Xiaoyuzhou:

Listen on WeChat:

This week we also tried recording and editing a video podcast — check it out on Koji's Xiaohongshu/Bilibili/WeChat Channels (video version goes live one week after this article).

Part 1: Quick-Fire Questions — Getting to Know Yan

👦🏻 Koji

How old are you, Yan?

👨🏻 Yan Wang

Starting with the depressing questions, huh? (laughs) I'm already an old man, past the age where I'd be first on the layoff list.

👦🏻 Koji

Where did you study?

👨🏻 Yan Wang

Undergrad at USTC. PhD at Columbia University.

👦🏻 Koji

What do you do now?

👨🏻 Yan Wang

Applied Scientist at Samsara, working on computer vision for dashcams. In my spare time, I keep exploring agent AI.

👦🏻 Koji

Your MBTI and zodiac sign?

👨🏻 Yan Wang

Forgot my MBTI, just remember it was a rare type. Scorpio.

👦🏻 Koji

Many people discovered you through your technical writing. Beyond AI and code, what does your life look like?

👨🏻 Yan Wang

That's a broad question — I do have some unusual experiences. In 2008, I was a Beijing Olympics torchbearer. Later I got licenses for excavators, aircraft, and boats.

These sound cool, but they all share one motivation: I want to personally experience the boundaries of what humans can do.

I didn't get a pilot's license because I'm rich and want a private plane. I can't fly, and I've always been curious — what's it like to fly above the clouds myself, looking at the scenery, taking photos? You can't take anything with you when you die, so why not experience as much of this world as possible?

I'm also an amateur photographer. I shoot astrophotography, microscopy, multi-spectral imaging. I shot enough that Leica noticed and invited me to a solo exhibition in Seattle.

These explorations are fundamentally the same thing:

Using tools to expand the boundaries of human perception — to see farther, smaller, even things invisible to the naked eye.

Photography has become a really important part of my life.

👦🏻 Koji

You strike me as someone who's both geeky and experimental in life, especially with all your interesting AI attempts. I first came across your blog when you started writing about ChatGPT — back when everyone thought it was sometimes genius, sometimes an idiot. Did you feel that way too?

👨🏻 Yan Wang

Oh, absolutely. In my early AI days, I constantly felt it couldn't handle the simplest tasks, totally stupid. But maybe because I'm fairly empathetic, I'd sometimes think from the AI's perspective: if I were it, and my boss asked me to do this, I probably couldn't pull it off either.

It's like hiring a fresh grad from Tsinghua or Peking University — they're smart, but if you don't explain the context, even a textbook-perfect solution won't land. The problem is we didn't clarify the requirements.

Later I found AI errors fall into two broad categories:

  1. It's genuinely not smart enough — like messing up basic arithmetic.
  2. It's actually quite smart, but we're terrible PMs, assigning tasks while constantly changing requirements. Suddenly this won't work, that won't work — the real issue is we never gave enough context from the start.

So looking back, many "artificial idiot" moments weren't the AI failing — it was us failing as its product manager. The blame might be on us.

👦🏻 Koji

You actually empathize with AI.

👨🏻 Yan Wang

This way when AI takes over the world, maybe I'll land a decent position (laughs).

👧🏻 Ronghui

Do you deliberately adjust yourself to become better at using AI?

👨🏻 Yan Wang

I've made quite a few adjustments.

AI used to give irrelevant answers all the time. But I found that once I explained the background properly, its responses improved dramatically. The problem wasn't the AI — it was me not providing enough context.

But typing is exhausting, and I'm lazy (laughs). So later I used GPT-4o to build a real-time voice input system that transcribes, understands, and responds in real time. I rely on it to quickly "dictate prompts" — hundreds of words per minute — and the bandwidth between me and AI suddenly opened up.

The results were clear: AI output quality improved, and I became more willing to hand it complex tasks.

I used to think of AI as my little assistant, but with enough context, it became my big brother (laughs).

Now when I write prompts, I'll tell it project background, lessons from past failures, even my boss and colleagues' preferences.

It can even give you VP-level perspective: "If you pitch it this way, your boss can take it to his boss and brag about it." This kind of senior-level insight was completely inaccessible to me before.

And this advice actually works. My proposals land better with bosses, persuade colleagues more easily, and things move forward more smoothly.

This entire transformation started from something as small as "dictating prompts with voice."

Once you lower the barrier to input, AI's real capabilities can finally be unleashed.

👦🏻 Koji

I recently got access to ChatGPT's new voice input feature — it's incredibly easy to trigger, like "pull to refresh" but "swipe up to start voice input."

— Feels like this is also about reducing typing friction, encouraging more natural and frequent input, thereby providing more context to AI.

👨🏻 Yan Wang

I got access too.

I think this shows two things:

First, ChatGPT is moving toward being a meeting assistant. You can record meetings directly, then have it summarize and answer follow-up questions afterward. Voice is clearly becoming central to its product form.

Second, they had Whisper speech recognition before, but it wasn't as good as GPT-4o's real-time voice. Mainly because Whisper isn't backed by a large language model, while GPT-4o is an LLM — the gap in understanding and feedback capability is substantial.

👧🏻 Ronghui

Typing has such low information density. Even people who love typing struggle to express things completely.

👨🏻 Yan Wang

Especially on phones — it's genuinely painful (laughs).

Part 2: How to Truly Integrate AI Into Daily Life

👦🏻 Koji

So once voice input made communicating with AI and providing context easier, what did you do next?

👨🏻 Yan Wang

After I started talking to AI more frequently, I quickly ran into a major problem: AI has no memory.

Every LLM inference is context-independent unless the product itself "maintains memory" for you, like ChatGPT's memory feature. But either many products don't have this, or ChatGPT doesn't do it well.

That meant I had to repeat context every time — family preferences, project goals — which was incredibly inefficient.

I tried writing a standard context block and copy-pasting it for fixed scenarios, but with constantly changing work content, this got annoying fast. It felt like writing documents, too much friction.

Then I realized: every time we write a prompt, we're essentially telling the AI "who I am." So instead of teaching it from scratch each time, what if I integrated it into my life long-term, so it naturally "remembered"?

That's when I started recording with my Apple Watch.

Since I work from home, privacy concerns were minimal. I use the built-in Voice Memos app, which auto-syncs to iCloud, gets transcribed, and feeds into my database.

The recording system delivered an immediate surprise: once while driving, I zoned out and nearly crashed. Luckily my Apple Watch was recording, so I verbally debriefed right then and set a reminder, then summarized that evening. The voice recognition automatically organized these to-dos. Having the lesson captured in the moment, with follow-up review, genuinely improved my driving.

👦🏻 Koji

So now you can just talk to yourself anytime, thinking out loud — like reminding yourself what to do tonight.

👨🏻 Yan Wang

Exactly. This further reduced the friction of use.

👦🏻 Koji

Do you find yourself becoming more careful, censoring what you say?

👨🏻 Yan Wang

Not really. Since the entire system is self-built, doesn't go through any commercial platform, the data is completely under my control.

But if I were using a third-party system, then yes, I'd probably hold back.

👦🏻 Koji

What other gains have you seen from this system? Anything unexpected?

👨🏻 Yan Wang

Two particularly surprising ones.

The first was that driving incident I mentioned — the recording helped me debrief the whole thing.

The other unexpected gain: voice recognition funneled my daily conversations and thought processes with AI into a concentrated record. The information density was high, and it unexpectedly became a highly efficient way to capture information.

But recordings alone weren't enough. After transcription, they were just piles of TXT files sitting on a hard drive — how to actually use them?

So I built my own "bootleg ChatGPT." It had two features I really needed:

  1. It could connect to multiple models (GPT, Gemini, DeepSeek, Qwen, etc.).
  2. It used an agentic AI approach — it didn't just answer questions, it could call tools, look up information, do searches on its own.

I also hooked it up to a retrieval engine that could access my voice transcription database. If I asked it to find background on a project I'd mentioned last week, it would search it out itself.

And it wasn't traditional RAG (that's a static pipeline), but an agentic workflow that could dynamically decide how to search, whether to switch keywords, how many rounds to run — all autonomously determined by the AI.

The results were genuinely effective, and validated all the recording and organizing work I'd done.

👦🏻 Koji

How long have you kept this recording habit going?

👨🏻 Yan Wang

Audio for over two months. Video just started two weeks ago, already taken over 20,000 photos.

👦🏻 Koji

Wow, tell me about this video system? How are you shooting with the Insta360?

👨🏻 Yan Wang

Right, the Insta360 is tiny, has a magnet on the back. I wear the magnetic necklace under my shirt, and the camera snaps onto my chest — almost unnoticeable. I use vlog mode, shooting 15 seconds every two minutes, battery lasts about 4 hours. But I still find it less than ideal, needs frequent charging. So I built my own small device: microcontroller + Micro SD card + CMOS camera module, compact, deep sleep capable, a small lithium battery lasts 1–3 days. Still in development, but I find it pretty interesting.

👦🏻 Koji

I have to say, Yan Wang's hands-on ability is incredible!

At last month's "AI + Hardware" offline salon by Crossing, there was a startup called Looki that presented something similar: an AI hardware device that magnetically attaches to your chest, takes photos or video every few minutes, records your day, and can even auto-generate vlogs later.

An entire startup's product, and you DIY'd it yourself.

👨🏻 Yan Wang

Sounds similar, but mine really has no technical sophistication — it's a Huaqiangbei approach. All off-the-shelf parts, I just assembled them. Most of the code was written by AI too.

Embedded development used to be painful. Now with O3 (GPT-4o) helping me debug, efficiency is way higher. Paste an error, it immediately tells me how to fix it. O3 has become my big brother (laughs).

👦🏻 Koji

Shooting 15 seconds every two minutes, after recording for a while, what have you discovered?

👨🏻 Yan Wang

Yes, and the process itself has been quite interesting.

I use Qwen 2.5 VL, a local model, to process these video frames, mainly doing three things:

  1. Privacy filtering: For example, automatically detecting and deleting footage when I enter a restroom.
  2. Generating search keywords: So I can quickly find experiences later through keyword search.
  3. Assessing image quality: Only clear, well-composed frames with human faces get kept for the future "memory system."

I plan to turn these images into an image search engine — looking back years later, like a digital In Search of Lost Time.

I also used machine learning to cluster these 20,000 images, selecting the 200 most representative ones, then fed them to Gemini for analysis.

The results were remarkably accurate. It not only spotted health issues (stress, poor posture) but also roughly inferred my profession and interests.

This experiment convinced me that a picture is worth a thousand words. It generated lots of insights, and I plan to continue.

👦🏻 Koji

How did it judge your posture? It's chest-mounted, how could it tell?

👨🏻 Yan Wang

Probably from seeing me hunched over, hand propping up my face — it inferred from those postures.

👧🏻 Ronghui

You asked it to pick "most representative images," but "representative" itself is pretty abstract, right?

👨🏻 Yan Wang

Right, that step wasn't done by AI — I used traditional machine learning. For highly repetitive scenes, like me programming at my computer, similar frames would be merged first, then the most divergent ones selected as representatives.

👧🏻 Ronghui

Which analyses were expected, which surprised you? The health issues one sounds unexpected.

👨🏻 Yan Wang

Expected: it identified objects in my life. For example, I really like a small model shaped like a cinema camera. Gemini not only recognized it but accurately named the brand and model, even inferring my interest in photography. These were details I hoped it would notice, but the precision still surprised me.

👧🏻 Ronghui

Through sound and image, you're helping AI understand you more comprehensively — actively eliminating the information gap between us and AI. What do you think happens when AI has more information about you?

👨🏻 Yan Wang

The information gap between humans and AI is genuinely enormous.

Before starting a new project, our first instinct isn't to write a document or prompt — it's to grab coffee with a colleague and talk through the framework in conversation.

In meetings, when the boss frowns, we immediately know to adjust. But AI has no eyes or ears, can't perceive these subtle signals. If the future becomes more AI-friendly, AI capabilities could amplify tenfold or even hundredfold, with huge impact on research and daily life.

So lately I've been researching one direction: making AI "proactively intervene." Right now we manually trigger AI tools, or say "Hey Siri." But could AI provide real-time feedback? Like if you say "The capital of France is London," it immediately pops up: "It's Paris." Much more useful than correcting after the fact.

The future I envision: everyone wearing XR glasses, AI like an external brain, providing real-time prompts, strategy, even helping you organize your words — not just summarizing, but assisting live. This is a direction I find deeply worth exploring.

👦🏻 Koji

I recently met a team working on "proactive AI" called Proactor AI[3]. Their product listens all day, judges needs and offers suggestions before you even speak. But the challenges are twofold: one, processing massive user data; two, intervening precisely without being intrusive. I believe this proactive form is AI's future.

👨🏻 Yan Wang

Agreed. Will buy on release.

I'm also thinking about another approach: "retrospective invocation."

Currently using AI usually requires pressing a button first, then speaking — like ChatGPT. But if a device is continuously recording, like a camera's "pre-record" function: once you press the button, it sends the previous few seconds of audio to AI, letting it judge what you want. This lowers the barrier to use and can complement the "proactive intervention" model.

👦🏻 Koji

The proactive AI you just mentioned reminds me of what the Proactor team said: sometimes AI needs context that isn't just about the past, but also what you're about to do next. That's why they went with an always-listening, proactively-interrupting service-style AI.

So once you've given AI "eyes" and "ears," how does it actually help you handle daily tasks?

👨🏻 Yan Wang

I used to interact with AI mostly through "verbal sparring" — I'd give commands, it'd offer suggestions, and I'd still be the one executing everything. Then I started testing ChatGPT's Operator feature, letting it actually "get its hands dirty" for me.

I had AI do two things for me, and both worked out well.

The first was online grocery shopping. Before, I'd have to manually search, select items, add to cart — often spending twenty to thirty minutes. Now I just tell AI by voice what I want, and it automatically searches, adds to cart, even checks my order history to know my preferred brands. Sometimes it predicts what I'm running low on and proactively restocks. At the end I just tap an execute button, everything's added, I quickly review and place the order. The whole process is compressed to five minutes.

The second was shipping packages. In the United States, shipping involves tons of manual form-filling, very time-consuming. Now I just tell AI the package details and address, and it operates in the background — selects USPS, skips insurance, and five minutes later tells me it's done. I only need to check out.

Part 3 Thoughts on How People Should Use AI

👦🏻 Koji

Feels like everything you're doing is pushing efficiency to the limit.

👨🏻 Yan Wang

Exactly. Behind this is a concept I've been thinking about: cyber longevity. Not the fantasy of uploading consciousness and living forever, but rather: in the same amount of time, can you get more done?

Take grocery shopping. Going to a physical supermarket might cost only 100 yuan, but burns an hour. Online groceries cost 20 yuan more, but I can get it done in five minutes. Essentially I'm buying back 55 minutes with that 20 yuan.

If you frame it as "20 yuan for 55 minutes of life," everyone would say yes. But say "it's 20 yuan more expensive," and suddenly people don't want to spend. That's what I mean by cyber longevity: not living more years, but making each day more efficient.

👦🏻 Koji

Some people just love wandering through wet markets. For them, that's the joy of living. They'd rather spend their saved time doing exactly that.

👨🏻 Yan Wang

Absolutely, that's a completely valid choice. But if you don't enjoy browsing markets, you could spend that time with family, gaming, or doing nothing at all. In a sense, that's also a form of longevity — or time arbitrage.

👦🏻 Koji

In our last episode, we interviewed Vakee from RockFlow. He said after launching their financial agent Bobby, the original app became almost unnecessary — the future might just be conversational agents. Recently Sam Altman also mentioned that GUIs will likely be replaced. What do you think of this view?

👨🏻 Yan Wang

GUIs were originally created to lower the barrier to using computers, so people who couldn't code could still operate them. But in scenarios like grocery shopping or shipping packages, they've become obstacles to efficiency. Now I use natural language to have AI operate the GUI for me — like sending a proxy — and efficiency is actually higher.

This suggests that GUIs, an interaction paradigm from decades ago, may really be due for a rethink.

👧🏻 Ronghui

Are you seeing any companies exploring alternatives to GUIs?

👨🏻 Yan Wang

Quite a few. The "double tap" gesture Apple Watch recently introduced is one example. If glasses devices become widespread in the future, things like subtle hand movements, wrist flicks, directional changes — these IMU signals could all become new interaction methods. Lots of room, but it's still hard to tell which form will become mainstream.

👦🏻 Koji

I feel like you embrace new tech faster than most people — an early adopter among early adopters.

👨🏻 Yan Wang

Yeah, I'm pretty geeky.

👦🏻 Koji

Have you come across any new products recently that you think might blow up?

👨🏻 Yan Wang

Nothing that's completely blown me away as a full product, but many new features excite me. ChatGPT's weekly updates — Codex, Operator, and so on. I'm mostly using their components, combining them with my own ideas to make a "bootleg ChatGPT." Much of the code is also generated by Cursor or Trae.

I think AI's biggest change is that we no longer have to wait for manufacturers to create the ideal "hit product." Now we can pour our own ideas into existing products and assemble tools that truly fit our needs.

👦🏻 Koji

When you came on our show last time, it was right when Manus launched. It's been three or four months — are you still using Manus or similar agent products? In what scenarios?

👨🏻 Yan Wang

I've been using Manus consistently, especially when I'm out and can't sit at a computer but need to quickly calculate something or do research.

For example, I recently followed a coffee bean auction, Best of Panama. They had a sample box, and I wanted to know if those three sample boxes were worth the price. The most direct calculation would be last year's auction unit price times 100 grams, but manually looking that up and calculating is tedious.

I spent 30 seconds explaining my needs to Manus, it spent ten minutes generating results with intermediate steps. The price was significantly cheaper than this year's sales, so I decided to buy.

These kinds of agent tools are quite practical in daily life.

👦🏻 Koji

Wow, I feel like you pay attention to so many things every day.

👨🏻 Yan Wang

Yeah, this is exactly what I mean by cyber longevity.

Without these tools, you'd either buy blindly or diligently look up information and calculate — half an hour gone. With Manus, it's like gaining an extra half hour of life.

👧🏻 Ronghui

I feel like you have such high energy — maybe because AI is handling so much for you, saving you a lot of mental bandwidth.

👨🏻 Yan Wang

It definitely helps, but it's not all AI. I think high energy comes partly from personality, partly from what you do every day. If you're doing interesting, challenging things rather than mechanical, repetitive work cleaning up after others, you'll naturally be more motivated. Like I mentioned — delegating research to AI, leaving me to just make decisions. That sense of accomplishment and happiness is completely different. What I'm talking about isn't just increasing density, but increasing the "quality" of life. In that sense, it's also extending lifespan.

👧🏻 Ronghui

I feel like the happiest people tend to be two types: people like you who use AI very proficiently, and people who completely don't care about this technology with no information anxiety. It's the people in the middle who feel like there's something new to learn every day, so much they haven't caught up on, constantly anxious and torn.

👨🏻 Yan Wang

True, that's a kind of fate's blessing I suppose.

👧🏻 Ronghui

What's the biggest joy in building these projects yourself? You really seem to enjoy the process.

👨🏻 Yan Wang

The biggest joy is "making it work." Things I couldn't do before, now I can — like learning to fly.

For me, the breakthrough in capability itself is the best reward.

Part 4 When AI Knows You Better: New Ethics in the Human-Technology Relationship

👧🏻 Ronghui

I recently read a novel you wrote — didn't expect someone researching AI to also write fiction.

👨🏻 Yan Wang

Actually AI wrote it; strictly speaking I just wrote the prompts.

👧🏻 Ronghui

There's one scene that stuck with me: someone comes home exhausted from work, their wife is in a bad mood, and AI prompts him how to respond. Is this part of your cyber longevity concept?

👨🏻 Yan Wang

Yes, that's exactly what I wanted to express. But thinking more carefully, it's also a bit absurd: if I need AI prompts to interact with my partner, am I still "me"? Yet the reality is, AI's advice is often more rational, more precise, even more considerate than ours. This creates a lot of conflicts like those in Black Mirror.

👧🏻 Ronghui

Does your novel also express your imagination of future "cyber lifestyles"?

👨🏻 Yan Wang

Yes, if I had to sum it up in one word: "awkward." One story is about someone coming home to find his wife crying from work stress, and AI proposes several comfort options: tell a joke, give him a script to read, or execute a full "emotional repair action package." But each option costs AI budget points. Because he has an important meeting next week, he can only choose the cheapest option, and it doesn't work.

People often have to make difficult, even absurd choices in the face of technology.

I'm optimistic about technology itself, but I also think it will bring social problems we may not be equipped to handle.

👦🏻 Koji

Do you ever feel like current AI understands you better, and helps you more, than any friend or family member? When you think about this, do you feel happy, disappointed, or a bit terrified?

👨🏻 Yan Wang

It's actually pretty terrifying. Once you realize AI really can improve your efficiency and improve your relationships with others, it's easy to become dependent on it, even a bit addicted. But the question is: is that still "me" living? Or is AI living through my body? That boundary can get really blurry.

👦🏻 Koji

Do you ever feel like AI is living through your body?

👨🏻 Yan Wang

Not that extreme — after all, we don't have those AR glasses from the novels yet. But I do consult AI on many decisions. Whether it's where to go tomorrow or how to handle family conflicts, its solutions are often more rational, more mature, and more effective than mine. But in a sense, those decisions are made by it, and I'm just executing. That feeling is pretty terrifying.

👧🏻 Ronghui

Do you think there are many people living a highly AI-integrated lifestyle like yours right now?

👨🏻 Yan Wang

Not many. I don't see many around me.

Partly because there just aren't that many people using AI deeply in the first place. And partly because the field is so new and so fragmented — even among early adopters, people might be on completely different paths with almost no overlap.

👧🏻 Ronghui

I once watched a Bradley Cooper movie called Limitless. He takes this smart drug and becomes energetic, brilliant, and successful, but ends up desperately dependent on it — even wanting to manufacture it himself.

That state reminds me of how using AI feels now. There's something similar about it.

👨🏻 Yan Wang

Yeah. What's even scarier is that my first reaction to your description was: if only I had AR glasses right now that could search up the movie title based on what you just said. That recursive Black Mirror feeling — I really do need that smart drug.

👧🏻 Ronghui

You're an AI-enhanced human (laughs).

👨🏻 Yan Wang

Yes, it's my external brain.

👧🏻 Ronghui

I want to ask something that might be a bit offensive: do you feel like you've lost anything by using AI?

👨🏻 Yan Wang

I haven't really felt like I've lost anything, at least not yet. I record what I do every day, and it's gotten even easier with the Apple Watch's voice memo feature. I also feed these records to AI and ask for its advice. It told me I'm too goal-oriented, that I don't have any genuine downtime. If there's something I've lost, maybe it's that idle, do-nothing leisure. But looking back, even without AI, I'd fill my time with something — learning to fly, tinkering with other stuff. I probably wouldn't be idle anyway.

👦🏻 Koji

Have you ever thought about this: one day you pass away, but AI still remembers everything about you, and others can still "talk" to you through it. Would you want that?

👨🏻 Yan Wang

Why not? I'd already be gone, so what happens after is beyond my control. In a way, that's a form of cyber immortality. What matters is whether you've left any impact on the world, even after you're gone.

👦🏻 Koji

As a parent, what are your thoughts on teaching children to coexist with AI?

👨🏻 Yan Wang

I think this is incredibly important. My child is still young, but I believe that rather than counting or memorizing Tang poetry, what's more crucial is cultivating an intuition from an early age: what can be handed off to AI, and what are core competencies that shouldn't be easily surrendered; how to judge AI's performance, how to be a good "boss" to AI. This capability is more meaningful than knowing hundreds of poems or learning multiplication two years early. I believe that in the future world, getting exposed to and learning to use AI early will be critical.

👦🏻 Koji

How old is he now? How does he interact with AI?

👨🏻 Yan Wang

Just over three. He doesn't interact with AI much — mainly listens to AI tell him stories. We use ChatGPT's real-time voice mode and have tried guiding him to chat with AI, but he's not very interested yet.

But we did something pretty fun with Manus: when telling the Snow White story at bedtime, we'd sneak in little messages like "you should eat your meals properly," and he was surprisingly willing to listen.

Part 5 Future Visions: From Using AI Well to Managing AI Well

👦🏻 Koji

To wrap up the show, we'd love for you to share what changes you've observed in recent months that feel most exciting and might signal massive commercial value.

👨🏻 Yan Wang

I have two observations.

First, AI's pace of evolution hasn't slowed at all. Just like we couldn't have imagined today's ChatGPT two or three years ago, six months ago I wouldn't have predicted that Claude Code and various tools would mature this quickly, or that Facebook would offer a $100 million signing bonus. AI is really sprinting toward its limits.

Second, Agentic AI is becoming mainstream. I've always believed this is the right direction for AI, and I'm glad to see it gaining momentum. The Agentic Workbench I'm building now is in this space, so I'm particularly excited.

👦🏻 Koji

Agentic Workbench?

👨🏻 Yan Wang

I developed a small tool similar to ChatGPT that I call Agentic Workbench.

Its core capability is connecting to my local personal database. I implemented an Agentic Retrieval System — unlike RAG, which follows a fixed workflow, this system activates a search tool only when the AI autonomously determines it needs more background information to complete a task. It constructs its own keywords, retrieves relevant information from my memory database, and uses that to support its analysis and response, or iterates on keywords for further search when necessary.

Through this approach, I've turned information utilization into a low-friction, AI-autonomously-driven process.

👦🏻 Koji

You mentioned having many favorite products and plenty of your own ideas. If you could participate in building them, what would you do differently? Can you give two or three examples?

👨🏻 Yan Wang

My favorite product is ChatGPT — I personally think it's at least two or three body lengths ahead of Gemini and Claude. I'd be thrilled to join OpenAI and add new features to GPT. Conversely, I particularly dislike the Gemini app — I could rant about it for half an hour (laughs) — but if Google asked me to improve it, I'd do that too.

There are also startups like Manus, Cursor, and Trae. I use these products too and have plenty of ideas to contribute.

👧🏻 Ronghui

You mentioned you're also building a product. Can you tell us about it?

👨🏻 Yan Wang

It's basically a "knockoff ChatGPT," but with two improvements: one, making it more transparent so you can see how the AI is thinking and have the chance to intervene — something ChatGPT and Gemini can't do yet. Two, adding lots of tools, like converting YouTube audio to text, connecting to my own database, etc. — overall, more tailored to my own usage habits.

👧🏻 Ronghui

How do you typically allocate your time?

👨🏻 Yan Wang

I actually only work 2–4 hours a day. The rest I hand off to AI. It's very efficient at writing code — I'm still in the top four for code commits at my company. Yes, top four means fourth (laughs).

AI saves me massive amounts of time, letting me freely pursue more things I'm interested in.

👧🏻 Ronghui

If you had to give efficiency advice to an engineer in a similar role, what would you say?

👨🏻 Yan Wang

Learn to use AI.

👧🏻 Ronghui

What if he's already using it?

👨🏻 Yan Wang

Then I'd suggest: don't just see AI as a tool.

AI is something like a person. We should treat it as a subordinate, not a tool.

You need to use it like a good boss — clarify the task, provide sufficient information, check deliverable quality, help it get unstuck when it's blocked.

With this mindset, many common failure patterns can be avoided.

👧🏻 Ronghui

I've had a similar feeling recently. Just yesterday I caught myself saying "ChatGPT told me something" — it felt like I was talking about a person, not a tool.

👨🏻 Yan Wang

Yes, because it's so powerful that we have to entrust it with so much background information. Take driving — it's actually quite complex, but we automatically filter out so many things: traffic police, signals, pedestrians, so driving feels simple. But AI is different. We're asking it to handle more and harder tasks, so naturally we need new ways to collaborate with it.

At this point, management theory becomes useful. We need to manage AI like we manage a person.

👦🏻 Koji

Can you expand on how to apply management principles to working with AI?

👨🏻 Yan Wang

Sure. Simply put, the essence of management is ensuring that when you assign a task, the other party understands what you're saying, knows how to execute, and ultimately delivers results that meet standards. You handle communication, technical risks, outcome verification — the whole process.

But managing AI is very different from managing people. For example, people need motivation, need one-on-ones — AI doesn't need any of that.

What AI really needs is context. You have to maintain its context well for it to perform well. So collaborating with AI is actually a new management skill — similar to managing people, but requiring entirely new methods.

👧🏻 Ronghui

I think context is really crucial. Misunderstandings happen all the time in human communication, and this problem is even more obvious when talking to AI. It doesn't know your background, so it might give irrelevant answers. But if you provide enough information, its responses become excellent.

👨🏻 Yan Wang

Yes. Without sufficient information, AI easily hallucinates. Because it's trained to be helpful to users, but when it encounters a situation with no information yet still needs to help, it has no choice but to make things up. That's where hallucinations come from.

👧🏻 Ronghui

What AI-related information sources do you follow?

👨🏻 Yan Wang

I don't really visit fixed websites. Every week I have ChatGPT do a Deep Research on the AI field, then I read the report it compiles.

👦🏻 Koji

Thank you so much, Yan Wang, for joining us at Crossing. While many are still figuring out how to use AI well, you're already using it as a second brain — even helping you understand the world and understand yourself.

We look forward to continuing to observe AI's evolution with you. Welcome back anytime!

👨🏻 Yan Wang

Sure!

👧🏻 Ronghui

Thank you.

References

[1] Computing Life: https://grapeot.me/

[2] From User to Builder: /2238e8f58ed18087a91acf9fcf252e9e

[3] Proactor AI: http://proactor.ai/