Charging Into the "AI Startup Battleground": What's on Medeo's Mind? | Chatting with Chenran on AI Video Products and Opportunities for Young People

Find your own "aha moment."

Last week, AI video generation tool Medeo[1] launched. Users input a text description, and it automatically handles shot breakdown, script generation, music addition, and video creation — triggering a minor "viral moment" in our WeChat Moments that same day.

The product comes from ONE2X, founded in early 2024 by Guan Wang, former head of large model products at Moonshot AI, with a focus on AI video. For this podcast episode, we invited Chenran, Medeo's product lead, to share the process of building Medeo, his thoughts on Agent + video, and his reflections on current industry consensus.

In our conversation, Chenran kept bringing up Cursor — besides sharing how it inspired his product thinking, we also discussed his feelings about current AI products and his habit of constantly building demos to experience the latest models or AI applications. What struck us was Chenran saying that in the process of making these demos, he had no particular goal; he just did it for fun.

"Fun" gave him lots of positive reinforcement. As he put it, "fun" is a rare and precious quality in this era. We believe there are many young people like Chenran, and increasingly more of them, and through this podcast we just want to say: enjoy it.

Listen on WeChat:

Listen on Xiaoyuzhou:

👦🏻 Koji

This week's Crossing guest is Chenran. His team is called One2X, and they recently launched an AI video editing product — Medeo. The product received considerable praise after its minor viral debut. I've known Chenran for over a year now, from a PM who hand-built demos to someone who can now independently lead a highly complex AI project. I've seen bamboo-shoot-like growth speed and vitality in him, embodying the "active doer in the AI era" spirit that Crossing advocates. He's a quintessential representative of this philosophy.

In this episode, we'll chat with Chenran about two things: first, how Medeo breaks through in the fiercely competitive AI video track; second, his personal growth — how this generation of young people can seize the opportunities AI brings. Let's start with a rapid-fire Q&A.

Rapid-Fire Q&A with Chenran, Medeo Product Lead

👦🏻 Koji

First, Chenran, how old are you?

👦🏻 Chenran

👦🏻 Koji

Where did you study?

👦🏻 Chenran

Undergrad in computer science at Fudan, grad school in computer science at Cornell.

👦🏻 Koji

What did you do before Medeo?

👦🏻 Chenran

After graduating, I was at a big tech company for less than 9 months before joining One2X. I'd done quite a few full-stack projects before, like the "AI Will" project with The Fair, and also tried some small Agent-direction projects on my own.

👦🏻 Koji

Right, Chenran and I worked together on a project called "AI Will" — we'll get to that later.

👦🏻 Koji

MBTI and zodiac sign?

👦🏻 Chenran

ENFP, Gemini.

👦🏻 Koji

If you had to pitch Medeo in one sentence, what would you say?

👦🏻 Chenran: It's an AI studio that lets both beginners and pro users generate professional videos from a single sentence.

Thoughts on Medeo's Minor Viral Moment After Launch

👦🏻 Koji

Of all the videos generated by Medeo so far, what's the biggest hit you've seen?

👦🏻 Chenran

Nothing I'd call a viral hit, but there's one I really love: someone took a text story I'd posted on Twitter and turned it into a video with Medeo, and @'d me. I found it pretty interesting — I totally didn't expect them to turn my story into a video.

"Medeo Video Example"[2]

👦🏻 Koji

The last time we met was at AI Hacker House, when Flowith was launching their new product Neo, and you guys had just released Medeo. Your servers got overwhelmed and you apparently pulled an all-nighter fixing them. But you still decided to come to AI Hacker House to check out the Agent Neo launch event. So how has this week been in terms of user feedback and overall feeling?

👦🏻 Chenran

Actually, this launch didn't give me a particularly strong sense of reality. Even though it was our official launch, the whole process had been planned for a long time — it felt like completing a milestone task. What pleasantly surprised me was users saying our features and section divisions were very clear, which showed that months of effort weren't wasted.

👦🏻 Koji

How does this feedback and user spread compare to your expectations?

👦🏻 Chenran

We actually didn't set explicit virality goals. This launch was more of a prototype-stage experiment. We chose this timing to keep pace with the rhythm. We didn't expect some official accounts and KOLs to repost it organically — that exceeded expectations, and I felt pretty gratified.

Launch Event + Invite Codes + KOL Hype = AI Product Launch Template?

👦🏻 Koji

These days many AI product launches follow the formula of "launch event + invite codes + KOL hype," but it seems Medeo didn't take this path — you didn't treat it as a launch campaign. What was the thinking there?

👦🏻 Chenran

Mainly two reasons. First, AI tool products inherently need long polishing time, and we haven't been at it that long. I believe tool products need to be refined over years. Medeo is still in early stages — I wouldn't call it a "highly complete" product yet. Second, our team style isn't to chase traffic with gimmicky launches. We'd rather stay heads-down polishing the product, building up user word-of-mouth, rather than prematurely focusing on metrics like DAU.

Why Choose to Build Medeo?

👦🏻 Koji

So why did you choose to build Medeo in the first place?

👦🏻 Chenran

Let me start with how I joined One2X. The team was initially formed by Guan Wang, Yaoxing, and another technical co-founder. Guan was previously product lead at Moonshot AI, focused on model training results. He really wanted to explore the application layer and brought up AI video with me.

My background is pretty cross-disciplinary — I'm a programmer, but I've also worked as a director and in content creation. The AI video direction happened to let me maximize both strengths (content creation and technical development), combining them to build a tool product of my own.

As for why Medeo now — on one hand, video generation is maturing. Though last year we judged that automatic video generation was still constrained by cost and stability, this year we've seen enough new possibilities to indicate market timing was about right. Actually, Medeo's original positioning was video editing, and it still is — we just emphasize end-to-end generation of deliverable videos, while also providing project files for post-production edits. It's not simply "editing" or "generation" — we'll keep exploring both directions.

It Had to Be Video: Modal Elevation Brings Greater Economic Returns

👩🏻 Ronghui

Did you choose this among many directions, or were you set on video from the start?

👦🏻 Chenran

Set on video from the start — it had to be video. We judged that as long as you elevate the modality — converting text or voice information into video form — it brings higher economic returns, lets you get to a financial outcome faster. And people's video-watching habits are only getting stronger; actually, for a WeChat official account article, if there's a video version, many people would rather watch the video. So we had this very intuitive conviction from the beginning.

Once we settled on video, the scope is huge, so we initially chose "auto-editing" as our track — specifically "URL to video": taking an article with images and text and converting it into an auto-edited video.

Later, seeing video generation technology maturing, we decided to make "one-click video generation" our flagship feature. But I actually think these two directions aren't in conflict — we want to explore video production technology across the board. Whether it's generated, edited, or retrieved, as long as viewers want to watch it and it's good video, how it's made is handled by the AI behind Medeo — users don't need to worry about it.

👩🏻 Ronghui

You mentioned "economic returns" as a judgment — can you expand on that? Any specific calculations or assessments?

👦🏻 Chenran

We did run some experiments at the time. For example, converting articles to video — we found these videos did get watched. Even if the original article only had a few hundred reads, the video version might get more plays. We also observed something: video actually has lower information density than articles, but spreads faster and viewers are more willing to watch.

So we've always held this philosophy: take high-information-density text and elevate it once into video. We even built a small MVP to validate that this was valuable. We also ran a news account focused on AI news ourselves. I can't say which account, but it was aimed at general users — you might watch it and think it's pretty silly, but the data feedback was decent.

It's hard to find precise data to back this up — it's more of a gut feeling. Now Crossing is doing video too, and everyone's willing to watch video. My sense is that modality elevation is a "crushing" economic upgrade. Take a piece of text: once you convert it to video, its "unit price" might be ten times higher. Of course, I'm guessing here — can't verify it precisely.

With AI, can articles be easily turned into video? No

👩🏻 Ronghui

Did you run into any challenges when you were testing "article-to-video" early on? Because anyone who's written scripts knows, scripts and articles are actually quite different. Could you share what you learned from the MVP that was particularly valuable?

👦🏻 Chenran

You have a sharp eye. Going through this process — and I have some directing experience myself, though not in short video — I felt this very clearly: a verbatim script and video grammar are completely different. They don't need to be tightly coupled to the article.

Video has its own language system: how to write the hook, how to close, how to deliver emotional value in the middle and keep people watching... These are what video cares about, not the information density that articles emphasize most. Data, tables — in video these are actually "point deductions." The moment you start reading them, viewers bounce. But if you can give them a sense of substance, that's more useful than substance itself. Viewers need to feel "I watched this and gained something," not a pile of cold numbers.

👦🏻 Koji

So what do you do when you encounter an article with tables?

👦🏻 Chenran

The ideal case is when the table itself is already an image — we can just insert it into the video. Or we're experimenting with more visual ways to make tables dynamic, make the presentation more interesting.

No video templates designed for the model yet — no bandwidth, and also a belief in giving the model freedom

👩🏻 Ronghui

So do you need to break it down into different script templates?

👦🏻 Chenran

We haven't gotten that refined yet. We're still pretty rough-and-ready, no templates written for different categories. Though I agree that different types should have different SOPs — we just haven't done it.

👦🏻 Koji

Is this from lack of bandwidth, or intentional? Do you believe the model will handle it itself later, so you don't need to impose templates and workflows now?

👦🏻 Chenran

One, the biggest reason is we genuinely don't have the bandwidth — not enough time to get into the weeds. On the other hand, I do believe that as models upgrade, they'll be able to handle script-writing and video tasks without leaning too heavily on structured guidance.

I think "less structure is key." Starting with heavy templates actually restricts the model's potential. Giving the model freedom — the results might be better. We've felt this deeply ourselves.

Medeo time investment: watching videos, understanding production flow, imitating

👦🏻 Koji

You said you don't have enough bandwidth — where does most of your time go?

👦🏻 Chenran

Although I tell people I'm Medeo's product lead, I'm mainly responsible for optimizing video quality. My core work is researching video directions, predicting future possible play patterns — like AI one-click video generation, AI auto-editing, these technical paths. I think about what "new video categories" are, define how this video category can be made through AI. Then I might manually cut a few demos, then develop the automated version. Another part is prompt engineering, workflow or agent building — this video generation algorithm framework, I basically built it myself.

What I spend the most time on: watching videos. I immerse myself in Douyin, Xiaohongshu, YouTube Shorts, all kinds of video. Only when you've watched enough, sunk into the vibe, can you really get how emotion is transmitted in video. Then I'll manually imitate their editing, analyze how they do it, and finally crystallize it into code. Since I'm a full-stack engineer, once I understand the content I can develop it.

So I'm half content, half code — content is感性, code is理性. I need to structure and deconstruct what it should be; I also need to feel my way through, see how these videos transmit the first point, what the rhythm is like. Both are essential.

👩🏻 Ronghui

You said you watch a lot of short videos, feeling that storytelling, emotion-transmitting way. What types do you mainly watch? How short are the videos? Have you distilled their formulas?

👦🏻 Chenran

I don't watch with particular direction — basically everything. But I do have some observations.

For example, Douyin videos have a very consistent rhythm, because anything that survives in this swipe-up-swipe-down feed basically has similar tonality. I've been thinking: can this tonality be structurally expressed, or actually should it not be expressed structurally — I can't quite feel this yet, but there is indeed a shared emotional transmission. Now I have an intuitive judgment: can this video survive on Douyin — this judgment comes from watching.

Another thing I directionally watch is science popularization or news videos on YouTube Shorts — like hosts talking to whiteboards about the universe, about concepts. They're almost stylistically identical. Though they have their own speaking speeds and accents, since the style is uniform, that means it can be structured. Their videos made me feel: "emotional value is more important than knowledge information." The more I watched, the more I felt video isn't transmitting "knowledge density," it's transmitting "emotional experience." Feeling good after watching — that's what matters.

👦🏻 Koji

Today we wanted to chat with the Medeo team. One2X has two co-founders, Wang Guan and Yaoxing — we both know them, really like them, they're very special founders. But after Medeo's launch, Ronghui and I actually most wanted to invite Chenran. On one hand we're closer, we meet almost every month; on the other, our podcast has invited many founders, but people on the front lines doing product, design, R&D are relatively few. Today we also want to talk with you about frontline实战 feelings. You just mentioned you're both doing content and R&D. I'm curious — how do you define your role in the team? How do you coordinate with others? Because it sounds like you're a one-person team, you can do everything.

👦🏻 Chenran

Indeed, I spend most of my time on pre-research and investigation work, basically completing it alone. The team is my solid backup — they handle engineering, product落地, and other parts. I generally push work to the demo stage, then hand it off. My role is more like "prophet" — I like playing with various AI tools, every new generative video tool that comes out I'll try. Like yesterday Claude 4.0 was released, I watched the livestream, started testing as soon as it came out, and after testing thought "indeed牛逼."

I spend a lot of time on demos, because only after you've run things through, made the first video, ten videos, a hundred videos, can you really understand what this thing is. This process feels like "painting" to me — very creative, very fun.

👦🏻 Koji

Previously Chenran and I collaborated on a project called "AI Will" — we encouraged every young person to spend 10 minutes writing a will guided by AI, thinking about what really matters in life. Chenran not only did R&D, but also handled product and design — that very aesthetically pleasing poster was his. Seeing that poster made me realize: in the AI era, aesthetic sense remains critically important. We all know how to use AI to generate images, but how to write prompts, how to choose what looks good after generation — these ultimately determine the product's feel.

Back to the topic — your team is remote. Wang Guan and Yao Qin are both in Beijing, but you don't meet every day. This organizational form is quite special. Can you talk about your team structure? How do you coordinate with them?

👦🏻 Chenran

We're fully remote. I experienced remote work during my TikTok internship in the United States before, so it's not new to me. But this is my first time experiencing an organization where everyone is remote.

First, meeting efficiency improved. We deliberately control time — daily standups capped at 30 minutes, all-hands resolved within an hour. More time is spent on independent thinking and problem-solving. This demands high initiative, but our team has磨合ed very well. When we have autonomy to allocate our time, everyone can actually focus better on doing things well. We don't set KPIs — the key is hiring the right people, giving them an innovation-friendly environment, and they can make good things. This is also part of One2X's cultural DNA.

As for my coordination with Wang Guan and Yaoxing, it's actually not that much. I spend most of my time coding, tuning models, researching tools — with them it's more strategic communication, like deciding when to do what, directional judgment when seeing new papers. Sometimes a paper lets us validate previous technical assumptions, thereby determining product direction. But execution-level detailed work, I can handle myself.

👦🏻 Koji

You just mentioned you judge technical direction based on new papers. Can you give an example — like recently, what paper came out that made you feel "yes," that the technical direction was validated?

👦🏻 Chenran

We've been paying attention to reinforcement learning last year, believing this path would have major breakthroughs. Until DeepSeek appeared at that critical node — we were especially excited, felt we'd bet on the right direction. Many of our product hypotheses were actually based on this technical assumption.

Yaoxing summarizes recent paper progress at weekly meetings. He uses O3 to help read papers, goes through all the new papers in the AI circle this week, then screens out the most important technical paths. We've also been following Google's "unified multimodal" route — this is too strong. We were betting on this direction last year, and recently Google released new results, which further validated our judgment. Medeo's own technical path is also based on these judgments.

Medeo's vision: let people without training use tools to make content they want to express

👦🏻 Koji

In your view, what are your biggest competitors? What do you think of outsiders saying Medeo is doing what Veed.io, Invideo could already do a year ago? How do you view these voices of homogenization? Or is there some differentiation that people haven't seen, which is why such negative voices appear?

👦🏻 Chenran

That kind of criticism definitely exists. We saw these more established competitors when we chose this direction, but our company's founding mission wasn't to go down that path. The initial product form does look similar. Basic AI video editing features are convergent — you can't escape the core functionality that video expression itself demands. Editors end up looking alike; there are established SOPs for this. Given that premise, what a product should do is find that 3% of innovation. What we care more about is how to provide higher-quality information products in a smarter way. We want the model to express stronger "assertiveness," producing higher-value video with fewer tokens. That means better visual effects and content quality. Our goal isn't to build another CapCut or Invideo, but to build something that lets people who "can't express themselves" do so through video.

Looking similar at this stage is completely normal. Future iterations will diverge significantly.

👦🏻 Koji

You mentioned earlier that while you look similar to them, your missions are different. So in your view, what's your mission? What's theirs?

👦🏻 Chenran

I may not have expressed myself clearly just now. What I meant is: people are trained in language — everyone is familiar with grammatical rules, we've all studied Chinese; but people haven't naturally received training in video grammar. When you use a video editing tool to cut a video, it's not just about learning the tool itself. More importantly, you don't know how to use "video language" to express something. For example, when you shoot a vlog, how do you tell a story, how do you narrate, how do you arrange the editing sequence — these are things untrained people simply don't know.

The problem we want to solve is making this expression process less painful, letting untrained beginners, and even experienced professional users, quickly express their ideas through tools — whether it's news, science popularization, stories, or ads. Actually, we haven't clearly defined which video category Medeo will serve yet; that depends on market feedback and technological direction. One2X believes video will become an important information expression medium, just like text, and we hope Medeo can let everyone fully express themselves through video.

We also believe One2X is helping people express information. The phrase we identify with most is "one piece of information, multiple expressions." Because with AI-powered processing methods, people in the future probably won't care about the intermediate editing and generation process — they'll focus on the result. And this result can have multiple expressions. For example, with the same piece of information, you might want to show it to elderly people, children, young people, or people with different ideologies — the expression would definitely differ. But essentially it's still "multiple expressions of one piece of information." What we want to do is minimize the pain in the middle, from one piece of information to multiple expressions. And video is the first expression medium we've chosen.

👦🏻 Koji

I'm still a bit confused. You said the products look the same but the missions are different, so you're not worried about homogenization. But the mission you described — Veed.io and Invideo might endorse it too. They might have been doing what you're doing now a year ago. So how should we understand this competition?

👦🏻 Chenran

That's a fair challenge. I think video is an extremely broad track, and its language expression system is very complex. For example, the tools needed for editing marketing videos versus podcast videos are completely different. Podcasts may rely more on voice and text information, while marketing videos depend more on emotion, sound effects, transitions, and product showcases — so the tools built for them will also be completely different.

The category we've currently chosen is news, science popularization, knowledge-based, or story-based videos. They're different from short videos with talking heads and subtitles. The development direction of video editing tools depends on which category you want to serve. Once the direction is set, the tools will differentiate. At the beginning everyone looks like a general-purpose editing tool, but we won't build another CapCut. CapCut is like an IDE that can write all kinds of code — it can do anything, and there are vertical scenarios waiting to be explored. But it's too early now; we can hardly solidify which direction we'll definitely go.

👦🏻 Koji

So you haven't solidified your direction yet, which is why the products look similar now. But teams like Veed.io and Invideo have been at it for a long time and haven't moved toward a clear direction either. Is it possible they ultimately won't verticalize, and instead it'll be winner-take-all?

👦🏻 Chenran

I don't really think this video track will be dominated by one tool. You said Veed.io hasn't solidified its direction, but I think it has. Look at all its marketing points — they're basically around "talking-head video plus subtitles," solving the problem of videos with people on camera plus subtitles, and doing it very well. Their stylized subtitles even went viral. I think every team will eventually choose the expression domain they're good at.

Video expression itself is complex enough: you can go to extremes with subtitles, or dig deep into sound effects or generation. But wanting to master everything is impossible, unless you're a major player like CapCut. But even leading video editing tools like CapCut or Premiere Pro may become so bloated that people want a lightweight tool focused on a specific type of editing.

For example, a podcast-specific editing tool that only needs to add subtitles — there's no need for such a heavy system. Vertical tools can deliver an extreme experience instead, like training a model for a specific task and optimizing every detail. Users are willing to pay for this kind of product.

As for Agent, there are also many Agent teams now saying they can do video too. Users will definitely ask what's different about you. I think the difference still exists — for example, when users want to edit subtitles, they still need an editor, a GUI, because there are fundamental details in editing products that you can't get around. Agent can generate content, but it's hard to edit. That's its drawback.

The final products may look similar, but the user paths will be very different, and the production process paths for users to complete a product will also be very different.

"I believe more that the video domain should be 80% workflow and 20% agentic"

👦🏻 Koji

LoveArt, as a Design Agent, can already generate video now. Although it can't edit yet, it might add that later. What do you think about everyone doing AI Agent now? How is your thinking different?

👦🏻 Chenran

We've discussed the pros and cons of Agent and Workflow internally too. When to use Agent? When to use Workflow? At least in Medeo's product, the editing domain itself has SOPs. For example, when directors or editors work, they usually organize footage first, understand the footage, conceptualize the script, then do rough cuts, fine cuts, and finally add subtitles and packaging. This is a process-oriented workflow.

Since human thinking is already process-oriented, I think Agent may not be that suitable in this scenario. I'd rather use another term: Agentic. Agent is more like an entity, a worker. Agentic is a technical approach. If the problem you want to solve is open-ended — you don't know what users will do — then Agentic might be more suitable, because open problems need open solutions. But video creation SOPs are already very mature, so we think this is better solved with Workflow, with Agentic as just an auxiliary technical approach. "We believe more that the video domain is 80% Workflow, 20% Agentic" — only with this ratio can results be delivered stably.

Cautious on MCP: it's a protocol, but doesn't solve real implementation problems

👦🏻 Koji

Have you seen any consensus in AI circles lately that everyone's hyped about, but you particularly disagree with?

👦🏻 Chenran

I'm cautious about MCP. It's essentially a protocol that solves ecosystem and portability problems. It doesn't bring any era-defining technology. For example, interoperability between models and tools doesn't have much to do with technological breakthroughs. At its core, it's still a prompt management method, but prompts can be managed in many other ways. You can decide for yourself how to manage this context, how to handle model outputs, how to structure them, how to save your costs.

The market discussion around MCP is very hot right now, and media coverage is extensive. But when we do vertical scenarios, we don't really need ecosystem expansion that much, so MCP doesn't help us particularly. Of course, if you're building a platform like Coze, or Agent, then using MCP to connect plugins has value. But if you want to build a vertical product, you don't need to prioritize MCP, because it adds engineering complexity without solving your actual technical problems.

Another insight: solving user needs doesn't necessarily require the optimal technology. For example, you don't necessarily need the strongest model. Many AI products now achieve functionality through multi-model mixing and system composition. Each node in the middle may not need the best model. I've even found that sometimes switching from 3.5 Sonnet to 3.7 Sonnet makes results worse. These differences can be quite subtle — different models vary greatly in instruction following and alignment. While the Claude series improved coding capabilities, it might develop strange little bugs on other tasks.

These issues often go undiscovered unless you test them yourself — you won't see them in the news. Only people tuning prompts on the front lines can feel them. And these insights generally don't circulate in the market. They might appear in some corner of some forum, but such information is hard to find secondary verification for. You only know when you see it with your own eyes. So my view is: don't blindly trust so-called "best technology." The technology that solves problems is good technology.

Most-used tool: Cursor

👦🏻 Koji

I'd like to shift to something lighter. I've noticed lately that people use all kinds of models — not just for programming, but daily Q&A model choices are becoming more fragmented too. Everyone's preferences are different. For example, I've mainly been using ChatGPT lately. One reason is it added Memory functionality, another is that a friend works at OpenAI and gave me a free Pro membership. With these two factors combined, I'm increasingly "locked in" to the entire ChatGPT ecosystem — basically chatting with it every day. Chenran, what have you been using most lately?

👦🏻 Chenran

I spend most of my time in Cursor. I basically try every model that can be connected in Cursor. For coding, I most commonly use the Claude series. My favorites are Gemini and Claude — each has its own personality.

I think Gemini's output has better taste, probably because it puts more effort into multimodal training. But its output is relatively "restrained" — it doesn't overcomplicate things. Claude 3.7, on the other hand, I've always said it's like a kid with ADHD. You've already finished a task, and it'll proactively refactor your code, optimize the logic, add a few more files, create a few more functions — always "adding scenes." Especially in Cursor, it can come across as overly eager. But Claude's instruction-following ability is actually worse than Claude 3.5's. I use both models frequently, and with enough use you can really feel their different "personalities."

By comparison, I use OpenAI's models less now — only for image processing. OpenAI feels very "serious," very "mathematical" to me. Its output is extremely structured, often giving me entire tables, which I don't really like. But my friend Yaoxing loves OpenAI's models — he's a heavy GPT user because it outputs substance, no fluff, very orderly — like a standard J-type person on the MBTI. I'm probably more P-type; I need a bit more emotional value.

👦🏻 Koji

Yeah, same here. Some people just see a table and feel happy, thinking this chaotic information has been organized into a clear structure and suddenly the world makes sense.

👦🏻 Chenran

But tables make me dizzy — I hate them. So my preference for OpenAI's models isn't that strong. Grok has been getting a lot of use lately too; I've seen many users on Xiaohongshu trying out its text model. It feels like everyone has their own preferences, especially for everyday Q&A scenarios where the gap between models isn't really that big anymore.

Questions I Often Think About as a PM

👩🏻 Ronghui

So Chenran, as a PM, what have you thought about most, or what do you consider the most important question, in the process of building Medeo?

👦🏻 Chenran

When building Medeo, I've actually been constantly looking at another product I really like — Cursor, which we've mentioned many times. I believe anyone building AI products has been inspired by Cursor. Cursor solved the "three-party relationship" problem. Previously when we built products, we considered the interaction between the traditional tool itself and the user — this was a "two-party, one-side" relationship. The user journey could be fully mapped out before the user even entered the product; it was highly deterministic, with branches and decision trees.

But with AI, the product structure becomes a triangle with three vertices: "human, AI, tool," forming three edges. For example, with our AI editing product, we first need editing software, then consider how "humans" interact with this software — this part is relatively easy. What's truly complex is what AI should "do" in this process, how it intervenes in human creation, and how humans direct AI to work and then revise its output.

Almost all AI products right now are grappling with this three-party relationship. What Cursor does well is that its role division for AI is extremely clear. It sets different functional tiers: from the most basic inline hint, up to edit mode, then chat mode, and finally agent mode. Each tier's capabilities and trigger logic are very clearly defined, and the entire interaction path is explicit. It solves the three-party relationship very well.

The question I think about most with Medeo is exactly this: how do we handle this three-party relationship? The interaction logic between traditional editing software and humans is relatively clear, but what AI should do within this software is hard to define. Because you're essentially embedding what was originally the user's creative process into the tool itself — something traditional products never had to consider. You wouldn't think about the tool producing something before the user even arrives. But with AI products, you often have to "become the creator" first, letting AI produce initial content, then waiting for the user to modify it. Under this model, the product path is no longer deterministic, because AI's output has infinite possibilities — so constructing this three-party relationship is what I find most difficult.

Another question I often think about concerns AI creativity. I've found that AI struggles to generate content beyond human imagination. For example, with scripts or storyboards it writes, after enough testing I can roughly guess what it'll output — genuine "aha moments" are actually rare. That's why I keep following new models, wanting to know if they bring new "aha moments."

But overall, I think AI is better suited to play an "execution role with understanding capabilities" rather than replacing human creativity.

Tasks with SOPs that are repetitive are best suited for AI, such as:

  • Batch writing prompts;
  • Gacha (material recommendations or auto-selecting clips);
  • Optimizing prompts;
  • Writing shot descriptions under existing structures;
  • 24-hour automated execution workflows, etc.

These are things AI excels at: it has both understanding capabilities and strong computational power and time advantages. These capabilities should be maximized.

But if you expect it to deliver stunningly creative ideas, you'll often hit a wall. So one question I frequently ponder is: how do we build SOPs suitable for AI execution, letting it demonstrate advantages within controllable workflows, rather than pushing it toward a creative role it can't fulfill?

How Does an AI Product Lead Spend Their Day?

👩🏻 Ronghui

You just mentioned Cursor and said you really like this product. Let's shift angles: in your daily product work, what products do you personally love and use frequently? Besides Cursor, which products have given you particularly good user experiences? And what does your typical day look like? We're recording on Friday today, so we'd also love to hear if among the new models or products you've recently tried, there's anything you'd especially recommend to everyone.

👦🏻 Chenran

Let me start with how I spend my day. We just talked about what I mainly handle at work. Actually, during my rest time, I invest more energy in "input." For example, watching movies, TV shows, listening to music, researching film trends and cultural movements, reading books, doing "pulling films" (analyzing shot structures, etc.). Recently I've been catching up on screenwriting knowledge, because I never attended film school, but I feel this knowledge is important for my product work, so I'm currently studying screenwriting principles.

I also watch film analysis content while eating, and read before bed. In short, I'm constantly exploring how to become a creator. I think preserving "creative desire" outside of work is very difficult — you all create too, and I remember you said before coming that you don't need to express creative desire right away, it will naturally emerge. But this "creative desire" is really very fragile. I'm still figuring out how to maximize protecting and stimulating it.

As for your other question about recently tried products — I still want to mention Cursor again. Although it launched last year, I still consider it one of the best AI products. It truly achieved "production deployment," solving the interaction problem between AI and humans, clearly defining AI's functional boundaries, while also delivering stable results in actual use. Many AI products today are still at the concept or prototype stage, including our own product which I'd also call a prototype — achieving truly stable delivery is actually very difficult. But Cursor did it. It could do this because it focused on "AI writing code," fully leveraging the model's current strongest capabilities. This is actually counterintuitive: "Stable deployment should be a basic product requirement, but in AI products it's instead the hardest thing to achieve."

AI products are essentially about finding "determinism" within "randomness." And to extract determinism from a high-entropy (uncertain) state requires enormous systematic effort — this is really very difficult. Recently I've also been testing some Agent products, like Manus and LovArt, which have also given me quite a few surprises — I'm sure you've all seen related content.

Stable Delivery Is More of a Product Problem; Currently Hard to Solve Because Engineering Is Difficult

👩🏻 Ronghui

You just mentioned the keyword "stable delivery." I want to follow up: what changes do you think will happen in the short term, say within a year, or in the relatively longer term, two to three years, that could make stable delivery more common?

👦🏻 Chenran

I believe stable delivery is more of a product-level problem to solve, not something that must rely on new technology. So-called stability means that when users perform the same operation, they hope to get roughly the same result. Even if there's randomness, that randomness should be an "adjustable parameter," like the Randomness slider in Midjourney — the higher you slide it, the more random the expected results.

But the problem now is that most products still can't achieve this "controllable randomness." For example, with current Agent products, your first use might be stunning, but you'd hope to reproduce the same experience next time. Yet the possibility of reproduction is extremely low. Even just asking it to make a webpage, next time the tech stack it uses might be completely different.

Users' psychological expectation is "I want stable results," but current products struggle to satisfy this. This isn't because technology is insufficient, but because "engineering" is too difficult. So I don't think we need to wait for the next technological breakthrough — we need to give products time to mature.

👩🏻 Ronghui

Indeed, everyone knows "stability" is important, especially PMs. But as you said, there's always a gap between this and user expectations. What do you think this gap is?

👦🏻 Chenran

I believe all PMs understand that stable paths and predictable experiences are what users need, and ideally they should "exceed expectations." But this is exactly what's hard about AI products: you have to extract "determinism" from "randomness." In traditional product development, you could draw the roadmap before users even arrived. Where they clicked first, what it triggered, where they went next — all could be designed crystal clear. But AI products are different; every step has randomness, and there are many steps. If it's a 100-step AI Q&A process, the final output is completely beyond human control. We certainly want to preserve AI's "random charm," but whether this "randomness" is controllable is extremely difficult.

I often use the Zelda (Breath of the Wild, Tears of the Kingdom) example. They build open-world games not by letting you wander completely unrestricted, but by giving freedom to explore within limited boundaries. They set big goals, like towers, and small reward points, like Korok seeds. During exploration, you're drawn to these points, thus deviating from straight lines, taking detours, forming unique journeys.

AI products are the same — we want each user's path to be unique, but this "randomness" needs to be guided by design, making people feel it's "worthwhile" and "surprising." This requires enormous time and refinement behind the scenes. The industry isn't mature enough; products need more trial and error and accumulation. They need a certain amount of time to incubate a better product.

👩🏻 Ronghui

From another perspective, does this also mean there's huge opportunity?

👦🏻 Chenran

I do think so. But this opportunity can only be "waited for" — either we create it, or someone else does.

Just like the moment The Legend of Zelda launched, everyone saw its "aha moment." Similarly, some AI product in the future might deliver that same aha moment. Every user could have their unique use case, their own open-world experience. AI products are fundamentally an open world — everyone's path is different. But only when that product actually appears will we know "this is it." Everyone's trying in that direction, and so are we.

AI Has Broken Traditional Learning Paths — Everyone Has More Opportunity in This New Era

👩🏻 Ronghui

Have you seen people around you actively seizing this kind of opportunity?

👦🏻 Chenran

Most people around me are AI entrepreneurs. For example, Hyacinth and A Wen, who you've had on your podcast, and "Ni Hao," who won the Xiaohongshu Independent Developer Gold Award. They're all incredibly strong super-individuals who've carved out their own unique paths in both aesthetic sensibility and development capability. They didn't get there through school-taught knowledge — they were driven by interest, studying step by step. This shows how the AI era has broken the constraints of traditional learning paths.

Before, if you wanted to learn frontend, you had to go step by step through HTML, CSS, and JavaScript. But now with AI, you can choose what you want to do, then use AI to learn the relevant skills. Everyone can build their own tech stack based on their interests. I think this is a huge opportunity for young people. Take me — born in 1999, with little market experience and lacking business sense. But that doesn't matter, because we're actually less likely to be constrained by traditional experience.

AI is an entirely new field. Many experienced people are instead limited by old ways of thinking, while young people are bold and willing to experiment — they can create lots of interesting things. When I was at a big company, I deeply felt that: nobody understood it, and I反而 became the one who understood it most. "In this new AI era, everyone truly starts from equal footing."

👩🏻 Ronghui

This reminds me of our previous guest Tong Chao, who said his biggest realization was that in the AI era, no one's understanding is ahead of anyone else's — what matters is "iteration speed."

👦🏻 Chenran

I'll admit I'm young, so of course I'd say that. But iteration speed really is crucial — it's created huge opportunities for me. I wouldn't claim my experience is richer than others', but if everyone's on the same level doing this, having the courage to try making demos is extremely important. I spend enormous amounts of time making demos now. Even outside work, I'm making all kinds of small demos, little toys, experimenting. I don't necessarily need to chase output — I do it to gain "physical intuition." Many demos are just about finding the feeling, because before an AI product is truly made, you have no idea what that "feeling" is. And understanding is built through rapid iteration of that physical intuition, again and again.

For example, when Claude 4.0 dropped yesterday, it sparked new product forms. Or when GPT-4o's image-to-image capability first achieved consistency in image generation, tools like LoveArt suddenly blew up — because they essentially used GPT-4o to ensure that consistency. Only by quickly trying these technologies at the first moment can you build new perceptual understanding at the first moment.

Many Things Are Done Randomly, Just for Fun — and Fun Is a Very Rare Quality

👩🏻 Ronghui

You said earlier that only when a demo is made can you truly experience a product's "feeling." If we limit it to just the past month, roughly how many demos have you made? What concrete gains have they brought you? Or rather, what has that "feeling" internalized into?

👦🏻 Chenran

When I make demos or test tools, I basically have no clear purpose. I'm a typical P-type person — many things happen randomly. I don't plan "I must make a demo this week." Usually it's a sudden flash of inspiration, and I just follow that idea. In this process, I'm enjoying the act of "doing," not necessarily needing a concrete result.

Frankly, most demos I've made look meaningless in retrospect. But this experience suddenly becomes useful at some moment. When you look back and try to positively derive "I did this for future use," you can't actually do it. It's better to just touch whatever interests you. For example, I've been testing Gemini 2.5 Pro's auto-editing capabilities recently. I simply think video is a field AI should take over, and I want to see the upper limits of its capabilities. Not because of work needs, but out of pure passion.

I've always had deep feelings about "fun." The earliest was when I was still studying in the United States, the day ChatGPT launched I started using it. Nobody around me was discussing it, but my eyes immediately lit up. So the NLP and computer vision I'd been studying could be this fun! Before that I'd always felt this major was just about face recognition — pretty boring. But GPT's conversation just struck me. I started staying up late playing with it, debating with it, playing Black Stories, doing emotional companionship — I tried every kind of interaction.

At that time I didn't realize this was a "track" — I just thought it was incredibly fun. Later I realized: if you find something fun, users will probably find it fun too. A product that can truly convey emotion must have an emotionally full, passionately engaged creator behind it. I've always hoped to maintain this emotionally abundant state, randomly doing some emotionally rich exploration. If I can find that "aha moment," I believe users will definitely feel it too.

The Reward of Constantly Making Demos: Feeling the Fun Is Positive Feedback

👩🏻 Ronghui

My question actually wasn't about whether you're purposeful — I wanted to know more about, after doing many things, if you stop to summarize, have these demos subtly become some kind of habit or understanding for you? Like when I started this podcast, it wasn't to improve my expression skills, but at some point this year I suddenly realized: I'm not as nervous when speaking as I used to be.

👦🏻 Chenran

Yes, I understand what you're asking. I think the core reward is still "fun" itself. Because much of my emotional positive feedback does come from making some fun demos. The process of making demos is very much like building with LEGO for me — a very pure form of entertainment.

And now among this group of friends doing AI entrepreneurship, we'll share our demos with each other. For example, Hyacinth started demoInn, and we'll all show off new things we made "out of passion" in the group. There's no particular technical improvement in this process, but it brings a lifestyle, a state of continuous exploration. You enjoy it, and I'm not trying to gain anything from it.

👦🏻 Koji

What you said reminds me of Paul Graham's essay How to Do Great Work. He said to find things that come easily to you, that you enjoy, that others find difficult. Take them seriously, and perhaps you can achieve great rewards. Maybe your demos happen to hit the hottest AI track, but for someone else it might be fishing, birdwatching — as long as it's something you're good at and find easy to do, there will be a sense of accomplishment. Finding such a thing is actually a person's great fortune.

👦🏻 Chenran

Yes, at the end of the day, I think "enjoyment" sounds very "small-self," but it's actually hard to talk about "superheroism" anymore. I deeply feel the downward cycle this generation faces. Jobs are harder and harder to find, many people can't even maintain emotional stability as a luxury, let alone have any grand dreams.

But I've always kept my dreams. Even if the economy is bad, even if everyone is just looking out for themselves, I still believe creation has meaning. What I'm doing with Medeo now is essentially my college dream. I started teaching myself to shoot video and direct in my freshman year, never imagining I'd later combine it with AI, connecting my major and my passion. What I'm doing now is exactly what I most wanted to do.

It's just that in this era, the hardest thing to preserve is actually "creative desire" and "passion." In a depressed environment, maintaining creative enthusiasm is too difficult. When I talk to friends about my dreams, at most they can say "that's nice," but they won't feel it's relevant to them. That's the state now. So I think preserving that passion and creative impulse is crucial. I'm constantly searching for new creative desire myself. Maybe my experience isn't enough yet, but at least I'm on this path, hoping my journey can bring a little inspiration to others.

In College I Thought My Major Was Useless, Then ChatGPT Appeared

👩🏻 Ronghui

I deeply empathize with what you said about "wanting to preserve creative desire." Because the challenges really are quite big now — on one hand, AI can replace many things; on the other, expression on social media is becoming increasingly formulaic, increasingly "standard answer" — those tired clichés you don't even need to think about, you can automatically recite. You just mentioned something I particularly want to follow up on — you said in college you once felt your major was useless. What was the situation then? For example, what was the overall industry atmosphere? And what was your first reaction seeing ChatGPT?

👦🏻 Chenran

I graduated undergrad in 2021. Looking at it overall, that was still the internet industry's peak period — people weren't finding jobs difficult, salaries were pretty good. But looking back now, that was already the apex; we just didn't realize what "economic cycles" meant then, what "the end of an era" meant. My biggest feeling at that time was: everything was too mature. All kinds of applications, all kinds of commercialization paths had already been figured out. There were no particularly major technological breakthroughs. In this established order, what was most valued were experienced veterans — they knew the rules, knew how to optimize. But for us fresh graduates, everything had to start from zero, learning honestly, with little room for creativity. I really felt then that this major — computer science — was just a tool to me, purely for "making money."

Then GPT appeared. I knew at first glance it was completely different. This thing wasn't optimizing existing frameworks — it created an entirely new framework. Its appearance directly pushed the line of "what our major can do" much further out. It could converse, generate text, even imitate human style. I was playing with it, hacking it every day back then. Actually I had no idea what "Prompt Engineer" meant at the time, nor imagined these play patterns would become a "track." I just purely thought: this is so interesting.

The Story Behind Koji's AI Hacker House: "Finding Your Kind"

👦🏻 Koji

Listening to you talk about this resonates so much. You said you were excited seeing ChatGPT then, but people around you were unmoved. I had that experience too. I recently started an AI Hacker House, and everyone's been asking me why. Actually at first I couldn't quite say why either — I just had a strong impulse to do it.

Then one day it suddenly clicked — the best period of my life was around 2010, my "Wudaokou era." I'd just graduated, and all my classmates were heading to places like Microsoft Research Asia and IBM. Nobody cared about Web 2.0. But I was obsessed with Twitter and Facebook — I thought user-generated content was incredible. There was just no one at school who shared that interest, so I was pretty lonely.

But once I got to Wudaokou, I immediately found my people — a group willing to grind from 10 a.m. to 10 p.m., six days a week, pouring all their time into products they genuinely loved. We understood each other, encouraged each other, and broadened each other's horizons.

Looking back now, I'm building the AI Hacker House for young people today — for people like 2010 me, like present-day you — to create a space where you can "find your kind" and feel less alone.

👦🏻 Chenran

Yeah, very similar feeling.

👦🏻 Koji

Alright, another shameless plug.

👩🏻 Ronghui

Chenran, you strike me as so young and full of energy.

👦🏻 Chenran

Do I? I'm always worried, like, so many guests on this show have such substantial things to say. But honestly, it was only last year that I even understood what entrepreneurs do, through listening to podcasts. I'd listen to founder stories every day to "de-brainwash" myself, and now here I am on a podcast myself.

Protecting the Drive to Create: Live Somewhere Quiet, Reduce Socializing, Quit Devices, Read Books

👩🏻 Ronghui

You mentioned "creative drive" earlier — I'm really curious, do you deliberately do anything to protect it? Like making more demos might be one way, but are there others?

👦🏻 Chenran

I've thought about this for a really long time. Probably the thing I've most repeatedly reflected on these past few years. My creative drive was actually strongest in college — making films, writing, running channels. But after graduating, it gradually disappeared. I didn't know why, and I never managed to get it back. I knew I still had the desire to express myself, but I was always missing that bit of energy, that catalyst.

Now I've started using very physical methods to "protect" it. For example, I moved to Pudong, somewhere quieter. It's not that I don't socialize, but I think socializing should be "actively chosen," not "obligatory." I've also started "quitting devices." I have a phone lockbox — I literally lock my phone away. I got so sick of scrolling, to the point where I felt Douyin was toxic, and I didn't want to be held hostage by it anymore. Instead, I started reading. Reading really works — it quiets you down and gives you new inputs.

I'm someone who needs massive input to produce output. Right now is probably a recharging phase. But honestly, I haven't fully found my "creative trigger" yet. I have lots of ideas in my head, but they're all only 30% formed — I can't execute them. When my creative drive was strongest, it was always when my emotions were running high. Especially pain. A few days ago on the high-speed rail, I suddenly got emotional and started crying out of nowhere, cried the whole way. But that kind of crying felt so good — like I'd finally found some emotional channel in myself. Back when I made films, it was also because some period was so painful that I wanted to communicate that pain. For me, that's creation.

👩🏻 Ronghui

I don't think that's strange at all. Look at so many artists, filmmakers, writers — their peak creative periods are rarely their happiest times.

👦🏻 Koji

Well, thank you so much today, Chenran. We talked about Medeo, and about your team's "counter-mainstream" way of working. I found it really inspiring myself. The latter half, about your personal journey — I think that'll give a lot of people strength too. We covered creation, growth, and how you live. I believe this episode won't just be enjoyable to listen to, it'll move a lot of people. Thank you for your time, and we look forward to having you back at Crossing. Bye.

👩🏻 Ronghui

Bye.

👦🏻 Chenran

Bye.

🚥

References

[1] Medeo: https://www.medeo.app/

[2] "Medeo Video Example": https://x.com/huangyun_122/status/1924166263926132884

[3] AI Will: https://pagen.so/page/qnbud4m04yiwqj

[4] How to Do Great Work: https://paulgraham.com/greatwork.html