20 Questions to Decode the Year of AI Video | A Conversation with Luma AI Product Manager Barkley
From Sora's Stunning Debut to a Year of Fierce Rivalry
Last week, our "20 Questions" segment launched, and we're incredibly grateful for the support — it gave us the motivation to keep pushing forward! This week, we're continuing with 20 questions to map out the progress in a specific domain: AI video models.
On February 15, 2024, Sora made its stunning debut, capturing the industry's full attention. Video models and video generation applications quickly became the focal point of the AI field. 2024 saw a fiercely competitive landscape: in Silicon Valley, there was Pika, Runway, and Google's DeepMind; domestically, players included Hailuo AI, Keling AI, Vidu, PixVerse, as well as Tencent's Hunyuan and ByteDance's Doubao.
For this episode of Crossing, we've invited Barkley, a product manager at Luma.ai[1], a leading video model startup based in Silicon Valley (and the sole PM on Luma's team, which has raised $160 million).
Through 20 questions, we'll explore the innovations and transformations in AI video models over the past year, and understand the movements of key players. He'll also share his observations as an industry participant on the year since Sora's launch, his analysis of engineering and management capabilities versus algorithmic breakthroughs, and firsthand insights on what people are discussing in Silicon Valley and how the PM role is evolving.
During our conversation, Barkley also told us about meeting Sam Altman at an after party and discussing whether vision is a necessary path to AGI. We hope this content — part observation, part analysis, and part industry gossip — proves helpful.
ps: Barkley was previously Koji's overseas marketing intern at Tangdao. We're all typical "crossover" professionals — from brand marketing and consumer goods to tech, internet, and AI. This interdisciplinary experience perfectly embodies the theme of Crossing: cross-domain thinking often brings unique perspectives and insights, and in this rapidly iterating AI era, such diverse backgrounds are actually a source of differentiated competitiveness.
Listen on WeChat:
Listen on Xiaoyuzhou:


Luma AI's Development and the Ray2 Model
🚥 Koji
For this week's Crossing, we've invited Barkley, product manager at Luma AI.
Luma AI is a globally leading AI video model company. They've raised $160 million, roughly over 1 billion RMB, so their every move draws industry-wide attention, and they've been hailed as one of OpenAI Sora's top rivals.
What surprised us: "A company that's raised nearly 1 billion RMB has only one product manager." We're thrilled to have Barkley, that sole PM, join us at Crossing.
This week is also Crossing's "20 Questions" column. We've prepared 20 heuristic questions for Barkley. We hope to help everyone cut through the information noise and build a clear, systematic understanding of the latest changes and progress in the AI video model industry. First, let's have Barkley introduce himself.
👦🏻 Barkley
Thanks to Koji and Ronghui for the invitation. I've actually been listening to Crossing for many episodes, so it's an honor to come on this podcast and chat about some developments and observations in the AI video space.
My name is Barkley, and my Chinese name is Dai Gaole. I'm the PM at Luma working on the video model layer, mainly responsible for data and model evaluation.
After graduating from undergrad in the US, I worked as a PM at TikTok. I was on TikTok's effects team, where I first got exposed to video generation and image generation. Since then, I've been working on applying CV technology and Diffusion-related tech to AI effects.
I joined Luma in June 2023. At the time, Luma was still a 3D generation company — we were doing 3D reconstruction and 3D generation. Around late 2023, we started pivoting to video. I followed the company as we shifted from 3D-related data and features to video.
I started with evaluation, then gradually took on data and fine-tuning work. That's a brief intro about me.
🚥 Koji
Let's begin our 20 questions.
Our first question — we're recording this right at the one-year anniversary of Sora's release. But it feels like five, eight, even ten years have passed. Yet thinking back to this time last year, when Sora had just dropped, it was the second or third day of Chinese New Year. I woke up in the middle of the night, checked my phone, and my entire WeChat Moments was exploding with shock.
A year after its release, the world feels like it's changing by the day, especially in the AI model space. So my first question for Barkley: from the front lines, do you feel there have been paradigm-level innovations in video models in the AI field since Sora's one-year anniversary?
👦🏻 Barkley
I think it depends on how you define paradigm innovation. At the model layer and architecture level, there hasn't been much fundamental change.
Because Sora's release really validated the DiT architecture for video models — Diffusion Transformer. This architecture replaced the previous pure UNet plus Diffusion-based video architectures, creating a massive leap in video generation quality and model capability.
Virtually all video models since then have followed the DiT route. But on the product and feature side, I think there have been many developments that may not count as major paradigm innovations, but represent gradual iteration. For example, understanding of physical world dynamics, consistency preservation research and its product manifestations, and human figure and motion generation.
It's actually been a process of gradual iteration, with new model architectures continuously improving and technical updates bringing ongoing changes.
So I think while this represents a paradigm innovation, it may not fully qualify as a revolutionary paradigm shift — that kind of fundamental transformation.
🚥 Koji
Second question — people are naturally curious about Luma AI.
Luma AI has raised so much money. What have you mainly been working on these past few months? How much can you share?
👦🏻 Barkley
We just released our next-generation Ray2 model this past month. We initially launched Ray2 with text-to-video, and gradually rolled out image-to-video last week as well. The feedback from the creator community has been very positive.
On one hand, it accurately understands many physical world dynamics — precise physical simulations that previous models couldn't demonstrate, like a ball rolling down a staircase. On the other hand, we've done fine-tuning and data processing in certain vertical domains, such as anime. This allows our model to perform well in these vertical areas, rather than just being a general-purpose model that only excelled at generating realistic video.
On top of this, we do a lot of research work. Because we position ourselves more as a research lab, and this research lab is still research-focused. So we do frontier research on things like real-time video generation and video understanding models.
The ultimate goal is to help video models better understand the laws of our physical world. Even though we feel Ray2 is already quite good, we see that scaling laws remain effective for video models, so we can still push this to the next level.

Evaluation Standards for Video Models
🚥 Koji
I'm actually quite curious — you build your own video models, and you also pay attention to other video models. When a new model comes out, how do you evaluate whether it's actually good?
Because it doesn't seem like language models, where there are many standard benchmarks and correct answers — solving math problems, coding challenges. In the video model space, how do you evaluate whether a model's output is good? What are these benchmarks?
👦🏻 Barkley
I think there genuinely aren't many public benchmarks available right now. For us, we define our own.
Through user interviews and understanding the creator community, we identify metrics that seem reasonable. One such metric might be aesthetics — esthetics, or aesthetic quality. Of course, this is subjective. In these cases, if the video model has an API, we'll batch-generate a set of videos. Then we rely on a global crowdsourced network to evaluate these videos and determine which might be better aesthetically.
Others include adherence to real-world physics. Google actually has a benchmark for this — I can't recall the exact name — but they selected a set of prompts they felt represented how the physical world operates, then tested how different models performed on them. So sometimes we'll also create a dedicated batch of prompts to test, say, how well a model simulates the real world. Beyond that, there's consistency, prompt alignment — how well it follows instructions — these standards are admittedly quite subjective, but they're criteria we've established based on creator research and our understanding of video model use cases.
🚥 Koji
Here's a very straightforward question: who do you think is the strongest in the world right now?
👦🏻 Barkley
Objectively speaking, based on what we've tested, Veo 2 is probably the strongest right now — Google's model.
Of course, I think all models involve some degree of trade-off. For instance, there's often a trade-off between motion and consistency.
If a model produces very large motion, consistency becomes relatively harder to maintain. If its aesthetics are strong, diversity may be harder to preserve — trade-offs like this.
Then there's inference: model size. A model with better results might take longer to run inference. For Veo 2, we feel its output quality on pure generated video clips is probably considered the best in the industry, but its generation time may also be relatively long.

Major Industry Players and Their Positioning
🚥 Koji
Our fifth question: could Barkley help us map out the landscape? We just heard what Luma AI is doing, but everyone also wants to know what the major players are up to.
For example, in Silicon Valley besides Sora, there's Pika, Runway, and Google DeepMind's Veo 2. Domestically there's Hailuo AI, Keling AI, Vidu, Pixverse — actually quite a lot, very competitive. Could you walk us through what each of them is doing?
👦🏻 Barkley
Sure, though I may not be entirely accurate, so don't blame me if I get something wrong — this is just my personal understanding (laughs).
First, the more big-tech-aligned players. We'd already consider OpenAI a big tech company. On the overseas side, that's DeepMind and OpenAI. DeepMind has been continuously pushing their Veo model. And I think DeepMind's scope is fairly broad, combining various multimodal capabilities — for example, they recently hired Tim Brooks from OpenAI's Sora team to work on this so-called world model concept. They may be thinking about how to push Veo 2 to greater extremes while also figuring out multimodal input and output.
OpenAI, after releasing Sora — the actual product — received feedback that felt somewhat underwhelming, like expectations weren't met. But from what I understand, Sora is still iterating on their next-generation model, and combined with OpenAI's existing capabilities in multimodal visual understanding, they're likely aiming toward more of a world model direction, more toward AGI.
Runway is probably more focused on the film and television space — they do a lot of professional editing, collaborate with film studios, and want to achieve the best video generation results in that domain.
Our current positioning leans more prosumer. We don't necessarily want to go directly after high-end film production or partner with these big companies. Instead, we're probably looking more for small-to-medium independent video creators.
Our definition of prosumer: the money they save using our product far exceeds our current subscription price. We believe this creates a very strong retention level — they'll be willing to continue paying for it.
Pika's difference from us may be that they're more focused on the consumer end, because Pika is currently doing a lot of AI effects, creating viral content through video, and much of this viral content targets novice users using AI for entertainment purposes. Pika is using this to break into the consumer market.
That's my understanding of the US players.
As for domestic players, I may have less information. My external impression is that Hailuo AI leans more toward pursuing global growth at scale — I feel they have very large global user numbers, but may not be as focused on profitability, mainly wanting to explore what usage looks like in different countries and regions in more consumer-facing scenarios.
Keling AI, my impression is they're more focused on commercialization metrics. They want to figure out how to improve the model while ensuring the business has positive revenue and growth. And they care more about commercial revenue in key countries and regions, and what their gross margin is per video inference generation.
Pixverse feels more like Pika's positioning in the US — more effects-driven and consumer-facing. Others like Vidu and Tencent's Hunyuan, I actually don't know much about, so I don't really know their specific positioning and direction.
Based on a simple gut feeling, Hunyuan is an open-source model, so I feel they're more about building their own ecosystem, and Vidu may still be in more of a research and prosumer positioning.
🚥 Ronghui
Sixth question: based on your frontline experience. How do people in Silicon Valley talk about domestic video models and applications? Was there a difference before and after DeepSeek?
👦🏻 Barkley
I think it splits into two groups. On the practitioner side, we actually cover domestic video model companies, especially when I'm responsible for model evaluation. So we maintain continuous attention and understanding of domestic video companies' results. I do feel that in video models specifically, many domestic companies are very strong. And over the past year, we've found a general pattern: "whoever released their model last probably has the best results." Because naturally that model has been trained longer, refined longer, seen more data, had more optimization, and accumulated learnings from previous models.
But on the other hand, creators in Silicon Valley — or more broadly in the US — I feel may not have known much about domestic video models before. And they may have been more inclined by habit to use American domestic ones, like Runway, like us, and including Sora which still attracted many artists after launch.
These creators may have previously only seen some information about Keling AI and Hailuo AI on Twitter. And among these more high-end creators, some may have tried them. But I feel that after DeepSeek broke through, there were many Twitter posts — saying while you're paying attention to DeepSeek, also pay attention to these Chinese video model companies, their results are actually quite good — and then there was this organic "water army" promoting Keling AI and Hailuo AI's results. Of course, especially Keling AI during this DeepSeek period also continuously released new model versions, which attracted more attention.
🚥 Koji
Our seventh question: earlier when we discussed various companies, I heard what seemed like two broad directions. One is more user-facing, the other more research-facing.
In your view, how do you understand these different directional choices, and after choosing directions, who chose what, and do work priorities become noticeably different?
👦🏻 Barkley
I think actually these differences aren't obvious at the start, especially at this stage. But I think this depends heavily on the founder's vision and thinking.
For us, we've been firm in saying we want to pursue a larger visual understanding world model. We believe this is an indispensable part of the path to AGI. So in our research we won't focus solely on video generation itself, but will simultaneously do a lot of research on visual understanding models. And we may also work in frontier areas where the probability of success currently looks small, but we feel that if it does succeed, it would be a new breakthrough.
I think this requires significant vision to sustain, plus continuous investment.
Because a defining feature of research is that you never know what it will produce. It's very possible that nine out of ten research ideas fail and turn out to be unworkable.
But if one of them works, scaling it up can produce surprising results — I think Sora's paradigm innovation is exactly that. But this does require a certain cost investment, and a company willing to pursue it long-term. So I think since we define ourselves more as a Research Lab, we will continue to persistently invest in this area. And I think for the others, like DeepMind, like OpenAI, these big players have also been continuously pursuing AGI.
They also believe that multimodality, including video understanding and video generation, is a key piece of the puzzle toward AGI. They continue to study differences between different models and iterate on the technology. Runway previously proposed the concept of world models, and I think they probably also have some research focused in this direction.
But for some more application-layer companies, of course they will continue to iterate their models too, but I think they may focus more on video generation itself — how to apply this video generation, how to fit current consumption scenarios, and what new forms it might create in the future. I don't think these two approaches represent a particularly clear-cut path choice. Because right now we feel that video models are still at a very early stage, not even at the GPT-3 stage of language models, so these path choices aren't especially clear. But I think in the coming years, these differences will gradually become apparent.
Balancing Research Roadmap and Commercialization
🚥 Ronghui
I'd like to add a question here. You just mentioned that your company has a relatively research-oriented positioning and direction.
Could you share how your company views the balance between research investment and commercialization as a company? Because for example, OpenAI was discussed for a long time regarding the difficulty of balancing this issue and its massive early investments.
Do you think the balancing problem faced by video model companies is similar to that of companies like OpenAI focused on text models? Or could there be a completely different development path?
👦🏻 Barkley
Let me answer the first question first — on the choice between investing in research and commercialization, I think we actually lean more toward investing in research today. Commercialization is indeed a relatively important metric for us, but not that important.
We currently rely more on funding to continue next-generation research. But we also ensure that our costs for inference and research remain relatively controllable.
On this point, the VCs in the United States still show a lot of long-term trust — they say invest this money, and even in the latest funding round, there were no explicit requirements for our commercialization data. Instead, they mainly want to see how we achieve in the visual domain, whether it's AGI or this world model definition, what this direction looks like.
So I think on this point, we've always defined ourselves as a Research Lab. And it's also the understanding that Silicon Valley VCs have of us, and the trust they give us.
🚥 Ronghui
Is the strategy for video models different from text models?
👦🏻 Barkley
Overall we believe the scaling law will be the same.
That is, the same development we've seen in the text domain over the past two years will replay in video models.
Everyone keeps scaling up model size until the model possesses basic general capabilities, then possibly developing larger scales than GPT-4's base model, and developing corresponding reasoning capabilities — this is more about understanding the real world and reasoning about its objective laws. This development path won't differ much from text models, because both are based on the Transformer architecture. The core of Transformer lies in continuously refining training data, expecting new capabilities to emerge from the model.
Where video models differ from language models is that video data is orders of magnitude larger and noisier. Not all information in a video is useful, but the model typically receives all of it. Getting the model to understand the relationships and patterns between this information is more challenging than simply scaling up language model data. Therefore, in engineering practice, video models may require completely different training approaches from language models.
🚥 Koji
Our eighth question is, many people online mention that the necessary path to AGI might not be text but vision — what do you think of this?
👦🏻 Barkley
There is a debate in the Silicon Valley AI research community, divided between the language model camp and the world model/visual model camp.
On the language model side, Anthropic (Claude's parent company) firmly believes that as long as language models continue to scale, they can understand world relationships through the human language corpus, so Claude has never developed multimodal models. In contrast, Meta's chief scientist Yann LeCun and Fei-Fei Li believe that humans primarily learn world patterns through vision, and visual feedback is an intuitive process, so visual models are indispensable.
After OpenAI's dev day last year, I ran into Sam Altman and asked him about Sora (not yet released at the time), and whether video generation is a necessary path to AGI. He asked me how I learned world patterns — was it through observation? I answered: "Yes."
Sam Altman said to me: "We can't expect a model that only 'reads books' to learn all the laws of the world, so OpenAI will definitely pursue visual understanding."
This made me feel that although OpenAI may not invest heavily in the Sora direction, they will still increase investment in visual and multimodal research.

Definition and Vision of World Models
🚥 Koji
Since we happened to mention Fei-Fei Li, our ninth question is actually to ask you to explain to everyone — what exactly is Fei-Fei Li's world model?
👦🏻 Barkley
I think different people may have different definitions. So my understanding of world models may come from including Fei-Fei Li's speeches I've seen, and also including some of Yann LeCun's previous public speeches.
But I think world models, in the Silicon Valley understanding, have two parts. One is understanding of the world, that is, 'all the physical laws of the world.' For example, if I'm holding a cup in my hand right now, and when I let go the cup falls, what shape will the cup break into on the ground, the effect of gravity, the effect of ground friction, the effect of different materials — can a visual model understand the physical laws that would actually occur in the world? This is the first level.
The second level is, after understanding, can it simulate future events that haven't happened yet — this is more on the generation side. For example, if I give it a photo of my hand holding a cup, and tell it to simulate my hand releasing the cup and it falling, what will happen — can it understand precisely? So we feel that for world models, the objective laws of the world, physical laws, and understanding and generation of all visual information in accordance with physical laws are two sides of the same coin. When you achieve a world model, it can simultaneously achieve precise understanding and precise generation and simulation of our current physical world.
And applying this to ultimate AGI, if it needs to handle any vision-related task — for example, if we imagine a robot in the future that needs to pick up a cup with its hand and deliver it in front of you for you to drink — then it must simultaneously have the capability to understand and simulate the entire process.
🚥 Ronghui
So what kind of inspiration or impact do you think this brings to your entire field?
👦🏻 Barkley
Regarding the concept of world models, we feel the inspiration and impact is that we won't limit ourselves to just generating video. We feel that all multimodal information should become input and output for this model. So our ultimate goal may be that to achieve this world model, to achieve visual AGI, it would more likely be an anything-to-anything model.
That is, video, images, sound, including various human voices, including sound effects, including common knowledge and know-how about the world. For example, as humans we know how to pick up something broken from the ground — this is also information that the ultimate world model may need to know. When all this information ultimately comes together, it can achieve multimodal input and multimodal output. This is what we feel, when imagining the capabilities needed for the model from the ultimate end goal, this is what we need to do from the research side now.
🚥 Ronghui
Can I understand this as actually raising the difficulty significantly?
👦🏻 Barkley
Yes, I think it also raises by an order of magnitude what's needed, whether in terms of data or in terms of research work. It's not just focusing on video input and output as a single modality.
🚥 Ronghui
Right, because it raises the dimensionality of information tremendously.
👦🏻 Barkley
Yes, and ultimately it may need some form of combination with language models. In fact, many current visual understanding models rely on a base language model as the pathway for understanding such condensed information.
🚥 Ronghui
In the direction that Fei-Fei Li is pursuing, besides them, who else is there currently?
👦🏻 Barkley
I think they're taking a more 3D-oriented expansion, so the path they've chosen is just one of the possible paths.
Because Luma was previously focused on 3D reconstruction and 3D generation. In fact, some of the directions World Labs is pursuing share a lot of similarities with our earlier work. But the reason we ultimately chose video as our path was that we felt through understanding video, through scaling up massive amounts of data, perhaps we don't necessarily need to understand the physical laws of this world through 3D. So I think this might be the difference in approach between us and World Labs — though we're both striving toward world models, we've made different choices on the path.
As for DeepMind, I think their world model also stems more from the video generation domain. Like Genie 2 they released last year, which can simulate various games where you can do 360-degree perspective shifts and see real-time generated scenes — but that too was more based on a video generation path rather than 3D reconstruction.
🚥 Ronghui
You mentioned this, and it made me wonder — did you abandon the 3D route entirely?
👦🏻 Barkley
I wouldn't say abandoned. We feel there are stages to this path that should be approached progressively. We don't think now is the time to scale up or do 3D at large scale.

The Critical Role of Engineering and Management
🚥 Koji
Our tenth question — last time we spoke with Barkley, you mentioned a view that to continue breaking through, a company's engineering and management capabilities likely create more value than algorithmic innovation.
Could you expand on that?
👦🏻 Barkley
I think this applies more to engineering and management around data. Since I'm more involved in data and evaluation, I'm not as familiar with engineering issues on the model side.
But for example with data, we often find that if you have a system that can quickly inject and output data, it dramatically improves model training speed. Because ultimately, following our understanding of scaling laws, the more data a model sees, the broader its understanding and generation capabilities become.
At that point, it's not about some architectural breakthrough in research — it's about how quickly I can get the model to understand this video data. All video might need some compression, but how do I preserve as much information as possible during compression? That's more of an engineering problem than a pure research problem.
And then there's what the data pipeline should look like — this too is more about management, how a company decides to run this whole process from data collection to annotation to finally segmenting usable clips for the model. It feels like an industrial kitchen to me. If data is the ingredient, you need a complete pipeline — one person chopping, one person washing, one person sorting everything into categories. Then finally deciding how to cut these ingredients, what proportions to toss into the wok.
There's really no research innovation here, but it's about doing things more efficiently through engineering and management — efforts that can significantly improve model capabilities.
🚥 Koji
Our eleventh question — let's talk about algorithmic breakthroughs. Have you seen any companies making new meaningful attempts recently?
👦🏻 Barkley
I may not be the most qualified to answer this, since I look at things more from a PM perspective.
For example, with Sora's release last year, everyone saw that DiT was viable at large-scale data scaling. On top of DiT, people made various modifications to the architecture itself, which may ultimately show up in different companies' models.
Beyond that, I think there are some functional points — like how to do video editing, even image editing. In these areas there are some new methods evolving from existing algorithms, which I feel is a pattern where research, the academic community, proposes some interesting hypotheses.
As a company or startup capable of training larger models, we scale up data and explore whether these can find broad application at larger scale, then ultimately decide if this is a meaningful attempt. I think it leans more toward this kind of effort — taking novel ideas, small innovation points, scaling them with data, and finally applying them to products.
🚥 Ronghui
You just spoke about the importance of engineering and management contributions. That made me think — isn't the challenge really that no one's done this before?
👦🏻 Barkley
Yes.
🚥 Ronghui
There's no reference sample. Do you have any takeaways from your personal experience that you feel particularly strongly about, that you could share with peers?
Or based on your understanding of your company and others doing this kind of thing without a reference template — what kind of atmosphere, what kind of incentive structure does a company create to push things forward more efficiently?
👦🏻 Barkley
I think it's mostly about being bold in trying things on these kinds of problems — essentially, brute force miracles (laughs).
We don't have standard answers anyway, so we just try. For example, with evaluation there's no unified standard. For third-party evaluators, they haven't been trained for this either — they don't know how to evaluate. For a video, what does aesthetic quality mean? So we establish different examples and standards. If it contains these elements, we might consider it more aesthetically pleasing, or less so.
We look at whether this ultimately aligns with our expectations. We first try various standards at scale with these annotators, see which ones better match what we ourselves see, including what the community expects.
Including these evaluations, we'll later have many ways to pass references. But a lot of this started from communicating with the creator community — they'd tell us how they felt looking at samples after evaluating them. We'd continuously adjust our standards based on their feedback. But I think a lot of it is just boldly trying, a process of trial and error.
🚥 Ronghui
Sounds like you need to build a train, but there's no train station yet, so you start by building the train station itself (laughs).
👦🏻 Barkley
It's like drawing a horse — whatever you draw, as long as it can run, that's fine. Whether it has the most scientifically accurate body structure, that may not matter at our current stage.
🚥 Ronghui
Do you think your peers are doing similar things? They need to do similar things too, right?
👦🏻 Barkley
I feel like everyone's crossing the river by feeling the stones, and I think in general, this is characteristic of a new field for startups. It might even feed back into our hiring standards.
Our hiring standard has always been: "This is a problem that's never been solved before — how would you approach it?" Our CEO, like us, often asks candidates this question.
🚥 Koji
What's the excellent answer?
👦🏻 Barkley
That depends on what the specific thing is, and what people's approaches might be.

AI Video Applications and Future Trends
🚥 Koji
We've talked quite a bit about the industry and various technical breakthroughs. Let's move on to products, and try to keep it light with some gossip (laughs).
Our twelfth question for Barkley — you're a PM, so you probably pay close attention to various applications. In recent months, have you seen any AI video applications that particularly impressed you or caught your eye?
👦🏻 Barkley
For impressive AI video applications, I can think of two particular examples. Not necessarily specific apps or products, but rather use cases I found impressive.
First was the application that emerged when Luma released its first-generation video model. Initially people just uploaded photos of two people trying to make them hug, which evolved into combining old photos of deceased relatives with modern photos. For example, a grandfather and granddaughter's photos — first arranged side by side, then through the prompt "let them hug," creating a time-spanning embrace: "an old photo and a modern photo, perfectly fused together, hugging." This was something I found very moving and human at the time — the feeling of reconnecting with departed loved ones.
Beyond that, there are some interesting video applications, including transformations of different things that emerged last year. There was a trend at the end of last year called apple dog — a dog holding an apple in its mouth, then you'd see the apple, the dog with the apple suddenly disappear, followed by all sorts of interesting transformed scenes. I thought that was pretty fun too.
🚥 Koji
So for our 13th question, let's make a prediction: in 2025, video models will keep evolving.
What kinds of new entrepreneurial opportunities or application scenarios do you think these innovations and breakthroughs might unlock?
👦🏻 Barkley
For example, we think that in 2025, video models will be able to maintain character consistency — at least for human characters — very well.
Previously, if you wanted to generate a continuous story, you had to put enormous effort into getting the model to learn, or constantly "roll the dice" to get the model to reliably produce scenes featuring the same character.
From what I've seen in research breakthroughs so far, I think this character consistency problem will improve dramatically in 2025.
At that point, you could genuinely use it to easily produce films with coherent narratives, or adapt scenes from text-based novels. A lot of fan-created content, for instance, could become a new video format that spreads online.
Another direction I'm personally quite interested in is real-time video generation. This may not be fully realized in 2025, but if we can reduce video generation latency to very low levels, then it's possible that while watching content, I could modify the video in real time.
For example, I don't like a certain ending of Harry Potter, and I want to see a different possible scenario play out. So while watching Harry Potter, I might have a dialogue with the video model and say, this is the ending I want to see, or in this scene, here's another possibility I'd like to see happen. And the model would respond immediately, generating a different ending.
The application scenarios that real-time video generation could enable — I'm more excited about this becoming a new form of content consumption.
Going forward, the boundary between producers and consumers may blur. Everyone could edit videos, and all video content would be customized toward them.
These are what I think could give rise to new application scenarios and even new entertainment opportunities. Though a lot depends on research progress, so it's hard to say whether it'll happen in 2025.
🚥 Ronghui
Is there anything you think will definitely be realized in the short term, that will happen right away?
👦🏻 Barkley
I think character consistency is something that should be achieved in the short term. Because you can see that many AI companies, including us, have already made some very good progress at the model level.

Juicy Gossip from the Video Industry
🚥 Koji
Speaking of gossip, you know when every company loves to chat about gossip the most — it's when everyone's having lunch together.
I'm curious: lately, when you've been having lunch with colleagues, what have you been talking about? What industry news and developments have come up that left an impression, that you'd want to share with everyone?
👦🏻 Barkley
We actually talk about gossip from other companies, including how we're constantly recruiting AI talent globally.
Sometimes we see their past experiences at previous companies. As a startup, we sometimes gossip about these big tech companies — what their management and AI research is actually like. Because we feel like research at a lot of big tech companies is in a very conflicted state.
Under this kind of multi-layered management, where research doesn't necessarily make the final decisions, but researchers still need to maintain a certain degree of autonomy and independence — you end up seeing some internal political struggles that can happen at big companies.
We treat this as lunch table gossip. Including why, at these big companies, a lot of researchers don't necessarily feel they can do their best work — this is feedback we get from many researchers who've come from places like Google DeepMind, from Meta.
Often, when a non-researcher manager is weighing whether to pursue cutting-edge AI research or keep their team producing consistent output, most managers will choose the latter, because it's the safer path.
But this incentive structure exists fundamentally because AI research is so uncertain. Under big tech evaluation systems, if you don't produce results, it likely means no promotion path and no chance for your team to survive. So I think sometimes these problems actually hinder innovation. We've been chatting about this a lot at lunch recently — found it pretty interesting.
🚥 Ronghui
Has your company discussed DeepSeek at all? And as a Chinese person, you might also be someone your non-Chinese colleagues ask about this.
👦🏻 Barkley
Yeah, I remember after DeepSeek came out, our CEO asked me a question: what's China's innovation and economic environment actually like? Because he was hearing very contradictory information — on one hand, it seemed like most Chinese companies weren't doing large model technology research and were focused on the application layer. But on the other hand, you have impressive companies like DeepSeek emerging.
As a PM here, I still have a lot of exchange with China, so I felt that after DeepSeek came out, it was a bit of a shock to Silicon Valley. That a Chinese company could achieve a breakthrough in pure foundational model technology, with such good results, and then grow so fast at the application layer — I don't think any Chinese company has achieved that in the global market before. So for us, it also means being more focused on recruiting talent from China.
For companies like DeepSeek, a lot of it is homegrown Chinese talent. We think about how to attract that kind of talent to come work with us and create more possibilities for AGI. On the other hand, from what I sense of the atmosphere in China, it's more like a shot in the arm for China's AI sector. If you believe in long-termism, if you truly believe in the vision, it will eventually be realized.
I actually think we can still feel a lot of that atmosphere here in Silicon Valley. For us, it's also a kind of reaffirmation — to keep pursuing AGI in the video domain, to keep scaling up the model, to keep doing foundational research.
🚥 Ronghui
Our next question — you just mentioned that your company takes pursuing AGI as its goal. And Runway's CEO had a fairly well-known article / speech before. He said they no longer see themselves as an AI company.
The whole piece was really emphasizing technology and finding good applications. I think you two are developing in different directions. When we chatted before, you also mentioned that the two CEOs had a spat on Twitter.
👦🏻 Barkley
Yeah. That was Cristóbal, Runway's CEO — a post he pinned to the top of his Twitter.
The gist was that Runway isn't an AI company, Runway is a media entertainment company. He said anyone still calling themselves an AI company, that era is over, wake up. AI will become a foundational thing, like water and electricity. So calling yourself an AI company today is actually meaningless. Because it'll eventually become something everyone uses. So you really need to think about what the application scenarios are.
After he posted this, our CEO retweeted it, quoting: "Anyone who stumbled into AI without truly understanding AI would say something like this." And attached a video of a frog shooting out its tongue that our Ray2 generated.
I don't think either is strictly right or wrong — they're really two sides of the same coin. Though it could also be a matter of timing. From our CEO's perspective, and what our company believes more broadly, AI at its current stage won't become something foundational like water and electricity. That is, frontier AI research itself will bring new paradigms, new application scenarios and breakthroughs.
This is also what we continuously observe in the industry: any improvement in a model can significantly broaden application scenarios.
So we still believe more strongly in continuing to focus on foundational model research, and the application scenarios will come naturally. Though that doesn't mean we don't focus on applications, or don't listen to what our users actually want. But relatively speaking, I think Runway may focus more on the media entertainment industry. Especially with their collaborations with many film studios — they probably want to hear a lot of feedback from these studios about what application scenarios they want, and then make corresponding model improvements. I think that's also a path choice, and it's not necessarily possible to see absolute right or wrong at this stage.
🚥 Ronghui
Different companies have different strategies and choices, so they'll have different viewpoints and ideas. But quite interesting.
👦🏻 Barkley
Yeah, what struck me is there's a saying I often hear: "everything in the bay area happens on Twitter."
All these company CEOs will just directly trash each other on Twitter, very distinctive personalities. I find it one of the most entertaining things about being in Silicon Valley — just eating popcorn and watching.

The Product Manager's Evolving Role in the AI Era
🚥 Koji
Actually, Crossing previously did a very popular piece — the AI-era product manager playbook. We covered a lot of ground in that episode. For instance, how PMs need to redefine themselves, what new skills they need to learn to deliver real value in an AI product.
So from your own hands-on experience — and this is our 16th question — what changes have you seen in the PM role at AI companies? And how did you transition from being a PM for AI effects at TikTok to becoming a model-layer PM at Luma AI? Could you share some stories and insights from that journey?
👦🏻 Barkley
My mindset has shifted dramatically over the past two years, largely from understanding how different the roles are at a model-layer startup versus a product-driven giant.
At ByteDance, I had strong ownership — I defined how effects should work, participated in research discussions, made requests, and the research team would assess feasibility. We'd ship effects on a predictable timeline, even though these projects might involve AI and carry uncertainty.
At Luma, as a model-layer PM, I found that the research lab is primarily researcher-directed, and I'm more in a supporting role.
Initially this positional gap felt uncomfortable. I gradually realized this might actually be a healthier model for research, since research itself is inherently uncertain.
The key difference between the AI era and the internet era: previously, PMs could clearly define requirements, features, target audiences, and metrics because engineers could guarantee functional delivery. Now everything is in a state of chaos — out of ten research ideas, nine might fail and only one succeeds. In this environment, PMs mostly help researchers determine the initial direction to try, rather than demanding that every idea must succeed and ship. That would be unrealistic.
When handling data and model evaluation, I serve as a connector between researchers, end consumers, and creator communities. Evaluation results feed back to researchers, identifying gaps and exploring how data collection and labeling can fill capability holes. But the specific execution and direction ultimately rest with the researchers. I genuinely don't have the ability to direct model iteration — I just try to provide firsthand user information to help them make better decisions.
🚥 Ronghui
I want to add two follow-ups. First, Barkley, you're relatively young — have you observed what senior PMs, those at higher levels, actually do?
Second, you mentioned the nature of your work — could some of this be specific to your company because research carries such outsized importance there? Have you talked to PMs at other companies about whether their work leans in different directions?
👦🏻 Barkley
On the first question — I think this is partly because our company only has one PM — but I believe even for more senior PMs, the entire arc of AI moving from research to product has really only just begun with ChatGPT. The industry is maybe only two and a half years old. All PMs need to re-adapt to this system, learning how to build application scenarios on top of it, or help models iterate better and drive deeper research.
On the second question, I actually talk more with PMs at other model-layer companies. I do feel that on projects like Sora and Veo, my peers are doing similar work — all focused on data, evaluation, and other tasks core to the model, which require user insight and understanding.
But I think PMs at model-layer companies differ significantly from those at application-layer companies. For example, I know that at application-layer companies like ByteDance's Dreamina, PMs are more focused on exploring how to apply models — regardless of whose model they use. They care about finding the best application scenarios, and how to make model capabilities more accessible to users through feature design and interaction patterns. Other application-layer companies search for the best application scenarios, interactions, and implementations for specific models based on their different contexts.
So I do think there's a substantial difference between model-layer PMs and application-layer PMs.
🚥 Ronghui
Question 17: At your company or others you've observed, what stands out as different or special about hiring requirements now compared to before?
👦🏻 Barkley
Based on our current PM hiring needs — or rather, what we look for as a model-layer company — we prefer PMs with prior experience in data or evaluation work at the model layer. This is still a relatively small pool right now.
Even if candidates lack that experience, we want them to be able to ramp up quickly — to figure out a method that hasn't previously been defined as standard, especially in a startup environment. Because no one can guide you; everyone expects you to adapt to the role and start producing immediately.
So I think the ability to quickly establish evaluation standards where no objective standards exist — that's probably what's different about our PM hiring compared to the past. We don't focus much on past specific experience, unless it's directly relevant. What we particularly care about is whether you can get up to speed and get things done.
🚥 Ronghui
We previously spoke with Li Leding, who said that PMs are more important now than ever before.
Also, a lot of what I know about Silicon Valley PMs actually comes from reading one person's newsletter. Since he used to be a PM himself, he focuses heavily on that angle and discusses many things from it. I'm curious — has there always been this much emphasis on PMs, whether in communities or content?
👦🏻 Barkley
I don't think Silicon Valley has emphasized PMs and PM communities to the same degree as China. The PM role really only became prominent with the rise of mobile internet, and mobile internet developed far more vibrantly in China than in the US.
In the US, many companies still lean engineer-driven, and now are gradually becoming more research-driven. Relatively few are fully PM-driven companies. Management approaches like ByteDance's or Tencent's would be considered quite unusual in Silicon Valley.
Regarding PM importance — I think with AI developing so rapidly, it's hard to definitively say what a PM role even is.
But perhaps the most critical thing is that the best PMs can quickly identify the essence of a problem, then find ways to solve it.
People with this ability, whether they become PMs, go into operations, or sales, will likely have strong prospects.
🚥 Ronghui
Question 18: You're in a very fast-moving industry. What do you do to keep learning and stay on top of new developments?
👦🏻 Barkley
I mostly talk with our researchers — they sometimes recommend papers they find interesting. Another important aspect of this industry is trying out various products.
As someone responsible for model evaluation, I use other video model products and our own product frequently. Also, for agent and LLM products, I'll try new ones when I see them. For example, I've been using Windsurf quite a bit lately to write small programs that interest me or could help me work more efficiently.
I think experiencing these products while simultaneously understanding the models that power them, how those models work, and what their boundaries might be — these are two very useful ways of learning for me as a model-layer PM.
🚥 Ronghui
Question 19: Have you observed what learning approaches are effective among people around you?
For example, you just mentioned that everyone's role feels somewhat mixed-up — I actually quite agree, there's this sense that in this era you're forced to learn everything. Is it like that around you too?
For instance, your colleagues or friends — are they in similar situations? And what valuable methods have you seen, like the approach you mentioned of regularly chatting with researchers and reading papers they recommend? I think that's a valuable way to learn — finding a high-quality information source and getting high-value information through their recommendations.
Have you seen other particularly effective, valuable methods?
👦🏻 Barkley
I had a friend who also worked as a PM at TikTok, and his approach to organizing AI-related papers and application information was excellent. He uses online diagramming software to map out all the products he's tried and papers he's read, attempting to find connections between them and build a large mind map.
Much of this also reflects what I've felt talking with researchers at my company. Their research process is essentially about finding connections between different methods and different models. For example, our researchers sometimes read language model papers, discover that certain approaches from language models might inspire us, and then try applying them to video models to see if the method works.
So I particularly admire my friend's approach — his ability to build connections between these papers and products, finding invariant themes that might be similar across different domains. These ultimately inspire what new products we might create, or what new application scenarios might emerge. I think it's a pretty excellent method.
🚥 Ronghui
Drawing connections across domains.
👦🏻 Barkley
Yes, AI — especially the Transformer architecture — really does give you this sense that there's actually some connection between everything in the world.
The way we use human brains to arrange, combine, and process these connections is incredibly inefficient. AI can discover hidden connections between things in massive amounts of data, giving rise to more powerful intelligence.
🚥 Koji
Our 20th question: Barkley, you're a PM in Silicon Valley.
We'd love to hear your perspective as a Chinese person in Silicon Valley — do you feel the AI era has brought different, perhaps better, or possibly worse, new career opportunities? What advice would you give?
How should people seize these opportunities?
👦🏻 Barkley
I think Chinese teams in China and Chinese-founded teams in the US do have some unique advantages.
For example, our understanding of both China and the US, including the tech markets in both countries. On the consumer side, there are very few product managers in the US who truly understand consumer psychology. The last consumer product that really took off in America was probably Snapchat, and after that it was TikTok — which was created by a Chinese team.
In terms of consumer insights and understanding of AI hardware, Chinese entrepreneurs have many unique advantages in these areas. Many Chinese products that went global have found success in the US, precisely because of our grasp of both markets and our command of the consumer ecosystem. I think there will be plenty of opportunities here, and we may see more teams founded by Chinese entrepreneurs emerging at the application layer in the future.
On the other hand, at the model layer, China's research capabilities are genuinely strong. This spirit of deep study and perseverance is a traditional virtue of the Chinese people. Here in Silicon Valley, you can also see that core researchers at many outstanding AI companies are actually Chinese — with different educational backgrounds, some having done PhDs in China, others in the US. These opportunities for Chinese people will continue to exist.
While geopolitics will have some impact, I believe more strongly that AI development should ultimately flow more freely globally, improving through a state that's somewhere between cooperation and competition. We'll also learn a lot from Chinese models and products, and think about how we can make improvements.
Finally, I'd like to mention that Luma AI is currently recruiting visual talent worldwide to join us in researching visual understanding and visual generation, working toward our vision of achieving AGI through the visual domain. We especially hope to recruit more Chinese talent. You can start working remotely, and we can also help with US work visas before joining our Bay Area office. If you're interested, feel free to reach out to me or apply through our careers page.
🚥 Koji
Alright, thank you Barkley. For those who want to get in touch with Barkley, you can check the comments section of our podcast. After we publish this, we'll have Barkley drop his contact info there. Let's wrap up for today then, thank you. Hope to have you back on Crossing sometime.
👦🏻 Barkley
Sure, thank you both.
Subscribe to the "Crossing" Podcast
🚦 We focus on the new industry shifts and entrepreneurial opportunities brought by the new wave of AI technology. "Crossing" was Steve Jobs's metaphor for Apple — standing at the intersection of technology and liberal arts, where great products are born. AI is transforming every industry, and we seek out, interview, and bring together "active doers" of the AI era. Together with them, we explore and embrace the new changes and new possibilities.
👦🏻 Host Koji: Co-founder of The Fair and Tangdao. I believe technology, especially AI, will fundamentally transform society and empower humanity in the future. Welcome to chat with me, bounce ideas around, and connect on the next possibility. Koji on Jike[2], Koji's website[3]
👧🏻 Host Ronghui: Works at a tech VC, former Silicon Valley correspondent for CBNweekly. Ronghui on Jike[4]
Join the "Crossing" Membership Community
☀️ First-hand AI news and insights
👫🏻 We encourage everyone to date/make friends/find future collaborators
🦀 Add our assistant on WeChat to join: Rwkfbcianvd, or scan the QR code below


References
[1] Luma.ai: http://luma.ai/
[2] Koji on Jike: https://okjk.co/0JSUes
[3] Koji's website: https://koji.super.site/
[4] Ronghui on Jike: https://okjk.co/0cbnYV