The Best Way for Humans and AI Agents to Work Together Hasn't Been Invented Yet | A Conversation with Paperboy

One look, and the Agent gets it.

👦🏻 Podcast Interview: Koji

🥷 Edited by: Crossing

🧑‍🎨 Layout: Zeoooo

🚥 This week on Crossing, our guests are the Paperboy team (https://www.paperboy.com). John Yang, 21, CEO. Jett Chen, 19, freshman at CMU, founding engineer. The Paperboy team has 12 people, 10 engineers, and has raised $4.7 million.

John believes: The best way for humans and AI agents to work together probably hasn't been invented yet. Though we already have Claude Code, Codex, Manus, and OpenClaw, they're all essentially session-based + prompt-based. The user opens a window, types a prompt, waits for completion, closes it. Next time, starting from zero.

Paperboy is trying to find a more natural, continuous, and collaborative agent interface and memory structure — agents should learn by observing how you use your computer, organize conversations via IM rather than sessions, and proactively reach out to you instead of waiting for your prompt.

If you're building AI products, AI infra, or thinking about how agents enter team workflows, we hope this episode gives you something to think about.

Listen on WeChat:

Listen on Xiaoyuzhou:

🎬 The video podcast is also live on Koji's WeChat Channels, Xiaohongshu, Bilibili, YouTube, and other platforms.

Lightning Round

👦🏻 Koji

Let's start with Crossing tradition — lightning round. How old are you two?

🧑🏻‍💻 John Yang

👨🏻‍💻 Jett Chen

👦🏻 Koji

What schools did you graduate from?

🧑🏻‍💻 John Yang

Didn't graduate. I was studying architecture at Pratt Institute.

👨🏻‍💻 Jett Chen

I graduated from Shanghai Starriver Bilingual School, and I'm a freshman at CMU, just finished my first year.

👦🏻 Koji

What's your MBTI and zodiac sign?

🧑🏻‍💻 John Yang

ISTJ, Gemini.

👨🏻‍💻 Jett Chen

INTJ, Virgo.

👦🏻 Koji

What were you doing before starting this company?

🧑🏻‍💻 John Yang

Paperboy is my second company. My first was called Million. We built a lot of open-source dev tools in the React ecosystem, then made a product called Same.Dev that let ordinary people recreate any website's UI just by entering a URL.

Million was in YC Winter 24.

👨🏻‍💻 Jett Chen

Before Paperboy I was just a high schooler who liked doing open-source projects and playing CTF. I built something called EarthKit that could use multimodal tech to guess a photo's location based on the image itself, working better than traditional pure neural network models.

👦🏻 Koji

When was that?

👨🏻‍💻 Jett Chen

About a year or two ago.

Starting Point: I'm Not Happy With Today's AI Products

👦🏻 Koji

Let's introduce what kind of product Paperboy is.

🧑🏻‍💻 John Yang

Paperboy is a very early-stage company with the mission of exploring the best way for me to collaborate with AI.

Last year, after I finished Same and used Manus, I always felt something was off about AI products on the market. Paperboy started from me trying various methods, constantly exploring different paths. We're trying to solve some technical problems, and also problems of product form.

For example, I shouldn't need to dump all my files, emails, and personal information into a chat box. If I want to collaborate with someone while also talking to an agent, there should be a very simple way for us to do that in the same context window.

Or, after an agent knows a lot about me, it should proactively do things for me ahead of time, but current chat windows can't do that at all. And right now all products are session-based — once you have too many sessions, you can't find previous conversation context anymore.

These problems are generally the gap between model capabilities and real-world application. I think there's still huge room for exploration and innovation in product experience, so our company is called "Paperboy Products" — products, plural.

👦🏻 Koji

We'll get into that later. Let's finish lightning round first — what's your funding situation?

🧑🏻‍💻 John Yang

We raised $4.7 million in '25.

👦🏻 Koji

Cool. Revenue and profit? Product hasn't launched yet, right?

🧑🏻‍💻 John Yang

Gross margin is zero or even negative, we're losing money every day, haha.

👦🏻 Koji

When can people roughly start using the product?

🧑🏻‍💻 John Yang

We've sent a prototype agent that can learn from OS activities to some friends.

But it's too expensive and doesn't run very well. We're working hard to finish the next generation this month, then push it out again.

👦🏻 Koji

When we publish this episode, we'll put the link below — interested friends can go sign up for the waitlist.

🧑🏻‍💻 John Yang

Yeah, I think by the time this episode comes out, people should be able to see it.

👦🏻 Koji

Got it. Current team size?

🧑🏻‍💻 John Yang

12 full-time employees, 10 of whom are engineers.

👦🏻 Koji

The first time I met John, you showed me a document you prepared for an internal team meeting. The first line was: "The best way for humans to collaborate with AI probably hasn't been invented yet."

What did you see when you wrote that? Has your view changed since then?

🧑🏻‍💻 John Yang

Yeah, that was the doc I prepared for our first all-hands. At the time the company was just me, Du, Chen, and Jett — four people.

We started from a core belief: the best way to collaborate with AI hasn't been invented yet, and we have the opportunity to be the team that finds that answer. Cursor was the first to really start trying to find the best way to do AI programming, and they had massive success. They were the first company to truly focus on that goal, and proved how important it is to get there first.

You asked what new insights I've had since then. I think the cool thing is, it's actually a constantly moving target.

You can never truly reach market expectations — you can only keep getting better. Because whenever you make something new, everyone else sees it. If other teams have taste, if users have taste, they can discover new pain points. Pain points always exist, so the only thing you can do is continuously improve.

It's a constantly moving target. This feeling has gotten stronger since OpenClaw and Anthropic's Claude Cowork launched.

👦🏻 Koji

Recently many founders have felt somewhat despairing, because "bombshells" keep dropping during the entrepreneurial process. In the six months since your company started, the industry has also changed dramatically — from Claude Code to OpenClaw, then Hermes.

How do you feel? Did what you originally wanted to build change dramatically because of these giants?

Under the Claude Code Bombardment

🧑🏻‍💻 John Yang

Not really. I think the space for exploration comes from humans.

From a technical and product perspective, problems can be divided into three categories:

First, technically, an Agent must be able to truly learn from the user's environment. It has to integrate into the user's existing workflows — where data is actually produced — like files on their computer and various software applications.

Second, it must be personalized. Personalization means you don't need to prompt it frequently, and you can trust it with more complex, important tasks and decisions. This also means it needs to be more reliable, running continuously over longer time periods.

Third, in terms of design, the experience must be extremely intuitive — users shouldn't have to learn it like a new tool. If your Agent is proactive enough to generate its own ideas, what form should carry those proactive outputs? It needs a complete environment, it needs personalization, and it needs to collaborate well with your existing team.

So when you look at new tools on the market, you won't find any new dimensions emerging. The development across these three dimensions is still constrained by human teams themselves.

The Two Core Problems of Agents

Cursor and Manus are currently the most successful agent forms, but John says they have two fundamental problems — and this directly defines what Paperboy is building.

👦🏻 Koji

Could you briefly introduce Paperboy to everyone? Many of our podcast listeners are probably already heavy users of Claude Code, Manus, or other Agents — why should they give Paperboy a chance?

🧑🏻‍💻 John Yang

Currently, Claude Code and Manus are the most successful Agent forms, but they're session-based, one-on-one, and prompt-driven.

This creates two important problems. First, being session-based means in their sidebar, you have multiple workspaces (projects), and under each project there's a pile of sessions. Every time you want the model to do something new, you have to start a new session.

Second, your interaction with the model is: you input a prompt, wait, send another message, it replies.

The problems with this approach are:

First, the Agent is passive. You have to describe things very specifically. You can create skill documents (like agent.md) to tell it what to do, but you have to actively maintain them, and it's very hard to translate your taste, judgment, and way of working into pure text.

Second, sessions are discontinuous. Having hundreds or even thousands of sessions is a terrible experience. I know that in some past sessions, the context window contained very valuable insights, but if I didn't deliberately save them at the time, that information is lost forever.

Paperboy directly addresses both of these problems.

First, the Agent must learn by observing how you use your computer. This includes your screenshots, keystrokes, mouse movements, meeting audio and video, browsing history, iMessage, and so on — of course, only if you authorize Paperboy to access this information.

Second, interaction should exist in a continuous chat stream, with a history much longer than a single context window, and it should be searchable. The standard product form should be like iMessage or WeChat: you have a bunch of chats, you tap in and continue the conversation with whoever's in there.

👨🏻‍💻 Jett Chen

To add on the session and context window problem. With current products like Claude Code and Manus, you could argue they have infinite context windows because they have compaction mechanisms.

Some newer products, like Poke from Interaction, Zo Computer, and even OpenClaw, have adopted similar forms — no sessions, your interaction with the Agent is a continuous conversation stream.

One major differentiator for Paperboy is the source of context. Their context mainly comes from the user's historical chat history with the Agent, or user-provided emails, messages, etc. We initially tried exporting users' WeChat or iMessage chat data, but quickly realized this wasn't a scalable approach.

The most scalable way is actually through the operating system level, observing the user's daily computer use to collect data. We found this gives a very comprehensive understanding of what the user does every day.

And from an information density perspective, the information density of daily computer use is extremely high. Observing 60 minutes of computer use teaches you far more than observing 60 minutes of WeChat chat.

So we decided very early on to achieve user adaptation through OS-level context.

After Screen Data Became Industry Consensus

"Collecting user screen data to build a Context Layer has, to some degree, become industry consensus."

👦🏻 Koji

A previous podcast guest, the founder of AirJelly, also built a desktop client to capture as much user context as possible. Recently OpenAI's Chronicle has a similar idea.

Is your approach similar to everyone else's? Or is there something different?

🧑🏻‍💻 John Yang

Grabbing raw data from computers and processing it into memory will become a universal trend. Not just startups like AirJelly that are focused specifically on this, but products like Codex, Claude Cowork, and Claude Code will all eventually do this — it's the next most obvious frontier for context.

Of course, different teams will process these raw stream data very differently. How you select and structure this information directly relates to the Agent's specific application. For example, a company focused on studying how users reply to emails will have a completely different memory structure and way of compressing raw data streams than ours.

There's enough room to maneuver here, and there's still enough opportunity to become the first company that truly understands who the user is across all applications, all relationships, and models the user's day.

This capability alone obviously attracts everyone to pile in. But depending on the specific application, there's still huge room for algorithm customization and variation.

👨🏻‍💻 Jett Chen

I also think that collecting user screen or computer usage data to build context has, to some extent, become industry consensus. What's more important is what you specifically do within this paradigm.

Currently, products like Codex or Littlebird treat screen data as a context layer. For example, Codex Chronicle's use case is: by collecting screen data, learn how users typically develop an application. If your use case is different, the final pipeline will also be different.

This is a very new field. Beyond collecting user data, there's actually a lot more you can do, which requires significant engineering and research. For example, how do you build the best proactive Agent? Is it predicting the user's next keystroke, or predicting what they'll do in the next hour? These are relatively underexplored problem spaces.

I don't think anyone has found an absolutely best solution yet, so for a company, exploring this area is still a great choice.

👦🏻 Koji

If a user installs Paperboy today, what's the biggest value they can feel in the first hour, even the first 5 minutes? What highlight do you want users to feel immediately?

🧑🏻‍💻 John Yang

We'll probably start with meeting prep.

Within an hour, it's very important to show users the framework of what your product can do, setting expectations. With memory, one characteristic is: the longer you use it, the better it gets. So you need an initial phase where users trust it to learn.

When you open it, you'll see a real chat window, not just a simple prompt box. Once you authorize it to access your calendar and email, it will start reading the information you give it, then ask some small questions about who you are, and start giving suggestions like: "Hey, I see you have a meeting coming up, want me to look at relevant materials?"

In this regard, I think Poke from Interaction does it best — they've found the secret: connect to context the user already has, and show them you're an agent that can truly interact, adapt to their personality, and help in a proactive but not annoying way.

This way, you can establish an expectation with users: we're an Agent that can proactively message you in a reasonable way.

MiniVivian & AutoJohn

👦🏻 Koji

How long has your team been using Paperboy yourselves? In this process, is there any "Aha Moment" you can share?

🧑🏻‍💻 John Yang

Our team member Vivian, she previously worked at Xiaohongshu and HSG YUE. We have a Vivian Paperboy called MiniVivian. My Paperboy is called AutoJohn. In our Slack, team members constantly ask AutoJohn questions directly — it handles all incoming inquiries, helping the product and design teams find the help they need.

Take MiniVivian as an example. Vivian does a lot of recruiting work, and MiniVivian is like her recruiting intern on the team. Because it understands all the judgments and taste I've shared with it about what kind of people we want to hire and where to find them — this information comes from our meetings and Slack communications. It can more accurately help Vivian source candidates on GitHub, Xiaohongshu, and Twitter, saving her massive amounts of time.

I think Vivian hasn't used Claude since February this year. She can't use it, because Claude doesn't have this background — you can't have it help with candidate reference checks, there are too many things about judgment criteria that you'd have to tell it from scratch.

👦🏻 Koji

Because it has more context, so when you prompt, you can almost not prompt at all.

🧑🏻‍💻 John Yang

Yes, I hate prompting. Since I started Same, I didn't want to write prompts. Of course you need to communicate, but we humans don't think in prompts — we send messages and expect the other person to know what we're talking about. We enjoy that high-bandwidth communication relationship.

I think smart people are happy to be told what they don't know, especially the things "we don't know we don't know."

Today's models are smarter than us, so frankly, I look forward to the day when I can just lie back and let AutoJohn become something smarter than me, with a higher IQ.

👦🏻 Koji

What about you, Jett? Any "aha moments" using Paperboy?

👨🏻‍💻 Jett Chen

First off, it's true that a lot of the time, chatting with John's "AutoJohn" is better than chatting with John himself.

👦🏻 Koji

When you're talking to his agent替身, do you worry that its intentions don't fully represent John, leading to misunderstandings?

🧑🏻‍💻 John Yang

My view is that users ultimately have to take responsibility for their own agents. The AutoJohn setup process isn't something that happens overnight — it's not like one day you suddenly have an entity you can pull into Slack.

There's an onboarding flow in between, where the agent asks you questions like: "How much information am I allowed to share with this person?" By default, it mimics my behavior — I share a lot with Jett, less with a new engineer joining the team, and the agent figures this out by observing all my chat history.

👨🏻‍💻 Jett Chen

Interacting with a person actually breaks down into many situations. There are things I wouldn't go to AutoJohn for, like permission issues that need John's personal approval. But in work scenarios, a lot of communication is information-based. John has the full context of Paperboy as a company, and as an engineer, I need to know how to create the most value for it.

In these cases, because AutoJohn on one hand factually possesses most of John's context, and on the other hand has developed heuristics very similar to John's through observation, I find AutoJohn extremely useful when making decisions about context and heuristics.

My other "aha moments" came earlier. When we built the text completion feature, I immediately found it useful for daily programming.

There are lots of AI command-line tools trending now, but they're not as smooth as traditional ones. And traditional tools often lack AI integration, which is annoying when writing scripts. With text completion, after I write a bunch of code and need to send a git commit, I can just type "@pb commit" in the command line, and it automatically writes the entire commit message for me — I hit enter and it's sent.

👦🏻 Koji

Can you expand on how this works?

👨🏻‍💻 Jett Chen

Our development process starts with building a system that can collect user data from the operating system and form effective memories. On top of that system, we have a framework that generates and updates a Markdown document about the user in real time.

This document includes the user's profession, activities over the past few days, even what they were doing seconds or minutes ago. The closer to the present moment, the finer the granularity of information.

So the Paperboy agent always has this context. With that foundation, we started looking for application scenarios. The first good one we found was implementing autocomplete anywhere in the operating system.

For example, when you're messaging on WeChat, you can type "@pb" as an activation word in the input box, followed by a brief instruction, or nothing at all.

👦🏻 Koji

If you don't input anything, it guesses what you're trying to do right now?

👨🏻‍💻 Jett Chen

Yes, because it has context.

👦🏻 Koji

Like working with a colleague who gets you — sometimes you don't need to speak, one look and they understand. Point at the screen: "Look here."

👨🏻‍💻 Jett Chen

Right, they immediately know which issue you're talking about. Paperboy was basically at that level. So whatever you're doing, it can provide context at the right moment.

For me, one "aha moment" was in the command line or GitHub. When I send a PR, it can directly write the entire PR description for me.

I found that what it writes is better than what Cursor or Claude Code would produce. Because when I'm developing a feature, I might be interacting with Claude Code one moment, messaging with John on WeChat the next, then doing research in the browser. Paperboy can aggregate all this cross-application context to generate a PR draft.

I was amazed by how deeply it understood what I was working on — far beyond AI tools that only have context from a single application.

👦🏻 Koji

Anything else?

👨🏻‍💻 Jett Chen

I suspect I have ADHD — I get distracted easily while working. I'm often using Claude Code while browsing Hacker News, clicking on interesting articles and reading for half an hour, then coming back to find Claude Code finished running ten minutes ago.

For this, I used to have Paperboy give me a report every evening on how productive I was. I even had it "scold" me when I was inefficient or distracted. I feel like knowing Paperboy is "watching" me work makes me more productive.

👦🏻 Koji

Not "Big brother is watching you."

🧑🏻‍💻 John Yang

A large part of my work now is using different products, doing research, and talking to people — unfortunately I can't write as much code as I used to. But Paperboy is great because it can keep all these different sources of information, at different granularities, in its head.

For example, when I want to learn about WeChat's history, or case studies of successes and failures in various network-effects business models, after doing the research I need to connect these dots and relate them to Paperboy's development.

At that point I have many ideas colliding in my head, and I need a tool to help me organize my thoughts, even remind me: "Wait, you tried a similar product format months ago and discovered these issues." I know these things, but when doing this kind of cross-level thinking, while I could manually draw it out on paper, having a smart model that really knows me makes the process much faster.

WeChat Group Chats Inspired the Interface Design

👦🏻 Koji

Any recent examples?

🧑🏻‍💻 John Yang

We started designing Paperboy's latest interface two weeks ago. We had a problem: if we want to sell the product to VCs, they want personal CRM modules, meeting reminder modules, deal tracking modules. But the product also works for founders, real estate salespeople, who need completely different modules. What do we do? Make these skills, plugins, or recipes for users to pick from?

That didn't feel right for Paperboy. Then Paperboy pointed me in a direction: What's something in the apps we use daily that can hold an infinite list without making people feel overwhelmed? The answer is instant messaging.

Think about our iMessage and WeChat, especially WeChat — you have contacts, and a ton of group chats. Three people might create four different groups for different topics. This is actually a very intuitive way of organizing conversations.

👦🏻 Koji

True, we create a new group for every podcast episode, even with the same members. Because multiple episodes are in post-production simultaneously, having everything in one group would be chaotic.

🧑🏻‍💻 John Yang

Right. The beauty of IM is that when a group becomes inactive, it sinks down, and you can hide it too. So WeChat itself is like an inbox.

That was our inspiration: we knew the product needed an inbox feature, so why not design it as something that is an inbox? IM is. That's where the latest interface came from.

👦🏻 Koji

From what I've heard, Paperboy has many distinctive features and interaction concepts. But as we discussed earlier, capturing all context is becoming consensus. If your interaction paradigm proves effective, others will catch up quickly.

How do you view competition? Who are your main competitors? How do you maintain sustainable advantage?

🧑🏻‍💻 John Yang

Only two companies really matter today: OpenAI and Anthropic. Everyone else is behind them. They have all the advantages: they have the models, Anthropic just got more compute, and they have excellent distribution channels. All we can hope for is to have some taste advantage like Cursor, and be one step ahead in exploring new interaction interfaces.

This is all about the individual user product experience, and that part won't disappear. But the other piece is the enterprise market, which they'll definitely want to enter because that's where the money is.

Paperboy actually wanted to build AI Slack as its first product. That was our initial answer to "if the best way to collaborate with agents hasn't been invented yet, what should we build?"

👦🏻 Koji

That's also a hot space — many products are exploring what tools we should use when agents enter teams. What do you think of products like Slock, Multica, and Moxt in the market?

🧑🏻‍💻 John Yang

First, I really like Slock — Richard is smart, he and his co-founder have great backgrounds, and the team is excellent. But the difficulty of building an AI version of Slack is: how do you get enterprise customers to switch platforms? This is purely my personal view.

For AI Slack's business model to work, you have to sell to teams. But for enterprises, unless there's a way to seamlessly migrate all data from their existing Slack with a one-to-one experience, the switching cost is too high. And you still need to build a ton of connectors to integrate with all the other data systems enterprises already have. You can't skip this foundational work.

Second, to be frank, Slack isn't a great product. Its channels and threads interface isn't pleasant to use. At a certain scale you have to use it because there's no better alternative, but it's definitely not a well-designed product.

And that's just on the user and acquisition side. Then there's the Agent problem. As Jett said, we don't think an Agent can learn very fast if the only information source it has is the messages you send it.

People with jobs are busy. They don't have time to care about training an Agent. They want a product that works out of the box and keeps getting better on its own. One reason we gave up on building a Slack replacement is that we felt Agents inside Slack would still need more context to learn naturally.

👨🏻‍💻 Jett Chen

I think Slack has strong network effects that prevent it from being displaced by new players. For example, it has a feature called Connect — a lot of inter-company communication now happens through that instead of email, and that's very hard to replace. Combined with all its other integrations, Slack has become somewhat like WeChat: indispensable.

By contrast, if you position the Agent at a higher level — at the OS layer, as something complementary to Slack — you actually get a much better path to landing in enterprises.

👦🏻 Koji

So is that what you're doing? Trying to complement Slack?

🧑🏻‍💻 John Yang

Yes. Our assumption is: if you're going after the enterprise market, they're already using Slack for communication. If you build something like Slack, you'll have a very hard time selling it to any team above 50 people, because the cost of getting them to stop using Slack is too high.

So bringing the Agent into workflows people already have, rather than trying to replace everything and build from scratch, is the more sensible approach.

The Last Interface and Five Kinds of Speed

👦🏻 Koji

I noticed a blog post on Paperboy's website titled The Last Interface, which mentions "five kinds of speed." Could you expand on that? I understand this reflects your thinking on the context layer and memory system.

🧑🏻‍💻 John Yang

Sure, though I won't go too deep into the details. The choice of "five" isn't because we literally have only five layers — it's just a reasonable way to categorize things.

The underlying idea is important. There's a nonprofit in San Francisco called the Long Now Foundation, founded by Stewart Brand, who published a theory called "Pace Layers." He divides the world into six levels: fashion, commerce, governance, infrastructure, culture, and nature.

These layers change at different speeds. Fashion is fastest; nature (like the laws of physics) barely changes at all.

The world is made up of these layers with different rhythms, and their interconnection constitutes our society.

I've been fascinated by this theory since high school, and it's shaped my worldview. When I think about how to efficiently and scalably represent the real world we live in within the world of Agents, I feel we need these same kinds of pace layers.

👦🏻 Koji

Interesting. So what corresponds to unchanging "nature," and what corresponds to ever-shifting "fashion"?

👨🏻‍💻 Jett Chen

For products, if you classify user tasks by duration, you find many different kinds. For example, replying to a message on WeChat might take 10 seconds — that's one end of the spectrum.

On the other end, a user might spend several hours reading 10 long reports and then make a business decision.

🧑🏻‍💻 John Yang

That might still be on the shorter side — long-term tasks are measured in months.

👨🏻‍💻 Jett Chen

We believe every segment of this spectrum can be augmented or automated by some kind of AI Agent, but each segment requires different things.

If you want to automate "replying on WeChat," the best product form is probably an autocomplete suggestion that pops up when you tap the input box — something we're already doing. But as the time span gets longer, the product form becomes increasingly uncertain. How do you automate a task that spans several hours? What's the best form for that? That's still a very worthwhile area to explore.

👦🏻 Koji

Before recording this podcast, several members of your team mentioned that when communicating with John, they don't feel like he's from '04 — that he has a maturity beyond his years. I felt the same. John, where do you think this comes from?

🧑🏻‍💻 John Yang

I think deep down I'm still quite childish, especially when I'm with Jett — we have a lot of fun together.

👨🏻‍💻 Jett Chen

I'd say you're an old soul.

🧑🏻‍💻 John Yang

At least with the team, I just have a lot of passion for business. We want to do something very hard in a very short amount of time, so clear thinking, efficiency, and focus are essential.

👦🏻 Koji

How do you achieve that? Easier said than done.

🧑🏻‍💻 John Yang

It has to start with myself. My top priority is defining what success means for the team, and making sure everyone has a clear, accurate understanding of the context behind their goals and how to execute them.

For me, the best part of founding a company is getting to learn many things I've always wanted to explore and understand — this curiosity is my biggest motivation. From high school through college, I did several internships, working on different types of tasks at different companies.

Later, when we founded Million, we started with B2B, selling web performance analytics tools to large enterprises. We quickly learned that was a very hard business to scale — as 20-year-olds, we weren't good at building enterprise relationships.

From then on, I've been exploring what kinds of markets actually make sense for what kinds of players. The way I like to learn is by spanning different business scales, not just staying within the startup world.

👦🏻 Koji

So it's like founder-market fit?

🧑🏻‍💻 John Yang

Yes, and more than that — it's the entire team's fit with the market. I think market selection is one of the most important things for early-stage founders. You have to make sure you understand the space you're entering, know where all the leverage points are, and where the challenges will come from.

Only then can you build the team, and only after that come product experience, fundraising, customers, and technology.

👦🏻 Koji

Any advice on how to choose the right market?

🧑🏻‍💻 John Yang

I'm probably going to sound like I'm stating a bunch of obvious things.

The market needs to be large enough, and not just in dollar terms — you need to look at sustainability and whether it can keep growing, so your team can keep iterating on the product and solidifying market position over the next decade.

Looking at the most successful products in history, they've never been single products — they're product matrices. You have to expect that what you're building might take ten years to truly become a mass consumer product.

Also, you have to start with problems around you. For me, that's a way to understand the market, but then you need to understand a lot of history to see whether business dynamics and economic patterns are in your favor, and whether there are successful precedents from the past.

There are so many details here — I feel like I'm still just scratching the surface. Paperboy, for me, is an opportunity to test what I've learned.

👦🏻 Koji

Paperboy is probably your first time as CEO — at your last company you were a co-founder. What's the biggest change since becoming CEO?

🧑🏻‍💻 John Yang

It's incredibly hard, really. You're responsible for everything.

As a co-founder you feel some of that too, but as CEO, I'm the person most likely to lose the company a lot of money, because my decision leverage is so high. If I'm wrong about direction, the cost is extremely high.

As the team grows, I'm also responsible for hiring, and for helping the team and individuals keep growing, making sure everyone is in the right role. These are basic management skills, and this is basically my first time managing, so a lot of it is learning on the fly.

For example, how do you do performance reviews? How do you do one-on-ones? How do you make sure sometimes I only need to communicate with team leads, rather than micromanaging everyone.

👦🏻 Koji

Where did you learn these management skills?

🧑🏻‍💻 John Yang

Before Paperboy, I worked at Manus for a while. Their CTO, Pan Pan, is an outstanding engineering manager. I had several one-on-ones with him where I just asked him everything I wanted to know: "How do you keep everything so organized?"

I've also spent time reading. I still think the best-written management book is Andy Grove's High Output Management, plus Ben Horowitz's The Hard Thing About Hard Things, and Bill Campbell's biography Trillion Dollar Coach.

I also got myself a CEO coach — we talk for an hour every week, which has been enormously helpful, especially early on. It gave me a space to talk through all the specific problems I was facing.

👦🏻 Koji

How did you find her?

🧑🏻‍💻 John Yang

An investor introduced us — she was also that investor's coach. She was previously a VC executive and had worked at many large companies; she's an expert at coaching startup CEOs.

I've never tried seeing a therapist, because I don't really believe in psychotherapy — it feels like you just go in and talk about emotions. But coaching is much better. You can certainly talk about emotions, but more importantly you can talk about business, about everything happening in the company, and it's all confidential.

It's also interesting to use Paperboy during my coaching calls — it listens in on our conversation, then helps me follow up on commitments I made during the call and organizes everything.

Two Kinds of Engineers, One Book, One Coach

👦🏻 Koji

Your team is now 12 people — not small. With so many AI tools and Agents available today, how has building a team changed compared to the past?

🧑🏻‍💻 John Yang

If you want to build serious infrastructure, you need people who understand infra. We've hired two kinds of people on the engineering team: one kind is like Jett — young, high-IQ, creative, able to quickly build prototypes for every hard problem.

The other kind are people with rock-solid fundamentals in a specific domain. You have to deeply understand systems and the underlying OS. We hired an engineer from AWS who had previously done something akin to Windows kernel development. I think for the foreseeable future, these kinds of domain specialists are indispensable.

👦🏻 Koji

Jett is a freshman now. We've seen some Silicon Valley companies like Palantir directly offering high schoolers jobs, believing that college is no longer necessary in the AI era. As a founding engineer at Paperboy, why are you still committed to college?

👨🏻‍💻 Jett Chen

For most people, college still makes a lot of sense. At least at CMU, most of my peers don't actually know what they want to do in the future. When you're uncertain, improving your technical skills — college is a pretty good option for that. Schoolwork only takes up part of your time. Most of it is yours to decide how to spend, giving you four years to explore.

But at the same time, if you already know exactly what you want to do when you get to college, and you genuinely have the opportunity, I think dropping out is also a rational choice.

👦🏻 Koji

John, as a young, first-time CEO-founder, how did you attract talented people to join you?

👨🏻‍💻 Jett Chen

I can talk about what drew me to John. A lot of the time when you choose a founder, you look at their slope, not their intercept — slope matters way more than intercept. And I think John is a high-slope founder.

Look at his track record: he's pulled off a lot of impressive things in a very short time. That shows he's highly agentic, able to learn fast and execute. For founders, that's arguably the most important quality. You need to think clearly and adapt quickly to whatever situation comes up.

First, I think he's a good founder. I've worked with him long enough to have high confidence that we collaborate smoothly. Second, he has good taste across research, engineering, and even product.

I feel like working with him, I can learn a lot of lessons and grow personally. I personally find John to be a very charismatic person.

👦🏻 Koji

Major confession session here. John, what about you? Do you have a system for convincing candidates to join?

🧑🏻‍💻 John Yang

I treat every candidate as an individual — no fixed playbook. It has to be mutual. I've never liked hard-selling candidates on the company.

Recruiting is one thing, but what matters more is my reliability. That's something I care about deeply: after I hire someone, can I deliver on what I promised them? I have to make sure I can consistently bring the team what they need.

You want to pick exceptional people — you can't work with mediocre ones. You can tell from the conversation, from their past work, from the life decisions they've made.

If they're experienced, there's decades of accumulated decision-making behind them. You can learn a lot from that. That kind of thing can't be faked.

👦🏻 Koji

Besides Paperboy, what's your favorite AI product to use?

🧑🏻‍💻 John Yang

I still like Cursor, but I still don't use Codex.

👨🏻‍💻 Jett Chen

I especially like Codex.

🧑🏻‍💻 John Yang

Cursor has been a huge inspiration for me. I'm happy for them, and I hope their acquisition goes smoothly.

Turning Down Cognition, Vercel, Sentry — Then What?

👦🏻 Koji

John, when you first started Million, I heard you received an acquisition offer from Cognition, the parent company of Devin?

🧑🏻‍💻 John Yang

Yes, and from other companies too, like Vercel and Sentry.

👦🏻 Koji

You didn't accept any of them at the time. How did you think through and decide on those acquisition offers? Looking back, are you glad about that decision, or do you have regrets?

🧑🏻‍💻 John Yang

First, these were all acqui-hires, so the offers weren't for much money. And joining these companies, you're somewhat just another employee working on someone else's ideas.

I've never been that interested in developer tools, so asking me to join Sentry, Vercel, or Cognition — I wouldn't want to do that. So it didn't make sense. If you can stay independent and pursue your own vision, of course you stay independent. That's what every founder wants.

👦🏻 Koji

Let's go back to Codex, which you like. Why Codex and not Claude Code?

👨🏻‍💻 Jett Chen

Claude Code has a different philosophy. It represents the most ambitious vision of how software engineers will work in the future. I think its most important function is demonstrating how good the Opus model is, not how good its CLI tool is. I appreciate the Opus model, but I don't really like Claude Code's CLI itself.

I like Codex for a few reasons.

First, its core agent is open source — the CLI and agent loop are both on GitHub.

Second, if you use their desktop app, you'll find the polish is far beyond Claude Code's desktop experience. They previously acquired a team called Sky that worked on Apple Shortcuts. OpenAI hires people like that to build really refined products. For example, their recent Codex Pets is both polished and fun. Including Codex's computer and browser use, everything is very polished.

I also think Codex was the first product to make high-parallelism, multi-agent simultaneous work its primary UI.

Finally, OpenAI's infrastructure advantage over Anthropic is much larger. So with Codex, the model quality is similar, but it's much more stable.

OpenAI can subsidize more compute, while Anthropic doesn't have enough investment to subsidize, leading to some strange moves — like if your codebase contains OpenClaw or Hermes agent, they might charge you 10x.

👦🏻 Koji

If you could buy their secondary market stock right now, would you buy OpenAI or Anthropic?

👨🏻‍💻 Jett Chen

I think I'd buy OpenAI.

👦🏻 Koji

All in on OpenAI?

👨🏻‍💻 Jett Chen

Maybe 80% OpenAI, 20% Anthropic.

👦🏻 Koji

Why?

👨🏻‍💻 Jett Chen

Both companies will do well, but I prefer what OpenAI is doing. I think Anthropic does best in interpretability and social impact.

But Anthropic is too opinionated about many models, both morally and in terms of specific usage patterns.

By contrast, OpenAI's philosophy is: with minimal restrictions, users should be able to use the model however they want. I personally lean more toward free will, so I prefer OpenAI's less prescriptive values. Also, I think OpenAI will have the compute advantage.

🧑🏻‍💻 John Yang

If I were making an investment decision, I wouldn't buy either at current valuations. But personally I actually prefer Anthropic, because of their commitment to safety.

👦🏻 Koji

What do you think of the idea that "models eat everything"?

👨🏻‍💻 Jett Chen

You could argue that in the future, many companies will offer roughly equivalent models, and product differentiation will still come down to product companies. That won't necessarily happen, but it's a possibility.

Cursor is a great example. They started as a product company, and now they're both a product company and a model company, using XAI's compute to train frontier coding models.

So I think companies shouldn't limit themselves to being "model companies" or "product companies" — usually they're both.

👦🏻 Koji

If we gave you $3 million in virtual money to invest in three teams you know, who would you pick?

🧑🏻‍💻 John Yang

Slock is great — I'd bet on Richard. Then robotics companies, I think that's absolutely going to be a massive category.

Finally, I believe the next company as great as Cursor will be born in the consumer space, not enterprise AI. Paperboy is of course in that market because I follow this conviction, but the consumer market is so large, with infinite opportunities.

👨🏻‍💻 Jett Chen

I'd also invest in robotics. After the wave of automating knowledge work, the next biggest opportunity is likely using model capabilities to transform the physical world.

Another one is security. AI plus security is an excellent space — the demand is basically unlimited. The stronger models get, the stronger both attack and defense capabilities become. That market is massive.

🧑🏻‍💻 John Yang

The follow-up question is, how do you compete with Anthropic, who will also get into this business?

👨🏻‍💻 Jett Chen

Anthropic and OpenAI will definitely have the best models. The question is, once you have the best model, what's your next optimization layer? Is it building better toolchains? Is it teaching agents the various heuristic experiences of human security researchers? Or is it building infrastructure, like constructing a massive multi-agent system to continuously find vulnerabilities? The exploration space here is enormous.

It's a very elite market — if your product finds vulnerabilities even 10% more efficiently than others, it's extremely valuable.

🧑🏻‍💻 John Yang

We haven't talked about long-horizon models yet, but Harvey just released long-horizon legal AI. And another John Yang recently released something similar to SWE-bench that can reproduce and complete entire codebases from a series of spec documents. That trend isn't going away.

👨🏻‍💻 Jett Chen

Security capabilities are critical to national security. For startups, it's also a great opportunity.

Because the country with the strongest model will have an incentive to say, "This model is only for us." So what do other countries do? Many different nations or interest groups will need this kind of model, so the market is enormous.

👦🏻 Koji

One last question: a year from now, what do you hope Paperboy looks like? What are you looking forward to, or what are you afraid of getting wrong?

🧑🏻‍💻 John Yang

Hopefully we're not losing money every day. I want to keep hiring the best people I can find, so talent density has to be even higher than it is now.

There are too many details we could discuss, but from a team and business perspective, we need positive cash flow, and we need to keep building a stronger team.

👦🏻 Koji

Alright, thank you both.

🧑🏻‍💻 John Yang

Thank you.

👨🏻‍💻 Jett Chen

Thank you.

Crossing is looking for independent writers to produce AI product and model reviews.

If you've written articles like: "Hands-on with PixVerse C1" or "Hands-on with LibTV", please contact zeo0811@gmail.com. Your email should include: ① a brief bio, ② AI review articles you've written.

We offer competitive compensation. Looking forward to observing and documenting the AI era together 🎪