云启资本

云启资本

@yunqipartners

科技常新,寻找未来开创者

534 articles18 episodes

Articles

The Pros and Cons of OpenAI's New o1 Model, Explained | Yunqi Tech π --- On September 13, Beijing time, OpenAI released its long-awaited new model series, o1-preview and o1-mini — the first fruits of its "Strawberry" project. This represents a significant departure from the GPT series, with o1 built on an entirely new training paradigm. ## What Makes o1 Different? The core innovation is **test-time compute scaling** — essentially, giving the model more time to "think" before responding. o1 spends seconds to minutes internally reasoning through problems, testing different approaches, and correcting its own mistakes before delivering an answer. This mimics how humans tackle complex problems: we don't always blurt out the first thing that comes to mind. We pause, consider alternatives, backtrack when stuck. o1 does something analogous, using a **chain-of-thought** process that's hidden from the user. ## Where o1 Shines **STEM domains, particularly math and coding.** OpenAI's benchmarks show dramatic gains: - **International Mathematical Olympiad (IMO) qualifying exam**: GPT-4o solved 13% of problems; o1 reached **83%** - **Codeforces competitive programming**:

The Pros and Cons of OpenAI's New o1 Model, Explained | Yunqi Tech π --- On September 13, Beijing time, OpenAI released its long-awaited new model series, o1-preview and o1-mini — the first fruits of its "Strawberry" project. This represents a significant departure from the GPT series, with o1 built on an entirely new training paradigm. ## What Makes o1 Different? The core innovation is **test-time compute scaling** — essentially, giving the model more time to "think" before responding. o1 spends seconds to minutes internally reasoning through problems, testing different approaches, and correcting its own mistakes before delivering an answer. This mimics how humans tackle complex problems: we don't always blurt out the first thing that comes to mind. We pause, consider alternatives, backtrack when stuck. o1 does something analogous, using a **chain-of-thought** process that's hidden from the user. ## Where o1 Shines **STEM domains, particularly math and coding.** OpenAI's benchmarks show dramatic gains: - **International Mathematical Olympiad (IMO) qualifying exam**: GPT-4o solved 13% of problems; o1 reached **83%** - **Codeforces competitive programming**:

Where's the ceiling for large language models?

Podcasts

Vol.17 48-Hour Xiaohongshu Hackathon Hit: How an AI-Native Product That Broke the "Retention Curse" Was Built --- Two weekends ago, I participated in a 48-hour hackathon hosted by Xiaohongshu. Our team of four built an AI-native product from scratch — no code, no design background between us — and ended up winning the "Most Popular" award. The product? A voice diary app called **"Echo"** that uses AI to turn fragmented daily moments into serialized, episodic "life podcasts." Think *This American Life*, but starring you. What surprised me wasn't that we won. It was that people kept using it *after* the demo. Here's the dirty secret of AI hackathons: most projects die the moment judges stop clapping. The "retention curse" is real — users try your GPT wrapper once, say "neat," and never return. We broke that pattern. Our daily active user rate among beta testers hit 34% in week one, which for a hackathon product is basically unheard of. How? Three deliberate choices we made against hackathon orthodoxy. **First, we refused to build a chatbot.** The default AI product in 2024 is still "talk to a large language model." We explicitly rejected this. Chat interfaces create *performance anxiety* — users feel pressure to ask the "right
Vol.17 48-Hour Xiaohongshu Hackathon Hit: How an AI-Native Product That Broke the "Retention Curse" Was Built

---

Two weekends ago, I participated in a 48-hour hackathon hosted by Xiaohongshu. Our team of four built an AI-native product from scratch — no code, no design background between us — and ended up winning the "Most Popular" award.

The product? A voice diary app called **"Echo"** that uses AI to turn fragmented daily moments into serialized, episodic "life podcasts." Think *This American Life*, but starring you.

What surprised me wasn't that we won. It was that people kept using it *after* the demo.

Here's the dirty secret of AI hackathons: most projects die the moment judges stop clapping. The "retention curse" is real — users try your GPT wrapper once, say "neat," and never return. We broke that pattern. Our daily active user rate among beta testers hit 34% in week one, which for a hackathon product is basically unheard of.

How? Three deliberate choices we made against hackathon orthodoxy.

**First, we refused to build a chatbot.**

The default AI product in 2024 is still "talk to a large language model." We explicitly rejected this. Chat interfaces create *performance anxiety* — users feel pressure to ask the "right

Vol.17 48-Hour Xiaohongshu Hackathon Hit: How an AI-Native Product That Broke the "Retention Curse" Was Built --- Two weekends ago, I participated in a 48-hour hackathon hosted by Xiaohongshu. Our team of four built an AI-native product from scratch — no code, no design background between us — and ended up winning the "Most Popular" award. The product? A voice diary app called **"Echo"** that uses AI to turn fragmented daily moments into serialized, episodic "life podcasts." Think *This American Life*, but starring you. What surprised me wasn't that we won. It was that people kept using it *after* the demo. Here's the dirty secret of AI hackathons: most projects die the moment judges stop clapping. The "retention curse" is real — users try your GPT wrapper once, say "neat," and never return. We broke that pattern. Our daily active user rate among beta testers hit 34% in week one, which for a hackathon product is basically unheard of. How? Three deliberate choices we made against hackathon orthodoxy. **First, we refused to build a chatbot.** The default AI product in 2024 is still "talk to a large language model." We explicitly rejected this. Chat interfaces create *performance anxiety* — users feel pressure to ask the "right