云启资本

云启资本

@yunqipartners

科技常新,寻找未来开创者

534 articles18 episodes

Articles

The Pros and Cons of OpenAI's New o1 Model, Explained | Yunqi Tech π --- On September 13, Beijing time, OpenAI released its long-awaited new model series, o1-preview and o1-mini — the first fruits of its "Strawberry" project. This represents a significant departure from the GPT series, with o1 built on an entirely new training paradigm. ## What Makes o1 Different? The core innovation is **test-time compute scaling** — essentially, giving the model more time to "think" before responding. o1 spends seconds to minutes internally reasoning through problems, testing different approaches, and correcting its own mistakes before delivering an answer. This mimics how humans tackle complex problems: we don't always blurt out the first thing that comes to mind. We pause, consider alternatives, backtrack when stuck. o1 does something analogous, using a **chain-of-thought** process that's hidden from the user. ## Where o1 Shines **STEM domains, particularly math and coding.** OpenAI's benchmarks show dramatic gains: - **International Mathematical Olympiad (IMO) qualifying exam**: GPT-4o solved 13% of problems; o1 reached **83%** - **Codeforces competitive programming**:

The Pros and Cons of OpenAI's New o1 Model, Explained | Yunqi Tech π --- On September 13, Beijing time, OpenAI released its long-awaited new model series, o1-preview and o1-mini — the first fruits of its "Strawberry" project. This represents a significant departure from the GPT series, with o1 built on an entirely new training paradigm. ## What Makes o1 Different? The core innovation is **test-time compute scaling** — essentially, giving the model more time to "think" before responding. o1 spends seconds to minutes internally reasoning through problems, testing different approaches, and correcting its own mistakes before delivering an answer. This mimics how humans tackle complex problems: we don't always blurt out the first thing that comes to mind. We pause, consider alternatives, backtrack when stuck. o1 does something analogous, using a **chain-of-thought** process that's hidden from the user. ## Where o1 Shines **STEM domains, particularly math and coding.** OpenAI's benchmarks show dramatic gains: - **International Mathematical Olympiad (IMO) qualifying exam**: GPT-4o solved 13% of problems; o1 reached **83%** - **Codeforces competitive programming**:

Where's the ceiling for large language models?

Podcasts