OpenAI's Unprecedented Product Blitz: A Complete Breakdown of Its 12-Day Launch Marathon
The most noteworthy thing isn't just o3.
This week on Crossing, we're joined by Guizang and Dacongming to recap everything OpenAI announced during its 12 days of launches. Beyond the world-stunning o3, what other new features, technologies, and highlights deserve attention?
For instance, Dacongming believes Day 9's release — which barely made a splash — is especially noteworthy: the OpenAI API update is as significant as o3 itself, because it provides critical infrastructure for future AI application development. The continuous iteration of structured output capabilities deserves attention (from 36% to 100% success rate improvement), which will greatly accelerate AI agents and projects connecting AI to the real world.
From Day 1 through Day 12, we'll not only comprehensively cover each day's announcements, but also share our personal experiences and insights.

Illustration generated by Recraft
🚥 OpenAI 12-Day Launch Record 🟢 Day 1: Full-power o1, ChatGPT Pro $200 membership, o1 Pro
🟢 Day 2: Reinforcement Fine-Tuning (RFT) based on o1
🟢 Day 3: Sora
🟢 Day 4: ChatGPT Canvas
🟢 Day 5: Full Apple ecosystem GPT integration
🟢 Day 6: 4o real-time video calls, video understanding, screen understanding, Santa voice
🟢 Day 7: ChatGPT Projects
🟢 Day 8: ChatGPT Search fully open with optimized experience, free users eligible
🟢 Day 9: o1 API (supports Function Call, with Function Call web access), real-time voice API updates/price cuts & SDK release, new model support: PFT preference fine-tuning
🟢 Day 10: ChatGPT's 800 number, WhatsApp
🟢 Day 11: ChatGPT desktop can read other apps, supports o1 and 4o advanced voice
🟢 Day 12: OpenAI o3 officially released!

👬🏻 Guest Introductions Guizang is the creator of "AIGC Weekly[1]" Newsletter and the "Guizang's AI Toolbox" WeChat public account — in my opinion, the most worthwhile AI newsletter to subscribe to in the entire Chinese internet. I've followed it for two years; it's practically a weekend ritual, and I've benefited immensely.
Dacongming is the creator of the "CyberZen" WeChat public account, and this is his second appearance on Crossing.
In my circle of friends, both tracked all 12 days of launch developments throughout. I'm confident they can not only provide timely updates but ensure high-quality content as well.
Listen on WeChat:
Listen on Xiaoyuzhou:

The Stunning o3 Release: Technical Breakthroughs and Impact of the New Model
🚥 Koji
Alright, our opening question for both of you: "What do you think was the most noteworthy highlight of these 12 days of launches?"
👦🏻 Dacongming
Hello everyone, I'm Dacongming. I'll take this one first. In my view, the most noteworthy aspect isn't one highlight, but two.
The first is undoubtedly the o3 release. It delivered a completely, far-and-away leading model. Though it's expensive — answering one question might cost $3,500, a figure I "measured with a ruler."
The second is a detail hidden during the release period. Around Day 9, there was a developer update. This included Realtime API updates and Go language support. But the core was that it allows structured output during o1 and Realtime operations, which lays groundwork for next year's AI agent explosion. These two points are what I consider extremely important.
🚥 Koji
Good, we'll expand on both shortly: the o3 release and the series of API updates for developers on Day 9.
Master Guizang, what do you see as the most noteworthy highlight?
👦🏼 Guizang
I'm Guizang. For me, it's also o3 — unquestionably the most noteworthy.
Honestly, OpenAI has been leading the entire industry's direction. While it may not always do everything best, when the industry hits a wall, it always charts a new path.
Not long ago, people kept saying pre-training had hit its ceiling, right? With o3, we saw breakthrough results. Though it wasn't so obvious with o1 — it didn't make people so convinced about this reasoning evolution direction — with o3 we saw clear progress and advancement. I think this is hugely significant for confidence across the industry, including investment confidence.
🚥 Koji
Is there any way to help people grasp just how powerful o3 is?
👦🏻 Dacongming
The most direct example is this programmer leaderboard. On Codeforces, a more hardcore competitive programming platform than LeetCode.
Many exceptionally skilled programmers compete there. For instance, OpenAI's current chief scientist has a Codeforces rating of 2655. This time, o3 scored 2727, surpassing OpenAI's chief scientist by a wide margin. On the current leaderboard, that ranks 175th among humans — which is absolutely absurd.
🚥 Koji
There's another staggering figure about o3: completing a single task costs roughly $3,500, equivalent to 20,000 RMB.
I saw Dacongming also wrote on his public account that if you ask o3: "Which is larger, 9.09 or 9.11?" — that question alone burns 20,000 RMB. Does this also show that computing power can still work miracles?
👦🏻 Dacongming
There's actually a small detail here. When comparing o3 to o1, there are two versions.
One is the low-compute version, where a single task costs roughly $20 — probably what we'll use in the future. The other is the high-compute, more thorough mode, whose compute requirement is 170x that of the low-compute version. That works out to $3,500.
From $3.50 to $3,500 — roughly a 1,000x increase.
👦🏼 Guizang
So that low-compute mode — I saw it achieved over 75% on the ARC test set, and this version is $20 per run. That's actually manageable.
👦🏻 Dacongming
When we looked at that performance chart, something interesting emerged:
The percentage of accuracy and the exponent of compute consumed have a linear relationship. We can draw an almost straight line: every tenfold increase in compute yields a fixed percentage gain in accuracy.
🚥 Koji
Between 10% and 20%.
👦🏻 Dacongming
What this means is that if we wanted 100% accuracy on this test set, the compute cost would be astronomical.
And that's not all. On the new test set, we saw o3's high-compute mode reach 88% accuracy, but on the second version of the ARC leaderboard, its accuracy drops to just 30% — it gets compressed further.
If we used ARC test set standards to achieve AGI, even at current compute costs, we'd likely be talking millions of dollars.
🚥 Koji
I also saw Master Guizang posted a long piece on Jike about what o3 made him feel. You mentioned a striking claim: "In the coming years, we may remember the moment o3 was released last night as vividly as we remember when ChatGPT launched."
What made you so excited about the o3 release, seeing it as a milestone?
👦🏼 Guizang
Actually, these are summaries of what various big names have said.
Terence Tao mentioned that tech people could probably hold out against large language models for several more years, but now they've been yanked to 25% success rate in one go. Including those programmer competition leaderboards mentioned earlier — this represents a deeply inspiring future.
From o1 to o3, just three months to reach this level of progress. If this scaling law continues, will we see o4, o5 in the first half of next year?
When o4 and o5 release, setting aside other fields and looking only at math and code — will humans be completely unable to catch up? Honestly, code is the foundation of our entire software world, so this will bring enormous change.
👦🏻 Dacongming
To add some AGI context here — this was from Mark's share at the last OpenAI offline event I attended.
At this o3 launch, the opening was Mark and Sam Altman presenting together. Mark offered an interesting perspective:
When we reach AGI depends on how we define AGI. We'll soon reach whatever we define as AGI, and then we'll have a new definition of AGI, and keep chasing.
OpenAI selected ARC as its AGI evaluation partner. ARC cited a mainstream AGI formulation: a system capable of automating most economically valuable work. By this standard, we could consider o3 as having nearly reached AGI. But soon, as we hit this AGI standard, we'll have an even higher, newer standard.
🚥 Koji
This is interesting — it's about how AGI should actually be defined. Previously, there was never any consensus. In ARC's definition, true intelligence means being able to do economically valuable work.
This also means that AI comforting your emotions or empathizing with your feelings — these don't fall within their definition of AI.
👦🏻 Dacong
So a new definition was offered:
AGI isn't about how many skills you have, because skills can be acquired through training. It's about how well you can learn.
A baby, by our most intuitive understanding, we naturally consider him or her to be AGI. But they have no skills whatsoever. They can't code, let alone reach the top 175 human level. But they're very good at learning. They can master language from scratch. They can learn to use chopsticks. They can express needs through crying.
So should our definition of AGI shift from "how many skills" to "how much can it autonomously learn afterward."
12 Days of OpenAI: From Full o1 to Real-Time Video Calling
🚥 Koji
Let's quickly review what was released over these 12 days, then we'll dive into each day in detail.
Day one: the full version of o1 launched. ChatGPT also introduced its controversial Pro membership at $200 per month. And o1 Pro was released that same day.
Day two: Reinforcement Fine-Tuning (RFT) was released.
Day three: the official version of Sora finally launched.
Day four: Canvas was introduced, a feature competing with Claude Artifacts — a change in interaction design.
Day five: relatively quiet, mainly announcing full Apple ecosystem integration with GPT.
Day six: with Christmas approaching, real-time video calling and video understanding for 4o was released. It can understand real-time video streams and your previously shared screens, answering questions based on video feed and screen content. And with Christmas coming, you could even call Santa Claus.
Day seven: Projects feature launched — something Claude already had.
Day eight: ChatGPT search fully opened to all users, including free tier. Lots of detail improvements: search directly from browser address bar, video search, and integrating 4o's real-time voice into search.
Day nine: o1 API released, a series of developer-facing interfaces. We'll have Dacong go into detail on this later, since he considers this equally noteworthy as o3.
Day ten: somewhat quiet — support for calling ChatGPT by phone, plus a WhatsApp chatbot.
Day eleventh: reiterating the previously released ChatGPT desktop app that can read content from other applications. So instead of always screenshotting for ChatGPT, you can directly let it see your screen content and answer questions. This supports calling the o1 model, and also 4o real-time voice conversations.
Day twelve: the heavyweight release we just discussed — o3, which shocked the entire industry.
Let's go back to day one. Many people were very excited then. I'm guessing Master Cang and Dacong, you probably stayed up watching the launch. Can you share your reactions to seeing o1 Pro and the $200 ChatGPT Pro membership?
👦🏻 Dacong
My biggest reaction was: first, are they crazy? $200 is far beyond normal payment habits — would anyone actually be dumb enough to buy this? Second, I bought it (the dumb one), then used o1 Pro, and discovered it's genuinely amazing.
I often think through things with AI or ChatGPT — how to approach projects, how to plan things. When I converse with 4o, basically I say something and it just fills in along my lines, sometimes sloppily, and I have to correct it many times. But with o1 Pro, it can break down everything I need to do very clearly in a single conversation. This saves me an hour of repeated revision, making it feel totally worth it.
🚥 Koji
I saw another take: the $200 Pro membership is worth it because it's somewhat like an unlimited, 7x24 always-online "Her" — like that sci-fi film — because you can start unlimited real-time voice conversations with 4o anytime.
Master Cang, did you buy this membership on day one to try it out?
👦🏼 Guizang
I didn't buy the $200 membership on day one. At the time I genuinely thought only suckers would buy it.
For o1 Pro, I saw they used a lot of reasoning approaches in testing. I think this was also a marketing problem — the cases they chose, of course to test intelligence, using reasoning for math and physics, that's fine.
But you need to intersperse some cases that ordinary users would actually use, to experience how powerful it is. They missed this, leading to my perception being: okay, your physics and math are strong, but useless to me, because I don't know how much improvement it actually has in real open-domain intelligence.
But later I subscribed for Sora. Only after using it did I discover that for open-domain questions, as Dacong said, it gives very comprehensive and novel perspectives when discussing problems, and the answers are very structured, so this is indeed quite worth it.
🚥 Koji
Can you share a specific example? What did you use o1 Pro for?
👦🏼 Guizang
The first time I tried it yesterday, I wanted to write an annual summary of me and AI. Because I had so much to say, I wanted it to give me an outline or some directions to write about. The directions it gave were very worth referencing. We know the problem when writing: when you go to 4o or Claude, as Dacong mentioned, it either repeats what you said, or states the obvious, or gives content completely unrelated to your profession and experience.
But o1 Pro doesn't do this. It really gave very constructive suggestions, and you could completely follow its outline to write step by step. This is very impressive, but this impressiveness is a very subjective result — you can't describe in words how impressive it is. Only when you see its response yourself do you think: this is exactly what I wanted.
👦🏻 Dacong
Let me add one more detail here. As mentioned, if you're a Pro member, you get unlimited access to advanced voice mode. If you use it via API, advanced voice mode consumes about $50 per hour on average.
If you particularly enjoy chatting with AI, you only need to chat for 4 hours to make back that $200.
🚥 Koji
Honestly, chatting with 4o really does feel like talking to a real person sometimes.
👦🏼 Guizang
I think 4o's problems are mainly two: one, the response isn't fast enough, and two, it's too expensive. And once you turn it on, the phone gets very hot — there might be issues with its implementation. By comparison, I have none of these burdens with Google's Gemini.
I sometimes feel burdened talking with 4o — partly because it's expensive, partly because it feels very heavy. But chatting with Gemini has no such burden. Although it only speaks English for now, I can chat very casually, and its response is much faster than ours, probably because the model is smaller. This is also something I found very strong in my usage.
🚥 Koji
Actually during these 12 days, Gemini also released 2.0. Although it definitely didn't get the PR attention that OpenAI did, I feel its word-of-mouth is very good. We'll also share our experiences using Gemini 2.0 with everyone later.
Alright, let's look at day two. Day two's release was a reinforcement fine-tuning based on o1, called RFT. Can you introduce what RFT is?
👦🏻 Dacong
For example, if you want GPT-4o to speak very concisely, but it can't do so on its own, you need to fine-tune it, giving it many examples to learn from on top of its existing foundation.
o1 actually doesn't fully fall into the traditional large language model category — it's a polymer combining a large model with Agent, except it built the Agent part inside the large model itself, enabling autonomous reflection.
Traditional fine-tuning no longer applies. If you want o1's output to have certain tendencies, whether in thinking style or output format, you need a new kind of fine-tuning. Hence RFT, a variant of the original FT. Its target shifted from the original large model to this Agent-form large model that o1 represents.
🚥 Koji
Understood. So this release also didn't attract particularly much attention that day. Because the experience it brings to C-end users isn't that direct.
👦🏻 Dacong
It's not just indirect for C-end users — it's indirect for B-end and developers too. Because o1 is too expensive; normally you wouldn't put it in your product, the costs don't work out. And fine-tuning costs even more than using o1 directly. So for projects, the vast majority of cases won't consider using it for now.
But from another perspective, we know models keep getting cheaper. If its cost drops to a more accessible level, and you have similar needs, I believe many developers would fine-tune it.
Sora Officially Launches: Breakthroughs and Shortcomings in Video Generation
🚥 Koji
By day three, 12 hours before the launch, rumors were flying everywhere that Sora would officially release that night. Many people did stay up waiting. When Sora finally launched, the reception was mixed, with skeptical voices gradually growing louder.
Master Cang, you mentioned that when o1 and o1 Pro launched, you didn't pay the $200 for membership. But Sora made you pay for membership. Can you share your experience using Sora after subscribing?
👦🏼 Guizang
If you're a Plus member ($20), you can only generate videos up to 720P, and you're limited to generating about a dozen videos before running out of quota. If you want to use it for serious video creation, you must upgrade to the $200 membership. So I ultimately chose to pay.
After subscribing, I found it has two main aspects. On one hand, its features are indeed very refined. For example, the storyboard feature allows you to input multiple videos in sequence, and it will use first/last frames or other methods to do transitions, connecting all clips into a complete video. This is indeed very well done in terms of interaction and functionality.
Speaking of the model itself, let's first look at its foundational capabilities. Taking text-to-video as an example, at its best quality it is indeed excellent, but this high-quality output is very limited. You could say it's only slightly better than current best video models, reaching top-tier standard.
The training process for video models is actually similar to language models — you start with a text-to-video foundation, then fine-tune on images for image-to-video generation. But when it comes to image-to-video, their fine-tuning is clearly insufficient. It feels rushed. If they had trained it properly, it wouldn't be in this state. The most basic requirement we have for image-to-video is simply that it moves, regardless of quality. But the reality is that 90% of the time, you input an image, wait several minutes, spend a few dollars worth of credits, and the output is still a static image.
This is incredibly frustrating. It's no longer a matter of service quality or model performance — it's predatory business conduct. You're advertising functionality that's completely unusable while charging exorbitant fees. That's fundamentally defrauding users.
🚥 Koji
Wow, that's a scathing critique.
👦🏼 Guizang
Yes, it's genuinely a matter of integrity. You paid 1,500 RMB for a monthly membership to use this feature, but in practice it's almost entirely non-functional.
🚥 Koji
Dacongming, anything to add?
👦🏻 Dacongming
While I'm not a professional video creator, the infinite loop and storyboard features were genuinely pleasant surprises for me.
🚥 Koji
Speaking of this Sora release, there's another notable detail. When we recorded a podcast with Monica Founder Red Xiao a few days ago, he mentioned that Sora launched without an API, which is quite rare in OpenAI's history.
This seems to indicate that for OpenAI this year, developing applications has taken precedence over providing APIs.
👦🏼 Guizang
I think there are two core objectives: acquiring data and capturing market share — occupying user mindshare.
So for other companies, developing applications has always been the most important direction. Because we all know that simply releasing APIs and selling tokens creates no moat and generates no scale effects. You have to build products, retain users through distinctive features, expand your user base, and make users dependent on your product. That's the correct path.
Canvas vs. Artifact: Two Different Product Design Philosophies
🚥 Koji
Speaking of how major model vendors need to develop applications and add features to increase user stickiness, we can talk about the Canvas feature released on day four. Though worth noting, Claude launched Artifact six months prior, which received considerable praise and genuinely boosted productivity.
On this point, could you two explain what Canvas is, and if possible, compare it with Artifact?
👦🏻 Dacongming
Let me explain. Starting with Artifact: when the model generates HTML or JS-enabled frontend code, it can render that page directly within the Claude interface, letting you see the results in real time. Similarly, if it generates Markdown content, that too can be rendered and previewed directly in the browser. This is extremely helpful for checking frontend code output.
ChatGPT's Canvas feature originally evolved from its code interpreter functionality. For example, if you ask it to write an algorithm for the chicken-and-rabbit cage problem, it not only displays the code in a code block but can also run it and show results — essentially running a Python server in the background. This capability was later expanded beyond code execution to display various text content and support editing and modification of that text.
🚥 Koji
I saw an interesting use case online: a user asked ChatGPT to annotate his paper, specifically requesting the style of a philosophy professor. The final presentation in Canvas resembled Word's comment format — the original text displayed in the main area, with annotations in a sidebar accurately pointing to specific passages being commented on.
This was genuinely surprising. Compared to simply asking a large model to revise an article, this interaction represents a significant leap in experience.
👦🏻 Dacongming
OpenAI has indeed quietly released many features recently — no launch events, no press coverage. That's interesting in itself.
For instance, the article annotation feature you just mentioned actually builds on OpenAI's quietly launched Predicted API last month (also called prediction mode). This lets you input an article, specify revision requirements, and it quickly marks places needing modification or correction with suggested changes.
I suspect the annotation feature in Canvas likely employs this technology that's been live but not officially announced.
🚥 Koji
Right, this is actually quite useful. I'd been using Notion AI, asking it to directly make changes within Notion. But it would just change things — unlike when you ask a colleague or lawyer to review a document, where they preserve revision history and you decide whether to accept or reject each point. Now OpenAI can do this too.
👦🏻 Dacongming
And there's another interesting point here: because it's revision rather than rewriting, it can quickly process lengthy content while preserving your main structure. Beyond article revision, this is also extremely useful for code modification.
Often when you ask it to modify code, your code interacts with other legacy code, and once you alter the structure, things can get very messy. If it only modifies certain parameters while handling the relationships between them, that becomes highly practical. This is another application of predictive output.
🚥 Koji
Master Cang, anything to add?
👦🏼 Guizang
The person who developed this feature shared some thoughts on October 4th, discussing her core thinking about it.
She mentioned two key points: first, minimize the need for users to think about when to trigger features or which ones to use — let the AI decide. This is a presentational approach, displaying content that doesn't fit well in conversation through friendlier means, such as long texts, copy, and external renders.
The Canvas author's vision differs — she wants to build the ultimate interface for AGI. In her imagination, this ultimate interface is a blank canvas that users can freely adjust.
Her core philosophy is to create a creative partner that assists and guides creation. This explains why the annotation feature mentioned earlier matters so much — it perfectly aligns with the creative partner positioning. We can reference real-world colleague collaboration: colleagues comment on your work, offer suggestions, and you choose to accept or reject. The same applies in code review, with annotations and comments for you to decide on adoption.
It's fundamentally designed as a creative partner, which is a fundamentally different approach from the presentational scheme mentioned earlier, and thus gives rise to many different features. For instance, Canvas is actually quite heavy, with many features simulating what a creative partner should do with content. Artifact's vision is simpler: provide more appropriate display formats for content that doesn't show well in conversation. This is the core reason for the functional design differences between the two.
Real-Time Video and Project Management: Innovation and Evolution in AI Interaction Methods
🚥 Koji
I think this actually reflects different product philosophies. Speaking of which, one very exciting prospect for 2025 is that beyond traditional chatbot interactions, we'll see many new interaction methods emerging. This innovation is already sprouting everywhere — not just in AI coding with Cursor, not just the agent mode Davin brought, but also Canvas, Recraft's text-to-image and image-to-image, and Recraft's infinite whiteboard for image editing. There are so many product innovations they're becoming hard to count.
When recording the podcast with Monica's Red Xiao last week, he mentioned that 2024 felt a bit boring, because it still felt like linear extrapolation of the chatbot interaction form from ChatGPT 3.0's launch. But one reason 2025 is particularly worth anticipating is that various user experience methods for interacting with AI are springing up like mushrooms after rain.
By day five, it felt like a PR event for Apple — essentially a press release announcing you can use ChatGPT on Apple devices. Nothing particularly noteworthy there.
Day six was about 4o's real-time video calling and video understanding capabilities, including calls with Santa Claus. This caused some ripples on social media, as many creators used ChatGPT to chat and joke with Santa.
What were your reactions or thoughts after seeing the day six release?
👦🏼 Guizang
Advanced real-time voice is an extremely polished feature, and the most effective way to make ordinary users feel AI's intelligence.
Average users may not understand what o1 or o3 mean, thinking "I don't need that." But for real-time voice calls, ordinary users feel "this is genuinely amazing" because it simulates scenarios that only existed in science fiction films. So on Xiaohongshu or Douyin, any content using real-time voice — whether it's graduate students using it to identify chemical reagents and guide experiment preparation, or people "dating" GPT — easily resonates with ordinary users. It really hits home.
🚥 Koji
Right, including practicing spoken language, mock interviews — these functions have become quite practical. I tried it myself too, because Gemini 2.0 launched similar functionality around the same time. You can turn on the camera, hold something up and ask "what's this," and the recognition rate is quite accurate. I even pointed at a poster on my wall and asked, "This is a film festival poster, can you tell me which year and which film festival?" And it could make reasonable guesses.
👦🏻 Dacongming
Let me add some information. The two main selling points released on this day were video calling and screen sharing.
Starting with video calling: if we look back at OpenAI's external investments and partnerships over the past year, we'll find the company has been involved in many offline and hardware-related scenarios. If ChatGPT can smoothly teach you to brew coffee or conduct chemistry experiments, this capability can be migrated to those hardware products they previously invested in, becoming a genuinely "killer" feature. We'll find that the skill sets and technical approaches are consistent throughout.
Take chemistry experiments: currently you point a camera at chemical equipment. If this camera and GPT were built directly into chemistry instruments, then combined with robotic arms, it could become an automated workflow.
Then there's screen sharing. You might remember that last year Microsoft launched a brand called Copilot. One interesting feature was that you could have a back-and-forth conversation with your computer, letting it autonomously complete tasks. This required feeding page information to the assistant, a feature that reportedly may have been shelved. But in this ChatGPT release, it can monitor information from other apps. I'm not sure how deep this monitoring goes, but it's likely through a partnership with Apple that allows access to fairly deep-level information. On mobile devices, this becomes an additional assistive tool. For example, Hearthstone players could ask for advice on their next move while playing.
Later on Day 11, they're also releasing a desktop client feature along similar lines. It can understand what's on your screen — whether you're programming, gaming, or chatting — and theoretically offer reply suggestions and guidance.
🚥 Koji
We'll get to this later, but on Day 11 they're also releasing a desktop client feature along similar lines. It can read and understand what's on your screen — whether you're programming, what game you're playing, or even conversations you're having with others — and give you guidance on how to respond, all theoretically possible.
👦🏻 Dacong
This approach, frankly, kills off a lot of Copilot products.
🚥 Koji
This brings to mind the classic dilemma of AI entrepreneurship: when a large model company like OpenAI releases a new feature, you feel excitement and despair simultaneously.
Day 7 brought the Projects feature. You can put all kinds of files from a project into one folder and have conversations with that folder. This gives the model a knowledge base and context for better responses. Claude actually had this feature six months ago; OpenAI is just catching up now.
After this launched, did you two see any interesting use cases?
👦🏼 Guizang
I may not know the specific pre-training or model training details. But this feature shares a characteristic with the artifact feature mentioned earlier: during inference, or during model training, we need to analyze and categorize corpora, identify high-quality data, and then use this synthetic data for retraining.
The core issue here is that much content is open-ended, and the conversational value of language model outputs is difficult to verify. If you want to use it for retraining, there can be problems. These two features address this to some extent.
For example, with Projects — all the files I put in a project and the conversations around them are basically on one topic. If many people are conversing, we can filter through other data screening methods. This solves the problem of categorizing quality conversations, while also附带 some real-world, potentially non-synthetic data. This is very helpful for model training or data collection.
The same applies to artifacts. Claude's artifact feature is essentially about sharing. I only need to analyze sharing counts and click-through rates to judge the quality of large language model-generated code, which positively correlates with code quality or conversation quality. Then at the code level or long-text level, I can screen and select these as corpora, reducing screening costs. This has very positive effects for model training itself or data collection. We'll likely see increasingly more of this design in many other excellent AI projects.
🚥 Koji
I saw a great official example — putting a job seeker's various resumes, social media links, and other materials into the same project. This way, the model can better understand who you are. With this information, you can have OpenAI provide career advice or conduct mock interviews with you.
👦🏻 Dacong
At the end of last year, OpenAI updated its privacy policy, stating that as a ChatGPT user, all your interactions with OpenAI — whether in ChatGPT or social media interactions with ChatGPT — may be used by OpenAI as training data. The subsequent release of GPTs (what we called OpenAI agents at the time) also reflected this.
As Teacher Guizang said, this makes it more convenient for users to use GPT, while in the process of enjoying that convenience, they're also doing data labeling for OpenAI. This is a very clever approach that doesn't provoke much backlash.
🚥 Koji
Everyone's chasing the data flywheel. When tool applications have no moat and it's difficult to build a social flywheel, how to increase user stickiness becomes crucial.
On Day 8, ChatGPT fully opened up its search feature, with multiple optimizations to the search interface and experience. How did it feel to use?
👦🏼 Guizang
I don't have much of an impression of ChatGPT's search feature (laughs). Its search quality and results aren't particularly outstanding among mobile AI search products. If I had other options, I'd just use Google directly.
Day 9's Key Update: Structured Outputs and a Major API Breakthrough
🚥 Koji
Let's jump to Day 9. This day OpenAI released various APIs for developers. Dacong specifically highlighted this at the beginning. Dacong, please tell us what was released on Day 9 and why you think it's so important.
👦🏻 Dacong
Overall, from the official announcement, they released:
- OpenAI's official API (previously in preview).
- Realtime API (advanced voice interaction API) with price cuts and an SDK, so you don't have to write compatibility code yourself.
- A new fine-tuning method called "preference fine-tuning."
Why is this important? In 2023 we had agents; this year Bloomberg predicted AI agents would explode, and we're gradually sensing this, including Coze's growth. Behind this agent growth is an important technical innovation — structured outputs. For example, telling your home light to dim to half brightness: the light can only receive structured information in a format like JSON (e.g., "Light 19, brightness 50%"). AI can serve as a translator here.
Last year with GPT-4's 0613 version, there was still no official standard method for structured outputs. Using prompt engineering to achieve structured outputs, the success rate for adjusting a light from 78% to 50% was only 35.9%. This April that success rate improved to 75.3%, and in May reached 86.4%.
The August 6 update brought a standard structured output interface, achieving 100% success rate in strict mode. This is why after August 6, we saw agent tools like the Cursor agent version springing up everywhere.
o1 is a powerful reasoning tool. If you want its output to control machinery or IoT, you need structured outputs. Before the Day 9 release, o1 didn't have structured output functionality, or required prompt engineering to achieve it, which was unstable. Now it supports standard structured outputs, enabling 100% application of high-quality reasoning to device control.
Realtime API also now supports structured outputs. o1 needs longer reasoning time, but many scenarios (like turning off lights) don't require that. Realtime API latency is under 300 milliseconds — the light turns off within 0.3 seconds of speaking. Additionally, Realtime API costs $50 per hour, meaning when productizing, you need to find application scenarios that can earn more than $200 per hour.
Online scenarios earning over $200 per hour, and only through voice chat — that's genuinely hard to imagine. However, in Realtime API, they've distilled a mini model, dropping costs to $5 per hour. While you can't find products earning $200 per hour, scenarios earning $20 per hour do exist — for example, providing online homework tutoring for overseas students. This is precisely why Realtime API now has commercial viability.
The newly released SDK is also important. Not all developers are skilled at handling voice models, especially since the previous WebSocket solution wasn't familiar to many. With the new SDK, you can call the model directly, and it supports the WebRTC solution that many are familiar with, making commercial use of Realtime API much easier.
This update also hides an unannounced feature. Previously when we said "end-to-end model," we meant voice-to-voice without text in between. This update brings a "multi-to-multi" model. It can simultaneously receive your file information, text, voice, video, and other multimodal inputs, while outputs can include text, function calls, and voice. Interestingly, the text and voice outputs are related but not necessarily identical, meaning they're not constructed sequentially but simultaneously.
For example, if I ask the AI "why do three monks have no water to drink," it can do three things simultaneously: show an animation, point with a mouse at the eldest monk and say "this is the eldest monk, he doesn't want to carry water and wants the youngest to do it," then point at the youngest monk and say "this is the youngest monk, he doesn't want to carry water and wants the eldest to do it," while also narrating the story's background.
Before the Day 9 release, this kind of interaction was impossible. The official presentation didn't go into detail about this, but if you read the documentation carefully, you'll find this is actually the core of the Day 9 release.
🚥 Koji
In reviewing these 12 days of content, a reminder: OpenAI is very good at marketing. Much of this 12-day content was released for marketing purposes and doesn't necessarily represent the most important technological advances and core capabilities. On the other hand, OpenAI also operates in a fiercely competitive environment, so some of the most powerful features may not have been made public. They may also be using these 12 days of releases to influence competitors' thinking and pace.
Therefore, beyond focusing on publicly released content, we should also pay attention to what's not being disclosed — you might discover some valuable insights.
👦🏻 Dacong
Another release was preference fine-tuning. Preference fine-tuning means I can define the AI's output preferences, telling it what kind of expression style I like. This is a more advanced feature — I can not only tell the AI what I like, but also what I dislike. It's somewhat like setting blacklists and whitelists. I don't need to specify in prompts "don't do this," "don't be verbose," "don't say redundant things," "use what kind of language" — I can directly fine-tune these preferences into the model, improving its stability.
These improvements working together lay the groundwork for the possibility of an agent explosion in the coming year.
🚥 Koji
So 2025 is a year very much worth looking forward to. Various industries should see agents achieving better real-world applications. Previously many applications were difficult to land, with final results not ideal enough to replace sufficient human work.
Although Day 9 was a low-key release, after Dacong's analysis, we can see its value to the entire application ecosystem is enormous.
👦🏻 Dacong
There's another interesting phenomenon here. Before optimized outputs came along, all our interactions with AI happened through chatbots — even when AI completed many tasks, the results were ultimately presented as conversation. But if it's equipped with function calling, combined with various IoT devices and other technologies, it can establish very tight connections with offline devices and the commercial world.
New Real-Time Interaction Experiences: Voice Calls and Screen Reading
🚥 Koji
Day 9 was a very hardcore day. Day 10, on the other hand, became a very fun day — ChatGPT launched a phone service, releasing an 800 number for users to call. However, this service only offered 15 minutes of experience time, letting users get a simple taste of what it feels like to talk with AI.
What was released on Day 11 was actually a feature that had already been live for some time, not something new — the ChatGPT desktop app can read the screen content of other applications and interact with users based on that content.
👦🏼 Guizang
Due to specific issues for mainland users, I haven't been able to experience this feature, and I've been trying to avoid using the client. But I have a question — since I haven't used it, I'm not sure whether it reads the entire screen or only specific content. For example, when using Xcode or VS Code, does it read all the content in the entire editor window, or only the selected portion on screen? The significance of these two approaches differs greatly.
🚥 Koji
My understanding is that it can read the content.
👦🏻 Dacong
It can read information at three levels:
- First, screenshot content — this it can definitely read.
- Second, it can directly read content within software.
- Third, during the reading process, it pays extra attention to sections highlighted or selected by the user with their mouse, as well as the context around these selections. This means it not only knows what the user selected, but can also understand the complete context of the selected content.
In ChatGPT's Mac client, when you hover your mouse over the banner, you can see exactly what content will be sent to ChatGPT. For example, when writing code, if you select a certain part, when you hover over it, you can see it will send information about specific files in VS Code, while also marking which information needs priority attention. These permission request details can be viewed.
🚥 Koji
Moving to Day 12, which was the first part we discussed at the opening — the stunning release of o3. Looking back at these 12 days of content, there are two aspects most worth anticipating:
- First, the release of o3. Currently o3 is still in internal testing, you can apply but the pass rate is low. It's predicted that a streamlined o3 mini might launch around January next year.
- Second, the series of APIs released for developers, which could have unimaginably significant impact on application development and the prosperity of the agent ecosystem. Engineers and entrepreneurs should pay special attention to the new opportunities this brings.
After reviewing these 12 days of content, I'd like to ask everyone present: at this launch event, were there any other details worth mentioning, or small things that most people didn't notice?
👦🏻 Dacong
There were two interesting details about this launch event:
The first was about the display arrangement — on day one of the launch, there would be one plush toy on the table or shelf, two on day two, and so on, up to twelve on the final day. It was a rather playful arrangement.
The second detail was that each release came with some thought-provoking information, such as timelines for when AGI might arrive. These pieces of information were more like悬念 left for the audience, sparking reflection and speculation.
🚥 Koji
Seems like they were leaving headlines for the media (laughs).
👦🏻 Dacong
Right, it's quite interesting how OpenAI creates viral moments through this approach. I want to share some extra information with you, though I'm not sure how much you already know — you can judge for yourself. I deliberately held back some content that looked like internal documents.
🚥 Koji
I think one small detail worth noting is the importance of Chinese people within OpenAI. At the o3 launch, a new Chinese researcher emerged: Hongyu Ren, a Peking University alumnus. Additionally, it's rumored that on the o1 mini project, there were three main Chinese leads — besides Hongyu Ren, there were Kevin and Jiahui.
Master Guizang, did you notice any other details worth adding?
👦🏼 Guizang
Chinese people stood out prominently during these 12 days. I feel like the proportion of Chinese people even exceeded the sum of white people and other ethnicities combined. This is indeed a notable change at OpenAI right now.
Also, yesterday I saw someone raise an interesting question: why does it seem that Indians' insights in the AI field are less prominent? More precisely, it's not about quantity, but that there's almost no Indian voice at all.
👦🏻 Dacong
Some time ago I attended an offline OpenAI event in Singapore, where I met Mark from the launch event in person, as well as many new and old friends from OpenAI. In conversation, I discussed with some people a question: who might be OpenAI's strongest competitor? I originally thought it would be Claude, because domestically everyone was saying Claude had beaten OpenAI, but I got a surprising answer.
The answer was Google.
While this doesn't represent OpenAI's view, why Google? Mainly for two reasons:
First, every model has its lifecycle, and whether training costs can be recovered within six months to a year is a crucial question. To recover costs, you need enough customers willing to pay. Google has its own full suite of office software, with a deeply integrated ecosystem — they don't worry about distribution.
Meanwhile, Claude is currently tied to Amazon Cloud, which primarily provides cloud services and struggles to quickly expand its market. Therefore, even if Claude develops faster, it may not recover costs in time. By comparison, both Google and OpenAI have this capability.
2024 AI Breakthroughs and 2025 Outlook: The Evolution from Tool to Agent
🚥 Koji
This reminds me of how different perspectives collide. When Guangmi was recently asked which of the seven giants he favored most, he mentioned Amazon. Because he believes the collaboration between Anthropic and Amazon is very healthy. You can also see from Amazon's financial reports that AI-driven revenue achieved 100% growth. The combination of Anthropic plus AWS cloud services creates strong synergies. So Amazon's future development is quite promising. Overall, 2025 may see many dramatic shifts and exciting developments.
Since this episode is likely our last of the year, I'd particularly like to ask you two: looking back at 2024 from the end of the year, what were the most impressive AI technological or product breakthroughs for you? Let's have Master Guizang answer first.
👦🏼 Guizang
I think the two most important breakthroughs were: first, Claude 3.5's code capability breakthrough, especially for frontend code; and second, Sora's release and the combination of multimodal input and output.
Thanks to OpenAI sharing detailed architecture details when releasing Sora, letting us see the development path, this led to a series of subsequent advances, including FLUX and other image models, as well as better video models like Hailuo AI and Pika. Additionally, multimodal output made it possible to produce multimodal content like video and audio at the Agent level.
Combined, these two breakthroughs foreshadow that next year we'll see more automated content generation. AI products have previously been constrained by their tool-like nature, making it difficult to build moats or reach ordinary people. Next year may bring major changes — in content production, more ordinary people will be able to enjoy personalized content produced by AI.
Regarding Claude's second breakthrough in code capability, particularly the breakthrough progress in frontend code capabilities. For example, when we mentioned Cursor or Davin, why they performed better after October was because of the structured output capabilities Dacong just mentioned. Also, Claude 3.5's code capabilities on metrics like GSM8k truly became practically reliable.
I had a profound realization:
A designer friend of mine who completely doesn't understand development, after I showed him tools like bolt and v0, underwent an astonishing transformation. He used to be intimidated by development, believing he could never master it. But the day after my demonstration, he showed me an app he had developed himself — a Mandarin-to-Cantonese conversion tool.
The tool was well-made and achieved all his design goals, and he truly had zero development experience. This kind of change is deeply meaningful for ordinary people or creative individuals. Next year we'll see more similar cases, just like this year there was Huashu's Kitten Fill Light, and Chunxiang Zhao's works. More cases may emerge next year, truly liberating human creativity.
🚥 Koji
As 2024 draws to a close, people's feelings about AI progress vary greatly. Some believe AI has advanced astonishingly, already able to automate more than half of daily work; others feel AI hasn't made particularly major progress, remaining stuck in the chat interface form.
I think both views have merit, but if you only use the web versions of Kimi, Doubao, ChatGPT, and Claude, you might indeed not feel much change. However, if you've used new tools like Cursor, Devin, or Recraft, I believe you can genuinely feel AI's tremendous progress over the past year.
At Crossing, we've always emphasized one key phrase: "active actors in the AI era." An important measure of active engagement is proactively trying various new tools. So I especially recommend everyone take time to experience these new tools and personally feel AI's rapid development.
Speaking of the most impressive AI breakthrough in 2024, for me it was encountering Devin at year-end. It showed me what the Agent we've been talking about should look like — for the first time, AI truly felt like a colleague-like Agent, and one with excellent intelligence quotient, emotional intelligence, management ability, and project planning capabilities. So I'm very much looking forward to next year seeing not just continued progress in AI programming, but also hoping to see similar Agent interaction patterns expand to various fields. Dacong also mentioned that behind Agent progress is the improvement in Function Calling success rates.
Speaking of which, I'd also like to ask Dacong: what was the most impressive AI breakthrough for you in 2024?
👦🏻 Dacong
My perspective leans more toward the project side. Whenever I get my hands on a new AI product, whether it's Cursor or something else, I think about which APIs it's calling, how it's calling them serially or in parallel, then deconstruct it to see what kind of wrapper it's ultimately wearing.
Actually, all the flashy AI applications we see can be broken down into combinations of a few OpenAI APIs.
At this point, to anticipate what new tricks will emerge in the coming months, a clever method is to pay attention to OpenAI API updates every week. For example, the Function Calling success rate I just mentioned jumping from 30% to 100% — what changes can that breakthrough bring? I make a habit of carefully reading through the documentation every week.
Summarizing all of OpenAI's API changes and applications this year, they revolve around one core theme: structured output. In March last year, OpenAI released its first structured output solution. It wasn't provided as an API at the time, but as a private beta — you gave OpenAI a sample file, and under specific calling conditions, it could return structured responses.
By June last year, OpenAI had identified Agent as a viable scenario to land on. They found many partners developing Agents and indicated they would further iterate on the structured output solution. On November 6 last year, OpenAI quietly released JSON Mode, signaling that structured output was becoming a mainstream priority.
This year we found that whether it's real-time interaction APIs or reasoning APIs and so on, everything has grown around structured output. Every time a product lands, it corresponds to structured output reaching a new standard.
In the current paradigm, structured output has transformed from "you give me information, I give you a file" to "you give me a pile of information, I simultaneously give you a pile of files," enabling you to process multiple tasks at once, with the success rate for each task jumping from 30% to 100%. This allows AI to handle larger and more complex interactions.
Therefore, the most impressive breakthrough for me in 2024 is:
Structured output has transformed from a clever toy into a core factor that genuinely impacts the real world, developer ecosystems, and project ecosystems. This factor, though hidden behind the scenes, remains largely unknown to the general public.
🚥 Koji
After discussing OpenAI's 12-day launch event, we also have to mention that during this period, Google dropped its own bombshell with the release of Gemini 2.0. I was genuinely stunned after using it myself. Whether it's the feedback quality of its Flash Thinking model version or the way it displays its entire reasoning process, both are deeply impressive.
The text volume of the reasoning process even exceeds that of the final answer. You can see how intelligently and seriously this agent approaches and breaks down each problem, thinking step by step to arrive at an answer. And it's fast — responding within seconds, much quicker than when o1 was released. Its multimodal experience is also remarkably smooth and seamless.
I'd like to ask you both: when experiencing Gemini 2.0, what particular impressions or information would you like to share?
👦🏼 Guizang
My profound impression is its multimodal input and output capabilities. Dacong just mentioned that OpenAI also has this feature, but there's nowhere to truly experience it with OpenAI yet. Gemini has indeed gone further in video understanding. When I tested it, I gave it a 20-minute video without subtitles, asked it to transcribe the content and organize it into an article. The model completed the entire process in one go, even automatically polishing the text and outputting the finished product directly — this capability is astonishing.
Additionally, I previously referenced Hyacinth's work to do a breakdown test: I gave it a one-minute AI-generated video, which was multiple shot clips created by an AI creator — possibly containing a dozen or so shots within that minute. It could output the start and end times of each shot, along with the specific content of each shot. This way I could quickly replicate this video based on mature DiT video models.
Using the two models in combination, I can almost replicate any video. Just click confirm, put the materials into the video generation model, and it automatically completes editing and output. If there's native audio, its built-in voice model can generate that directly too.
This represents a major leap in video content production efficiency.
🚥 Koji
That's all for today's episode. I'm very happy, and my anticipation for 2025 keeps climbing. I've gained a lot of positive conviction, believing that 2025 will bring many exciting new developments. Thank you both for sharing so many unique insights today, showing us the details, perspectives, and future implications behind the press conferences.
Thank you to both our guests. Let's look forward to 2025 together, and welcome back to "Crossing" next year.
Subscribe to the "Crossing" Podcast
🚦 We follow the industry transformations and new entrepreneurial opportunities brought by the new wave of AI technology. "Crossing" is Steve Jobs' metaphor for Apple — describing it as standing at the crossroads of technology and liberal arts, where great products are often born. AI is bringing change to all industries. We seek out, interview, and unite "active doers" of the AI era, exploring and embracing new changes and new possibilities together with them.
👦🏻 Host Koji: Co-founder of The Fair and Tangdao. I believe technology, especially AI, will fundamentally transform society and empower humanity in the future. Welcome to chat with me, bounce ideas around, and connect on the next possibility. Koji's Jike[2], Koji's website[3]
👧🏻 Host Ronghui: Works at a tech VC, former Silicon Valley correspondent for CBNweekly. Ronghui's Jike[4]
Join the "Crossing" Membership Group
☀️ First-hand AI news and insights
👫🏻 We encourage everyone to date/make friends/find future kindred spirits
🦀 Add our assistant on WeChat to join: Rwkfbcianvd, or scan the QR code below
References [1] AIGC Weekly: https://quail.ink/op7418
[2] Koji's Jike: https://okjk.co/0JSUes
[3] Koji's website: https://koji.super.site/
[4] Ronghui's Jike: https://okjk.co/0cbnYV