Code Brain | Is the GPT-4 Era Over? Global Users Put Claude 3 to the Test, and the Results Are Stunning

Significantly outperforming GPT-4. Has pure-text LLM development hit its ceiling?
Last night, Anthropic — OpenAI's biggest rival — unveiled its next-generation AI model family: Claude 3.
The lineup consists of three models, ranked from weakest to strongest: Claude 3 Haiku, Claude 3 Sonnet, and Claude 3 Opus. The most capable of the three, Opus, surpassed both GPT-4 and Gemini 1.0 Ultra on multiple benchmarks, setting new industry standards across mathematics, coding, multilingual understanding, and vision.
Anthropic claims Claude 3 Opus possesses knowledge equivalent to an undergraduate human.

With this release, Claude introduced multimodal capabilities for the first time (Opus scored 59.4% on MMMU, beating GPT-4V and matching Gemini 1.0 Ultra). Users can now upload photos, charts, documents, and other unstructured data for AI analysis and interpretation. The three models also continue Claude's signature strength — long context windows. The initial rollout supports 200K tokens, though Anthropic notes all three models can handle up to 1 million token inputs (available to select customers), roughly the length of the English editions of Moby-Dick or Harry Potter and the Deathly Hallows.
However, the most capable Claude 3 comes at a steep premium over GPT-4 Turbo: GPT-4 Turbo charges $10/$30 per million tokens for input/output, while Claude 3 Opus runs $15/$75.

Opus and Sonnet are now available on claude.ai and through the Claude API, with Haiku launching soon. Amazon Web Services also announced immediate availability of the new models on Amazon Bedrock. Here's Anthropic's official demo:
Following the announcement, researchers with early access began sharing their experiences. Some reported that Claude 3 Sonnet solved puzzles previously only cracked by GPT-4.

Others noted that in real-world use, Claude 3 hasn't completely dethroned GPT-4.

Hands-On Testing: Claude 3

URL: https://claude.ai/
Does Claude 3 truly live up to its claimed performance gains over GPT-4? Most early testers seem to think it's genuinely competitive. Here's what we found:
First, a trick question: Which month has 28 days? The actual answer is every month. Turns out Claude 3 still struggles with this type of riddle.

Next we tested areas where Claude 3 supposedly excels. According to Anthropic's official description, Claude is strong at "understanding and processing images," including extracting text from images, converting UIs to front-end code, interpreting complex equations, and transcribing handwritten notes. LLMs famously struggle to distinguish fried chicken from poodles; when we fed it an image containing both, Claude 3 responded: "This image is a collage featuring dogs and fried chicken pieces or nuggets that bear a striking resemblance to the dogs themselves..." — passing that test.

When asked how many people were in another image, Claude 3 correctly answered: "This illustration depicts seven small cartoon characters."

Claude 3 can extract text from photos, correctly handling even vertical Chinese and Japanese text:

What about internet memes? When shown an optical illusion image, GPT-4 and Claude 3 gave opposite interpretations:

Which one got it right?
Beyond image understanding, Claude's long-text processing is also notably strong. The full lineup offers 200K context windows and accepts inputs exceeding 1 million tokens. How does it perform? We fed it a recent paper from Microsoft and the University of Chinese Academy of Sciences, "The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits," and asked for a numbered summary. The complete response took roughly 15 seconds. This was using Claude 3 Sonnet; the Claude Pro version is faster but costs $20 per month.

Note that Claude currently limits uploaded documents to 10MB; exceeding this triggers a warning:

In its Claude 3 blog post, Anthropic highlighted significant coding improvements. One tester fed it raw ASCII art, and Claude handled it effortlessly:

We can now confirm that Claude 3 has stronger coding capabilities than GPT-4. Not long ago, Andrej Karpathy, who had just left OpenAI, issued a "tokenizer" challenge. Specifically, he asked an LLM to take his 2-hour-13-minute tutorial video and translate it into a book chapter or blog post format about tokenizers. Claude 3 rose to the task. Here's what Anthropic research engineer Emmanuel Ameisen shared:


Perhaps no longer conflicted by professional ties, Karpathy gave a fairly thorough and objective assessment: "Stylistically, it's actually pretty nice! If you look closely, you'll find some subtle issues/hallucinations. Either way, the near out-of-the-box system is impressive. I'm looking forward to playing with Claude 3 more, it looks like a strong model." If there's one related thing he felt compelled to say, it's that "people should be extra careful when doing eval comparisons, not only because the evals are worse than you think, but also because many evals are overfit in undefined ways, and also because the comparisons being made can be misleading. GPT-4 coding (HumanEval) is not 67%. My eye starts twitching whenever I see this comparison used as a proxy for coding."
Based on all these demanding test results, some have already declared "Anthropic is so back."
Finally, Anthropic also launched a prompt library covering multiple categories of prompts. If you want to explore Claude 3's new capabilities in depth, it's worth trying out.

Link: https://docs.anthropic.com/claude/prompt-library
The Claude 3 Model Family
The Claude 3 family consists of three versions: Claude 3 Opus, Claude 3 Sonnet, and Claude 3 Haiku.

Claude 3 Opus is the most intelligent model, supporting a 200k-token context window and achieving state-of-the-art performance on highly complex tasks. It handles open-ended prompts and novel scenarios with exceptional fluency and human-like comprehension. Opus shows us the outer limits of what generative AI can achieve.

Claude 3 Sonnet strikes an ideal balance between intelligence and speed, especially for enterprise workloads. Compared to peers in its class, it delivers strong performance at lower cost and is designed for high durability in large-scale AI deployments. Sonnet also supports a 200k-token context window.

Claude 3 Haiku is the fastest and most compact model, with near-instant responsiveness. Notably, it too supports a 200k-token context window. It answers simple queries and requests with unmatched speed, enabling developers to build seamless AI experiences that feel human-like.

Let's now take a closer look at the features and performance of the Claude 3 family.
Surpassing GPT-4 Across the Board, Setting a New SOTA for Intelligence
As the most intelligent model in the Claude 3 family, Opus outperforms competitors on most AI evaluation benchmarks, including undergraduate-level expert knowledge (MMLU), graduate-level expert reasoning (GPQA), grade-school math (GSM8K), and more. Moreover, Opus demonstrates near-human-level comprehension and fluency on complex tasks, pushing the frontier of general intelligence.
Additionally, all Claude 3 models — Opus included — show enhanced capabilities in analysis and forecasting, nuanced content creation, code generation, and non-English dialogue in languages such as Spanish, Japanese, and French.
The chart below compares Claude 3 models against competitors across multiple performance benchmarks. As shown, the strongest model, Opus, comprehensively outperforms OpenAI's GPT-4.

Near-Instant Responsiveness
Claude 3 models can power real-time customer chat, auto-completion, and data extraction tasks where responses must be immediate.
Haiku is the fastest and most cost-effective model in its intelligence tier. It can read through an arXiv paper with dense charts and figures (~10k tokens) in under three seconds.
For the vast majority of workloads, Sonnet is twice as fast as Claude 2 and Claude 2.1 while delivering higher intelligence. It excels at tasks requiring quick turnaround, such as knowledge retrieval or sales automation. Opus matches the speed of Claude 2 and 2.1 but with substantially greater intelligence.
Strong Vision Capabilities
Claude 3 offers sophisticated vision capabilities on par with other leading models. They can process a wide range of visual formats, including photos, charts, graphs, and technical diagrams.
Anthropic notes that some of their customers have over 50% of their knowledge bases encoded in various data formats such as PDFs, flowcharts, or presentation slides. The new models' robust vision capabilities are therefore especially valuable.

Fewer Refusals
Previous Claude models often made unnecessary refusals, suggesting a lack of contextual understanding. Anthropic has made meaningful progress in this area: compared to earlier generations, Opus, Sonnet, and Haiku are significantly less likely to refuse to answer even when user prompts approach system boundaries. As shown below, Claude 3 models demonstrate more nuanced understanding of requests, recognizing genuinely harmful prompts while refusing harmless ones far less frequently.

Improved Accuracy
To assess model accuracy, Anthropic used a large set of complex, factual questions targeting known weaknesses in current models. They categorized responses as correct, incorrect (or hallucinated), and uncertain — where the model admits it doesn't know rather than providing false information. Compared to Claude 2.1, Opus doubled its accuracy (correct answers) on these challenging open-ended questions while also reducing incorrect responses. Beyond producing more trustworthy outputs, Anthropic will also enable citations in Claude 3 models, allowing them to point to precise sentences in reference materials to substantiate their answers.

Long Context and Near-Perfect Recall
The Claude 3 family launched with a 200K context window, though Anthropic noted that all three models are capable of accepting inputs exceeding 1 million tokens — a capability offered to select users requiring enhanced processing power. To handle long-context prompts effectively, a model needs strong recall. The Needle In A Haystack (NIAH) evaluation measures how accurately a model can recall information from vast amounts of data. Anthropic strengthened this benchmark's robustness by testing across diverse crowdsourced document corpora, using 30 random needle/question pairs per prompt. Claude 3 Opus not only achieved near-perfect recall with over 99% accuracy — in some cases, it even identified limitations in the evaluation itself, recognizing that the "needle" sentences appeared artificially inserted into the original text.

Safety and Ease of Use
Anthropic stated that it has established dedicated teams to track and mitigate safety risks. The company is also developing approaches like Constitutional AI to improve model safety and transparency, while addressing privacy concerns that new modalities might raise. While the Claude 3 family showed advancement on key metrics for biological knowledge, cyber-related knowledge, and autonomy compared to previous models, the research places the new models within AI Safety Level 2 (ASL-2).
In terms of user experience, Claude 3 is better than previous models at following complex, multi-step instructions and adhering to brand and response guidelines, enabling more trustworthy application development. Additionally, Anthropic said Claude 3 models are now better at generating popular structured outputs in formats like JSON, making it easier to guide Claude for use cases such as natural language classification and sentiment analysis. 4
What's in the Technical Report
Anthropic has released a 42-page technical report, The Claude 3 Model Family: Opus, Sonnet, Haiku.

Report: https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_Card_Claude_3.pdf
The report details the Claude 3 family's training data, evaluation criteria, and more granular experimental results.
On training data: the Claude 3 models were trained on a proprietary mix of publicly available internet data up to August 2023, alongside non-public data from third parties, data labeling services and paid contractors, and internal Anthropic data.
The Claude 3 family underwent extensive evaluation across multiple metrics including:
- Reasoning capabilities
- Multilingual capabilities
- Long context
- Reliability / factuality
- Multimodal capabilities
First, on reasoning, coding, and Q&A tasks: the Claude 3 family was benchmarked against competitors across industry-standard tests for reasoning, reading comprehension, math, science, and programming. The results show they not only surpassed previous Claude models but also achieved new state-of-the-art results in most cases.

Anthropic evaluated the Claude 3 family on the LSAT, Multistate Bar Examination (MBE), 2023 American Invitational Mathematics Examination, and GRE General Test. Specific results are shown in Table 2 below.

The Claude 3 family features multimodal capabilities (image and video frame input) and has made significant progress on complex multimodal reasoning challenges beyond simple text understanding. One telling example: Claude 3's performance on the AI2D science diagram benchmark, a visual Q&A evaluation involving chart parsing and answering corresponding multiple-choice questions. Claude 3 Sonnet reached SOTA at 89.2% in a 0-shot setting, followed by Claude 3 Opus (88.3%) and Claude 3 Haiku (80.6%). Specific results are shown in Table 3 below.

University of Edinburgh PhD student Yao Fu offered his own analysis of the technical report. First, in his view, the evaluated models showed virtually no differentiation on metrics like MMLU, GSM8K, and HumanEval — what really matters is why even the best models still have a 5% error rate on GSM8K.

He believes that what truly distinguishes models are MATH and GPQA — these super-challenging problems are what AI models should target next.

Compared to previous Claude models, the areas showing the biggest improvements are finance and medicine.

On vision, Claude 3's demonstrated OCR capabilities suggest enormous potential for data collection.

He also identified several other trends:


Based on current benchmarks and hands-on experience, Claude 3 has made substantial strides in intelligence, multimodal capabilities, and speed. As the new family continues to be optimized and deployed, we may see an increasingly diversified large model ecosystem.
Blog: https://www.anthropic.com/news/claude-3-family
Source: Synced Original: https://mp.weixin.qq.com/s/zNX_7JoE9XRyAg_GCy85nA
References: https://www.cnbc.com/2024/03/04/google-backed-anthropic-debuts-claude-3-its-most-powerful-chatbot-yet.html https://www.aboutamazon.com/news/aws/amazon-bedrock-anthropic-ai-claude-3
