Product

DeepSeek-V4-Pro

DeepSeek V4 Pro

DeepSeek-V4-Pro is a large language model released by DeepSeek, positioned as a 1.6-trillion-parameter flagship in the company's model family . In benchmark testing by 葬AI, it averaged 7.9 minutes per task — faster than Zhipu's GLM 5.3 (12.7 minutes) but slower than xAI's Grok 4.6 (5.4 minutes) . The model sits alongside DeepSeek-V4-Flash, a lighter variant that has demonstrated competitive performance against larger rivals at roughly one-third the parameter scale .

The V4 generation builds on architectural innovations first established in DeepSeek-V3, including the Mixture-of-Experts (MoE) design that activates only a fraction of total parameters per token, multi-head latent attention (MLA) for reduced memory overhead, and multi-token prediction (MTP) for faster training and generation . These engineering choices have defined DeepSeek's reputation for achieving strong performance at lower compute costs .

As of mid-August 2026, 葬AI described the V4 Pro正式版 as "神一阵鬼一阵" — capable of impressive results but inconsistent — while noting that DeepSeek's track record with the Flash variant suggested the company could potentially close the gap with leading post-trained models through further refinement .

AI-generated — may contain errors, please verify.

DeepSeek-V4-ProProduct
DeepSeek V4 Pro
No graph yet
Mentioned in 7 articles

Coverage