Product

GPT-4 Vision

GPT-4V

GPT-4 Vision is OpenAI's multimodal large language model that processes visual input alongside text . First introduced as "GPT-4V" in late 2023, it represented the industry's shift from simple modality stacking toward genuine multimodal fusion, accepting images as inputs through Chat Completions API to generate captions, analyze real-world images, and read documents containing graphics .

The model's architecture follows what MiniMax's Junjie Yan described as the dominant industry pattern: a massive MoE-based transformer backbone with vision capabilities aligned to it, enabling strong visual understanding without native integration of other modalities like audio or video generation . Pricing for visual inputs scaled with image size—for instance, a 1080×1080 pixel image cost $0.00765 to process at launch .

GPT-4 Vision was subsequently superseded by GPT-4o in May 2024, which OpenAI positioned as a milestone for natively handling text, audio, and image inputs together , and later by the GPT-4.1 series in April 2025 with expanded context windows and improved coding performance .

AI-generated — may contain errors, please verify.

GPT-4 VisionProduct
GPT-4V
渲染中…
Mentioned in 4 articles

Coverage