3D Foundation Model Startup VAST Closes Hundreds of Millions of RMB Funding Round, Using Scaling Law to Build Next-Generation 3D Content Platform | Led by Oasis Capital

Counselor Vitality (Note: This appears to be a standalone phrase or title fragment. "参赞" can mean "counselor" (diplomatic title) or "participate in planning/advising"; "生命力" means "vitality" or "life force." Without additional context, this is the most direct translation. If this refers to a specific column or section title in a publication, it might alternatively be rendered as "The Vitality of Counselors" or "Counselor on Vitality" depending on intended meaning.)

VAST announced today that it has completed angel and Pre-A funding rounds totaling several hundred million RMB. The angel round was led exclusively by Oasis Capital, while the Pre-A round was led by Fortune Venture Capital and Primavera Venture Partners, with Inno Angel Fund and Shuimu Tsinghua Alumni Fund participating. The financing represents the largest amount raised in the 3D large model sector to date.

An investment lead at Oasis Capital commented: "In 2023, the VAST team already recognized the importance of 3D and committed firmly to building in this direction. They launched the world's first 3D generation model in 2023 and have continuously iterated, recently releasing the foundational 3D model Tripo 2.0, which opens a door to free creation. As VAST's angel investor, Oasis has witnessed the company's growth and transformation, as well as the team's endless creativity. The evolution of technology is full of challenges and requires long-term dedication and sustained investment. Oasis believes in the vitality of innovative technology and will continue to support the team as it grows."

The new 3D large model Tripo 2.0 is also officially launching today.

About VAST

Founded in March 2023, VAST is an AI company dedicated to developing general-purpose 3D large models. Its goal is to build consumer-grade 3D content creation tools and establish a 3D UGC platform, making 3D-based spaces a key element for user experience, content expression, and enhancing new quality productive forces.

In early 2024, VAST launched Tripo 1.0, a 3D large model with billions of parameters capable of generating 3D mesh models from images or text in 8 seconds. Since launch, users worldwide have generated over 5 million 3D models.

In March 2024, VAST partnered with leading global open-source community Stability AI to release the open-source 3D foundation model TripoSR. The model achieved top-tier performance of 0.5 seconds for single-image-to-3D generation and remains a popular project in the open-source 3D generation community.

Today, VAST introduces Tripo 2.0, validating the Scaling Law for 3D large models and pushing 3D generation to its next milestone.

Tripo 2.0 employs a hybrid architecture combining DiT and U-Net models, learning to capture geometric and material distributions in large-scale data to better ensure detail in 3D model geometry and output quality for materials.

Tripo 2.0 can generate shape geometry in 10 seconds and textures with PBR in another 10 seconds, setting a new standard for general model performance in 3D generation.

To our knowledge, Tripo leads globally in performance across all 3D generation tasks.

GPTEval3D: MLLM-based evaluation metrics (3D generation shape, texture quality, detail expression, input condition adherence, output diversity), designed to assess the semantic accuracy and quality of generated 3D content

Understanding Individual Objects Is the Beginning of Understanding the World

For users creating 3D content, text input offers the possibility of "speak it into existence, create worlds," while image input provides greater controllability during the creative process.

Unlike traditional 3D reconstruction applications, most pure creative concepts may exceed real-world physical constraints. Much 3D content in games, design projects, or virtual scenes has no physical counterpart in reality. Some environments are too extreme — even with substantial investment in advanced scanning equipment, they cannot be scanned, repaired, or reconstructed. Therefore, a 3D large model's ability to generate complex composite objects from text and its capacity for spatial understanding and reconstruction from single-image input become the most important criteria in any evaluation framework.

For Tripo, this means possessing the following capabilities:

First, precise language understanding — accurately reflecting user text input intentions in an object's geometric structure and compositional details, including spatial relationships between parts described in the text;

Second, deep and accurate spatial reconstruction ability — ensuring accurate inference of three-dimensional structure and depth information from a single image at any viewpoint, precisely reconstructing complex object geometry and texture details while maintaining overall consistency;

Third, understanding of physical laws and common sense — ensuring generated content aligns with user intent while maintaining logical consistency within basic physical constraints, striking a balance between creative freedom and physical plausibility.

This is Tripo's answer: seeing the big picture in small details, exploring the hidden sides of the world.

Examples include "a leaf spirit with teeth holding a leaf," "a basket with tomatoes, lettuce, and carrots," and "a flamingo standing on a glass sphere on water":

Results generated directly from www.tripo3d.ai; all are six-view renders of AI-generated 3D models

Take this jade fabric flower image as another example: Is the left flower cluster attached to or separate from the main bouquet? What is the overlapping relationship between leaves? What does the back of the bouquet look like?

Or this ship: What is the mast structure? How is the cabin designed?

Beyond refined text and visual input understanding, Tripo 2.0's generation results also achieve leading quality and fidelity, setting new industry standards (new state-of-the-art) in shape and texture quality and detail expression. Tripo can not only generate highly detailed and accurate 3D shapes capturing complex features and geometric structures, but also produce high-fidelity PBR (physically based rendering) materials with fine surface properties and realistic, rich visual effects.

Results generated directly from www.tripo3d.ai

Validating the Scaling Law for 3D Generation

VAST's algorithm team has been searching for the tokenizer in 3D generation, working to validate the Scaling Law in this domain.

Tripo 2.0 employs a sophisticated hybrid architecture fusing DiT and U-Net models. This fusion fully leverages the strengths of both architectures: DiT excels at capturing global context and long-range dependencies in 3D structures, while U-Net specializes in preserving fine details and local features. Combined with massive high-quality 3D data and multiple synthetic data augmentation techniques, this design not only significantly improves generation quality but also enhances model robustness, stability, and generalization capability.

On the engineering optimization side, the team improved efficiency through distillation, employing both guidance distillation and step distillation to substantially optimize performance without sacrificing quality (more algorithmic details to follow in VAST AI's subsequent technical reports).

Through over a year of exploration, the algorithm team continuously investigated the relationship between model scale and performance. Tripo 2.0 confirms that as model parameters increase and training data expands, generation quality improves predictably. By developing deep understanding of individual objects, Tripo 2.0 demonstrates reasoning capabilities from micro to macro — this ability to "see the big picture in small details" is the foundation for building complex 3D worlds.

Making Everyone a Super Creator

Tripo 2.0 can generate geometry in 10 seconds and textures with PBR materials in another 10 seconds. This means Tripo can not only reduce costs and increase efficiency in 3D industrial production pipelines, but also opens possibilities for real-time creation of more 3D content and interactive experiences.

Yachen Song, VAST's founder and CEO, stated: "We are now confident in announcing that VAST and Tripo 2.0 have reached the stage of Midjourney V4 in terms of quality. This represents a leap in user experience and enormous commercial potential. We appreciate the support from various investors, which enables us to continue exploring the future 3D ecosystem.

On the technology front, we will continue pursuing the Scaling Law for 3D generative AI, researching the fundamental principles governing the relationship between model scale, data volume, and generation quality, and seeking scalable paradigms for data, representation, and model architecture to push the boundaries of 3D generative AI. We will also explore more holistic 3D generation — not just individual assets, props, and characters, but also (dynamic) environments, motion, physics, and more. As an emerging frontier branch of large models, 3D generation is demonstrating unprecedented imagination for B2B and B2C applications across gaming, animation, film and television, 3D printing, internet and industrial product design, embodied intelligence, simulation, MR, education, spatial intelligence, and beyond. We believe that through more consumer-grade 3D creator tools and a universal 3D content platform, we can build a new and thriving generative 3D ecosystem."

Celebrating Vitality

What do you think vitality is?

Vitality is free creation, the expression of self-awareness.

—— Yachen Song, VAST Founder & CEO