From 30 Million to 3 Billion: Three Rounds, Wanwu Capital | Luzhou Shengmingli

VAST recently announced the completion of its Series B and B+ funding rounds.

VAST recently announced the completion of its Series B and B+ funding rounds, totaling approximately RMB 3 billion. In less than six months, the company has raised roughly RMB 5 billion in cumulative funding, setting a new record for the 3D-native AI sector.

Oasis Capital was the sole institutional investor in VAST's angel round. From the first round of RMB 30 million to this round of RMB 3 billion, we have remained consistently bullish on the company's trajectory and significantly increased our investment in this round. As 3D-native models evolve from "generation quality" to "production readiness," VAST continues to break through on native mesh generation, industrial pipeline compatibility, and world models, integrating into real production workflows across gaming, XR, 3D printing, film, and television. This round was led by Matrix Partners China, with participation from leading strategic investors including Perfect World, BlueFocus, Singularity Power Capital, Yanqu Games, Thundersoft, Guangdian Investment, and 37 Interactive Entertainment, as well as top financial investors including CDH Investments VGC, CICC Capital, CMC Capital, Hongtai APlus, Sanzheng Health Investment, Zhongping Capital, Fujian Industrial Investment Fund, Fujian Venture Capital, Zhuoyuan Asia, and Jiangxi Financial Holding. Existing shareholders Oasis Capital, Fortune Venture Capital, Primavera Venture Partners, 4399, MUHUA Ventures, Zhence Capital, and Huakong Fund all continued to significantly increase their stakes.

VAST's strategic investors now span every industry with concrete demand for 3D-native foundation models and world models: gaming, automotive, e-commerce, smart devices, XR, film and television, cultural tourism, advertising... They sit closest to end-user demand, understand the pain points of real production pipelines most intimately, and maintain the most rigorous standards for technology deployment.

The diversified backing from strategic industry capital signals that VAST's technical capabilities have crossed the threshold of "production readiness" and are becoming the default infrastructure for 3D content production across multiple sectors.

At the same time, VAST released the latest version of its flagship 3D-native foundation model, the Tripo P series: Tripo P2.0. As the world's first 3D-native foundation model to support native quad-dominant topology, P2.0 enables 3D generation to simultaneously satisfy three industrial-grade thresholds for the first time — "engine-ready, riggable, and editable" — serving professional pipelines in gaming, film, and real-time interactive content while providing creators in the Vibe Coding era with a 3D-native foundation that can be directly programmed, rigged, and iterated upon.

Proceeds from this round will be directed toward continuous iteration of 3D-native models and data capabilities, expansion of training and inference infrastructure, acceleration of product development and commercial deployment, and advancing AI 3D from "usable" to "delightful" as the 3D-native intelligent infrastructure for all industries.

Betting on non-consensus, until it becomes consensus

Tripo P2.0: The Most Vibe Coding-Ready 3D-Native Foundation Model

In March this year, VAST launched Tripo P1.0 — a new species built from the ground up with a fundamentally reconstructed technical framework.

Now, we present the Tripo P2.0 Preview. As the industry's first 3D-native foundation model to unlock native quad-dominant topology, P2.0 can directly generate 3D objects dominated by quads with more accurate shape reconstruction, controllable polygon counts, and more sensible part separation in mere seconds, achieving unprecedented pipeline compatibility.

We were pleasantly surprised to discover that P2.0 delivers not only higher-quality 3D generation but, thanks to its exceptional pipeline compatibility, is unlocking a long-anticipated new scenario: 3D Vibe Coding.

LLM coding capabilities have already significantly lowered the barrier for developing interactive 3D applications. Today, programming agents can generate Three.js pages, write partial interaction and animation logic, and continuously debug based on runtime results — but they lack 3D objects that can be rapidly generated, understood, and manipulated.

For natural language to truly drive complex 3D applications, generated objects must do more than "look right." They must be fast enough, accurate enough, and sufficiently controllable in complexity, with clear geometric and structural organization. Polygon count is merely one performance variable; actual runtime depends on file size, materials and textures, LOD, and target device. P2.0 bridges the critical gap from "agents can write code" to "agents can create and drive 3D objects":

  • Second-level generation enables agents to rapidly obtain 3D objects within a "generate-run-check-modify" loop, supporting rapid prototyping and multi-round iteration;
  • Target polygon range generation allows agents to select appropriate geometric complexity based on performance budgets for web, mobile, or real-time applications, reducing subsequent decimation and retopology work and satisfying requirements for "low poly, clean structure, load-and-go" — helping developers control performance overhead for web and real-time applications;
  • More faithful intent reconstruction reduces repeated generation and manual correction, improving stability and predictability when agents execute tasks continuously;
  • Quad-dominant, more regular topology facilitates subsequent editing, subdivision, and animation deformation;
  • More structurally semantic natural part segmentation (Semantic Segmentation) makes the 3D model a set of individually identifiable and manipulable components, creating conditions for agents to further name parts, locate objects in code, and write component-level interaction, animation, and behavior logic.

These characteristics make Tripo P2.0 the most Vibe Coding-ready 3D-native foundation model to date.

When code intelligence combines with 3D-native intelligence, Vibe Coding evolves from merely generating a webpage with a 3D model to directly creating complex, interactive 3D applications: humans express creative intent and judge results in natural language, agents handle planning, programming, and debugging, and P2.0 provides 3D objects and structures that can be further understood and manipulated.

Yanpei Cao, VAST's Chief Scientist, was among P2.0's first users and creators. He used P2.0 to generate a mechanical spider, then leveraged natural semantic part segmentation to have an LLM name individual components, write and debug motion logic, ultimately achieving component-level animation control for the entire mechanical spider.

Nexus Underlying Reconstruction: The Leap to Native Quad-Dominant Topology

Production-ready generation is the foundational principle running through VAST's algorithm and technology R&D. Along this path, models directly generate objects' geometry, topology, and materials, and further endow them with semantic parts, skeletal rigs, and physical properties — upgrading visual objects from merely "viewable" to "usable" functional objects whose generated results require no rework before entering professional production pipelines.

Quad-dominant (Quad) topology has been a major pain point preventing 3D generation from entering downstream pipelines. In computer graphics, quads are regarded as the universal mesh standard due to their stability, controllability, and predictability. However, traditional manual retopology often takes hours or even days, procedural retopology quality is rarely directly usable, and AI retopology is generally time-consuming with insufficient stability — while "generate high-poly first, then retopologize" introduces workflow redundancy and detail loss.

This is precisely the problem Tripo P2.0 sets out to solve.

The underlying algorithm of the Tripo P series, Nexus, is a generation framework designed for 3D-native meshes: while the industry-dominant autoregressive approach decomposes 3D structures into sequences and "writes" them out one by one, Nexus decouples vertex and topology generation, using hierarchical diffusion to globally generate vertices from coarse to fine, then encoding arbitrary topology through "spacetime interval" encoding — achieving fully order-free native mesh generation based on diffusion models for the first time. The related paper was accepted by the top international graphics conference SIGGRAPH, and the team's real-time demo based on Tripo P1.0 won the highest award, Best in Show, at SIGGRAPH 2026 Real-Time Live!

This underlying paradigm reconstruction made Tripo P1.0 the industry's first large model capable of directly generating professional modeler-grade 3D models in seconds, breaking the long-standing "impossible triangle" of quality, effect, and pipeline usability in AI 3D. On third-party evaluation leaderboards with over 200,000 users participating in blind voting, Tripo P1.0 held the top position until Tripo P2.0 took the baton.

VAST's algorithm team pioneered a route encoding arbitrary topology — including quads and non-manifold structures — into continuous features, enabling diffusion models to efficiently, stably, and end-to-end generate complete meshes. Tripo P2.0 thus became the industry's first 3D-native foundation model to unlock native quad-dominant topology, with comprehensive upgrades across four dimensions:

  • Precision: P2.0 supports higher shape resolution, with significantly enhanced overall visual performance and detail fidelity, particularly noticeable improvement in detail-sensitive areas such as character faces.
  • Broader polygon range: Adapts to differentiated polygon budget needs across scenarios, with maximum mesh polygon count expanded from P1.0's 20,000 tris to 50,000 tris / 25,000 quads. From lightweight mobile assets to high-precision hero assets, all can be output on demand.
  • Local editing support: The official release will support users to regenerate specific local regions individually while keeping unselected areas unchanged, satisfying more flexible modification needs and multi-round iteration while substantially reducing trial-and-error costs for model iteration.
  • More sensible part separation: Alongside precision improvements, the natural connected component granularity of models increases synchronously, with part separation becoming more reasonable and semantically coherent.

From Objects to Worlds: Full-Stack World Model Layout

Everything VAST does ultimately points toward the same vision — to democratize optimal interactive experience with intelligence. Algorithmic and model-layer iterations over the past half-year have delivered multiple milestones toward realizing this vision.

Project Eden, the general world model research project released this June, provides the underlying engine for next-generation interactive content. Distinct from industry-standard approaches such as "action-conditioned video generation" and "static 3D scene generation," Project Eden creatively decouples underlying state simulation from visual presentation natively, naturally unlocking capabilities including long-term environmental persistence, free scene reuse, and multi-user concurrent interaction.

VAST has already demonstrated a "text-to-interactive simulation scene" generation pipeline where generated assets come with colliders, joints, and physical properties, ready to plug into Nvidia Isaac Sim and Gazebo for scaled embodied intelligence team training; the generative renderer evolved from the project also demonstrated its application in embodied intelligence simulation scenes at SIGGRAPH 2026.

What truly determines a foundation model company's long-term value is its ability to continuously transform what the industry considers "impossible" or "not worth doing" into the next phase's default capability.

Today, the Tripo series is reshaping the underlying paradigm of native mesh generation through global diffusion generation, and driving the industry toward a new technical consensus: diffusion models are moving from exploratory solutions to the core technical route for native mesh generation. The starting point of this consensus was precisely the non-consensus that VAST held firm to: not aiming for demo effects but adhering to "production-ready" goals; seeking native representations for 3D rather than forcing it into language or image representation frameworks; not settling for the "false prosperity" of autoregressive routes in mesh generation, but exploring the higher ceiling of global generation.

Expanding the Boundary of "Production-Ready Generation," Becoming the 3D-Native Intelligent Infrastructure for All Industries

Production-ready generation is the foundational principle running through VAST's algorithm and technology R&D: let AI adapt to industrial standards accumulated over decades, rather than making humans clean up after AI. Along this path, models directly generate objects' geometry, topology, and materials, and further endow them with semantic parts, skeletal rigs, and physical properties — upgrading visual objects from merely "viewable" to "usable" functional objects whose generated results require no rework before entering professional production pipelines. When generated objects are no longer constrained by category or form, and evolution paths universally apply to arbitrary scenarios, 3D-native intelligence will truly have matured.

Today, the Tripo H series refreshes quality ceilings with sculptural geometric precision, while the P series outputs engine-ready clean topology meshes in seconds; 8K textures, intelligent splitting, and local editing make models "editable and producible." These capabilities are now landing in real workflows across gaming, XR, 3D printing, and film and television industries.

Gaming is one of the most important scenarios for AI 3D deployment

In July 2026, NetEase's Eggy Party launched a 3D model splitting feature in its Eggy Workshop, powered by Tripo: AI-generated complete models can be further disassembled into independently editable components, supporting local modification, replacement, and recombination — moving from "AI generation" to "AI editability." This marks the second collaboration between the two parties, following Tripo's native 3D models first integrating into Eggy Party's UGC creation system in September 2025. Previously, VAST partnered with NetEase's Where Winds Meet to launch the "Taiji of All Things" gameplay feature, where players upload a photo to generate interactive 3D props. Today, Tripo provides full-pipeline support from concept generation to animation rigging for multiple global top-tier game teams, integrated into art production pipelines for scene asset building, level prototype validation, and multi-option comparison; one game team completed a large-space scene that would traditionally require months of delivery using only 2.5 people in two weeks.

XR is another critical entry point for AI 3D

VAST has developed a native AI 3D generation application tailored for next-generation XR devices, requiring no controllers or professional software — language descriptions generate 3D content, manipulated through eye tracking and hand gestures. XR devices, constrained by mobile GPU compute, impose extremely demanding requirements on asset polygon counts and topology regularity, while the Tripo P series' support for native output at target polygon counts with clean topology, entering Unity, Unreal, and other engines without retopology, serves precisely as the technical foundation for such experiences. VAST's algorithm team recently open-sourced StereoWorld (CVPR 2026) with The University of Hong Kong, PICO, and other institutions — a binocular world model that simultaneously models left and right views from generation inception, solving spatial drift in monocular models when users move, and capable of converting existing monocular video to binocular stereoscopic content, lowering cold-start costs for XR content.

3D printing is the most direct consumer-grade outlet for 3D-native intelligence

Globally, consumer-grade 3D printing is moving from maker communities into home DIY, education, designer toys, and personalized manufacturing. The Tripo H series generates watertight models ready for direct slicing and printing without repair: Tripo was among the earliest AI 3D foundation models integrated into Bambu Lab's MakerWorld, and since the H2.5 series in late 2024 has helped users complete the full "upload image — generate model — slice and print" workflow, improving prototyping speed by approximately 90%. VAST now achieves full coverage of top 3D printing hardware manufacturers, with partners including Bambu Lab, Anycubic, Creality, and Elegoo; in July this year, VAST entered strategic partnership with HeyGears, whose full-color 3D printing platform HeyVerse integrated Tripo to jointly build an integrated desktop full-color printing closed-loop solution.

Film and animation is moving from "gacha" to precise control via 3D director's viewfinders

Intangible AI, an overseas creative platform founded by former Apple and Pixar members, has integrated Tripo as its scene asset supply engine — creators input text or images and receive characters, props, and architecture in seconds to drag into virtual scenes, selecting camera angles and positions before handing off to video models for rendering. Domestic video platform TapNow also launched "3D Studio" in August, using a "Tripo sets keyframes, video model outputs" combination that lets creators freely build white-box sets and position cameras, with every angle of the same performance ready for final output.

From gaming to XR, from 3D printing to film and television, models are no longer mere efficiency tools. As they are integrated into real production workflows across more industries, 3D evolves from the exclusive capability of professional creators into universal production infrastructure. VAST is accelerating native 3D intelligence into production scenarios across industries, and the investment from diverse strategic industry players represents direct recognition of this phase's achievements.

Let Everyone Create Worlds, Let the World See Every Creator

We have always believed that 3D is the native expression of world states. From creating objects to creating worlds, from tools to democratization, expressing oneself and building worlds will ultimately become everyone's instinct.

Today, the all-in-one AI 3D workspace Tripo Studio has gathered tens of millions of creators globally; the Tripo Education Program has covered over 350 universities across more than 30 countries, benefiting 300,000 students and professors, becoming a bridge connecting overseas creators and academia.

This year alone, VAST's footprint has spanned three continents and over 30 cities:

  • The inaugural Tripo AI 3D Game Jam received the most submissions of any global competition of its kind, with shortlisted works presented live at GDC;
  • The 3D creation workshop Tripo Donut Workshop crossed 15 time zones, landing in 7 cities including San Francisco, Tokyo, Seoul, Berlin, Paris, and London;
  • During SIGGRAPH 2026, VAST once again became the only Chinese company on the Keynote stage, with 5 papers accepted to Technical Papers, and won the highest award Best in Show in the Real-Time Live! segment, co-hosting a hackathon and industry night with Fei-Fei Li's World Labs;
  • Top works from the 3rd Tripo AI 3D Rendering Challenge appeared on the Nasdaq screen in Times Square, New York, with submission volume growing 20x over two years to become one of the world's largest 3D rendering competitions;
  • This week, VAST will again co-host a hackathon with World Labs, Mint, Founders Inc, and others in San Francisco to explore the next steps for spatial intelligence and generative 3D.

Argentine writer Jorge Luis Borges tells a story about "how to express the world" in The Aleph.

The narrator "Borges" has always mourned the deceased Beatriz, and so visits her family annually. Her cousin Carlos is writing a long poem attempting to encompass the entire world, and claims an "Aleph" is hidden in his basement. When the house faces demolition, he invites Borges to witness it.

The "Aleph" is a tiny, iridescent sphere, approximately two or three centimeters in diameter. It contains all spatial locations; whoever gazes into it simultaneously perceives everything in the universe — every place, every angle, past memories, distant people, and intimate details — past, present, and future, far and near, neither overlapping nor compressed. All things exist simultaneously in this tiny sphere.

The story raises questions:

Whether finite language can express the infinite and simultaneous;

Whether the artistic ambition of "completely recording the world through description" is destined to fail;

The true paradox thus emerges: Borges perceived the entirety of the world, yet could only describe it linearly, item by item. The Aleph contains everything, yet any narrative about the Aleph necessarily omits everything.

The way we understand the world through language, images, and video today is often similar: they record slices of the world from certain angles, at certain moments.

3D comes closer to the world itself — where objects are, how they relate to each other, how their states change; even language, images, and video exist within the same spatial state.

This is also VAST's necessary path from 3D generation to world models: from generating an object to understanding a world; from describing the world to making the world itself a generable, understandable, and interactive object.

3D describes not merely how we talk about the world; it represents how the world exists.

Where language cannot reach, worlds can be born anew.