VAST Closes $50 Million Series A to Build World Models and UGC Interactive Platform | Oasis Vitality

To start a business, you have to "believe first" rather than "see first."

VAST, which Oasis Capital led exclusively at the angel round, today announced the completion of a $50 million Series A funding round.

The round was co-led by Alibaba and Hengxu Capital, with follow-on investments from Yuanhe Puhua, Baidu Venture, and Orient Capital, creating a comprehensive empowerment network spanning top-tier capital, industry giants, and well-known strategic investors. Existing shareholders Primavera Venture Partners and the Beijing Artificial Intelligence Industry Investment Fund made additional investments above their pro-rata rights.

Founded in 2023, VAST has become a global full-stack leader in the multimodal domain, spanning foundational model R&D to application ecosystem deployment — its self-developed 3D foundation models have maintained industry-leading performance, with ecosystem partnerships covering leading enterprises including Alibaba, Tencent, ByteDance, NetEase, SAIC, Bambu Lab, and UBTECH, as well as over 90,000 developers; its Tripo Studio platform has attracted more than 6.5 million creators, generating nearly 100 million 3D models cumulatively.

Simultaneously, VAST officially released its new AI 3D large model family, continuing to set new industry SOTA benchmarks: the updated Tripo H3.1 maintains industry-first position across core metrics including input alignment, structural precision, texture quality, and generation speed, continuously pushing the boundaries of high-resolution model generation; the newly architected Tripo P1.0 redefines the algorithmic paradigm for AI 3D, capable of outputting professional modeler-grade 3D assets in 2 seconds — over 100x faster than existing solutions on the market.

Building on its accumulated data, talent, and system-level research and engineering experience, VAST has prioritized world model development in 2025, with its first world model set for release in the near term.

Proceeds from this round will focus on recruiting top talent for world models, continuous iteration of core algorithms, and data accumulation, while aggressively advancing its UGC interactive content platform. VAST is committed to enabling everyone to create, experience, and share interactive content, making interactive content a new information medium connecting digital and physical worlds.

Redefining the AI 3D Algorithmic Paradigm

VAST was founded with a singular mission: to let everyone create their own interactive worlds. This vision breaks down into three keywords: UGC, interactivity, and 3D worlds.

At the outset, the team attempted to tackle "worlds" directly, but quickly encountered a fundamental contradiction: the creation cost and barrier for interactive content is extraordinarily high. Producing AAA-grade interactive content demands massive human resources, time, and capital. These creative constraints determine the ceiling for content consumption formats and diversity.

So VAST chose to address the underlying problem first, using AI to redefine how interactive content is produced.

Over the past three years, the algorithm team's self-developed AI 3D large models have achieved continuous breakthroughs, with each generation of the Tripo H series defining the global state-of-the-art at launch.

The team's proposed SparseFlex representation not only captures fine details, open surfaces, and internal geometric structures with precision, but also supports more efficient training strategies that dramatically reduce memory overhead — making AI-generated 3D models truly viable for user experience and large-scale commercial applications.

The updated Tripo H3.1 delivers higher fidelity and alignment with input reference images, while achieving top-tier expression of both overall structure and local details. It resolves long-standing industry pain points across character anatomy, facial features, geometric text, and other challenging cases.

In comprehensive benchmarking of this flagship model, whether for font embossing, human facial features, or complex mechanical structures, Tripo H3.1 achieves industry-first performance across core metrics including input alignment, structural precision, texture quality, and generation speed.

For UGC interactive content, precision alone is insufficient. Creators need "speed" and "out-of-the-box usability" — high quality, high efficiency, and real-time deployment within engine pipelines. Yet a structural tension has long existed between quality and speed: pursuing precision means longer generation times, while pursuing speed sacrifices structure and detail. Topological structure, geometric precision, texture quality, engine compatibility, editability — solving one problem often comes at the cost of another.

This has been the primary bottleneck preventing AI 3D from scaling from PGC industry applications to mass UGC interactive content.

Therefore, VAST's algorithm team made a bold choice: rebuild the AI 3D large model from the ground up — starting from first principles, not merely making marginal optimizations along the existing path, but rethinking how 3D should be represented and generated with entirely new thinking and algorithmic frameworks.

This is the origin of the Tripo P series.

Existing 3D mesh generation has been constrained by compromises of serialization: lengthy serialized data severely limits generation efficiency, while unidirectional causal bias stifles global spatial interaction. Tripo P1.0 is the first to reconstruct the underlying paradigm of spatial generation, with its Smart Mesh functionality already live on the Tripo Studio platform.

Tripo P1.0 abandons localized computation of 3D objects in favor of constructing a unified native probability space. Rather than "predicting" the next point, the model performs macroscopic probability collapse on the overall spatial structure. The model first precipitates the "substrate" that constitutes form from noise space; subsequently, complex topological relationships evolve within this unified probability space, building upon that substrate.

Through this new spatial modeling philosophy, Tripo P1.0 breaks through dimensional construction barriers. It can directly complete instant alignment of high-dimensional information within disordered noise space, enabling extremely complex 3D topological structures to simultaneously "converge" into form — achieving a fundamental leap from local stitching to global emergence.

The end result: Tripo P1.0 can directly generate professional modeler-grade 3D models in 2 seconds — with clean topology, stable wireframes, and engine-ready output. The team has also found that under this new approach, the model's editability and the scalability of its precision hold exceptionally high optimization potential.

The release of the Tripo P series formally announces: AI 3D large model algorithmic paradigms have entered the 2.0 era, where speed, quality, and engineering viability now hold simultaneously.

Accelerating UGC Platform Construction to Welcome an Explosion of Interactive Content

From day one, VAST has believed: when the barrier to creating interactive content drops low enough, UGC interactive content will enter a new phase.

Today, this thesis is beginning to be validated — and faster than anticipated.

  • Over 6.5 million creators have produced work on Tripo Studio, generating nearly 100 million 3D models cumulatively;
  • Over 90,000 enterprise developers and partners have integrated Tripo capabilities into their own products and workflows;
  • From Blender and Maya to Unity and Unreal, ecosystem plugins cover mainstream 3D creation tools and content engines;
  • In smart manufacturing, interactive entertainment, virtual reality, and embodied intelligence, Tripo is increasingly becoming the default capability for 3D content production.

The significance of these numbers lies not in scale itself, but in the shift in behavioral patterns.

More and more users previously unexposed to 3D modeling are beginning to use 3D generation tools in high-frequency, lightweight ways. The process of generating models is approaching everyday expressive habits — instant generation, instant expression, instant sharing. 3D is no longer confined to professional production pipelines, but beginning to enter mass creative contexts.

Centered on this shift, VAST continues building creator ecosystems and exploring the boundaries of interactive content formats together with the community:

  • The Tripo Ambassador Program spans over 30 countries and 50+ universities globally;
  • The Game Hub community has gathered over 100,000 active developers, producing 2,000+ AI interactive content pieces;
  • Top 25 works from the AI 3D Rendering Competition S2 were displayed on Times Square in New York and multiple core landmarks worldwide;
  • The inaugural Tripo AI 3D Game Jam received the most submissions of any global competition in its category upon launch, with shortlisted works to be exhibited on-site at GDC.

Over the past year, the team traveled across three continents and 30+ cities, repeatedly discussing the same question with developers, creators, and enterprise partners: when 3D production efficiency crosses the inflection point, how will interactive content evolve?

Meanwhile, Tripo has embedded itself into broader AI production systems. Early on, it was the first to achieve node-level integration with ComfyUI, incorporating 3D capabilities into multimodal generation pipelines; subsequently, it became the first to support MCP and other model context protocols, enabling 3D generation to be called and orchestrated; now, Tripo model capabilities access the next-generation Agent ecosystem as Skills, giving intelligent agents the ability to generate, call, and edit 3D assets.

This means 3D capabilities are no longer solely for human creators. As agent collaboration gradually becomes an emerging foundational production method, 3D generation is entering the standard capability set of intelligent agents. Whether for automated content generation, game building, virtual scene construction, or e-commerce and interactive system asset production, 3D is becoming the spatial expression module of AI systems.

When creation costs fall, when both human creators and intelligent agents possess spatial generation capabilities, when creation frequency rises and ecosystem density reaches threshold — new content formats will naturally emerge once conditions mature.

The company judges that the inflection point where 3D generation shifts from professional production tool to mass expressive language is near.

Therefore, in 2026 VAST will accelerate construction of its UGC interactive content platform, integrating generation capabilities, distribution mechanisms, and interactive systems into a complete closed loop — transforming 3D creation from producing individual assets to creating continuously evolving interactive content.

From Making Objects to Making Worlds

The progression from understanding and generating objects to understanding and generating entire worlds is the technical and product roadmap VAST has consistently pursued.

From its founding, VAST bet on 3D as the most primitive, most natural, and highest information-density content modality. The team believes world models represent the ultimate form of general models, and must be built on native understanding of three-dimensional space.

In 2025, VAST has already directed core R&D resources toward world models. Its accumulated database of 50 million high-quality 3D and world models, the world's highest-density cross-disciplinary team at the intersection of AI and computer graphics, and its system-level research and engineering experience accumulated as an industry pioneer in 3D understanding and generation constitute VAST's distinctive advantages along this path.

VAST is committed to building general world models that will not only generate interactive virtual worlds, but also possess perception, understanding, and physical reasoning capabilities — broadly serving next-generation interactive content, embodied intelligence, simulation, and other expansive scenarios.

A deep connection exists between AI 3D and world models: the former is an indispensable foundation for the latter, and 3D is the universal interface for both machines and humans. Whatever final form world models may take, native understanding and generation capabilities for three-dimensional space constitute an essential component.

In the view of VAST founder and CEO Yachen Song, choosing to "believe first" rather than "see first" is what distinguishes startups from large platforms in essence, and is key to establishing first-mover advantage.

AI 3D was once regarded as a high-investment, highly uncertain market. Three years ago, VAST planted its first seed in this nearly overlooked land. At that time, there was no convergent path to follow, no mature technical framework — most saw only a wasteland, and turned their gaze elsewhere.

For three years, VAST has remained focused on sowing, irrigating, and fertilizing this land. Every algorithmic iteration redefined the industry's highest standard; every product update lowered the barrier to creation somewhat, and widened the boundary of industrial viability somewhat.

Now, a small grove has grown from this wasteland. Over 6.5 million creators produce here, 90,000 enterprises and developers build applications here, and increasingly many believe — a real forest will indeed grow here.

The meaning of planting trees extends beyond the trees themselves. Every tree changes the soil, moisture, and air around it, creating better growing conditions for the next tree. Every VAST user makes the technology more mature, the product more refined, the barrier to creation lower, the imagination for content consumption richer, and the starting point for those who follow higher — when individual creation converges, it ultimately transforms the entire system ecosystem.

Looking back years from now, it may be hard to pinpoint which day marked the beginning of the turning point. One will only remember that the trees grew more numerous, streams gradually began to flow across the ground, the shade connected into stretches, and the wasteland acquired a new name.