VAST Nears $200 Million in Series A Funding, Unveils World Model in Parallel | Oasis Vitality
Project Eden: A General World Model Research Initiative

VAST, an artificial general intelligence company and Oasis Capital angel-round portfolio company, recently announced the completion of its Series A+ and A++ rounds, totaling nearly $200 million.
The new funding was led by Zhence Capital and China Life Yangtze River Delta Sci-Tech Fund, with participation from strategic industry investors including the Shenzhen AI Terminal Industry Fund (whose industry partner is global leading terminal manufacturer HONOR), a prominent strategic corporate investor, and Shanghai Semiconductor Industry Investment; alongside top-tier financial investors including Shenzhen Capital Group, YuanSheng Ventures, Wofu Venture Capital, and F&G Capital. This diverse coalition of leading market-oriented funds, state-owned platforms, and industry strategic investors creates a multi-dimensional empowerment system. Existing shareholders Primavera Venture Partners, Eminence Ventures, Baidu Venture, and Orient Capital also continued to significantly oversubscribe.
This marks VAST's second capital infusion in just two months, following its March 2025 funding round.
Concurrently, VAST unveiled its new world model initiative — Project Eden. Unlike conventional approaches such as "action-conditioned video generation" and "static 3D scene generation," Project Eden creatively decouples underlying state simulation from visual rendering at a native architectural level. This breakthrough makes it the world's first world model enabling autonomous maintenance and deterministic control of world states, naturally unlocking disruptive capabilities including long-term environmental persistence, free scene reusability, and multi-user concurrent interaction. Project Eden aims to become the foundational engine for next-generation low-barrier interactive content creation, while providing embodied intelligence and other agents with highly logically consistent training and evaluation environments.
The proceeds will primarily fund top-tier talent acquisition for AI 3D large models and general world models, core algorithm iteration and data accumulation, while accelerating global market expansion and industrial ecosystem development.

Project Eden
A World Model with Persistent Environments and Multi-User Interaction
While large language models predict the next token and video models focus on rendering the next frame, the core mission of a world model is to simulate the next state of the world — that is, to model all changes resulting from the current environment and user actions.
This defines the two fundamental challenges any qualified general world model must solve: first, defining the world's current objective state (State), and second, driving the world's continuous autonomous evolution (Transition).
For VAST, the ultimate goal of its technical roadmap has remained constant: enabling everyone to create and freely explore countless interactive worlds with their own hands. Achieving this requires solving several foundational problems: long-horizon environmental memory, concurrent multi-user and agent interaction, and engineering solutions that are both low-cost and scalable.
The two dominant technical approaches in the industry currently cannot simultaneously satisfy the complete user need of creating worlds and continuously interacting within them.
The first is action-conditioned video generation: this approach only performs short-term pixel-level prediction based on limited spatial input actions, implicitly compressing world state into a constrained sequence of frames. Once objects leave the camera's field of view, the model can only hallucinate and reconstruct, making long-term state retention impossible and precluding multi-user coexistence in the same world.
The second is static 3D scene generation: while capable of constructing explorable three-dimensional spaces, this approach strips away the temporal dimension and physical operating logic, lacking any state iteration mechanism and thus failing to support dynamic interaction.
In response, VAST's general world model research initiative Project Eden innovatively adopts a three-layer decoupled architecture that breaks free from the inherent constraints of pixel generation:
The bottom layer is the structured state layer: establishing an evolvable 3D foundational substrate that uniformly maintains scene geometric structure, object identity attributes, and global event logic, fully bearing the world's objective state and autonomous simulation.
The middle layer is the conditional interface layer: serving as the conversion hub between state and rendering, transforming complete underlying 3D states into semantic and geometric conditional constraints adapted to generation tasks, based on different camera perspectives. All perspective renderings derive from the same underlying world, fundamentally guaranteeing physical consistency across shots and viewpoints.
The top layer is the generative rendering layer: relying on underlying objective states and intermediate-layer constraints, it renders refined visual imagery on demand in real-time, supplementing dynamic details to deliver intuitive immersive experiences to users.
Through this natively decoupled architecture of state simulation and visual rendering, Project Eden becomes the first globally to transform world states into independently operating entities that can be persistently retained, repeatedly edited, and multi-user shared, thereby naturally unlocking three core capabilities that traditional approaches cannot simultaneously achieve:
-
Long-term environmental persistence: World states exist independently of camera perspectives and are permanently stored, unaffected by scene transitions or user departures. Underlying state queries guarantee spatiotemporal consistency, supporting users' extended continuous roaming within scenes and thoroughly resolving industry problems like object disappearance and scene distortion.
-
Free scene reusability: Supports reading, writing, and dynamic intervention on underlying world states, with all user actions within scenes being genuinely retained. For example, after a user performs destructive or transformative operations on scene objects, subsequent users entering that scene will see fully consistent modified results. Scenes need not be regenerated, achieving global state continuity and efficient reuse.
-
Multi-user concurrent interaction: With decoupled state evolution and rendering pipelines, a single underlying world can simultaneously承载 large numbers of human users and AI agents in concurrent online interaction. Unlike traditional approaches where computational costs grow exponentially with viewpoints/user counts, this architecture offers controllable computational costs, not only supporting large-scale social interaction and massive online content ecosystems, but also serving as a critical foundation for cluster-based embodied intelligence training and multi-agent collaborative research — with outstanding commercial and scientific value.
Project Eden is positioned as the foundational engine for next-generation interactive content creation, and as a high-quality simulation base adapted for embodied intelligence training, comprehensively covering two core scenarios: interactive content and scientific research.
For interactive content, it provides one-stop capabilities for environment generation and interactive logic construction, supporting both mass creators using natural language and simple actions to one-click create multi-user shared interactive worlds, and serving content production and interactive experience deployment in gaming, film and television, VR/AR, and digital twin industries.
For scientific research, it provides simulation environments with complete physical rules, long-term temporal consistency, and free intervention capabilities, empowering large-scale embodied intelligence training and multi-agent performance evaluation.
The team has established a more pragmatic, scalable industry research paradigm: refusing to degrade world models into video generation tasks, instead using evolvable structured states as the foundation and generative models to drive high-fidelity visual presentation — a path aligned with technical fundamentals and more amenable to scalable deployment.
VAST's exploration of general world models continues to iterate:
On one hand, strengthening high-complexity scene simulation capabilities, enriching physical dynamic effects, expanding free-viewpoint boundaries, and refining object interaction granularity;
On the other hand, building dedicated state transition models to achieve closed-loop autonomous world updates based on agent interaction behaviors, while continuously optimizing real-time rendering performance and reducing deployment costs, so that world models can benefit more creators and developers.

Continuously Refreshing 3D Foundation Model SOTA
Creating Generational Gaps with the Industry
Over the past three years, VAST has consistently maintained algorithmic SOTA in the AI 3D domain. Each iteration of its self-developed Tripo series 3D foundation models has set new global industry benchmarks.
The Tripo H3.1 and Tripo P1.0 models officially launched in March 2026 (NEXUS, SIGGRAPH 2026) continue to maintain industry-leading performance by a wide margin: the former refreshes the precision ceiling of AI 3D with sculptural-grade geometric detail; the latter is the world's only 3D foundation model capable of outputting production-grade meshes within seconds, achieving hundred-fold speed improvements over other available solutions with a generationally leading technical approach. These sustained breakthroughs at the model layer equip VAST with the underlying conditions to push 3D assets from "viewable" to "usable, interactive, and evolvable."
Recently, Tripo Studio launched two new algorithmic breakthroughs:
- 8K Textures: Every Detail Withstands Scrutiny
Tripo 8K Texture is the industry's first native 8K AI texture algorithm.
The new AI texture precision has surpassed the limits of human visual resolution, enabling full-distance lossless presentation of 3D assets: close-up inspection reveals no flaws, and extreme magnification remains crisp. With this algorithm, AAA rendering quality and cinematic detail can both be natively AI-generated.




For a long time, 8K textures have been the exclusive domain of high-end 3D assets. Manual creation by senior texture artists requires 3-5 days; material scanning with projection onto models takes 2-3 days, with stringent equipment and site requirements and per-texture costs of $500-2,000 — affordable only for top-tier projects.
VAST compresses the entire production pipeline to under 2 minutes, with near-zero marginal cost per texture. Independent creators and small studios can now easily access film and AAA-grade texture quality; for large teams, the production bottleneck for high-definition textures is completely eliminated, available on demand.
Technically, this feature employs native multi-channel synchronous generation, with all-dimensional materials reaching 8K resolution, revealing every fiber with intact detail even after magnification. Output assets directly integrate into professional pipelines including Unreal, Unity, and Blender, requiring no secondary repair.
- Segmentation V2: More Precise, More Controllable Intelligent Part Splitting
In May 2025, VAST launched the industry's first intelligent part splitting feature in Tripo Studio Beta, enabling AI-generated 3D assets to be automatically segmented and directly enter downstream pipelines.
Tripo Studio serves diverse industries including gaming, 3D printing, industrial design, and virtual reality, yet different scenarios demand significantly varying segmentation granularity. Users of previous versions often needed to expend additional effort on manual adjustments, reducing overall production efficiency. One year later, the team launched the upgraded Segmentation V2, building on enhanced multimodal 3D structural understanding models and part naming mapping mechanisms to deliver higher-precision, more controllable 3D asset splitting capabilities.
The upgraded V2 generates 2D pre-segmentation previews before executing 3D splitting, making results clearly visible; it also introduces three granularity levels corresponding to real downstream scenario needs for assembly detail:
- Low (3–6 parts): For scenarios focused on primary structure, such as 3D printing and concept presentation;
- Medium (6–15 parts): Corresponding to common assembly granularity in game development and film production pipelines;
- High (15+ parts): For highly segmented assets including fine modules, mechanical structures, and detachable toys.
For the 3D printing industry, combined with the concurrently launched Quick Cap feature, the full "generate—segment—cap—print" workflow is further compressed.
- Pioneering Frontier Research, Building Open-Source Ecosystems
At VAST, open-source is not an incidental byproduct of technology spillover. As the spatial language shared between humans and machines, 3D's underlying infrastructure should be built through open collaboration.
In March 2024, VAST partnered with Stability AI to jointly open-source TripoSR, pioneering the compression of single-image 3D generation speed to the 0.5-second level — the model rapidly became a mainstream choice for global creators.
In March 2025, VAST launched its second open-source season, successively releasing eight projects including TripoSG, TripoSF, UniRig, and HoloPart, covering the full core chain from foundation models to functional components. Multiple results have been integrated into mainstream creative tools including Blender and ComfyUI, with UniRig firmly established as the global benchmark for open-source 3D auto-rigging solutions.
Now, VAST's third open-source season has officially concluded. This season focused on dynamic interactive content, deeply exploring new possibilities in representation forms and application scenarios:
• Jointly open-sourced with Tsinghua University: TripoSplat (DeG, SIGGRAPH 2026): revolutionizing 3D Gaussian density control logic through a learnable probabilistic sampling mechanism, enabling models to autonomously complete dynamic computational allocation, so that 3D content is no longer limited to static resolution but becomes "dynamic resolution" adaptable to devices and application scenarios;


• Jointly open-sourced with The University of Hong Kong: AniGen (SIGGRAPH 2026): one-click generation of animatable 3D assets from single images, completing geometry, texture, skeleton, and skinning generation within a unified model, enabling 3D content to be dynamically interactive immediately upon generation;
• Jointly open-sourced with Tsinghua University: SkinTokens: the industry's first transformation of skinning weights into token form, achieving joint bone and skinning generation within the same autoregressive framework, pushing AI auto-rigging capabilities to animation and gaming industry industrial standards;
• LegoACE (SIGGRAPH Asia 2025): supporting dual text and image input, generating physically assemblable LEGO models through block-by-block autoregressive generation.
Through three years of deep cultivation, VAST has built a complete AI 3D and world model open-source algorithm ecosystem, with over 30 open-source projects released, covering the full technology stack from foundational representations to generation pipelines. Continuously opening core technologies to global researchers and developers, making frontier technology truly serve every creator.





