Skip to main content

One post tagged with "AI UGC"

View all tags

AI UGC 3D Quick? Generate Professional-Grade 3D Models in 2 Seconds, Alibaba and Baidu Bet on VAST

· 17 min read

This article is republished with permission from "Game Artist".

​Another Chinese AI startup has secured funding.

On March 5, AGI company VAST announced the completion of a $50 million Series A round. This comes less than 9 months after its tens-of-millions-of-dollars Pre-A+ round in June last year.

The round was co-led by Alibaba and Hengxu Capital, with participation from Yuanhe Puhua, BV Baidu Ventures, and Oriental Jiafu, among others. The investor lineup includes top-tier capital, industry giants, and prominent strategic investors. Additionally, existing shareholders Spring Capital and the Beijing AI Industry Investment Fund made super-proportional follow-on investments.

VAST founder and CEO Song Yachen previously worked at SenseTime, where he was responsible for strategic analysis and commercialization of multiple AI projects, and also participated in the founding of MiniMax. In 2023, Song Yachen founded VAST. Based on its self-developed AI 3D and world models, the company aims to create mass-market interactive content creation capabilities, lead a creator equality movement for everyone, and ultimately build an interactive world content platform. Currently, VAST has become a full-stack global leader in the multimodal domain, spanning from foundation model R&D to application ecosystem deployment.

According to official disclosures, its self-developed 3D foundational large model continues to lead the industry. Ecosystem partnerships cover leading companies such as Alibaba, Tencent, ByteDance, NetEase, SAIC, Bambu Lab, and UBTECH, as well as over 90,000 developers. Its Tripo Studio platform has gathered more than 6.5 million creators, generating nearly 100 million 3D models in total.

Regarding this investment, Hengxu Capital stated: “Our assessment of AI technology has always revolved around one core question: can it truly land and create value? VAST’s 3D generation solution impressed us with its technical maturity and product completeness — not only are the results leading, but more importantly, the solution is ‘usable.’”

BV Baidu Ventures also commented: “3D generation is a crucial foundation for the future virtual world and a key direction for multimodal generation and world models. Compared to video generation, 3D naturally carries physical laws, spatial relationships, and object interaction logic. It is an essential path for AI to understand and simulate the real world. In the past, the industry paid less attention to the 3D modality compared to video, but now 3D generation has reached a significant inflection point in both technical breakthroughs and commercial deployment. This is the golden window for investment.”

Along with the funding announcement, VAST released a new family of AI 3D large models.

The upgraded Tripo H3.1 maintains industry-leading performance across core metrics such as input alignment, structural accuracy, texture quality, and generation speed. The new architecture, Tripo P1.0, redefines the algorithmic paradigm of AI 3D, capable of outputting professional-grade 3D assets in just 2 seconds — a speed increase of over 100x compared to existing solutions.

Furthermore, leveraging its accumulated data, talent, and systems-level research and engineering experience, VAST has focused on developing world models since 2025, with the first world model expected to be released soon.

It is understood that the new funds will be primarily used to recruit top talent for world models, iteratively improve core algorithms, accumulate data, and vigorously promote the construction of a UGC interactive content platform. This is to better realize the vision of “enabling everyone to create, experience, and share interactive content, making interactive content a new type of information carrier connecting the digital world and the physical world.”

3D-Generated UGC
About to Enter a More Mass-Market Phase

When founding VAST, Song Yachen believed that when the barriers to creating interactive content are low enough, UGC interactive content would enter a new phase.

Now, that judgment is being validated, and at a faster pace than expected.

• Over 6.5 million creators have already created on Tripo Studio, generating nearly 100 million 3D models in total. • Over 90,000 enterprise-level developers and partners have integrated Tripo’s capabilities into their own products and workflows.

• Ecosystem plugins cover mainstream 3D creation tools and content engines, from Blender and Maya to Unity and Unreal. • In areas such as smart manufacturing, interactive entertainment, virtual reality, and embodied intelligence, Tripo is gradually becoming the default capability for 3D content production.

The significance of these numbers lies not in their scale, but in the structural change of behavior.

More and more users who had no prior experience with 3D modeling are now using 3D generation tools frequently and in lightweight ways. The process of generating models is approaching everyday expression habits — generate instantly, express instantly, share instantly. 3D is no longer confined to professional production pipelines; it is entering the context of mass creation.

Around this change, VAST is also continuously building a creator ecosystem and exploring the boundaries of interactive content forms together with the community, including:

The Tripo Ambassador Program covering over 30 countries and 50+ universities; The Game Hub community with over 100,000 active developers, generating more than 2,000 AI interactive content pieces; The top 25 works of the AI 3D Rendering Contest S2 displayed on New York’s Times Square and key landmarks worldwide; The first Tripo AI 3D Game Jam, which received the highest number of entries among similar global events since its launch, with selected works exhibited at GDC.

Over the past year, VAST has been repeatedly discussing the same question with developers, creators, and enterprise partners across three continents and over 30 cities: when 3D production efficiency crosses a critical threshold, how will interactive content evolve?

Meanwhile, Tripo has also embedded itself into a broader AI production ecosystem. Initially, VAST was the first to achieve node-level integration with ComfyUI, incorporating 3D capabilities into the multimodal generation pipeline. Later, it became the first to support model context protocols like MCP, enabling 3D generation to be invoked and orchestrated. Now, Tripo’s model capability is integrated into the new generation of Agent ecosystems in the form of Skills, granting agents the ability to generate, invoke, and edit 3D assets.

This means 3D capability is no longer only for human creators.

As Agent collaboration gradually becomes an emerging fundamental production method, 3D generation is entering the standard capability set of agents. Whether for automated content generation, game building, virtual scene construction, or asset production for e-commerce and interactive systems, 3D is becoming the spatial expression module of AI systems.

When creation costs drop, when both human creators and agents possess spatial generation capabilities, when creation frequency increases, and when ecosystem density reaches a threshold, new content forms will naturally emerge under favorable conditions.

VAST thus concludes that the tipping point where 3D generation transforms from a professional production tool to a universal language of expression is near.

To this end, VAST announced that in 2026 it will accelerate the construction of a UGC interactive content platform, integrating generation capabilities, distribution mechanisms, and interactive systems into a complete closed loop, turning 3D creation from producing single assets into creating constantly evolving interactive content.

​Technical Moat: Redefining the AI 3D Algorithm Paradigm

In fact, as early as its founding, VAST established the vision of “enabling everyone to create their own interactive world.” The key to realizing this vision encompasses three dimensions: UGC, interaction, and the 3D world.

At the start, the team tried to directly tackle the “world” aspect, but soon discovered a fundamental contradiction: the cost and barrier to creating interactive content were extremely high. Producing AAA-level interactive content requires massive manpower, time, and cost. The barrier to creation sets the ceiling for content consumption forms and diversity.

Therefore, VAST chose to solve the underlying problem first, using AI to redefine the production method of interactive content.

First, the team proposed the SparseFlex representation method, which not only accurately captures model details, open surfaces, and internal geometry but also supports more efficient training strategies, significantly reducing memory usage, making AI-generated 3D models truly industrially usable for user experience and large-scale commercial applications.

The upgraded Tripo H3.1 goes further, with higher fidelity and alignment to input reference images, while also achieving top-tier expression for overall structure and local details. It resolves long-standing bottlenecks in cases involving character forms, faces, and geometric text.

VAST conducted a comprehensive benchmark test of this flagship model, and the results showed that across core indicators such as input alignment, structural accuracy, texture quality, and generation speed, Tripo H3.1 achieved industry first.

For UGC interactive content, accuracy alone is insufficient. Creators need “speed” and “out-of-the-box” usability — high quality and high efficiency, with real-time implementation within engine pipelines.

But there has long been a structural tension between quality and speed: pursuing precision means longer generation time, while pursuing speed often sacrifices structure and detail. Topology, geometric accuracy, texture quality, engine compatibility, editability — each problem solved often comes at the expense of another.

This became the key bottleneck for AI 3D transitioning from PGC industry applications to large-scale UGC interactive content.

Therefore, VAST’s algorithm team made a bold fundamental choice: rebuild the AI 3D large model from the ground up — starting from first principles, not just making marginal optimizations along existing routes, but thinking with a new mindset and a new algorithmic framework, rethinking how 3D should be expressed and generated.

This is the background for the Tripo P series. The release of this series also marks the entry of the AI 3D large model algorithm paradigm into Phase 2.0: speed, quality, and engineering usability begin to coexist.

Specifically, existing 3D mesh generation is often plagued by sequential compromises: lengthy sequential data severely limits generation efficiency, and unidirectional causal bias kills global spatial interaction. Tripo P1.0, for the first time, reconstructs the underlying paradigm of spatial generation. Its Smart Mesh feature is already available on the Tripo Studio platform.

Tripo P1.0 abandons local computation of 3D objects and instead constructs a unified native probabilistic space. The model no longer “predicts” the next point, but performs a macro-level probabilistic collapse on the overall structure of space. The model first precipitates a “substrate” for constructing the form in the noise space; then, complex topological relationships evolve on that substrate within the unified probabilistic space.

With this new spatial modeling philosophy, Tripo P1.0 breaks down dimensional barriers. It can directly complete high-dimensional information instant alignment in an unordered noise space, allowing extremely complex 3D topological structures to “converge” and take shape synchronously, achieving a fundamental leap from local assembly to global emergence.

The end result: Tripo P1.0 can generate a 3D model at the level of a professional modeler in just 2 seconds — with clean topology, stable edge flow, and engine-ready.

Moreover, VAST also discovered that under this new approach, the model’s editability and the scalability of its accuracy have very high optimization potential.

​From Creating Objects to Creating Worlds, The Future Evolution of AI 3D

From understanding and generating objects to understanding and generating entire worlds has always been VAST’s consistent technology and product roadmap.

At its inception, VAST bet on 3D — the most primitive, natural, and information-dense content modality. At that time, VAST made it clear: the world model is the ultimate form of a general model and must be built on a native understanding of three-dimensional space.

In 2025, VAST has already focused its core R&D resources on the world model direction. The previously accumulated database of 50 million high-quality 3D and world models, the team with the highest global talent density intersecting AI and graphics, and the systems-level research and engineering experience accumulated as an industry pioneer in 3D understanding and generation constitute VAST’s unique advantages on this path.

VAST is committed to building a general world model. It will not only be able to generate interactive virtual worlds but also possess the ability to perceive, understand, and physically simulate, serving a wide range of next-generation interactive content, embodied intelligence, simulation, and more.

This investment is also based on VAST’s belief in the deep connection between AI 3D and world models: the former is an indispensable foundation for the latter, and 3D is a universal interface for both machines and humans. No matter what form world models ultimately take, native understanding and generation of three-dimensional space will be a crucial component.

In the view of VAST founder and CEO Song Yachen, choosing to “believe first rather than see first” is what distinguishes a startup from a major platform and is key to establishing a first-mover advantage.

AI 3D was once seen as a market with high investment and high uncertainty. When VAST was founded in 2023, there were no convergent proven paths or mature technical frameworks in the industry, so many people did not choose this highly challenging path.

Over the past three years, VAST has continuously pushed the industry standard limits of algorithms through sustained investment, lowering the barrier for creators and helping to broaden the usable boundaries of the industry.

Now, over 6.5 million creators are creating on the ecosystem VAST has built, and 90,000 enterprises and developers are building applications there. Each VAST user, in turn, makes the technology more mature, the product more complete, the creation barrier lower, the imagination of content consumption richer, and the starting point for later entrants higher — when individual creations converge, they ultimately change the entire system ecosystem.

Investors' Judgment and Thoughts on AI 3D

​Hengxu Capital:

Our assessment of AI technology has always revolved around one core question: can it truly land and create value? VAST’s 3D generation solution impressed us with its technical maturity and product completeness — not only are the results leading, but more importantly, the solution is “usable.” From automotive design to embodied intelligence simulation, from industrial digital twins to smart manufacturing, efficient generation of 3D content is a key link in bridging the virtual and real. VAST’s pipeline is ready; the technology has been validated. It is precisely the ideal target we are looking for: “AI empowering a thousand industries.” We believe that breakthroughs in world models will redefine the way humans interact with the physical world, and VAST stands at the forefront of this transformation.

​Yuanhe Puhua:

Looking back at the evolution of major content platforms, from PGC to PUGC to UGC, each leap often requires fundamental changes in creation tools. ByteDance reshaped short video creation and distribution with recommendation algorithms and low-barrier tools. VAST is doing the same — lowering the threshold for 3D creation to an unprecedented level, which holds the potential to incubate a new UGC interactive content platform. We deeply recognize VAST’s original intention and vision of “everyone can create 3D,” and today they have already gathered a group of active creators, forming a unique community culture. We are particularly optimistic about VAST’s generalization ability in the toC direction — AI 3D creation will soon become as simple as shooting a short video, and the related content ecosystem will see explosive growth. We believe VAST can use the key of world models to open the door to the next generation of interactive content platforms.

​BV Baidu Ventures:

3D generation is an important foundation for the future virtual world and a key direction for multimodal generation and world models. Compared to video generation, 3D naturally carries physical laws, spatial relationships, and object interaction logic. It is an essential path for AI to understand and simulate the real world. In the past, the industry paid less attention to the 3D modality compared to video, but now 3D generation has reached a significant inflection point in both technical breakthroughs and commercial deployment. This is the golden window for investment.

From Fei-Fei Li’s World Labs to Luma, Google, and other companies, world models have become an important development direction for AI. We believe that VAST is a high-quality team with outstanding advantages and deep focus on world models in China — possessing both solid technical strength and productization capabilities, while simultaneously laying out creation tools and content platforms. It has formed a positive growth flywheel in data accumulation and commercial expansion.

We continue to be bullish on VAST leveraging its differentiated advantages in high-quality 3D data to maintain competitiveness and long-term growth potential in the global generative AI field.

​Spring Capital:

As an early investor in VAST, we have witnessed the team’s complete journey from technological breakthrough to commercial validation. Our decision to follow on in this round stems from the company’s business development far exceeding expectations over the past year: whether in revenue growth or user data, strong momentum is evident. What excites us even more is that the vision VAST anchored at its founding — to build a UGC interactive content platform — is turning from blueprint into reality. With continued breakthroughs in world model technology, the popularization of 3D creation is no longer a distant future but a tangible present. We firmly believe VAST will become the leader in this generational change.

​Beijing AI Industry Investment Fund:

After the previous round, we have increased our investment in VAST again, driven by our recognition of the team’s execution ability and technological leadership. Over the past six months, VAST’s commercialization progress has exceeded expectations, achieving large-scale applications in 3D printing, industrial design, cultural entertainment, and other fields, with revenue showing explosive growth. If Zhipu AI represents Chinese AI companies’ global competitiveness in foundational large models, then VAST is writing a similar story in the 3D and world model track. As a key AI model-layer enterprise supported by Beijing, VAST not only represents the forefront of technological innovation but also is a typical example of AI empowering new productive forces and promoting the deep integration of the digital economy and real economy. We look forward to VAST continuing to lead industry development and winning greater voice for China’s AI industry in global competition.