AI Music's Largest Single Funding Round: ACE Raises $40 Million in Pre-B Round | 5Y News
ACE Studio has accumulated nearly 100,000 paying creators and over 3 million professional production tracks.


The AI music generation space appears to have moved past the "type a sentence, get a song" phase. Competition now centers on controllability. While leading products like Suno and Udio can simultaneously generate lyrics, composition, arrangement, and vocals, they still grapple with unstable fine-grained control and murky data copyright issues.
Jing Guo has spent seven years working on precisely these problems.
In 2019, Guo founded ACE, initially exploring virtual singer products and accumulating expertise and data around localized control of vocals and instruments. In 2024, the team moved into professional music production, launching ACE Studio to capture more granular music production trajectories within real workflows. Starting in 2025, they began developing proprietary music models using this data, upgrading ACE Studio into an AI music creation platform supporting stem separation and localized editing — a direct improvement on musical controllability.
But Guo doesn't want to stop at the professional market. Building on this foundation, the team also developed Miya, a music content product for general consumers. In his view, "music is fundamentally about expression — it shouldn't be confined to the professional realm."
In his student days, Guo formed bands and wrote songs, yet complex arrangement and music theory knowledge once convinced him that music creation belonged to professionals. Now, he believes it was complicated software and mistaken assumptions that built music's mystique. What he ultimately wants to do is "remove the barriers to musical expression and let ordinary people participate in music production."
ACE's team now numbers nearly 40. Core members combine musical expertise with AI technical capabilities. Other team members include engineers from Tencent and ByteDance, as well as professional musicians and art educators. The company recently completed a Pre-B round, raising $40 million from CCV, Shunwei Capital, Alphaist, and other strategic investors, with CCV making consecutive large follow-on investments. Previous shareholders include 5Y Capital, Boyu Capital, and overseas investor HF0.
Before founding ACE, Guo worked on domestic growth for mass-market casual games at iDreamSky, including Fruit Ninja, Temple Run, and Angry Birds. Co-founder Wenxiao Zhao (CTO) brings expertise in machine learning, speech synthesis, music technology, and software systems architecture, having previously worked on game engine development at Tencent. Co-founder Conger Sheng (CPO) was among the youngest signed songwriter-producers at Warner Music China and has contributed to tracks for top-tier Chinese pop singers and idols.
Core algorithm researcher Dr. Ruibin Yuan graduated from HKUST and previously led development of Qwen's music model. The team has disclosed several academic achievements: lead author of MERT (Musical Representation learning with large-scale self-supervised Training), with over 7.4 million downloads on HuggingFace and adoption by leading AI music models including Suno; creator of a symbolic music language model integrating generation and understanding, shared and praised by Yann LeCun; and contributor to the MMMU benchmark used by OpenAI, Anthropic, and other major AI labs.

Building the Spotify of the AI Era
In 2008, music streaming platform Spotify emerged, transforming how the world consumed music. Previously, people bought CDs, downloaded individual tracks, or paid per song or album. Spotify let users access the entire catalog on demand, either by listening to ads or subscribing. Music shifted from a product purchased piece by piece to a service always available.
From ACE's perspective, music hasn't seen another breakthrough since that streaming revolution. Because production never transformed, subsequent platforms have merely expanded catalogs, improved playback experiences, and refined recommendations — still aggregating content, licensing music rights, then distributing and consuming.
They believe AI will finally deliver that breakthrough: the Spotify of the AI era will drive new content distribution and consumption through production-side transformation. ACE's roadmap to achieve this: first penetrate professional music production, then move toward mass content consumption. The team has largely completed step one and is now advancing into step two.
On the professional production side, core achievements center on ACE Studio, an AI production tool for professional musicians. Users can generate complete songs from a single prompt, remix existing works, generate individual instruments or vocal passages, input melodies and lyrics for AI singers to perform, or have different AI instruments play specified melodies. Once a song is complete, the platform can generate music videos or produce soundtracks and sound effects for video.
Its key distinction from "one-click song" products is that creators can split apart and adjust vocals, instruments, and localized segments within a song, then continue editing and refining — turning AI-generated music into workable production material that enters professional musicians' daily workflows. ACE Studio currently has nearly 100,000 paying creators, generating over $2 million in monthly revenue, 90% from overseas.
This product's early accumulation also significantly aided ACE's proprietary model development. The team's current models include the ACE Music Models and the open-source ACE-Step series, which can generate complete songs with vocals and accompaniment from text, lyrics, or reference audio, supporting genre control, cover versions, stem separation, accompaniment regeneration, and personal style training. Presently, model performance ranks just below Suno v5.5, but generation costs are one-fifth.

Caption: ACE's music models have trained for nearly a year, with performance approaching Suno v5.5
For mass content consumption, ACE relies primarily on Miya, an AI music creation and consumption platform for general users. People can generate songs simply by chatting about their experiences or expressing emotions — no music theory required — while also listening, sharing, and socializing within the platform. The product launched recently and remains in early stages.
Thus, ACE attempts to build outward from professional production tools, gradually forming a complete music platform covering models, creation, distribution, and consumption.
Currently, ACE uses its music models as foundational infrastructure, with ACE Studio serving as the production tool for professional musicians — generating subscription revenue while capturing high-quality creation trajectories. Miya pushes generative capabilities to ordinary people, letting users create songs through conversation and listen, share, and socialize on-platform. This user behavioral data in turn feeds back into model training, creating a closed loop from models to applications to data reinforcement.

Professional Music Creation Trajectories as Continuous Data Supply
A core competitive moat underlying ACE's models and products is granular, professional user data feedback.
When ACE first started with vocal synthesis and virtual singers, it accumulated a collection of vocal and single-instrument stem materials. During model pre-training, the team recombined and mixed these materials to batch-generate training songs with clear annotations. Compared to scraping complete songs from the internet, this approach reduces direct copyright music usage risks while helping the model more clearly learn internal song structures — the relationships between vocals, drums, bass, and other components.
Moreover, ACE Studio's nearly 100,000 paying creators include substantial numbers of professional musicians: pop music producers within the Hollywood and Grammy systems, Broadway, opera, and theater composers, independent musicians, arrangers, and commercial scoring professionals, plus numerous teachers and students from institutions like Berklee and CalArts. Their authorized creation trajectories serve as crucial training data for the models.
ACE initially attracted these professional users by addressing genuine music production pain points. Previously, after writing a melody, musicians needed to find suitable singers to record demos or complete harmonies within songs — costly, time-consuming work that music production constantly confronts. Starting in late 2022, ACE developed this functionality, transforming its early virtual singer technology into a professional production tool.
Later, during professional users' creative processes, ACE — with user consent — learned their production trajectories: how they programmed drum patterns, adjusted tempo, modified which notes and chords, added what instruments and strings, and ultimately how they stemmed, edited, and exported. Compared to crude prompt-to-song generation, these trajectories record professional musicians' judgments and production chains, substantially helping models learn professional production and musical aesthetics while improving the "gacha" problem of music generation.
In the post-training phase, the team also maintains continuous professional user engagement through weekly online training, professional user interviews, and university partnerships.
Senior producer SoulSpeak first establishes professional benchmark datasets. Then through connections at Berklee, CalArts, and other institutions, plus professional freelancer platforms, ACE continuously recruits musicians — organizing over a hundred freelance musicians monthly to evaluate outputs across melody, structure, arrangement, and orchestration. To ensure evaluator quality, ACE implements random problem assignment, timed responses, and anti-cheating exams to filter participants, maintaining data annotation and maintenance quality.
To date, ACE Studio has accumulated over 3 million professional creation trajectories. Throughout this process, ACE Studio has gradually added capabilities including full song generation, stem separation, localized editing, continuation writing, and fine-grained control, embedding itself into professional musicians' workflows.

The "Short Video Moment" for Music Creation
In Guo's view, AI music has been a consistently underestimated market. He believes that when the barrier to music creation is eliminated — when everyone can use music to express emotions, vent about life, tell stories — the logic of music production, distribution, and consumption will be completely rewritten.
And this potential remains largely unseen today.
Before short videos emerged, video production was similarly confined to professionals. When short videos took off, production barriers fell, and video shifted from a professional product to everyday expression. Platforms exploded with massive original content, distributed according to individual interests. Ultimately, short videos didn't merely create a new content format — they packaged live streaming, e-commerce, and local services into the same distribution system, reshaping video's production, consumption, and commercial ecosystem.
Guo believes music hasn't experienced equivalent transformation. Streaming changed how music is accessed, not who produces it. The entire market remains built on limited music supply and copyright constraints. Music has seen neither mass UGC nor true personalized distribution atop infinite supply.
When creation barriers drop, music supply will expand from limited daily new releases to massive, wildly diverse personalized content. Sufficiently rich supply will in turn drive genuine change on the consumption side: people will no longer merely chase stars, charts, and hit songs, but receive music closely matched to their own experiences, emotions, and immediate contexts.
Compared to AI video, Guo also believes AI music is more likely to produce an era-defining "AI Native" platform. On one hand, video creation already underwent short video's revolution; creation barriers have substantially dropped, and AI video mainly helps creators produce previously high-cost footage. On the other hand, AI-generated video still plays on short video platforms — an already near-"optimal" platform ecology that's difficult to disrupt again.
Music's ecological transformation space remains wide open. What AI music needs to do now, Guo believes, is "let people discover that music creation seems as simple as shooting short videos." This is why they developed Miya. Having built model foundations and completed the previous phase of data accumulation, they hope Miya will make music an everyday expression tool for ordinary people — turning casual conversation into songs, completely eliminating barriers of music theory and arrangement.
Another reason Guo is bullish on AI music's market prospects: "It's not a natural extension of large language models." Current general-purpose models can understand music, but pushing music to excellence sacrifices general intelligence — making music a domain where startups can still control the full stack from models to applications to data flywheels.
Guo now stands in this direction. What he ultimately wants to build is an AI-native platform simultaneously accommodating music production and consumption — bringing music its "short video moment."


