Skip to main content

NC AI to Debut ‘Monster Sound Generation and Conversion AI’ at INTERSPEECH 2025


Technology Promises to Transform How Game Studios Create the Roars of Their Fantastical Worlds

NC AI, one of Korea’s leading multimodal AI research organizations, will unveil a suite of advanced Monster Sound Generation and Conversion technologies at INTERSPEECH 2025, the world’s largest conference dedicated to speech and language technology.

The annual gathering, hosted by the International Speech Communication Association, will take place from August 17 to 21 in Rotterdam and this year centers on “Fair and Inclusive Speech Science and Technology,” a theme that underscores respect for the immense diversity of human and increasingly non-human voices.

At the conference, NC AI will present two research papers: one outlining the architecture and training methodology of a high-fidelity timbre conversion model tailored for monster sound production, and another detailing its real-time, web-based demo system. Attendees will be able to speak into the system and hear their own words transformed instantly into the roar or growl of a specific creature. An online demo will also be available for those following remotely.

The technology represents a significant leap forward for major MMORPG developers, where the creation of lifelike monster audio traditionally demands painstaking manual work. By analyzing sound at CD-quality resolution (44.1 kHz), NC AI’s model captures subtle textures of breath, metallic resonance, and vocal strain while preserving the underlying linguistic content. The system interprets both what is being said and how it is being expressed, enabling it to convert laughter, snarls, breathing, and other non-verbal cues with impressive naturalism. Intensity variations are processed at intervals of 0.005 seconds, producing an immediacy that makes in-game creatures feel vividly alive.

What once required sound designers to handcraft dozens of variations for each scenario can now be automated. The system expands the frequency range of human speech and reproduces the complex timbres associated with monstrous characters. Developers can finely tune attributes that convey aggression, intimidation, or playfulness, meaning the same monster can sound entirely different depending on its emotional or combat state.

The work rests on years of high-quality data construction. NC AI’s Audio AI team partnered with NCSOFT’s Sound Center to annotate an expansive game audio archive, categorizing acoustic properties such as timbre, spatial presence, noise characteristics, and atmospheric tone. The team also used specialized transformation tools like Dehumanizer to fabricate extreme vocal textures that cannot be captured through live recording, enabling the model to support an unusually wide range of non-human voices. NC AI presented its data augmentation methodology earlier this year at the Spring Conference of the Korean Acoustical Society, earning strong recognition from both academic and industry attendees.

Benchmark tests show the NC AI system outperforming several state-of-the-art models including DDDM VC, Diff HierVC, and Free VC, across measures of audio quality, naturalness, timbre similarity, and linguistic preservation. Researchers attribute the leap in performance to advances in high-resolution audio processing, optimized style controls, joint analysis of verbal and non-verbal elements, and improved texture and intensity reconstruction.

The technology also forms the backbone of NC AI’s generative sound effects tool Sound Palette, which allows creators to specify mood and timbre to instantly produce hundreds of variations. The tool is designed not only for gaming but for applications across film, advertising, XR experiences, and the broader digital content industry.

This research marks another milestone in NC AI’s effort to position itself as a central player in Korea’s push for AI sovereignty. The company was recently selected for the nation’s Independent AI Foundation Model Project, which aims to advance homegrown AI technologies for global competitiveness.

NC AI plans to release its INTERSPEECH research materials and demonstration videos through official channels in hopes of fostering new collaborations and accelerating the commercialization of AI-based audio tools. The company sees the conference as a timely opportunity to deepen partnerships with global researchers, studios, and platform providers.

“As a leading research institution in Korea’s multimodal AI field, we have merged years of game audio expertise with advanced modeling techniques to bring this technology to life,” said Nam-hyun Jo, Head of NC AI’s Audio AI Team. “Our goal is to use AI to turn creators’ imagination into reality and to redefine what immersive audio means for the digital content industry.”