Giving Humans and AI Complete Sonic Freedom: Noiz''s Vision for the Sound Layer

Noiz wants to own one layer of the AGI stack — the sound layer that hears, understands, and responds — built on three converging paths: real-time inference, unified modeling, and acoustic memory.

Giving Humans and AI Complete Sonic Freedom: Noiz''s Vision for the Sound Layer | Noiz

"Give humans and AI complete sonic freedom."

That line is written into Noiz's mission. It sounds ambitious, but behind it is a very specific evolution: from one person reproducing a voice-cloning paper from scratch, to an open-sourced MockingBird used by developers worldwide, to a product serving millions of creators and developers today, with audio understanding, generation, and editing folded into a single framework.

If AGI is a stack of layered capabilities, Noiz wants to own one specific layer: the sound layer — responsible for hearing, understanding, and responding, connecting content creation, embodied intelligence, and Agent interaction.

Contents

It started with a book and MockingBird

Founder Chen Weijia's path to starting Noiz was a deeply personal one. After Yufu Tech was acquired, he was looking for a new direction, and a dusty deep-learning book pulled him back toward voice. With no GPU on hand and his Python more than a little rusty, he still reproduced voice cloning from scratch and built the Chinese-language MockingBird, which gained serious traction on GitHub.

The MockingBird repo now points to noiz.ai: the open-source community was the starting line; the product is the long-term battlefield. That period left two lasting convictions: first, the real distance between engineering something that works and a research paper, and second, that a voice with genuine emotion and expressiveness was worth building a company around on its own.

In 2024, the team committed to "going deep on the AI sound layer," and Noiz AI took its current shape. The name NOIZ comes from "noise" — originally a term from inside the model, now the team's own symbol.

Three product pivots: generation, expression, creation

Noiz's product direction has gone through three growth phases, each one asking how to better serve what users actually need.

The earliest question was: can AI talk the way a person does?

Once cloning could produce speech, that turned out not to be enough: it also had to perform — emotion, pacing, and style all needed to be controllable.

After that: what creators actually wanted was full-track production — dubbing, scoring, sound effects, and editing, all inside the same studio.

The three pivots can be summarized as:

  1. Generation → get the words spoken aloud

  2. Control and expression → sound like the character, the brand, the scene

  3. Creation → a complete AI audio workstation

As the story goes at Tsinghua's x-lab "President's Cup" competition, the team stopped stacking feature counts and focused on one thing: does the expression actually land? Global registered users crossed a million within six months, with ARR nearing $4M — proof that this particular trade-off had real market feedback behind it.

What "the sound layer" means on the technical map

Noiz has spoken publicly about three main lines of work:

Real-time inference

Interactive drama, gaming, livestreaming, and Agent conversation all need low-latency audio. Work like AudioX-Turbo compresses multi-step diffusion down to very few steps — generating 10 seconds of audio in roughly 0.2 seconds on an RTX 4090 — laying the groundwork for reconstructing sound in real time.

Unified modeling across speech, music, and sound effects

Sound is physically the same phenomenon regardless of category. Unified modeling is what lets voice, ambience, and score stay coordinated within the same project, in service of "full-spectrum" listening.

Persistent acoustic memory

Agents and embodied systems need to remember "what this space sounds like" and "what timbre this user prefers." A sound layer built only on stateless TTS can't sustain long-term interaction — acoustic memory is the next piece of the puzzle.

From Audio-Omni (a unified framework for understanding, generation, and editing) to the AudioX series (joint audio-video generation), Noiz's research has never stopped moving forward. Only by pushing research forward can the product turn that capability into an interface creators can actually click.

Hear, understand, respond

The "sound layer" breaks down into a closed loop:

  • Hear: recognize the content, scene, emotion, and sound source

  • Understand: combine that with world knowledge to know what this audio is doing narratively

  • Respond: generate or edit the next appropriate piece of sound

That's different from a pipeline that's just "text in, audio out." A pipeline solves for throughput; a closed loop solves for interaction and creation. Chen Qian's read in interviews: voice will be one of the first interfaces to break out in the coming years, since the hardware can stay always-on at low power, and voice separation plus emotion recognition in complex soundfields remain largely unexplored territory.

Noiz's lead in TTS Skill installs on platforms like OpenClaw is one signal that developers already treat "an Agent that can talk" as table stakes; the next step is Agent-native audio capability, expanding from serving people to serving machines and software entities.

What creators can feel today

The vision translates into concrete product features:

  • text-to-speech: dubbing for Chinese short dramas, motion comics, and short-form video, with emotion control, pause markers, and saved voice presets

  • Voice Design: describe a sound the library doesn't have

  • Cloning: lock in a character from a 3–10 second sample

  • Sound effects and scoring: fill in ambience and action

  • API / Skills: embed into workflows and Agents

For individual creators, the sound layer means you don't need a recording studio to produce a professional audio track. For teams, it means voice assets that are reusable, localizable, and programmatically callable.

How this differs from "just another voice startup"

Voice AI is already its own major category, with leading companies valued in the tens of billions of dollars. Noiz's positioning isn't to repeat the "sounds more human" pitch — it's to go deeper on short-form emotion, full-spectrum sound, and model-product fusion:

  • Model layer: over a dozen audio models, with minor versions shipping as fast as every three days

  • Product layer: an all-in-one Studio, reducing the back-and-forth between editing software and outside sound-effect libraries

  • Ecosystem layer: open-source papers, GitHub projects, developer APIs

Chen Qian has described this as "model-product fusion": model iteration pulls the product forward, and real feedback pulls the model back in turn. Audio has an artistic dimension, and brand plus auditory taste will keep creating long-term differentiation — even as the underlying capability converges, workflow and scene understanding can still open real distance between competitors.


From reproducing a single paper to exporting audio tracks for millions of users every day, Noiz has stayed focused on one thing: giving sound the standing it deserves in the AI world. Expression needs warmth, scenes need texture, and interaction needs to actually hear and actually answer.

"Complete sonic freedom" isn't a slogan — it's the cumulative result of generation, expression, and creation stacked on top of each other. Once the sound layer is in place, picture, text, and action finally add up to the same world.

Next step: open Noiz AI's text-to-speech, save a character voice preset, and treat it as the first voice asset in your project.

Try Noiz AI Text-to-Speech ->

Frequently Asked Questions

What does "sonic freedom" actually mean?

On the creative side, any character, emotion, or environment can be expressed. On the technical side, generation, editing, and real-time interaction aren't constrained by legacy toolchains. On the interaction side, humans and AI communicate through natural sound.

What is Noiz AI?

Noiz AI is an AI sound creation toolkit for creators, brands, and developers. Core capabilities include text-to-speech, voice cloning, voice design, and sound effect generation, used across short-form video, podcasts, audiobooks, ads, education, game characters, and multilingual localization.

How is Noiz's "sound layer" practically different from plain TTS?

TTS only solves "read this text aloud." The sound layer covers the full chain of understanding, generation, and editing, and supports ambience, scoring, real-time interaction, and Agent integration, suited to producing complete audio content, not a single line.

Try Noiz for free