Audio-Omni: From Regenerate-and-Pray to One Framework for Understanding, Generation, and Editing

Audio-Omni folds audio understanding, generation, and editing into a single framework, so creators stop rerolling generations and start editing them instead.

Audio-Omni: From Regenerate-and-Pray to One Framework for Understanding, Generation, and Editing | Noiz

Anyone who's done AI voiceover or sound design knows the drill: type a prompt, generate a take, it's wrong, generate again, still wrong, keep rolling the dice.

Image generation has already moved past this stage — from "hope for a good roll" to editing workflows where you point at what's wrong and fix exactly that. Audio is still stuck in the earlier phase. Understanding, generation, and editing usually mean three separate tools, three exported files, and three timelines to reconcile by hand.

Audio-Omni is built to change that: understand it, make it, fix it — all inside one framework. The work has been accepted to SIGGRAPH 2026, with contributors from HKUST, Tencent WeChat Vision, Peking University, and others. Noiz's head of algorithms, Zeyue Tian, is among the paper's authors, and Noiz is steadily turning the underlying capability into shipped product features.

Contents

Why creators need all three at once

Noiz CEO Chen Qian put the real production pipeline plainly in an interview:

Short-form video, short dramas, and ad voiceovers are rarely "generate once and you're done." You first need to understand what's happening in the source material and where the emotional beats land. Then you generate the dialogue or ambient sound. Then, on the timeline, you tweak a single word, strip out a layer of noise, or drop in one impact hit.

Split across three separate models, something gets lost at every handoff: the understanding model doesn't know what's editable downstream; the generation model doesn't know what was analyzed upstream; the editing model never sees the earlier context at all. The creator, meanwhile, is the one ferrying files between three different interfaces.

Audio-Omni's approach is a unified framework: general ambience, music, and speech handled in one system, with instruction-level editing built in from the start.

The architecture: how the two halves work together

The paper's structure boils down to two components working together:

A frozen multimodal large language model (MLLM)

Handles the "thinking": understanding audio content, scene, and speakers, while drawing on real-world knowledge. Prompt it with "the instrument a Led Zeppelin drummer would play," and the model can reason its way to drums, then generate the matching sound. That kind of reasoning is nearly impossible to train from text-audio pairs alone.

A trainable diffusion generator (DiT)

Handles the "doing": turning that understanding into high-fidelity audio. High-level semantics flow in through cross-attention; low-level timing, timbre, and video alignment are handled through signal-level splicing, which keeps lip-sync and on-screen action tight.

The team also built AudioEdit, a dataset of more than one million instruction-editing examples, combining real-video mining with programmatic synthesis, so "add this, remove that, restyle it" finally has data to learn from.

One interesting finding: when connecting to the MLLM, the second-to-last layer often works better for audio synthesis than the final layer. The last layer leans too hard toward "predict the next token," which smooths away acoustic detail; the second-to-last layer keeps more of the information that generation actually needs. That's a useful data point for anyone combining large language models with audio generation.

What it can actually do: nine generation types, five editing types

Generation side (selected)

  • Text-to-ambience, text-to-music

  • Video-to-voiceover, video-to-score

  • Text-to-speech, zero-shot voice conversion

  • Cross-lingual instructions (Chinese, Japanese, French, and more all work as prompts)

  • Knowledge-grounded generation, style continuation from a reference clip

Editing side (selected)

  • Add: drop in a new element into an existing ambience — a skateboard, a lion's roar

  • Remove: strip out a specific component, like singing or a passing plane

  • Extract: pull one sound out of a mix, like an ambulance siren

  • Style transfer: turn a dog bark into a knock, while keeping the rhythmic shape

  • Speech editing: change a few words while keeping the voice as close to unchanged as possible

For short-drama teams, "fix one line without re-recording the whole scene" saves a huge amount of post-production time. For video creators, "the picture's locked, now add ambience that actually fits" becomes something you can count on instead of hope for.

How this differs from "one giant model does everything"

Multimodal LLMs are genuinely good at fusing audio and video, but sound has its own independent physical dimensions — spatiality, material texture, reverb, beat alignment — that don't live in the same data as pixels.

The Noiz team has consistently argued for the long-term value of an independent auditory engine: even if the product surface is a single entry point, the underlying system still needs to be trained specifically for sound and aligned specifically to acoustic laws. Audio-Omni is a concentrated showcase of that engine capability; Noiz's product layer turns it into something creators can actually click, not a paper demo.

What this means for Noiz users

In the short term, you'll feel the shift directly inside Noiz's all-in-one Studio:

  • When generating dialogue, the system has a clearer sense of "what scene is this line in"

  • Score, sound effects, and voice stay coordinated within the same project, with fewer app switches

  • As editing capability ships, the number of "reroll and hope" attempts drops, and so does the number of steps it takes to get something right

Longer term, Audio-Omni represents audio's shift from "generation tool" to "creation system." Image tools already proved that once editing matures, output capacity jumps another order of magnitude. Audio is at the same inflection point now.

Open source and further reading

Noiz is also pushing forward on AudioX-Turbo in parallel, compressing the number of steps in multi-step diffusion inference to make real-time interactive use cases actually viable. Understanding and editing solve "is this right"; inference efficiency solves "can this keep up with a live interaction." Both lines eventually converge in the product.


The next round of competition in audio creation isn't just about sounding more human — it's about who understands the source material better and requires fewer redos. Audio-Omni brings understanding, generation, and editing into one framework, laying the groundwork for that whole pipeline.

If you want a first feel for what a unified audio workflow looks like, start with Noiz AI's text-to-speech and try pairing dialogue with ambience inside the same project.

Try Noiz AI Text-to-Speech ->

Frequently Asked Questions

What is Noiz AI?

Noiz AI is an AI sound creation toolkit for creators, brands, and developers. Core capabilities include text-to-speech, voice cloning, voice design, and sound effect generation, used across short-form video, podcasts, audiobooks, ads, education, game characters, and multilingual localization.

How is this different from a plain TTS tool?

TTS solves "read this text aloud." Audio-Omni covers understanding and editing of ambience, music, and voice as well, built for producing a complete audio track, not a single line.

What was Noiz's role in the paper?

Core team members are paper authors, and Noiz is progressively engineering the underlying research into features creators can actually use in the product.

Try Noiz for free