Voice AI is one of the spaces Silicon Valley VCs have openly admitted to collectively missing.
The Noiz team has been at this longer than most: an open-sourced MockingBird, millions of users, ARR approaching $4M. Below are the ten calls we think are most worth remembering — for creators, investors, and peers alike.
Contents
Model-product fusion is a deliberate strategy, not a fallback
The category is still positive-sum — penetration is under 10%
Three near-term priorities: engineering, short-form, next-gen engine
Team DNA: algorithms and acoustic taste, in the same room
Noiz is nearly 30 people, mostly Gen Z, with close to half coming from Tsinghua and Peking University. Founder Chen Weijia was an early ASR engineer at Meta, later co-founded Yufu Tech (acquired by Tencent), and spent time deep in Agent work at TikTok. CEO Chen Qian brings nearly 20 years in music, having conducted for Peking University's China Music Society. Head of algorithms Zeyue Tian holds a PhD from HKUST and is first author on top-tier papers including AudioX and Audio-Omni.
Chen Qian made a point in the interview: a voice model needs more than algorithms — it needs auditory taste and artistic sensibility. Plenty of the team's engineers are former campus singing-competition winners or indie-band songwriters. The culture explicitly rejects pure grind-mode work — the belief is that you have to genuinely love life first, before you can develop a real feel for the emotion, texture, and rhythm of sound.
That explains why users keep bringing up Noiz's "taste": the product isn't built by translating paper metrics into buttons — it's built by putting does this actually sound right first.
No audio slot machines — closed-loop productivity instead
The industry is still stuck in a "pull the lever and hope" phase: regenerate repeatedly, roll the dice, low efficiency. Image AI has already reached a mature, precisely-editable stage; audio is inevitably heading toward the same understand → generate → edit trinity.
Noiz's product-layer goal is a Studio: users shouldn't have to jump between three tools — generation, editing, scoring, and multilingual localization should happen in one continuous flow. The model layer has already shipped over a dozen full-stack audio models, with minor versions shipping as fast as every three days. Audio-Omni getting accepted at SIGGRAPH 2026 is that "no slot machines" philosophy, written directly into a paper.
An auditory engine has to exist independently
Multimodal large models can package and output audio and video together, but Chen Qian holds a firm line: light waves and mechanical waves are physically two different things. Vision engines and audio engines are inherently separate at the architecture level. A visual world model has no audio data, so it can't generate a physically plausible impact sound; spatial hearing, 360° soundfields, and distance attenuation can't be trained from visual data either.
An independent audio model and auditory engine is a genuine long-term requirement. Lightweight fusion models are cheaper, but they can't support professional creative work or spatial audio.
Short-form is a structural variable, not a feature checkbox
ElevenLabs specializes further in long-form content: voice stability, cinematic quality. Noiz's other big bet is short dramas and short-form video — hooking a viewer in five seconds, high-density pacing, strong emotional intensity. That's not a "short-drama mode" bolted onto a menu — it requires retraining from the model layer up:
Data layer: heavy incorporation of short-video, short-drama, and e-commerce-native corpora
Output layer: full coverage from a 60 to a 95, paired with editing capability to fix flaws
Emotion layer: reinforced joy, anger, sorrow, and delight, tuned for dramatic tension and that satisfying, addictive quality
The choices of 400,000 short-video and motion-comic creators domestically have validated real demand for this path.
Match beats clarity
Film and TV dubbing wants studio-clean audio. E-commerce voiceover, on the other hand, shouldn't chase excessive clarity — a bit of ambient noise actually reads as more authentic. What users want isn't just "can I hear it clearly" — it's the right match: does the voice fit the scene, the platform, and the emotional register?
This judgment directly shapes product defaults. The same engine, serving an explainer video versus a tense short-drama standoff scene, should carry different parameter logic.
Model-product fusion is a deliberate strategy, not a fallback
Voice model vendors commonly run a dual track of model layer plus product layer. For Noiz, this is a deliberate choice: model iteration stays directly aligned with real user scenarios, and product feedback loops back to guide research, forming a high-signal flywheel. Model progress pushes the application forward; application data pulls the model forward — walking on two legs rather than waiting for the ecosystem to mature before adding a product layer.
The foundation can commoditize — taste never does
Will voice commoditize the way text models have? The team's read: the trend is real, but the art component sets a ceiling. Like music, voice can't fully turn into a standardized commodity. Brand, auditory taste, scenario-specific capability, tooling, and ecosystem all create lasting differentiation. Short-form video, film and TV, and hardware interaction will each end up with their own dedicated audio solutions.
The category is still positive-sum — penetration is under 10%
Lightspeed pegged ElevenLabs' share of the creator market at roughly 60%. Chen Qian's counterargument isn't about contesting that share head-on — it's about looking at the total market: AI voice tool penetration among creators is still under 10%. ElevenLabs adding $100M in ARR in a single quarter says the market is expanding, not that it's zero-sum.
Noiz has already crossed a million users globally, with ARR approaching $4M, closing seed and seed-plus rounds within roughly six months of formal incorporation, backed by investors including Northern Light Venture Capital and Innoangel Fund. The numbers say this category is still in an early expansion phase, not a fight over a fixed pie.
High-quality corpora are the long-term moat
Noiz puts high-quality corpora first. Beyond compliant sourcing, the real differentiator is the ability to curate and match data to specific scenarios. Architecture matters, but with top talent in the room, architecture alone rarely stays a permanent moat.
Three near-term priorities: engineering, short-form, next-gen engine
Noiz has three near-term priorities:
Engineering Audio-Omni into commercial-grade deployment
Doubling down on short-form scenarios to reinforce differentiation
Developing the next-generation AI audio engine, covering gaming, virtual spaces, and embodied and hardware auditory interaction
If you're building short dramas, motion comics, or high-density short-form video, the two calls easiest to put to use right away are: match beats clarity, and emotional intensity matters more than cinematic polish. Open Noiz Text-to-Speech, run the same confrontational line through both an "explainer" tone and a "dramatic" tone, and you'll hear exactly where the team has placed its bets.