We're releasing WorldSonus. While a world model is still drawing the next stretch of video, it can write stereo at the same time: a piece of picture, then a piece of sound.
Each chunk is 100 milliseconds of 48 kHz stereo. The paper timed one chunk at 41.2 milliseconds on a single NVIDIA H100. The real-time factor is 0.41. The paper, the code, and the weights are public. It is not in Noiz Studio yet.
World models are usually silent
A world model can keep generating an environment, but what it hands over is often picture only. WorldSonus sits after that and writes the sound. It looks at frames that have already arrived, plus the sound prompt you give it now, and writes the next left and right channels. Frames that are not out yet are invisible to it.
The authors are from the Hong Kong University of Science and Technology, Noiz, MetaX, and Shanghai Jiao Tong University. Corresponding authors are Zeyue Tian and Qifeng Chen. The full list is in the paper.
How the sound keeps up with the picture
It does not wait for the whole video. Sound follows the picture as far as the picture has gone.
One hundred milliseconds at a time
Each chunk is 100 milliseconds of 48 kHz stereo. That lines up with three frames of 30 fps video. The stereo codec inside also runs at 30 Hz, so one chunk of sound is three internal frames.
The model only keeps the last 5 seconds, which is 50 chunks. When the ring is full, the oldest chunk is dropped. A longer video does not keep growing the memory.
You can change the prompt halfway through
Want a different sound? Wait for the next 100 milliseconds and switch. What is already written stays. Earlier frames are not recomputed. In the 10-second training clips, 60% switch the prompt at the 5-second mark.
Left and right ears follow the camera
The output is a left and a right channel that follow a perspective camera. It is not mono copied twice. Training used 1,465 hours of audio: 999 hours from stereo video, 466 hours audio only. The stereo comes from two kinds of data: filtered real recordings, and panoramic surround folded into the camera view.
The video model in front only needs to pass frames it has already drawn. It does not open its internal state. SwanSphere does panoramic video and surround. WorldSonus does real-time stereo for an ordinary camera.
The picture is read two ways
DINOv3 reads the picture. One path makes a slightly longer summary, so the model can hold onto a sound event. The other path sends these three frames into the generation head, so a hit lands at the right time and place. The difference between neighboring frames is also used as motion.
ShiftNCE is training only. A frozen Synchformer scores whether the sound landed on the moment it should. After training, generation does not load it. In the pre-finetune ablations, taking ShiftNCE out moved DeSync from 0.831 to 1.067 on 4,096 clips of 10 seconds. Taking out the frames that go straight into the generation head jumped FAD from 2.42 to 15.82. Changing the chunk to 33 milliseconds or 200 milliseconds scored FAD 3.79 and 2.51. The release uses 100 milliseconds. Those numbers are the ablation, not the final table below.
Scores and samples
We mainly look at VGGish FAD: how far the generated sound sits from the reference sound. Lower is better. In Table 1 of the paper, WorldSonus scores 1.73 on 5-second clear-stereo VGGSound and 2.03 on the 30-second Interactive set.
Model | How it writes | FAD, 5 s clear-stereo VGGSound | FAD, 30 s interactive video |
|---|---|---|---|
WorldSonus | Watch and write, 100 ms chunks, 41.2 ms on one H100 | 1.73 | 2.03 |
PrismAudio | Sees the whole video and the text first | 2.09 | 6.08 |
ThinkSound | Sees the whole video and the text first | 2.94 | 7.00 |
AudioX | Sees the whole video and the text first | 3.00 | 5.52 |
V-AURA | Streaming mono, 640 ms chunks | 4.06 | 13.33 |
The 5-second set has 4,096 clips. The 30-second set has 1,024. FAD uses the mid channel. Numbers are from Table 1 of the WorldSonus paper.
Timing still isn't the strongest
It does not win every column. On 5-second VGGSound, WorldSonus DeSync is 0.686. ThinkSound is 0.481 and PrismAudio is 0.539. Both sit tighter on the picture. AudioX is 0.986. V-AURA is 1.287. Writing the sound right away, and pinning it to the frame, still pull against each other.
Left-right placement, and changing the prompt mid-stream
On 30-second Interactive, WorldSonus BiasSkill is 15.90. PrismAudio is 4.19, ThinkSound is 0.63, AudioX is −0.01. After the prompt switches, dual-caption match is 25.68% on VGGSound and 23.44% on Interactive. Against holding the original prompt, the gain is about 0.051 and 0.053.
How people hear it
Twenty listeners heard 40 clips of 10 seconds and made 400 pairwise comparisons. They preferred WorldSonus 58.8% of the time against AudioX, 80.0% against ThinkSound, and 65.0% against PrismAudio. Real recordings still win. All of this audio is generated. None of it is a field recording.
Seven samples
One clip from each group on the project page, the first in the group. Seven in all.
Paper, code, and weights
The paper, the inference code, and the weights are public. There is still no button for this in Studio. To run it yourself, follow the project setup. The paper and the model card do not list a price or a commercial service.
arXiv paper — arXiv:2610.08760, 6 October 2026. There is also an HTML edition.
Code — inference code. Add
--streamand it writes the next 100 milliseconds of audio as soon as three frames are ready.Weights —
worldsonus_150k.pt(final C3 no-H0, 150k EMA, 638 inference tensors),audio_codec.pt(causal 48 kHz stereo decoder, 30 Hz, 64 channels), andz_stats.pt. DINOv3 and T5Gemma 2 download separately under their own licenses.
The project code and these weights are CC BY-NC 4.0: attribution required, non-commercial use only. Ask about the license before any commercial use.
The base command on the model card is:
python -m worldsonus.infer --video input.mp4 --prompt "A train passes beside a river." --output output.wav
Add --stream and it writes 100 millisecond chunks. The README says the public code does not promise a fixed latency. 41.2 milliseconds is only the paper's measurement on one H100.
How far you can take it today
The model card is plain about this: it is a research inference release, not a live system already in production. Training code and training data are not in the drop. The audio codec is the public causal version from SoundReactor. The paper also says its reconstruction is weaker than SoundReactor's non-causal decoder.
Generated sound can miss events in the picture. It can also add artifacts. Do not treat it as a recording, and do not treat it as evidence. If a recognizable voice shows up, you need permission first.
Two things still unfinished
The paper names two gaps. One is a better causal audio codec: keep writing sound as it decodes, and keep more detail. The other is cleaner spatial stereo data. In public sets, some so-called stereo is mono copied twice, or dirtied when the channels were split. Sound that follows people and the camera still needs more reliable spatial audio.