WorldSonus: real-time stereo for world-model video

48 kHz stereo, one chunk every 100 milliseconds

阅读论文
WorldSonus method figure: watch-and-write, training-only ShiftNCE, stereo data, and a prompt change halfway through.

We're releasing WorldSonus. While a world model is still drawing the next stretch of video, it can write stereo at the same time: a piece of picture, then a piece of sound.

Each chunk is 100 milliseconds of 48 kHz stereo. The paper timed one chunk at 41.2 milliseconds on a single NVIDIA H100. The real-time factor is 0.41. The paper, the code, and the weights are public. It is not in Noiz Studio yet.

World models are usually silent

A world model can keep generating an environment, but what it hands over is often picture only. WorldSonus sits after that and writes the sound. It looks at frames that have already arrived, plus the sound prompt you give it now, and writes the next left and right channels. Frames that are not out yet are invisible to it.

The authors are from the Hong Kong University of Science and Technology, Noiz, MetaX, and Shanghai Jiao Tong University. Corresponding authors are Zeyue Tian and Qifeng Chen. The full list is in the paper.

How the sound keeps up with the picture

It does not wait for the whole video. Sound follows the picture as far as the picture has gone.

One hundred milliseconds at a time

Each chunk is 100 milliseconds of 48 kHz stereo. That lines up with three frames of 30 fps video. The stereo codec inside also runs at 30 Hz, so one chunk of sound is three internal frames.

The model only keeps the last 5 seconds, which is 50 chunks. When the ring is full, the oldest chunk is dropped. A longer video does not keep growing the memory.

You can change the prompt halfway through

Want a different sound? Wait for the next 100 milliseconds and switch. What is already written stays. Earlier frames are not recomputed. In the 10-second training clips, 60% switch the prompt at the 5-second mark.

Left and right ears follow the camera

The output is a left and a right channel that follow a perspective camera. It is not mono copied twice. Training used 1,465 hours of audio: 999 hours from stereo video, 466 hours audio only. The stereo comes from two kinds of data: filtered real recordings, and panoramic surround folded into the camera view.

The video model in front only needs to pass frames it has already drawn. It does not open its internal state. SwanSphere does panoramic video and surround. WorldSonus does real-time stereo for an ordinary camera.

The picture is read two ways

DINOv3 reads the picture. One path makes a slightly longer summary, so the model can hold onto a sound event. The other path sends these three frames into the generation head, so a hit lands at the right time and place. The difference between neighboring frames is also used as motion.

ShiftNCE is training only. A frozen Synchformer scores whether the sound landed on the moment it should. After training, generation does not load it. In the pre-finetune ablations, taking ShiftNCE out moved DeSync from 0.831 to 1.067 on 4,096 clips of 10 seconds. Taking out the frames that go straight into the generation head jumped FAD from 2.42 to 15.82. Changing the chunk to 33 milliseconds or 200 milliseconds scored FAD 3.79 and 2.51. The release uses 100 milliseconds. Those numbers are the ablation, not the final table below.

WorldSonus method figure: watch-and-write, training-only ShiftNCE, stereo data, and a prompt change halfway through.
Method figure · PDF

Scores and samples

We mainly look at VGGish FAD: how far the generated sound sits from the reference sound. Lower is better. In Table 1 of the paper, WorldSonus scores 1.73 on 5-second clear-stereo VGGSound and 2.03 on the 30-second Interactive set.

Model

How it writes

FAD, 5 s clear-stereo VGGSound

FAD, 30 s interactive video

WorldSonus

Watch and write, 100 ms chunks, 41.2 ms on one H100

1.73

2.03

PrismAudio

Sees the whole video and the text first

2.09

6.08

ThinkSound

Sees the whole video and the text first

2.94

7.00

AudioX

Sees the whole video and the text first

3.00

5.52

V-AURA

Streaming mono, 640 ms chunks

4.06

13.33

The 5-second set has 4,096 clips. The 30-second set has 1,024. FAD uses the mid channel. Numbers are from Table 1 of the WorldSonus paper.

Timing still isn't the strongest

It does not win every column. On 5-second VGGSound, WorldSonus DeSync is 0.686. ThinkSound is 0.481 and PrismAudio is 0.539. Both sit tighter on the picture. AudioX is 0.986. V-AURA is 1.287. Writing the sound right away, and pinning it to the frame, still pull against each other.

Left-right placement, and changing the prompt mid-stream

On 30-second Interactive, WorldSonus BiasSkill is 15.90. PrismAudio is 4.19, ThinkSound is 0.63, AudioX is −0.01. After the prompt switches, dual-caption match is 25.68% on VGGSound and 23.44% on Interactive. Against holding the original prompt, the gain is about 0.051 and 0.053.

How people hear it

Twenty listeners heard 40 clips of 10 seconds and made 400 pairwise comparisons. They preferred WorldSonus 58.8% of the time against AudioX, 80.0% against ThinkSound, and 65.0% against PrismAudio. Real recordings still win. All of this audio is generated. None of it is a field recording.

Seven samples

One clip from each group on the project page, the first in the group. Seven in all.

JoyAI-Echo · 16 s
Zing-0.5 · 16 s
Wan 2.7 · 16 s
Happy Oyster · 12 s
LTX-2.3 · 11 s
LingBot-World · 10 s
HunyuanVideo-1.5 · 16 s

Paper, code, and weights

The paper, the inference code, and the weights are public. There is still no button for this in Studio. To run it yourself, follow the project setup. The paper and the model card do not list a price or a commercial service.

  • arXiv paper — arXiv:2610.08760, 6 October 2026. There is also an HTML edition.

  • Code — inference code. Add --stream and it writes the next 100 milliseconds of audio as soon as three frames are ready.

  • Weights — worldsonus_150k.pt (final C3 no-H0, 150k EMA, 638 inference tensors), audio_codec.pt (causal 48 kHz stereo decoder, 30 Hz, 64 channels), and z_stats.pt. DINOv3 and T5Gemma 2 download separately under their own licenses.

The project code and these weights are CC BY-NC 4.0: attribution required, non-commercial use only. Ask about the license before any commercial use.

The base command on the model card is:

python -m worldsonus.infer --video input.mp4 --prompt "A train passes beside a river." --output output.wav

Add --stream and it writes 100 millisecond chunks. The README says the public code does not promise a fixed latency. 41.2 milliseconds is only the paper's measurement on one H100.

How far you can take it today

The model card is plain about this: it is a research inference release, not a live system already in production. Training code and training data are not in the drop. The audio codec is the public causal version from SoundReactor. The paper also says its reconstruction is weaker than SoundReactor's non-causal decoder.

Generated sound can miss events in the picture. It can also add artifacts. Do not treat it as a recording, and do not treat it as evidence. If a recognizable voice shows up, you need permission first.

Two things still unfinished

The paper names two gaps. One is a better causal audio codec: keep writing sound as it decodes, and keep more detail. The other is cleaner spatial stereo data. In public sets, some so-called stereo is mono copied twice, or dirtied when the channels were split. Sound that follows people and the camera still needs more reliable spatial audio.

FAQ

Can I use WorldSonus in Noiz Studio right now?

Not yet. This is a standalone research inference release. You run the paper, the code, and the weights yourself.

What is in each chunk of audio?

100 milliseconds of 48 kHz stereo, lined up with three frames of 30 fps video. The paper timed one chunk at 41.2 milliseconds on a single H100. The real-time factor is 0.41.

Can I change the sound prompt halfway through?

Yes. Switch at the next 100 millisecond boundary. Audio already written stays. Earlier frames are not recomputed.

Can I use the weights commercially?

The project code and these weights are CC BY-NC 4.0: attribution required, non-commercial use only. DINOv3 and T5Gemma 2 each keep their own licenses.

Is the audio on this page a field recording?

No. It is generated. It can miss events in the picture and it can add artifacts. Do not treat it as evidence.

Does the public code guarantee 41.2 milliseconds?

No. 41.2 milliseconds is the paper's measurement on one NVIDIA H100. The GitHub README does not promise that latency for the public script.

阅读论文