Audible World Models: Spatially Aware Sound Generation for 3D Worlds

Duowen Chen1, Jinjin He1, Gouthaman KV2, Sandeep Bangalore Venkatesh2, Bo Zhu1
1Georgia Institute of Technology 2Dolby Laboratories

A training-free framework that turns generated visual worlds into explicit audible world states with semantic sound sources, geometric grounding, and listener-dependent acoustic propagation.

NeurIPS 2026
Pipeline overview of Audible World Models.

From a single text prompt, Audible World Models builds a panoramic 3D proxy, identifies sound-producing objects and ambient regions, synthesizes dry audio for each source, anchors the sources to reconstructed geometry, and renders listener-dependent spatial audio with geometric acoustic propagation.

Abstract

Text- and image-conditioned world generators can create visually rich 3D environments, yet these worlds often remain silent or rely on soundtracks synthesized solely from text or rendered video. Although such audio can convey what should be heard, it lacks an explicit representation of where sound sources are located and how their perceived sound should vary with listener movement.

We introduce Audible World Models, a training-free framework that incorporates sound into the generated world state. Starting from a text prompt, our system constructs a panoramic 3D proxy, separates it into semantic layers, identifies sound-producing foreground objects and ambient background regions, and synthesizes dry audio for each sound label. It then anchors these sources to reconstructed geometry and renders listener-dependent spatial audio using geometric acoustic propagation. By explicitly linking semantics, geometry, and sound propagation, the framework maintains persistent source locations while adapting the rendered audio to changes in listener viewpoint and motion.

Experiments across 80 generated scenes demonstrate substantial gains in spatial consistency over text-, video-, and panorama-conditioned baselines, while preserving competitive semantic alignment. VLM-based assessments and human evaluations further indicate that our soundtracks are preferred for their audio–visual consistency, spatial plausibility, and motion-dependent behavior.

System Overview

Stage 1

Panoramic World Proxy

A text prompt is converted into a global panorama and layered 3D proxy, giving the system access to audible sources that may sit outside the current rendered view.

Stage 2

Semantic Audio Inventory

VLM agents parse foreground objects and broader background regions, assign concise sound descriptions, and mark each source as localized or diffuse for later spatialization.

Stage 3

Source Synthesis and Grounding

Text-to-audio synthesis produces dry clips for each label. Localized sounds are placed on object meshes or world sheets, and diffuse sounds are distributed over semantic background regions.

Stage 4

Acoustic Rendering

An acoustic parameter agent selects stable scene and source settings, then GSound renders binaural, listener-dependent audio through geometry-aware propagation along the trajectory.

Representative Video Examples

Each scene compares three soundtracks for the same camera trajectory: Stable Audio Generation (text-conditioned), MMAudio (video-conditioned), and Ours.

Headphones recommended to hear the spatial (binaural) effects. Videos and audio are compressed for the web.

Scene Prompt

A watercolor panorama of a windy beach concert, distant stage lights and crowd shapes, foreground as broad wet sand wash, no text.

Stable Audio Generation

MMAudio

Ours

Scene Prompt

A metal foundry with a glowing ladle edge and heat shimmer in the foreground, workers midground in protective gear, and towering furnaces with overhead cranes in the background, photoreal, no text.

Stable Audio Generation

MMAudio

Ours

Scene Prompt

A distant panorama of a crowded beach boardwalk, dense movement and kiosks in the distance, foreground kept as open sand strip, photoreal, no text.

Stable Audio Generation

MMAudio

Ours

Scene Prompt

A Monet-inspired impressionist scene of a wide meadow with distant poplars and a pale sky, foreground kept as soft color fields, no text.

Stable Audio Generation

MMAudio

Ours

Scene Prompt

A distant view of a large outdoor concert in a park, stage lights and crowd concentrated midground, foreground kept as empty grass and sky, photoreal, no text.

Stable Audio Generation

MMAudio

Ours

Scene Prompt

A wide view of a riverfront festival at night, a dense band of lights and crowds across the midground, foreground kept as dark open water, photoreal, no text.

Stable Audio Generation

MMAudio

Ours

Show 20 more scenes
Scene Prompt

A wide desert highway at night, a long line of distant headlights and taillights, foreground kept as empty sand and scrub, cinematic realism, no text.

Stable Audio Generation

MMAudio

Ours

Scene Prompt

A broad view of a factory yard, forklifts and loading docks active in the distance with flashing lights, foreground kept as an empty paved apron, photoreal, no text.

Stable Audio Generation

MMAudio

Ours

Scene Prompt

A temple courtyard in rain: a close stone lantern and wet tiles in the foreground, visitors midground under umbrellas, and a gate with hanging bells in the background with mist, photoreal, no text.

Stable Audio Generation

MMAudio

Ours

Scene Prompt

A school hallway between classes: a close row of lockers and posters in the foreground, students midground, and a bright stairwell and doors in the background, photoreal, no text.

Stable Audio Generation

MMAudio

Ours

Scene Prompt

A rural grain elevator site: a close gravel shoulder and road sign in the foreground, trucks midground, and tall silos with conveyor structures in the background under a wide sky, photoreal, no text.

Stable Audio Generation

MMAudio

Ours

Scene Prompt

A Monet-inspired impressionist view of a crowded riverside fair, colorful dots of people and lanterns across the midground, foreground as soft open water and reflections, no text.

Stable Audio Generation

MMAudio

Ours

Scene Prompt

A large rail yard panorama, multiple freight trains and signal lights in the distance, foreground mostly empty ballast and track edges, photoreal, no text.

Stable Audio Generation

MMAudio

Ours

Scene Prompt

A wide panorama of a container port at dawn, cranes and forklifts active in the distance with flashing beacons, foreground mostly open water and empty pier edge, photoreal, no text.

Stable Audio Generation

MMAudio

Ours

Scene Prompt

A busy commercial kitchen: a stainless prep counter with utensils in the foreground, cooks moving midground, and stoves with vents and shelves in the background under harsh task lighting, photoreal, no text.

Stable Audio Generation

MMAudio

Ours

Scene Prompt

A train station platform at dusk: a luggage suitcase and yellow safety line in the foreground, commuters waiting midground, and an arriving train with signal lights in the background beneath a long canopy, photoreal, no text.

Stable Audio Generation

MMAudio

Ours

Scene Prompt

A low-poly stylized scene of an open prairie with a distant wind farm, simple shapes, big sky, minimal foreground objects, no text.

Stable Audio Generation

MMAudio

Ours

Scene Prompt

A watercolor illustration of a wide coastal horizon with soft waves and distant sailboats, lots of negative space, gentle washes, no text.

Stable Audio Generation

MMAudio

Ours

Scene Prompt

A high mountain plateau with distant snow peaks, broad open rock and sparse grass, foreground kept empty, wide-angle realism, no text.

Stable Audio Generation

MMAudio

Ours

Scene Prompt

A downtown street parade at noon, with marching performers, confetti in the air, storefront reflections, crowd barriers, and a cross street filled with sunlight, realistic wide-angle.

Stable Audio Generation

MMAudio

Ours

Scene Prompt

A small village square, with a fountain in the center, people chatting near the shops, and a bicycle passing along the edge of the square.

Stable Audio Generation

MMAudio

Ours

Scene Prompt

A subway tunnel platform, with a train approaching in the distance, footsteps along the platform, and scattered voices under the station lights.

Stable Audio Generation

MMAudio

Ours

Scene Prompt

A beach at golden hour, with waves breaking on the left, children playing near the water, and seagulls circling above.

Stable Audio Generation

MMAudio

Ours

Scene Prompt

A working orchard during harvest, with long rows of trees, crates and ladders, a small tractor on a dirt lane, and a gravel road leading to a shed, photoreal natural light.

Stable Audio Generation

MMAudio

Ours

Scene Prompt

A windswept coastal village quay with boats and fluttering flags, post-impressionist painting inspired by Vincent van Gogh, bold swirling strokes, dramatic sky, vivid color, no text.

Stable Audio Generation

MMAudio

Ours

Scene Prompt

A riverside summer festival at sunset, with a small live band on an open-air stage, food trucks and string lights along the path, crowds moving between stalls, and water reflecting the last light, cinematic wide shot.

Stable Audio Generation

MMAudio

Ours

BibTeX

@inproceedings{chen2026audible,
  title     = {Audible World Models: Spatially Aware Sound Generation for 3D Worlds},
  author    = {Chen, Duowen and He, Jinjin and KV, Gouthaman and Bangalore Venkatesh, Sandeep and Zhu, Bo},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2026}
}