SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation

Yunrui Cai1,2,*, Xu Li1,†, Yucheng Zhou1, Jinchao Li1, Dingdong Wang2, Dongchao Yang2, Xixin Wu2, Chen Zhang1, Zhiyong Wu3, Pengfei Wan1, Helen Meng2,†
1Kling Team, Kuaishou Technology    2The Chinese University of Hong Kong    3Shenzhen International Graduate School, Tsinghua University
*Work done during an internship at Kuaishou Technology.    Corresponding authors.

SonicWeave is a unified flow-matching model for speech, singing, music, sound effects, and their mixtures. Text-conditioned general audio generation is moving beyond isolated speech, music, and sound-effect synthesis toward a single model that can compose them into controllable, coherent audio scenes. This unified setting is particularly challenging: heterogeneous components impose conflicting structural requirements on a shared backbone, while a complex mixed scene may contain locally distinct or overlapping content that demands fine-grained adaptation within the same clip. CPE-MoE combines a global text-and-diffusion prior with local acoustic evidence to route temporally contiguous audio chunks. A learned conflict gate favors the prior when local states are unreliable, while allowing local evidence to influence routing when a region departs from the global scene context. SonicWeave supports broad task coverage and fine-grained composition within complex scenes using a single set of weights. Across TTS, TTA, and TTM benchmarks, SonicWeave consistently improves over controlled Dense and Base-MoE baselines, while complex-scene evaluation demonstrates improved compositional quality and routing analyses reveal content-dependent expert specialization across diffusion phases.

Overview and method

Unified audio generation with CPE-MoE

SonicWeave teaser
Teaser figure
Global priorStructured scene text and the diffusion-step embedding summarize what to generate and at which phase. This chunk-independent prior provides a stable expert preference when local acoustic states are still unreliable.
Local evidenceChunk states pooled after joint self-attention reveal the acoustic content currently emerging in each region. A learned conflict gate lets this evidence influence routing when local content departs from the global scene context.
Temporally coherent chunk routingOne top-K assignment is shared by neighboring frames, while the selected experts still transform the original frame-level states. The chunk is the routing unit—not the representation unit—preserving short-range continuity and fine-grained modeling.
SonicWeave architecture and CPE-MoE
SonicWeave architecture figure
Method. A stereo VAE encodes audio into continuous latents, which are jointly modeled with a structured-caption stream in a flow-matching DiT. CPE-MoE groups audio tokens into contiguous chunks and fuses a global text-and-phase prior with local acoustic evidence for sparse expert routing. The selected experts transform the original frame states, so the chunk is the routing unit rather than the representation unit; text and time tokens remain on a stable shared path.

Public-task gallery

Speech, environmental audio, music, and singing

Text-to-Speech

English and Mandarin examples under the same model.

English speech

SeedTTS · EN
TTS
Transcript

“He didn't consider mending the hole — the stones could fall through any time they wanted.”

Structured prompt
Plain English speech.
<type>speech</type>
<lang>en</lang>
<speech>He didn't consider mending the hole — the stones could fall through any time they wanted.</speech>
<speaker_count>1</speaker_count>
<vocal>adult English voice, neutral tone</vocal>

SonicWeave output

Mandarin speech

SeedTTS · ZH
TTS
Transcript

“尽管是用民间资金办学,但亚奥学校的性质并不单纯。”

Structured prompt
Plain Mandarin speech.
<type>speech</type>
<lang>zh</lang>
<speech>尽管是用民间资金办学,但亚奥学校的性质并不单纯。</speech>
<speaker_count>1</speaker_count>
<vocal>adult Mandarin voice, neutral tone</vocal>

SonicWeave output

Text-to-Audio

Multi-source environmental audio from AudioCaps captions.

Multi-source environmental scene

AudioCaps · 104323
TTA
Caption

Horses trotting while wood clanks several times as a woman talks and birds chirp in the background alongside wind blowing into a microphone.

Structured prompt
Horses trotting while wood clanks several times as a woman talks and birds chirp in the background alongside wind blowing into a microphone.
<type>sfx</type>
<speaker_count>1</speaker_count>
<vocal>a woman talking in the background</vocal>
<sfx>horses trotting, wood clanking several times, birds chirping, wind blowing into a microphone</sfx>
<texture>multi-source environmental scene with background birds and microphone wind noise</texture>

SonicWeave output

Duck, water, and laughter

AudioCaps · 103898
TTA
Caption

Water splashing as duck quacks followed by brief laughter

Structured prompt
Water splashing as duck quacks followed by brief laughter
<type>sfx</type>
<sfx>water splashing, duck quacking, human laughter</sfx>
<texture>natural field recording</texture>

SonicWeave output

Text-to-Music and Singing

Instrumental and vocal music generation share the same weights and conditioning interface.

Solo cello performance

MusicCaps · _3GnNMrDuCs
TTM
Caption

This audio contains someone playing a piece on cello ranging from the low register up into the higher register. This song may be playing during a live performance.

Structured prompt
This audio contains someone playing a piece on cello ranging from the low register up into the higher register. This song may be playing during a live performance.
<type>music</type>
<music>cello performance ranging from the low register to the higher register</music>
<texture>possible live-performance setting</texture>

SonicWeave output

Upbeat country-rock singing

Song Describer · 233
Singing
Caption

A rock song with a country vibe, it has male vocals, drums, and acoustic guitar. It is an upbeat song.

Structured prompt
A rock song with a country vibe, it has male vocals, drums, and acoustic guitar. It is an upbeat song.
<type>singing</type>
<vocal>male vocals</vocal>
<music>upbeat rock song with a country vibe, drums, and acoustic guitar</music>

SonicWeave output

Complex-scene showcase

Multi-element composition in one clip

Late-night diner soft-rock singing

Complex Scene · 110
Mixed
Natural-language prompt

In a late-night roadside diner, an adult female singer performs, “Coffee goes cold, but the jukebox stays, playing our old song while the night leans into day.” Clean electric guitar, piano, bass, and light drums form a soft-rock backing.

Structured prompt
English soft-rock singing in a late-night roadside diner
<type>mixed</type>
<lang>en</lang>
<lyrics>Coffee goes cold, but the jukebox stays, playing our old song while the night leans into day.</lyrics>
<speaker_count>1</speaker_count>
<vocal>adult English female singer, clear mellow soft-rock tone</vocal>
<music>clean electric guitar, piano, bass guitar, and light drum kit</music>
<ambience>quiet late-night diner room tone</ambience>

SonicWeave output

Race-engineer dialogue

Complex Scene
Mixed
Natural-language prompt

During a race, a controlled male engineer says over English radio, “Box this lap; the rear temperatures are climbing.” The strained male driver replies, “Copy, but the steering is vibrating under braking.” “Understood, stay off the inside curb.” Engine roar, wind, tires, and radio pops collide. Calls are somewhat distorted by noise.

Structured prompt
English race-engineer and driver radio dialogue at high speed.
<type>mixed</type>
<lang>en</lang>
<speech>Box this lap; the rear temperatures are climbing. / Copy, but the steering is vibrating under braking. / Understood, stay off the inside curb.</speech>
<speaker_count>2</speaker_count>
<relation>pit engineer advising a race driver</relation>
<vocal>controlled male engineer and strained male driver through clipped team radio</vocal>
<sfx>Engine roar, wind, tires, and radio pops collide</sfx>
<ambience>race car at speed with strong wind trackside</ambience>
<texture>driver transmission is noisier and more distorted than the engineer response</texture>

SonicWeave output

Rural pondside conversation

Complex Scene
Mixed
Natural-language prompt

Beside a rural pond, a man urgently asks in Mandarin, “你要干啥?别往那边走,水边滑得很。” A younger voice replies excitedly, “我就看看那条鱼,刚才它跳起来了。” The man warns, “小心点,鞋底都是泥,摔一下可不得了。” Cicadas chirp, a distant dog barks, and water ripples against the muddy bank beneath light wind.

Structured prompt
rural outdoor Mandarin conversation with sudden questioning near a pond
<type>mixed</type>
<lang>zh</lang>
<speech>你要干啥?别往那边走,水边滑得很 / 我就看看那条鱼,刚才它跳起来了 / 小心点,鞋底都是泥,摔一下可不得了</speech>
<speaker_count>2</speaker_count>
<vocal>rural Mandarin accents, clear questioning male voice and younger excited voice</vocal>
<sfx>cicada chirping, distant dog barking, water rippling against muddy bank</sfx>
<texture>open rural pond atmosphere, light wind across microphone, dialogue remains intelligible over natural ambience</texture>

SonicWeave output

Indoor basketball commentary

Complex Scene · 007
Mixed
Natural-language prompt

A Mandarin sports commentator shouts, “三分出手了!球进了!现在只差一分,比赛还剩十五秒!” Tens of thousands of fans roar, stomp, blow air horns, and whistle inside a huge indoor arena.

Structured prompt
Mandarin sports commentary in a roaring indoor arena
<type>mixed</type>
<lang>zh</lang>
<speech>三分出手了!球进了!现在只差一分,比赛还剩十五秒!</speech>
<speaker_count>1</speaker_count>
<vocal>Mandarin male sports commentator, shouting into a microphone</vocal>
<sfx>massive crowd screaming, stomping bleachers, air horns, whistles</sfx>
<ambience>huge reverberant indoor stadium</ambience>

SonicWeave output

Home stir-frying narration

Complex Scene · 008
Mixed
Natural-language prompt

While stir-frying at home, a warm Mandarin-speaking woman says, “现在把葱姜蒜放进去,听到这个滋滋声就对了,火不要太大。” Oil sizzles fiercely in a wok, a spatula scrapes metal, and a range hood and occasional plates are audible.

Structured prompt
Mandarin home cooking narration with active stir-frying sounds
<type>mixed</type>
<lang>zh</lang>
<speech>现在把葱姜蒜放进去,听到这个滋滋声就对了,火不要太大。</speech>
<speaker_count>1</speaker_count>
<vocal>warm Mandarin female voice</vocal>
<sfx>oil sizzling loudly in wok, spatula scraping metal, ingredients popping in hot oil</sfx>
<ambience>small home kitchen, faint range-hood hum</ambience>

SonicWeave output

Alpine ski finish commentary

Complex Scene · 044
Mixed
Natural-language prompt

At an alpine ski finish, an English commentator exclaims, “She is two tenths ahead at the final split—now she only has to hold the line through the last gate!” Wind crosses the microphone, spectators ring cowbells, PA echoes, and a helicopter layers through the valley.

Structured prompt
excited English ski commentary in a windy alpine finish area
<type>mixed</type>
<lang>en</lang>
<speech>She is two tenths ahead at the final split—now she only has to hold the line through the last gate!</speech>
<speaker_count>1</speaker_count>
<vocal>energetic sports commentator with rising intensity and outdoor PA coloration</vocal>
<sfx>cowbells, crowd cheers, ski edges scraping packed snow, helicopter rotors</sfx>
<ambience>windy mountain finish zone with reflected loudspeaker sound across the valley</ambience>

SonicWeave output

Moving emergency transfer

Complex Scene · 046
Mixed
Natural-language prompt

As an ambulance stops, a paramedic says while moving, “Keep the mask sealed and watch the monitor; we are moving him inside on three.” The sirens swirl around, stretcher wheels cross pavement seams accompanied by tense, nerve-wracking background music, a monitor beeps, and footsteps stay close to the voice.

Structured prompt
English paramedic instruction during a moving emergency transfer
<type>mixed</type>
<lang>en</lang>
<speech>Keep the mask sealed and watch the monitor; we are moving him inside on three.</speech>
<speaker_count>1</speaker_count>
<vocal>adult English-speaking paramedic, breath-controlled urgent command captured on body microphone</vocal>
<sfx>ambulance siren, stretcher wheel impacts, oxygen-mask hiss, monitor beeps, fast footsteps</sfx>
<music>tense, nerve-wracking background music</music>
<ambience>hospital ambulance bay with hard exterior reflections and opening doors</ambience>

SonicWeave output

Windy mountain trail dialogue

Complex Scene · 076
Mixed
Natural-language prompt

On a mountain trail, two hikers say, “The blue marker should be beyond that ridge.” “I think we missed the turn near the stream.” “Then let’s check the map before the clouds close in.” Wind hits their jackets as gravel steps, birds, and distant thunder continue.

Structured prompt
English hiking dialogue on a windy mountain trail
<type>mixed</type>
<lang>en</lang>
<speech>The blue marker should be beyond that ridge. / I think we missed the turn near the stream. / Then let us check the map before the clouds close in.</speech>
<speaker_count>2</speaker_count>
<vocal>two slightly breathless adult English voices at different walking distances</vocal>
<sfx>boots on loose gravel, map unfolding, jacket fabric, distant thunder</sfx>
<ambience>exposed mountain path with gusting wind, sparse birds, and a faint stream below</ambience>
<texture>wind and motion noise fluctuate while the two voices alternate naturally</texture>

SonicWeave output

Roadside survey-car sighting

Complex Scene · 033
Mixed
Natural-language prompt

A bus passes by and two men excitedly say in Mandarin, “快看,是 Google Map 的采集车!” “真的是,它刚从这里开过去。” Busy road traffic and urban road noise fill the background.

Structured prompt
excited Mandarin dialogue after a bus passes by over busy road traffic
<type>mixed</type>
<lang>zh</lang>
<speech>快看,是 Google Map 的采集车! / 真的是,它刚从这里开过去。</speech>
<speaker_count>2</speaker_count>
<vocal>two excited adult Mandarin male voices</vocal>
<sfx>passing vehicle engines</sfx>
<ambience>urban roadside traffic noise</ambience>
<texture>conversation competes with outdoor road sound</texture>

SonicWeave output

Controlled comparison

Dense, Base-MoE, and SonicWeave

Prompt
Dense
Base-MoE
SonicWeave
Airport gate

At a busy airport gate, a female Mandarin PA announcer says, “前往上海的旅客请注意,CA八三七航班现在开始登机。” then repeats in English, “Passengers traveling to Shanghai on flight CA eight three seven may now board through gate twenty-six.” Her formal, slightly nasal voices come through the public-address system. Rolling suitcases, distant boarding beeps, overlapping conversations, and footsteps fill the terminal. Bright reflections from the large hall and PA-system compression color the sound.

Structured prompt
busy bilingual airport gate announcement with overlapping travelers and rolling suitcases
<type>speech</type>
<lang>zh,en</lang>
<speech>前往上海的旅客请注意,CA八三七航班现在开始登机。Passengers traveling to Shanghai on flight CA eight three seven may now board through gate twenty-six.</speech>
<speaker_count>1</speaker_count>
<vocal>Mandarin female PA announcer followed by English, formal broadcast diction, slightly nasal loudspeaker tone</vocal>
<ambience>busy airport terminal, suitcase wheels, distant boarding beeps, overlapping conversations, footsteps on polished floor</ambience>
<texture>PA system compression, bright reflections from large hall</texture>
Dense
1B dense baseline
Base-MoE
Token-level routing
SonicWeave
Chunk-routed CPE-MoE
VHF storm rescue

Over VHF radio in a violent storm, a panicked sailor calls, “Mayday, mayday, this is Harbor Seven, we are taking on water!” A coast guard operator answers, “Harbor Seven, rescue boat launching now, hold your position!” Waves, rain, wind, static, and dropouts batter the transmission.

Structured prompt
A sailor and coast guard exchange an emergency radio call during a violent storm.
<type>mixed</type>
<lang>en</lang>
<speech>Mayday, mayday, this is Harbor Seven, we are taking on water! / Harbor Seven, rescue boat launching now, hold your position!</speech>
<speaker_count>2</speaker_count>
<relation>panicked sailor requesting rescue and coast guard operator responding</relation>
<vocal>panicked sailor and coast guard operator over VHF radio</vocal>
<sfx>waves, rain, wind, static, and dropouts</sfx>
<ambience>violent storm at sea</ambience>
<texture>radio transmission battered by static and dropouts</texture>
Dense
1B dense baseline
Base-MoE
Token-level routing
SonicWeave
Chunk-routed CPE-MoE
Morning market

At a morning market, a customer and vendor bargain in Mandarin: “老板,这个西红柿多少钱一斤?” “四块五,今天刚进的货,特别新鲜。” “能便宜点吗?我多买一点。” Plastic bags, vendor calls, scooters, and cooking from a nearby stall blend together.

Structured prompt
A customer and vendor bargain in Mandarin at a busy morning market.
<type>mixed</type>
<lang>zh</lang>
<speech>老板,这个西红柿多少钱一斤? / 四块五,今天刚进的货,特别新鲜。 / 能便宜点吗?我多买一点。</speech>
<speaker_count>2</speaker_count>
<relation>customer bargaining with a vendor</relation>
<sfx>plastic bags, vendor calls, scooters, and cooking from a nearby stall</sfx>
<ambience>morning market</ambience>
<texture>foreground dialogue blended with surrounding market activity</texture>
Dense
1B dense baseline
Base-MoE
Token-level routing
SonicWeave
Chunk-routed CPE-MoE
Matched inputs highlight cross-lingual speech, event responsiveness, speaker structure, and foreground–background composition.

External systems

Complex generation against strong public systems

Prompt
Higgs Audio V2
Dasheng AudioGen
SonicWeave
Busy gym

In a busy gym, a trainer shouts, “Three more reps, you’ve got this! Keep your core tight. Two more!” Weight plates clank, treadmills run, people breathe heavily, and battle ropes strike the floor around him.

Structured prompt
A trainer shouts encouragement over the sounds of a busy gym.
<type>mixed</type>
<lang>en</lang>
<speech>Three more reps, you’ve got this! Keep your core tight. Two more!</speech>
<speaker_count>1</speaker_count>
<relation>trainer addressing gym participants</relation>
<vocal>trainer shouting encouragement</vocal>
<sfx>weight plates clanking, treadmills running, heavy breathing, and battle ropes striking the floor</sfx>
<ambience>busy gym</ambience>
<texture>foreground speech over dense gym activity</texture>
Higgs Audio V2
Dasheng AudioGen
SonicWeave
Speech, music, beeps, and cat

A man says, “Let’s see if this works now.” Background music is playing; a series of electronic beeps is followed by a cat meow.

Structured prompt
A man speaks over background music before electronic beeps and a cat meow.
<type>mixed</type>
<lang>en</lang>
<speech>Let’s see if this works now.</speech>
<speaker_count>1</speaker_count>
<vocal>a man speaking</vocal>
<music>background music</music>
<sfx>a series of electronic beeps followed by a cat meow</sfx>
<texture>background music under foreground speech with ordered sound events</texture>
Higgs Audio V2
Dasheng AudioGen
SonicWeave
Dialogue with electronic music

Two women discuss game rules: “You move only after the timer starts.” “Got it, then I will wait for the signal.” Energetic electronic music with bouncy synthesizers and drums plays underneath.

Structured prompt
Two women discuss game rules over energetic electronic music.
<type>mixed</type>
<lang>en</lang>
<speech>You move only after the timer starts. / Got it, then I will wait for the signal.</speech>
<speaker_count>2</speaker_count>
<relation>two women discussing game rules</relation>
<vocal>two women speaking</vocal>
<music>energetic electronic music with bouncy synthesizers and drums</music>
<texture>music playing underneath the dialogue</texture>
Higgs Audio V2
Dasheng AudioGen
SonicWeave
Classroom rehearsal

In a large classroom, children rehearse in English: “Wait, you’re supposed to say the magic word first!” “Oh right, I keep forgetting. Let me try again.” “Okay, from the top!” Chairs scrape, pencils tap, classmates whisper, and pages turn around them.

Structured prompt
Children rehearse English dialogue in a large classroom.
<type>mixed</type>
<lang>en</lang>
<speech>Wait, you’re supposed to say the magic word first! / Oh right, I keep forgetting. Let me try again. / Okay, from the top!</speech>
<speaker_count>2</speaker_count>
<relation>children rehearsing a dialogue</relation>
<vocal>children speaking in English</vocal>
<sfx>chairs scraping, pencils tapping, classmates whispering, and pages turning</sfx>
<ambience>large classroom</ambience>
<texture>foreground dialogue surrounded by classroom activity</texture>
Higgs Audio V2
Dasheng AudioGen
SonicWeave