From raw waveforms to a speaking Gemma 4 — a guided course built from three sources
📚 SoundStream (2107.03312) · AudioLM (2209.03143) · Gemma 4 E4B speech head (Frisson Labs)
You know transformers: a sequence of discrete tokens goes in, next-token prediction comes out. Text has a natural tokenization (BPE): a sentence is ~15 tokens per second of reading. Audio is different in two brutal ways:
The whole field's answer, in one sentence: compress audio into discrete tokens at a manageable frame rate, then treat speech generation as pure language modeling — predict next token, then decode tokens back to a waveform. The two papers you gave me supply the two halves of that pipeline:
Read the chapters in order — each builds on the previous. Every chapter ends with a self-check quiz.
Paper: SoundStream: An End-to-End Neural Audio Codec
Authors: Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, Marco Tagliasacchi (Google)
Year: 2021 (arXiv:2107.03312; published in IEEE/ACM TASLP, Nov 2021)
---
Traditional audio codecs come in two flavors, and both fail at low bitrates:
SoundStream replaces the entire hand-engineered encode→quantize→synthesize pipeline with a single neural network trained end-to-end. It compresses speech, music, and general audio at 3–18 kbps, runs in real time on a smartphone CPU, is streamable (causal, low latency), and — uniquely — one single model covers all bitrates in that range via quantizer dropout. Headline result: SoundStream at 3 kbps beats Opus at 12 kbps and roughly matches EVS at 9.6 kbps in blind listening tests — i.e., 3–4× fewer bits for equal quality.
The reader is assumed to know transformers; the key conceptual shift is that this is a pure convolutional autoencoder + discrete bottleneck + GAN loss, no attention at all, because it must run causally in real time on audio samples.
Input: single-channel waveform x ∈ ℝᵀ sampled at f_s = 24 kHz. The model is G(x) = dec(Q(enc(x))): encoder → quantizer → decoder, producing reconstruction x̂. Everything is causal: convolutions are padded only on the past, so the model can stream and its latency is determined solely by the temporal down/up-sampling ratio.
So enc(x) ∈ ℝ^{S×D} with S = T/320.
Why not a single codebook? To hit 6 kbps with 75 frames/s, each frame gets r = 6000/75 = 80 bits. A single VQ codebook would need N = 2^80 entries — impossible.
Residual (multi-stage) VQ cascades N_q VQ layers, each with a small codebook of N vectors:
ŷ ← 0; residual ← y
for i = 1..N_q:
ŷ += Q_i(residual) # quantize the current residual
residual -= Q_i(residual) # pass on what's left
return ŷ
Each layer gets an equal slice of the budget: r_i = r/N_q = log₂ N. For 6 kbps with N_q = 8: each codebook has N = 2^(80/8) = 1024 entries. The sum of the layer outputs progressively refines the estimate while keeping shape S×D constant — crucial for bitrate scalability. Bits per frame = N_q · log₂ N.
Codebook training details (all learned end-to-end by backprop):
Quantizer dropout → one model, many bitrates. During training, for each example sample n_q ~ Uniform[1, N_q] and use only the first n_q quantizers (structured dropout over quantizer layers). At inference, pick n_q for the target bitrate. Because embeddings keep the same shape regardless of n_q, no architectural change or retraining is needed — unlike product quantization (wav2vec 2.0) or concatenated VQ outputs (Lyra), which require retraining per rate. Bonus: quantizer dropout acts as a regularizer — the scalable model marginally outperformed bitrate-specific models at 9 and 12 kbps.
Mirror of the encoder: 1D conv → 4 blocks, each = transposed (upsampling) conv + the same 3 dilated residual units, strides reversed (8, 5, 4, 2), channels halved at each upsample, final conv (kernel 7, stride 1, 1 filter) projects back to a 24 kHz waveform. Asymmetric capacities were ablated: a lighter encoder barely hurts quality (ViSQOL 3.96 → 3.94) while a lighter decoder hurts more (→ 3.84) — echoing asymmetric designs in neural image compression.
Two discriminator families, trained jointly with the generator:
1. Wave-based (MelGAN-style multi-scale): the same convolutional discriminator applied to the waveform at original, 2×-, and 4×-downsampled resolutions. Each scale: initial conv → 4 grouped convs (group size 4, stride 4, channel multiplier 4, capped at 1024 channels) → 2 plain convs → logits.
2. STFT-based: complex STFT as real+imaginary channels, window W=1024, hop H=256. 2D conv (7×7, 32 ch) → 6 residual blocks alternating stride (1,2)/(2,2) along (time, frequency); final 1×(F/2⁶) conv aggregates frequency bins into per-time-step logits.
Both are fully convolutional → per-time-step logits, giving dense adversarial signal.
With K+1 discriminators (k=0 STFT; k=1..3 wave scales), hinge losses:
L_D = E[ (1/K) Σ_k (1/T_k) Σ_t max(0, 1 − D_k,t(x)) + (1/K) Σ_k (1/T_k) Σ_t max(0, 1 + D_k,t(G(x))) ] — push logits of real audio above 1, generated below −1.L_adv = E[ (1/K) Σ_k (1/T_k) Σ_t max(0, 1 − D_k,t(G(x))) ].L_feat = E[ (1/KL) Σ_{k,l} (1/T_{k,l}) Σ_t |D^(l)_{k,t}(x) − D^(l)_{k,t}(G(x))| ]. Stabilizes GAN training and gives a perceptual mid-level reconstruction signal.S_s(x) with 64 mel bins, window s ∈ {2⁶,…,2¹¹} (64→2048), hop s/4; loss combines L1 plus a log-domain term with weight α_s = s/2:L_rec = Σ_s Σ_t [ |S_s,t(x) − S_s,t(x̂)|₁ + α_s · ||log S_s,t(x) − log S_s,t(x̂)||₂ ].
L_G = λ_adv·L_adv + λ_feat·L_feat + λ_rec·L_rec with λ_adv = 1, λ_feat = 100, λ_rec = 1.The perception–distortion trade-off framing: reconstruction losses buy fidelity; the adversarial loss buys perceptual quality (plausible fine detail the L1 losses would blur).
Training data comes as (input, target, denoise) triples: if denoise=false, target = input; if true, target = clean speech. A binary, potentially time-varying conditioning signal is injected via FiLM layers (feature-wise linear modulation): a'_{n,c} = γ_{n,c}·a_{n,c} + β_{n,c}, with γ, β produced by a linear layer from the one-hot mode. FiLM is applied at the bottleneck (encoder- or decoder-side). This enables denoising toggled at inference with zero extra latency; denoising before quantization (encoder-side) also lowers the bitrate needed (~7–20% headroom vs. entropy coding was measured overall).
SoundStream is the progenitor of the audio token vocabularies that speech-language models are built on. The connection chain is direct:
1. Discrete audio tokens. RVQ turns any waveform into a sequence of discrete symbols (N_q codebook indices per 75 Hz frame). LLMs are sequence predictors over discrete tokens; RVQ is the bridge that makes speech a "language" an LLM can model. AudioLM, VALL-E, MusicLM, MusicGen, and Spirit-LM all use RVQ codecs directly descended from SoundStream (its successor EnCodec, and Google's own SoundStream variants) as their tokenizer; VALL-E famously uses EnCodec tokens with a transformer decoder to do zero-shot TTS.
2. The residual structure is a design pattern for LMs. Coarse quantizers capture prosody/speaker/timbre; fine quantizers capture acoustic detail. Hierarchical speech LMs (AudioLM's semantic→coarse→fine stages) exploit exactly this ordering: model the first few codebooks with one transformer, the rest with another.
3. Quantizer dropout → bitrate-scalable tokens prefigures variable-token-budget audio generation: the same tokenizer serves cheap (low n_q) and high-fidelity (high n_q) generation without retraining.
4. Engineering constraints that shaped the field: causal convolutions for streaming, real-time smartphone CPU decode, a single model across rates and content types — these defined the practical envelope that later textless-NLP and spoken-dialogue systems inherited.
5. GAN-decoder recipe. The adversarial + feature-matching + multi-scale mel reconstruction loss mix (with λ_feat = 100) became the standard recipe for neural vocoders and codecs (HiFi-GAN lineage → EnCodec → DAC), which is what makes decoded speech from tokens sound natural rather than buzzy.
In short: before an LLM can speak, something must convert speech into tokens an LLM can predict, and something must convert predicted tokens back into audio. SoundStream (2021) is the paper that established both halves with one clean, end-to-end-trainable, real-time architecture.
Q1. Why is RVQ needed at all — why not one codebook per frame?
At 6 kbps and 75 frames/s each frame gets 80 bits. A single VQ codebook would need 2^80 entries — impossible. RVQ cascades 8 codebooks of 1024 entries each: each codes the previous one's residual, and the sum refines the estimate. Bits per frame = N_q · log₂ N.
Q2. What does "quantizer dropout" buy you?
One model that serves every bitrate in 3–18 kbps: during training, randomly use only the first n_q quantizers. At inference, choose n_q for the target bitrate — no retraining, because the embedding shape doesn't change.
Q3. Why the GAN losses instead of just L1/L2 reconstruction?
Pure reconstruction losses blur fine detail (perception–distortion trade-off). Adversarial + feature-matching losses push the decoder to generate plausible fine structure, which is what makes decoded speech sound natural instead of buzzy.
arXiv: 2209.03143 (September 2022, Google Research)
Authors: Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, Neil Zeghidour
Venue: Preprint (later appeared at IEEE TASLP-adjacent venues; foundational for VALL-E, MusicLM/SoundStorm, VALL-E 2, and the whole "audio-as-tokens" line)
---
Raw audio is sampled at 16,000–44,100 values per second. If you naively treat audio like text and train an autoregressive Transformer to predict one sample at a time, you hit two walls:
1. Sequence length. Self-attention is O(n²); one second of 16 kHz audio is 16,000 tokens. Even 10 seconds is beyond practical context.
2. Structure. Models trained purely on waveforms (e.g., WaveNet) generate locally plausible but globally meaningless audio — "babbling speech" — because they never see abstractions like words, sentences, speaker identity, or musical phrase structure.
The insight of AudioLM is a two-level answer:
This paper is the "GPT moment" for speech: it shows a decoder-only Transformer, trained only on raw audio with zero transcripts, zero labels, zero phonetic annotations, learns syntax, semantics, speaker consistency, and studio-quality audio.
---
---
Three frozen components wrap three trained Transformer LMs.
Even at 6000 bps semantic tokens cannot reconstruct audio (ViSQOL 1.4); even at 6000 bps acoustic tokens have terrible phonetic discriminability. A LM trained only on acoustic tokens preserves voice and recording conditions from a prompt but produces babbling — this negative result motivates the whole design.
Each stage is a separate decoder-only Transformer trained to predict next tokens; stages run sequentially at inference. RVQ's coarse-to-fine structure motivates the split.
Why three stages rather than one flat model?
1. Conditional-independence factorization: p(z_t | z_<t, y_<t) ≈ p(z_t | z_<t) — semantics don't depend on fine acoustics.
2. Shorter sequences per stage → much cheaper training/inference than one model over interleaved semantic+acoustic tokens.
3. Stage 3's chunking decouples its cost from total audio length.
Ratios: each semantic token (25 Hz) corresponds to 2 coarse acoustic frames (50 Hz), each with Q′ = 4 tokens → 2Q′ = 8 tokens per semantic token in stage 2, and 2(Q − Q′) = 16 tokens in stage 3.
---
---
---
1. It established the recipe every subsequent system uses. Audio → discrete tokens → next-token prediction → decode. VALL-E (Microsoft, Jan 2023) is essentially AudioLM's acoustic side with text conditioning; MusicLM and SoundStorm (Google, 2023) build directly on the AudioLM token hierarchy; VALL-E 2, SPEAR-TTS, and AudioPaLM all inherit the semantic/coarse/fine cascade.
2. Proved the "LLM of speech" is learnable without text. sWUGGY/sBLIMP scores near or above supervised toplines show that next-token prediction on semantic tokens acquires lexicon and grammar — the strongest pre-LLM-era evidence that speech can be a first-class language-modeling modality.
3. The semantic/acoustic split is a general design pattern. Disentangling what is said (low-bitrate, structure-bearing tokens) from how it sounds (high-bitrate, detail-bearing tokens) is exactly how later systems separate content from speaker/style, enable zero-shot voice cloning from 3-second prompts, and get controllability (swap semantics, keep voice).
4. RVQ as a coarse-to-fine hierarchy. The Q′ coarse/fine split and the flattening-with-offsets trick became the standard handling of RVQ tokens in VALL-E, SoundStorm (which later parallelized this with masked modeling), and MusicGen.
5. Showed the practical necessity of cascade modeling. One monolithic LM over all tokens is intractable at audio scale; staged conditioning (semantic → coarse → fine, with the finest stage chunked) is the compute-feasible factorization — a template echoed in later audio, image, and video token cascades.
6. Set the responsible-AI template for voice cloning. 51.2% human indistinguishability + 98.6% machine detectability framed the deepfake-discussion that still surrounds speech LMs.
---
Source: Borsos et al., "AudioLM: a Language Modeling Approach to Audio Generation," arXiv:2209.03143 (2022). Full text extracted from the PDF; all numbers above are from the paper's Tables I–IV and Section IV.
Q1. Why can't you just train one LM on SoundStream tokens?
Because the paper shows it babbles: acoustic tokens preserve voice and recording conditions but lose linguistic structure — an LM on them alone produces speech with consistent speaker but nonsense content (that's the motivating negative result). Semantic tokens carry structure but can't reconstruct audio.
Q2. Where do semantic tokens come from, mechanically?
Take the 7th layer of a pre-trained 0.6B w2v-BERT Conformer's MLM module, normalize each dim, run k-means with K=1024, use cluster indices as tokens at 25 Hz / 250 bps. They encode what is said but almost nothing about who (speaker-classifier accuracy drops to 3.2%).
Q3. Why split RVQ layers into coarse (4) and fine (8) stages?
RVQ's residual structure means early codebooks carry the "big picture" (speaker, timbre) and later ones carry detail. Stage 2 conditions on semantics and predicts the 4 coarse layers; stage 3 predicts the 8 fine layers conditioned only on coarse tokens (conditionally independent), run on 3-second chunks so cost doesn't scale with audio length.
Source: "Grafting a Speech Head onto Gemma 4 E4B" — Frisson Labs Blog.
Provenance per the article: Google's Gemma 4 E4B model card, the official E4B config, Google's audio-understanding docs, the authors' local MLX-VLM implementation inspection, and Kyutai's Moshi/Mimi paper.
Gemma 4 E4B is a small, fast, multimodal text-decoder model shaped like a typical vision/audio-language model rather than a native speech-to-speech model:
All modalities land in one 2560-wide shared sequence consumed by the 42-layer decoder.
Concrete numbers from the article (attributed to Google's audio-understanding docs):
Gemma inputs (text / audio embeddings in one 2560-wide stream)
→ Frozen Gemma 4 E4B (42 layers)
→ TAP POINT: learned mix of the LAST 6 decoder layers
(after the transformer stack, before the tied text-output head)
→ Trainable Gemma→Mimi audio head (only trained module)
→ Mimi codec tokens (8 codebooks @ 12.5 Hz)
→ Frozen Mimi decoder
→ 24 kHz waveform
Say this naturally as speech:\n{text} fed to Gemma; no TTS in the loop.Q1. Why tap the last 6 layers instead of the model's text output?
Reading hidden states before the tied vocabulary head means the speech branch is parallel to text decoding, not downstream of it — the head learns directly from Gemma's multimodal internal representation rather than from decoded text (which is what bolt-on TTS does).
Q2. What's the loss and target?
Cross-entropy over Mimi codec-token IDs (8 codebooks), with teacher speech Mimi-encoded from WAVs. Only the 152M-param Gemma→Mimi head trains; Gemma and the Mimi decoder are frozen.
Q3. What did the smoke test prove, and what's still missing?
It proved the wiring: Whisper recognized ≥2 target words in 5/7 text-driven clips and 6 words in the audio-in→audio-out run. Missing: speaking generated answers (needs AR decode + training on hidden states of generated tokens), generalization beyond memorized phrases, temporal/duration modeling, streaming, more data.
Putting the three sources together, here is the canonical recipe for making a text LLM speak, distilled from 2021→now:
Train (or take pretrained — EnCodec, Mimi, DAC) a codec: convolutional encoder → RVQ → GAN decoder. Decide your token budget: AudioLM used Q=12 codebooks × 1024 entries @ 50 Hz (600 tok/s); Mimi uses 8 codebooks @ 12.5 Hz. The RVQ hierarchy gives you a free coarse-to-fine ordering.
Audio without linguistic structure babbles. Extract tokens from a self-supervised speech model's intermediate layer (w2v-BERT layer 7, k-means K=1024 @ 25 Hz in AudioLM; HuBERT units in textless NLP). These carry what is said; acoustic tokens carry how it sounds.
Stage A: AR LM over semantic tokens (the "what"). Stage B: AR LM over coarse acoustic RVQ layers, conditioned on semantics (the "who/style"). Stage C: fine layers, conditioned on coarse ones (the "polish") — chunked to keep cost tractable. Modern variations: text conditioning on top of semantics (VALL-E, SPEAR-TTS), or parallel masked prediction instead of AR for the acoustic stage (SoundStorm).
If your LLM already hears (audio tower → placeholder tokens in the hidden stream), you don't need to retrain it: freeze everything, tap late decoder layers (last 6 of 42 in the Frisson experiment), and train a small head (152M params) to emit codec tokens; decode with the frozen codec decoder. Loss: cross-entropy on codec-token IDs with teacher-encoded speech.
Design check: You want Gemma-4-E4B to answer out loud in its own voice. Using only ideas from this course, sketch the minimal system.
Option 1 (simple, production): Gemma generates text normally → external low-latency TTS (which internally is exactly this course: text → semantic/acoustic tokens → LM → codec decoder). Option 2 (research, the Frisson path): freeze Gemma, tap the last-6-layer hidden mix, train a head to predict Mimi tokens — but train it (a) on hidden states of generated tokens, (b) with AR or duration-aware temporal modeling, (c) on far more than 128 phrases. The head then emits codec tokens during the normal decode loop and a frozen Mimi decoder yields the waveform.