Teaching a Language Model to Speak

From raw waveforms to a speaking Gemma 4 — a guided course built from three sources

📚 SoundStream (2107.03312) · AudioLM (2209.03143) · Gemma 4 E4B speech head (Frisson Labs)

Chapter 0 — Why is speech hard for an LLM?

You know transformers: a sequence of discrete tokens goes in, next-token prediction comes out. Text has a natural tokenization (BPE): a sentence is ~15 tokens per second of reading. Audio is different in two brutal ways:

The whole field's answer, in one sentence: compress audio into discrete tokens at a manageable frame rate, then treat speech generation as pure language modeling — predict next token, then decode tokens back to a waveform. The two papers you gave me supply the two halves of that pipeline:

🎙️ waveform
→
SoundStream encoder + RVQ
discrete acoustic tokens
→
Transformer LM(s)
next-token prediction
→
SoundStream decoder
tokens → waveform
→
🔊 audio out

Read the chapters in order — each builds on the previous. Every chapter ends with a self-check quiz.

Chapter 1 — SoundStream: the audio tokenizer

3 kbpsbeats Opus @ 12 kbps
8.4Mparams, real-time on a phone CPU
75 Hzframe rate of embeddings (24 kHz audio ÷ 320)
2021Google · Zeghidour et al.

Research Digest: SoundStream — An End-to-End Neural Audio Codec

Paper: SoundStream: An End-to-End Neural Audio Codec

Authors: Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, Marco Tagliasacchi (Google)

Year: 2021 (arXiv:2107.03312; published in IEEE/ACM TASLP, Nov 2021)

---

1. What problem does it solve?

Traditional audio codecs come in two flavors, and both fail at low bitrates:

SoundStream replaces the entire hand-engineered encode→quantize→synthesize pipeline with a single neural network trained end-to-end. It compresses speech, music, and general audio at 3–18 kbps, runs in real time on a smartphone CPU, is streamable (causal, low latency), and — uniquely — one single model covers all bitrates in that range via quantizer dropout. Headline result: SoundStream at 3 kbps beats Opus at 12 kbps and roughly matches EVS at 9.6 kbps in blind listening tests — i.e., 3–4× fewer bits for equal quality.

2. Architecture, from first principles

The reader is assumed to know transformers; the key conceptual shift is that this is a pure convolutional autoencoder + discrete bottleneck + GAN loss, no attention at all, because it must run causally in real time on audio samples.

2.1 Setup and notation

Input: single-channel waveform x ∈ ℝᵀ sampled at f_s = 24 kHz. The model is G(x) = dec(Q(enc(x))): encoder → quantizer → decoder, producing reconstruction x̂. Everything is causal: convolutions are padded only on the past, so the model can stream and its latency is determined solely by the temporal down/up-sampling ratio.

2.2 Encoder (SEANet-style, fully convolutional)

So enc(x) ∈ ℝ^{S×D} with S = T/320.

2.3 Residual Vector Quantizer (RVQ) — the core mechanism

Why not a single codebook? To hit 6 kbps with 75 frames/s, each frame gets r = 6000/75 = 80 bits. A single VQ codebook would need N = 2^80 entries — impossible.

Residual (multi-stage) VQ cascades N_q VQ layers, each with a small codebook of N vectors:

ŷ ← 0;  residual ← y
for i = 1..N_q:
    ŷ        += Q_i(residual)      # quantize the current residual
    residual -= Q_i(residual)      # pass on what's left
return ŷ

Each layer gets an equal slice of the budget: r_i = r/N_q = log₂ N. For 6 kbps with N_q = 8: each codebook has N = 2^(80/8) = 1024 entries. The sum of the layer outputs progressively refines the estimate while keeping shape S×D constant — crucial for bitrate scalability. Bits per frame = N_q · log₂ N.

Codebook training details (all learned end-to-end by backprop):

Quantizer dropout → one model, many bitrates. During training, for each example sample n_q ~ Uniform[1, N_q] and use only the first n_q quantizers (structured dropout over quantizer layers). At inference, pick n_q for the target bitrate. Because embeddings keep the same shape regardless of n_q, no architectural change or retraining is needed — unlike product quantization (wav2vec 2.0) or concatenated VQ outputs (Lyra), which require retraining per rate. Bonus: quantizer dropout acts as a regularizer — the scalable model marginally outperformed bitrate-specific models at 9 and 12 kbps.

2.4 Decoder

Mirror of the encoder: 1D conv → 4 blocks, each = transposed (upsampling) conv + the same 3 dilated residual units, strides reversed (8, 5, 4, 2), channels halved at each upsample, final conv (kernel 7, stride 1, 1 filter) projects back to a 24 kHz waveform. Asymmetric capacities were ablated: a lighter encoder barely hurts quality (ViSQOL 3.96 → 3.94) while a lighter decoder hurts more (→ 3.84) — echoing asymmetric designs in neural image compression.

2.5 Discriminators (GAN training)

Two discriminator families, trained jointly with the generator:

1. Wave-based (MelGAN-style multi-scale): the same convolutional discriminator applied to the waveform at original, 2×-, and 4×-downsampled resolutions. Each scale: initial conv → 4 grouped convs (group size 4, stride 4, channel multiplier 4, capped at 1024 channels) → 2 plain convs → logits.

2. STFT-based: complex STFT as real+imaginary channels, window W=1024, hop H=256. 2D conv (7×7, 32 ch) → 6 residual blocks alternating stride (1,2)/(2,2) along (time, frequency); final 1×(F/2⁶) conv aggregates frequency bins into per-time-step logits.

Both are fully convolutional → per-time-step logits, giving dense adversarial signal.

2.6 Losses (the GAN recipe from HiFi-GAN/MelGAN lineage)

With K+1 discriminators (k=0 STFT; k=1..3 wave scales), hinge losses:

L_rec = Σ_s Σ_t [ |S_s,t(x) − S_s,t(x̂)|₁ + α_s · ||log S_s,t(x) − log S_s,t(x̂)||₂ ].

The perception–distortion trade-off framing: reconstruction losses buy fidelity; the adversarial loss buys perceptual quality (plausible fine detail the L1 losses would blur).

2.7 Joint compression + enhancement (denoising)

Training data comes as (input, target, denoise) triples: if denoise=false, target = input; if true, target = clean speech. A binary, potentially time-varying conditioning signal is injected via FiLM layers (feature-wise linear modulation): a'_{n,c} = γ_{n,c}·a_{n,c} + β_{n,c}, with γ, β produced by a linear layer from the one-hot mode. FiLM is applied at the bottleneck (encoder- or decoder-side). This enables denoising toggled at inference with zero extra latency; denoising before quantization (encoder-side) also lowers the bitrate needed (~7–20% headroom vs. entropy coding was measured overall).

3. Training pipeline

4. Key results and numbers

ClaimNumber Low-bitrate qualitySoundStream @3 kbps ≫ Opus @6 kbps and EVS @5.9 kbps; matches EVS @9.6 kbps and Opus @12 kbps (3.2–4× bit savings) Medium bitrates@6 kbps: EVS/Opus need 2.2–2.6× more bits to match High bitrates@12 kbps: 1.3–1.6× more bits needed by baselines vs. LyraSoundStream @3 kbps beats Lyra @3 kbps (learned encoder vs. fixed mel features) ViSQOL≥ 3.7 across 3–18 kbps; 4.01 @6 kbps (default config) Learnable encoder ablationFixed mel-filterbank encoder (Lyra-style): ViSQOL 3.96 → 3.33 @6 kbps — worse than a learned encoder at half the bitrate (3.76 @3 kbps) Model size / speed8.4 M params (32/32 ch), RTF 2.3–2.4× on Pixel 4; 16/16 ch → 2.4 M params, RTF >7× RVQ depth/size trade-off @6 kbps(N_q=8, N=1024): 4.01 · (16, 32): 3.98 · (80, 2): 3.92 — 80 one-bit quantizers train fine Latency ablationsStrides (1,4,5,8)=7.5 ms, (2,4,5,8)=13 ms, (4,4,5,8)=26 ms — all ViSQOL ≈4.01 at 6 kbps Quantizer dropoutScalable model matches bitrate-specific models at 6 & 12 kbps, slightly worse at 3 kbps; even slightly better at 9/12 kbps Joint denoisingNear-parity with separate SoundStream + SEANet pipeline (e.g., ViSQOL 3.02 vs 3.05 @0 dB SNR on VCTK) at half the compute and no added latency Entropy coding headroom7–20% potential bitrate savings from empirical-entropy coding of VQ symbols (unexploited in the reported rates) MusicFirst codec shown to encode music well at ~3 kbps (better than Opus @12 kbps in MUSHRA)

5. Why this matters for speech-capable language models

SoundStream is the progenitor of the audio token vocabularies that speech-language models are built on. The connection chain is direct:

1. Discrete audio tokens. RVQ turns any waveform into a sequence of discrete symbols (N_q codebook indices per 75 Hz frame). LLMs are sequence predictors over discrete tokens; RVQ is the bridge that makes speech a "language" an LLM can model. AudioLM, VALL-E, MusicLM, MusicGen, and Spirit-LM all use RVQ codecs directly descended from SoundStream (its successor EnCodec, and Google's own SoundStream variants) as their tokenizer; VALL-E famously uses EnCodec tokens with a transformer decoder to do zero-shot TTS.

2. The residual structure is a design pattern for LMs. Coarse quantizers capture prosody/speaker/timbre; fine quantizers capture acoustic detail. Hierarchical speech LMs (AudioLM's semantic→coarse→fine stages) exploit exactly this ordering: model the first few codebooks with one transformer, the rest with another.

3. Quantizer dropout → bitrate-scalable tokens prefigures variable-token-budget audio generation: the same tokenizer serves cheap (low n_q) and high-fidelity (high n_q) generation without retraining.

4. Engineering constraints that shaped the field: causal convolutions for streaming, real-time smartphone CPU decode, a single model across rates and content types — these defined the practical envelope that later textless-NLP and spoken-dialogue systems inherited.

5. GAN-decoder recipe. The adversarial + feature-matching + multi-scale mel reconstruction loss mix (with λ_feat = 100) became the standard recipe for neural vocoders and codecs (HiFi-GAN lineage → EnCodec → DAC), which is what makes decoded speech from tokens sound natural rather than buzzy.

In short: before an LLM can speak, something must convert speech into tokens an LLM can predict, and something must convert predicted tokens back into audio. SoundStream (2021) is the paper that established both halves with one clean, end-to-end-trainable, real-time architecture.

6. Caveats / limitations

✅ Self-check

Q1. Why is RVQ needed at all — why not one codebook per frame?
Answer

At 6 kbps and 75 frames/s each frame gets 80 bits. A single VQ codebook would need 2^80 entries — impossible. RVQ cascades 8 codebooks of 1024 entries each: each codes the previous one's residual, and the sum refines the estimate. Bits per frame = N_q · log₂ N.

Q2. What does "quantizer dropout" buy you?
Answer

One model that serves every bitrate in 3–18 kbps: during training, randomly use only the first n_q quantizers. At inference, choose n_q for the target bitrate — no retraining, because the embedding shape doesn't change.

Q3. Why the GAN losses instead of just L1/L2 reconstruction?
Answer

Pure reconstruction losses blur fine detail (perception–distortion trade-off). Adversarial + feature-matching losses push the decoder to generate plausible fine structure, which is what makes decoded speech sound natural instead of buzzy.

Chapter 2 — AudioLM: a GPT for speech

0 tokenstranscripts, labels, or phonemes used
3 stagessemantic → coarse → fine
51.2%humans couldn't tell continuations from real speech (chance = 50%)
60k hrsLibri-Light, fully unlabelled

Paper Digest: AudioLM — a Language Modeling Approach to Audio Generation

arXiv: 2209.03143 (September 2022, Google Research)

Authors: Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, Neil Zeghidour

Venue: Preprint (later appeared at IEEE TASLP-adjacent venues; foundational for VALL-E, MusicLM/SoundStorm, VALL-E 2, and the whole "audio-as-tokens" line)

---

1. The core problem

Raw audio is sampled at 16,000–44,100 values per second. If you naively treat audio like text and train an autoregressive Transformer to predict one sample at a time, you hit two walls:

1. Sequence length. Self-attention is O(n²); one second of 16 kHz audio is 16,000 tokens. Even 10 seconds is beyond practical context.

2. Structure. Models trained purely on waveforms (e.g., WaveNet) generate locally plausible but globally meaningless audio — "babbling speech" — because they never see abstractions like words, sentences, speaker identity, or musical phrase structure.

The insight of AudioLM is a two-level answer:

This paper is the "GPT moment" for speech: it shows a decoder-only Transformer, trained only on raw audio with zero transcripts, zero labels, zero phonetic annotations, learns syntax, semantics, speaker consistency, and studio-quality audio.

---

2. Background concepts (for a reader who knows transformers but not speech)

---

3. Architecture: the full pipeline

Three frozen components wrap three trained Transformer LMs.

3.1 Tokenizer A — acoustic tokens (SoundStream)

3.2 Tokenizer B — semantic tokens (w2v-BERT + k-means)

3.3 Why both? The trade-off table (Table I, LibriSpeech dev-clean)

TokenizationBitrateABX within/across (↓)ViSQOL (↑) Semantic (w2v-BERT)250 bps6.7 / 7.61.1 Semantic (w2v-BERT)6000 bps5.6 / 6.21.4 Acoustic (SoundStream)2000 bps22.4 / 28.73.3 Acoustic (SoundStream)6000 bps17.8 / 26.63.9

Even at 6000 bps semantic tokens cannot reconstruct audio (ViSQOL 1.4); even at 6000 bps acoustic tokens have terrible phonetic discriminability. A LM trained only on acoustic tokens preserves voice and recording conditions from a prompt but produces babbling — this negative result motivates the whole design.

3.4 The three-stage hierarchical LM

Each stage is a separate decoder-only Transformer trained to predict next tokens; stages run sequentially at inference. RVQ's coarse-to-fine structure motivates the split.

Why three stages rather than one flat model?

1. Conditional-independence factorization: p(z_t | z_<t, y_<t) ≈ p(z_t | z_<t) — semantics don't depend on fine acoustics.

2. Shorter sequences per stage → much cheaper training/inference than one model over interleaved semantic+acoustic tokens.

3. Stage 3's chunking decouples its cost from total audio length.

Ratios: each semantic token (25 Hz) corresponds to 2 coarse acoustic frames (50 Hz), each with Q′ = 4 tokens → 2Q′ = 8 tokens per semantic token in stage 2, and 2(Q − Q′) = 16 tokens in stage 3.

3.5 Inference modes

---

4. Training setup (concrete numbers)

---

5. Key results

ExperimentResult Acoustic generation fidelity (ground-truth semantic tokens → resynthesize)CER 3.4 / WER 6.0 vs original audio's 0.8/2.5; comparable to GSLM resynthesis (2.9/6.6), but GSLM is single-speaker, clean-condition only Speaker determinism (Table III)SoundStream reconstruction: 100% speaker-classifier accuracy; acoustic generation from fixed semantic tokens: 3.2% (vs 0.3% chance) → semantic tokens carry almost no speaker info Continuation speaker consistency92.6% speaker-classifier agreement between 3-s prompt and 7-s continuation of unseen speakers sWUGGY (lexical, zero-resource 2021)83.7 overall / 71.5 in-vocab — best among all text-free systems; beats forced-alignment phoneme topline (92.2/— is text-based; AudioLM beats every causal, generative system incl. GSLM's 68.7) sBLIMP (syntactic)64.7 — +8% relative over previous SOTA (CPC-BERT, 59.9); beats the supervised phone topline (66.8 is close, but AudioLM uses no supervision) Human subjective testRaters given 10-s clips (3-s real prompt + 7-s continuation vs real audio): correct labeling 51.2% (p = 0.23 vs chance 50%) → continuations statistically indistinguishable from real speech Deepfake detectionSimple conv classifier detects AudioLM speech with 98.6% accuracy (vs SoundStream-compressed originals) — humans can't, machines trivially can Piano continuationRaters preferred AudioLM over acoustic-tokens-only LM in 83.3% of pairs — the semantic stage fixes melody/structure, not just audio quality

---

6. Why this matters for speech generation with LLMs

1. It established the recipe every subsequent system uses. Audio → discrete tokens → next-token prediction → decode. VALL-E (Microsoft, Jan 2023) is essentially AudioLM's acoustic side with text conditioning; MusicLM and SoundStorm (Google, 2023) build directly on the AudioLM token hierarchy; VALL-E 2, SPEAR-TTS, and AudioPaLM all inherit the semantic/coarse/fine cascade.

2. Proved the "LLM of speech" is learnable without text. sWUGGY/sBLIMP scores near or above supervised toplines show that next-token prediction on semantic tokens acquires lexicon and grammar — the strongest pre-LLM-era evidence that speech can be a first-class language-modeling modality.

3. The semantic/acoustic split is a general design pattern. Disentangling what is said (low-bitrate, structure-bearing tokens) from how it sounds (high-bitrate, detail-bearing tokens) is exactly how later systems separate content from speaker/style, enable zero-shot voice cloning from 3-second prompts, and get controllability (swap semantics, keep voice).

4. RVQ as a coarse-to-fine hierarchy. The Q′ coarse/fine split and the flattening-with-offsets trick became the standard handling of RVQ tokens in VALL-E, SoundStorm (which later parallelized this with masked modeling), and MusicGen.

5. Showed the practical necessity of cascade modeling. One monolithic LM over all tokens is intractable at audio scale; staged conditioning (semantic → coarse → fine, with the finest stage chunked) is the compute-feasible factorization — a template echoed in later audio, image, and video token cascades.

6. Set the responsible-AI template for voice cloning. 51.2% human indistinguishability + 98.6% machine detectability framed the deepfake-discussion that still surrounds speech LMs.

Limitations (as the paper itself notes)

---

Source: Borsos et al., "AudioLM: a Language Modeling Approach to Audio Generation," arXiv:2209.03143 (2022). Full text extracted from the PDF; all numbers above are from the paper's Tables I–IV and Section IV.

✅ Self-check

Q1. Why can't you just train one LM on SoundStream tokens?
Answer

Because the paper shows it babbles: acoustic tokens preserve voice and recording conditions but lose linguistic structure — an LM on them alone produces speech with consistent speaker but nonsense content (that's the motivating negative result). Semantic tokens carry structure but can't reconstruct audio.

Q2. Where do semantic tokens come from, mechanically?
Answer

Take the 7th layer of a pre-trained 0.6B w2v-BERT Conformer's MLM module, normalize each dim, run k-means with K=1024, use cluster indices as tokens at 25 Hz / 250 bps. They encode what is said but almost nothing about who (speaker-classifier accuracy drops to 3.2%).

Q3. Why split RVQ layers into coarse (4) and fine (8) stages?
Answer

RVQ's residual structure means early codebooks carry the "big picture" (speaker, timbre) and later ones carry detail. Stage 2 conditions on semantics and predicts the 4 coarse layers; stage 3 predicts the 8 fine layers conditioned only on coarse tokens (conditionally independent), run on 3-second chunks so cost doesn't scale with audio length.

Chapter 3 — Making Gemma 4 speak: the grafted speech head

4.5Beffective params, 42 layers, 262K vocab
152Mtrainable params — everything else frozen
25 tok/sGemma's audio-input budget (16 kHz, 30s max)
8 codebooksMimi tokens @ 12.5 Hz as the speech target

Gemma 4 E4B Architecture — Digest & Speech-Head Analysis

Source: "Grafting a Speech Head onto Gemma 4 E4B" — Frisson Labs Blog.
Provenance per the article: Google's Gemma 4 E4B model card, the official E4B config, Google's audio-understanding docs, the authors' local MLX-VLM implementation inspection, and Kyutai's Moshi/Mimi paper.

1. Overall design

Gemma 4 E4B is a small, fast, multimodal text-decoder model shaped like a typical vision/audio-language model rather than a native speech-to-speech model:

Key numbers (spec panel)

SpecValue Effective parameters4.5B Decoder layers42 Text vocab262K Context window128K InputsText / Image / Audio OutputText Hidden width (decoder embedding)2560

2. Attention & decoder

3. Input paths (per-modality)

All modalities land in one 2560-wide shared sequence consumed by the 42-layer decoder.

4. Audio input handling

Concrete numbers from the article (attributed to Google's audio-understanding docs):

5. Where Gemma stops (no native speech output)

6. The Frisson experiment: grafting a speech head

Architecture

Gemma inputs (text / audio embeddings in one 2560-wide stream)
   → Frozen Gemma 4 E4B (42 layers)
   → TAP POINT: learned mix of the LAST 6 decoder layers
       (after the transformer stack, before the tied text-output head)
   → Trainable Gemma→Mimi audio head (only trained module)
   → Mimi codec tokens (8 codebooks @ 12.5 Hz)
   → Frozen Mimi decoder
   → 24 kHz waveform

Inference discipline

Training setup

ItemValue Trainable params152M (head only) Data split128 train / 36 valid / 36 test Checkpointsstep 500 (text smoke) / step 850 (audio-only continuation)

Results (smoke test — overfit architecture proof, not polished TTS)

What it tests vs. what's missing

7. Practical takeaway (Discord/game-buddy framing)

Caveats

✅ Self-check

Q1. Why tap the last 6 layers instead of the model's text output?
Answer

Reading hidden states before the tied vocabulary head means the speech branch is parallel to text decoding, not downstream of it — the head learns directly from Gemma's multimodal internal representation rather than from decoded text (which is what bolt-on TTS does).

Q2. What's the loss and target?
Answer

Cross-entropy over Mimi codec-token IDs (8 codebooks), with teacher speech Mimi-encoded from WAVs. Only the 152M-param Gemma→Mimi head trains; Gemma and the Mimi decoder are frozen.

Q3. What did the smoke test prove, and what's still missing?
Answer

It proved the wiring: Whisper recognized ≥2 target words in 5/7 text-driven clips and 6 words in the audio-in→audio-out run. Missing: speaking generated answers (needs AR decode + training on hidden states of generated tokens), generalization beyond memorized phrases, temporal/duration modeling, streaming, more data.

Chapter 4 — The full picture: build your own speaking model

Putting the three sources together, here is the canonical recipe for making a text LLM speak, distilled from 2021→now:

Step 1 — Get an audio tokenizer (SoundStream lineage)

Train (or take pretrained — EnCodec, Mimi, DAC) a codec: convolutional encoder → RVQ → GAN decoder. Decide your token budget: AudioLM used Q=12 codebooks × 1024 entries @ 50 Hz (600 tok/s); Mimi uses 8 codebooks @ 12.5 Hz. The RVQ hierarchy gives you a free coarse-to-fine ordering.

Step 2 — Get semantic tokens for structure

Audio without linguistic structure babbles. Extract tokens from a self-supervised speech model's intermediate layer (w2v-BERT layer 7, k-means K=1024 @ 25 Hz in AudioLM; HuBERT units in textless NLP). These carry what is said; acoustic tokens carry how it sounds.

Step 3 — Train hierarchical LMs

Stage A: AR LM over semantic tokens (the "what"). Stage B: AR LM over coarse acoustic RVQ layers, conditioned on semantics (the "who/style"). Stage C: fine layers, conditioned on coarse ones (the "polish") — chunked to keep cost tractable. Modern variations: text conditioning on top of semantics (VALL-E, SPEAR-TTS), or parallel masked prediction instead of AR for the acoustic stage (SoundStorm).

Step 4 — Wire it into your LLM (the Gemma 4 way)

If your LLM already hears (audio tower → placeholder tokens in the hidden stream), you don't need to retrain it: freeze everything, tap late decoder layers (last 6 of 42 in the Frisson experiment), and train a small head (152M params) to emit codec tokens; decode with the frozen codec decoder. Loss: cross-entropy on codec-token IDs with teacher-encoded speech.

What's still open (2026 frontier)

🎓 Final challenge

Design check: You want Gemma-4-E4B to answer out loud in its own voice. Using only ideas from this course, sketch the minimal system.
One possible answer

Option 1 (simple, production): Gemma generates text normally → external low-latency TTS (which internally is exactly this course: text → semantic/acoustic tokens → LM → codec decoder). Option 2 (research, the Frisson path): freeze Gemma, tap the last-6-layer hidden mix, train a head to predict Mimi tokens — but train it (a) on hidden states of generated tokens, (b) with AR or duration-aware temporal modeling, (c) on far more than 128 phrases. The head then emits codec tokens during the normal decode loop and a frozen Mimi decoder yields the waveform.