Augmenting Continuous Diffusion Language Models with Discrete Token Guidance
Hui Ren1,
Zihan Li1,
Chang Liu1,
Huidong Liu2,
Alexander Schwing1
1University of Illinois Urbana-Champaign 2Amazon.com, Inc.
HC-DLM augments continuous diffusion language models with discrete token guidance:
a shared latent plans every token jointly, while tokens read out at each step keep it anchored to valid text.
The latent plans, the tokens guide, at every step (simulated example). Interactive version on the project page.
Each family of diffusion language models has a blind spot:
- Discrete diffusion decodes tokens in parallel, but samples each one independently from its marginal.
- Continuous diffusion plans all tokens in one shared latent, but never checks that plan against real tokens until the very end.
HC-DLM keeps both strengths: tokens are planned jointly in the latent, and the latent is guided by tokens at every step. The continuous latent is the only persistent generative state; at every step the model reads a token draft $k_t$ out of the latent $x_t$, and that draft guides the next latent update.
$$ p_{\theta,\phi}(k_{0:T},x_{0:T})=p(x_T),p_\theta(k_0\mid x_0)\prod_{t=1}^{T}\underbrace{p_\theta(k_t\mid x_t)}_{\text{read out tokens}};\underbrace{p_\phi(x_{t-1}\mid x_t,k_t)}_{\text{token-guided latent denoising}} $$
One reverse step, four designs. Discrete diffusion updates tokens directly, one marginal at a time. Continuous diffusion denoises a latent that is blind to tokens and decodes only at the end. Hybrid models attach a continuous signal to a self-contained discrete chain. In HC-DLM the two levels talk at every step: tokens are read out of the latent, then guide its next update.
| Discrete diffusion e.g. MDM, LLaDA |
Continuous diffusion e.g. Diffusion-LM, Plaid |
Hybrid discrete–continuous e.g. CADD, CCDD |
HC-DLM (ours) | |
|---|---|---|---|---|
| Token dependence within a step | ✕ independent marginals |
✓ | ◐ via conditioning only |
✓ |
| Tied to tokens at every step | ✓ | ✕ only at t = 0 |
✓ | ✓ |
| Tokens revisable at every step | ◐ uniform kernel only |
✓ | ◐ uniform kernel only |
✓ readout from xt |
Noise corrupts tokens and latent independently, so training stays simple. A single variational bound on the token likelihood splits into three terms and trains everything end to end: reconstruction for the token readout, token-guided flow matching for the latent denoiser, and an entropy term that keeps the encoder from collapsing.
One principled objective, trained end to end. An encoder maps clean tokens to a latent, independent forward kernels produce the noisy pair, the token-conditioned denoiser predicts the clean latent, and the token predictor decodes it back to tokens.
From Sudoku and Countdown to open-domain text, HC-DLM leads discrete, continuous and hybrid diffusion baselines of the same size on Hard Sudoku, Countdown and LM1B. Parameter counts exclude token embeddings.
Left: solutions emerge early: accuracy of the intermediate prediction decoded at each step, vs. latent diffusion without token guidance. Right: holds up with fewer steps: Hard Sudoku accuracy against step count; MDM degrades sharply as steps shrink.
📋 Full tables (accuracy %; generative perplexity, lower is better)
Sudoku & Countdown (6M parameters, accuracy %)
| Method | Sudoku Easy | Sudoku Hard | CD4 | CD5 |
|---|---|---|---|---|
| MDM (top-prob. margin) | 89.49 | 49.88 | 50.8 | 21.3 |
| CCDD | 94.65 | 70.73 | 81.18 | 25.35 |
| HC-DLM | 94.21 | 72.41 | 84.41 | 37.52 |
LM1B (generative perplexity)
| Method | Params | Gen. PPL ↓ |
|---|---|---|
| MDM | 116M | 103.9 |
| Duo | 116M | 97.6 |
| LangFlow | 117M | 92.2 |
| Plaid | 109M | 77.3 |
| HC-DLM | 118M | 75.5 |
Ablation: neither level works alone. Latent DM removes token guidance from the denoiser; MDM removes the latent altogether (Sudoku accuracy %).
| Method | Cont. latent | Token guidance | Easy | Hard |
|---|---|---|---|---|
| MDM (top-prob. margin) | ✕ | ✕ | 89.49 | 49.88 |
| Latent DM (w/o token guidance) | ✓ | ✕ | 50.46 | 24.74 |
| HC-DLM | ✓ | ✓ | 94.21 | 72.41 |
The code and the artifacts are coming soon. Stay tuned!
If you find this work useful, please consider citing:
@misc{ren2026hcdlm,
title={Hierarchical Continuous Diffusion Language Models},
author={Hui Ren and Zihan Li and Chang Liu and Huidong Liu and Alexander Schwing},
year={2026},
eprint={2610.02193},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2610.02193},
}



