Training a diffusion LM doesn't need every token,
just the few pivots that shape the rest of the generation.
Masked diffusion language models (dLMs) offer a promising parallel alternative to autoregressive models for complex reasoning. However, they face a unique credit-assignment challenge, since reasoning failures typically stem from a few critical, early commitments that cascade downstream. Existing post-training overlooks this dynamic: supervised fine-tuning (SFT) weights all tokens uniformly, while online reinforcement learning (RL) attributes a single trajectory-level reward to every position. We introduce Pivot-SD, an efficient offline self-distillation framework that supervises only these high-impact bottlenecks (pivots). Pivot-SD isolates pivots using an information-gain metric measuring uncertainty reduction over the remaining masked block. It applies bidirectional supervision: successful pivots are reinforced with cross-entropy, while critical errors in failed trajectories are suppressed via targeted unlikelihood, preserving valid surrounding sub-steps. Using only 200 questions and four rollouts each, Pivot-SD improves LLaDA-8B-Instruct over full-sequence SFT and budget-matched diffusion RL baselines across math and code benchmarks.
1. Find pivots. For each commitment of token y_p at position p and step t, we fix it and measure how much the entropy over the remaining masked positions drops. The top-K commitments per trajectory are the pivots.
2. Supervise only pivots. A verifier labels each trajectory. Pivots of correct trajectories are reinforced with cross-entropy; pivots of incorrect ones are suppressed with unlikelihood, conditioned on the partially masked state M_t. All other tokens receive no loss.
| LLaDA-8B-Instruct | MATH500 | GSM8K | HumanEval+ | MBPP+ |
|---|---|---|---|---|
| Base | 31.40 | 75.13 | 32.32 | 43.12 |
| SFT-GT | 35.47 | 76.94 | 33.94 | 46.32 |
| SFT-SD | 31.73 | 73.90 | 35.78 | 45.92 |
| diffu-GRPO | 33.47 | 76.04 | 32.52 | 43.39 |
| wd1++ | 33.27 | 76.93 | 33.54 | 43.12 |
| Pivot-SD (ours) | 37.47 | 79.51 | 40.43 | 47.12 |
Accuracy (%), 256-token generation. All methods use the same budget: 200 questions, 4 rollouts each, 1,000 steps. Mean of three runs.
@inproceedings{kim2026pivotsd,
title = {Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models},
author = {Kim, Seo Hyun and Hong, Sunwoo and Choi, Younwoo and Chao, Chen-Hao and Yun, Se-Young and Krishnan, Rahul G.},
booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
year = {2026}
}
Placeholder entry; we will replace it with the ACL Anthology one once the proceedings are published.