Pivot-SDEfficient Self-Distillation for
Masked Diffusion Language Models

EMNLP 2026 Main (Oral)
Seo Hyun Kim1,2,*‡ Sunwoo Hong1,2,*‡ Younwoo Choi2 Chen-Hao Chao2 Se-Young Yun1,† Rahul G. Krishnan2,†
1KAIST AI2University of Toronto & Vector Institute
*Equal contribution   †Corresponding authors   ‡Work done as visiting researchers at the University of Toronto and the Vector Institute

Training a diffusion LM doesn't need every token,
just the few pivots that shape the rest of the generation.

Correct trajectory
Illustrative example of pivot selection. Left: predictive entropy of each masked position. Right: the sum of entropy. Rows are denoising steps. The most confident mask is unmasked (outlined). If the next prediction becomes much more certain, that commitment is a pivot (★).

Abstract

Masked diffusion language models (dLMs) offer a promising parallel alternative to autoregressive models for complex reasoning. However, they face a unique credit-assignment challenge, since reasoning failures typically stem from a few critical, early commitments that cascade downstream. Existing post-training overlooks this dynamic: supervised fine-tuning (SFT) weights all tokens uniformly, while online reinforcement learning (RL) attributes a single trajectory-level reward to every position. We introduce Pivot-SD, an efficient offline self-distillation framework that supervises only these high-impact bottlenecks (pivots). Pivot-SD isolates pivots using an information-gain metric measuring uncertainty reduction over the remaining masked block. It applies bidirectional supervision: successful pivots are reinforced with cross-entropy, while critical errors in failed trajectories are suppressed via targeted unlikelihood, preserving valid surrounding sub-steps. Using only 200 questions and four rollouts each, Pivot-SD improves LLaDA-8B-Instruct over full-sequence SFT and budget-matched diffusion RL baselines across math and code benchmarks.

Method

1. Find pivots. For each commitment of token y_p at position p and step t, we fix it and measure how much the entropy over the remaining masked positions drops. The top-K commitments per trajectory are the pivots.

g(t) = (H_pre(t) − H_post(t)) / |A_t \ U_t|

2. Supervise only pivots. A verifier labels each trajectory. Pivots of correct trajectories are reinforced with cross-entropy; pivots of incorrect ones are suppressed with unlikelihood, conditioned on the partially masked state M_t. All other tokens receive no loss.

L = 1[correct]·(−log P(y_p|M_t)) + λ·1[incorrect]·(−log(1 − P(y_p|M_t)))
Where each method puts its training signal. Green raises probability, red lowers it, struck-through data is discarded. Pivot-SD leaves the valid prefix of the incorrect answer untouched.

Results

LLaDA-8B-InstructMATH500GSM8KHumanEval+MBPP+
Base31.4075.1332.3243.12
SFT-GT35.4776.9433.9446.32
SFT-SD31.7373.9035.7845.92
diffu-GRPO33.4776.0432.5243.39
wd1++33.2776.9333.5443.12
Pivot-SD (ours)37.4779.5140.4347.12

Accuracy (%), 256-token generation. All methods use the same budget: 200 questions, 4 rollouts each, 1,000 steps. Mean of three runs.

Efficiency against online RL

45464748495051520510152025Total compute (hours)Average accuracy (%)~8× faster> +3 pts1k steps5k stepsPivot-SD 51.1diffu-GRPOwd1++Pivot-SD (ours)
Average accuracy over MATH500, GSM8K, HumanEval+ and MBPP+ (256 tokens) against total wall-clock hours on 2× NVIDIA L40. RL baselines generate their rollouts online, inside training. The Pivot-SD figure includes its offline data generation: 800 rollouts sampled once (2.5 h) plus 0.3 h of training.

BibTeX

@inproceedings{kim2026pivotsd,
  title     = {Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models},
  author    = {Kim, Seo Hyun and Hong, Sunwoo and Choi, Younwoo and Chao, Chen-Hao and Yun, Se-Young and Krishnan, Rahul G.},
  booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
  year      = {2026}
}

Placeholder entry; we will replace it with the ACL Anthology one once the proceedings are published.