S2D2: Fast Decoding for Diffusion LLMs via Training-Free Self-Speculation
Rattachement africain : us. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Block-diffusion language models offer a promising path toward faster-than-autoregressive generation by combining block-wise autoregressive decoding with within-block parallel denoising. However, in the few-step regime needed for practical acceleration, standard confidence-thresholded decoding is often brittle: aggressive thresholds hurt quality, while conservative thresholds require unnecessary denoising steps. Existing approaches that address this issue either require additional training or incur extra test-time compute. We present S2D2, a training-free self-speculative decoding framework for block-diffusion language models. Our key observation is that a block-diffusion model becomes autoregressive when the block size is reduced to one, allowing the same pretrained model to act as both drafter and verifier. S2D2 inserts a speculative verification step into standard block-diffusion decoding and uses lightweight routing policies to decide when verification is worth its cost. This yields a hybrid decoding trajectory in which diffusion proposes tokens in parallel, while the autoregressive mode acts as a local sequence-level critic. Across three mainstream block-diffusion families, S2D2 consistently improves the accuracy-speed tradeoff over strong confidence-thresholding baselines. On SDAR, we observe up to $4.7\times$ speedup over autoregressive decoding, and up to $1.57\times$ over a tuned dynamic decoding baseline while improving accuracy by up to $4.5$ points. On LLaDA2.1-Mini, S2D2 remains complementary to built-in self-correction, including a conservative setting where it is $4.4\times$ faster than the static baseline with slightly higher accuracy.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
Où se fait cette recherche
-
Red Hat (United States) pays non établi dans la noticeEntreprise
-
IBM (United States) pays non établi dans la noticeEntreprise
-
Iowa State University pays non établi dans la noticeUniversité ou école supérieure
-
MIT-IBM Watson AI Lab pays non établi dans la noticeStructure de recherche
-
Red Hat AI Innovation pays non établi dans la noticeInstitution
-
Core AI pays non établi dans la noticeInstitution
Red Hat (United States), IBM (United States) et Iowa State University, avec 3 autres affiliations.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.