Induction in Both Directions: A Mechanistic Analysis of In-Context Learning in Masked Diffusion Language Models
arXiv:2607.15893v1 Announce Type: new Abstract: While the internal mechanisms of autoregressive (AR) transformers have been studied extensively, much less is known about diffusion language models (DLMs), an emerging alternative that generates text by iterative denoising. In this work, we study how DLMs implement induction, a mechanism behind in-context learning in which the model finds a repeated context and copies the token that followed it. Our analysis compares attention-only AR models and ab...
arXiv cs.CL
·Andy Catruna, Emilian Radoi
·
// relacionados
Leia também
Blog
Hugging Face says an AI agent hacked its infrastructure, and it used AI to fight back
Editorial
Cura 1T: um modelo que aprende medicina treinando a si mesmo, sob supervisão humana
Blog
Beyond grep: The case for a context-rich AI coding harness
Blog