Blog
LLMs & Texto
Beyond KV Reconstruction: Functional Reconstruction for MLA Draft Models in Speculative Decoding
arXiv:2607.27269v1 Announce Type: new Abstract: Multi-head latent attention (MLA) is increasingly important for long-context LLM inference because compact latent states replace the growing key-value (KV) cache and reduce decoding memory traffic. Yet most capable open checkpoints use multi-head or grouped-query attention (MHA/GQA), so conversion is needed to obtain MLA's cache efficiency without retraining from scratch. Speculative decoding offers complementary acceleration, but its speedup depen...
arXiv cs.LG
·Weiye Shi, Fanxu Meng, Muhan Zhang
·
// relacionados