VarRate: Training-Free Variable-Rate KV Cache Compression for Long-Context LLMs

arXiv:2607.15498v1 Announce Type: new Abstract: The key-value (KV) cache is the main memory bottleneck in long-context large language model (LLM) inference. Two leading training-free families are both structurally limited: token-selection methods (SnapKV, Ada-KV) score importance from an observation window and evict low-scoring tokens, but eviction is irreversible -- so when the importance signal degrades under query-agnostic reuse, accuracy collapses by 11-15 points; uniform low-rank coding kee...

arXiv cs.CL ·Shahrzad Esmat, Dhawal Shah, Ali Jannesari ·
compartilhar: