Blog
LLMs & Texto
VarRate: Training-Free Variable-Rate KV Cache Compression for Long-Context LLMs
arXiv:2607.15498v1 Announce Type: new Abstract: The key-value (KV) cache is the main memory bottleneck in long-context large language model (LLM) inference. Two leading training-free families are both structurally limited: token-selection methods (SnapKV, Ada-KV) score importance from an observation window and evict low-scoring tokens, but eviction is irreversible -- so when the importance signal degrades under query-agnostic reuse, accuracy collapses by 11-15 points; uniform low-rank coding kee...
arXiv cs.CL
·Shahrzad Esmat, Dhawal Shah, Ali Jannesari
·
// relacionados
Leia também
Blog
Hugging Face says an AI agent hacked its infrastructure, and it used AI to fight back
Editorial
Cura 1T: um modelo que aprende medicina treinando a si mesmo, sob supervisão humana
Blog
Beyond grep: The case for a context-rich AI coding harness
Blog