SeKV: Resolution-Adaptive KV Cache with Hierarchical Semantic Memory for Long-Context LLM Inference
arXiv:2606.31145v1 Announce Type: new Abstract: Large language models increasingly operate over long contexts, where the KV cache becomes a dominant memory bottleneck: its size grows linearly with sequence length and must be retained throughout decoding, making full GPU caching prohibitively expensive without compression. Existing KV cache compression methods struggle to balance efficiency with faithful context preservation. Token eviction discards information, while semantic grouping fixes comp...
arXiv cs.CL
·Amirhossein Abaskohi, Giuseppe Carenini, Peter West, Yuhang He
·
// relacionados
Leia também
Blog
Um Guia de Programação para a Programação de GPU Baseada em Tiles da NVIDIA: De cuTile e Kernels Triton até Flash Attention
Blog
OpenAI's GPT-5.6 Sol Ultra reportedly solves a 50-year-old math problem in under an hour
Blog
Grupos terroristas estão usando todos os principais chatbots de IA para planejamento de ataques e desenvolvimento de armas
Blog