OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching
Paper LLMs & Texto

OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching

OasisKV improves LLM inference throughput by storing full KV caches in lower memory tiers and prefetching only relevant entries into HBM using speculative-decoding lookahead predic…

Hugging Face · Daily Papers ·Can Xiao, Sukmin Cho · ·▲ 12 upvotes

Este artigo está em destaque na seleção diária de papers do Hugging Face, curada pela comunidade de pesquisa em IA.

Autores: Can Xiao, Sukmin Cho, Junbong We, Zhixiong Niu, Jianyi Cheng, Yiren Zhao

  • 12 upvotes da comunidade
  • Temas: large language model inference, KV cache, speculative decoding, lookahead tokens, attention sparsity, memory-centric inference

Resumo

Resumo original (em inglês), extraído do paper:

OasisKV improves LLM inference throughput by storing full KV caches in lower memory tiers and prefetching only relevant entries into HBM using speculative-decoding lookahead predictions.

Onde ler

compartilhar: