Blog
Visão Computacional
CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention
arXiv:2607.25291v1 Announce Type: new Abstract: The quadratic cost of self-attention makes long-context inference prohibitively expensive, and proxy-based block-sparse attention has become a practical remedy. Existing methods typically rely on a proxy to predict a binary sparse mask and a kernel to consume this mask and perform sparse attention computation. Such an approach is effective under moderate budgets. However, as the budget tightens, the estimated proxy inevitably drops some salient blo...
arXiv cs.CL
·Yufei Xue, Lin Niu, Hong Liu, Siran Liu, Hanyong Shao, Wei Liu, Guanghua Yu, Jianchen Zhu, Jun Zhang
·
// relacionados
Leia também
Editorial
Evidence-RL: como obrigar um modelo de visão a de fato olhar para a imagem
Blog
Colaboração entre drone (VANT) e veículo terrestre não tripulado (VTNT) para navegação autônoma em terreno coberto de neve
Blog
XEns-CKD: Uma abordagem explicável baseada em ensemble para detecção do estágio de doença renal crônica
Blog