Blog
LLMs & Texto
Hierarchical Global Attention (HGA)
arXiv:2606.30709v1 Announce Type: new Abstract: Hierarchical Global Attention (HGA) is a drop-in replacement for dense causal attention in pretrained long-context transformers. HGA preserves the original checkpoint parameters: the pretrained $W_Q$, $W_K$, $W_V$, and $W_O$ projections remain unchanged, no calibration parameters are introduced, and no retraining is required. Applied to Qwen3-30B-A3B-Instruct-2507-FP8 on a single RTX~5090 (32GB), the patched model runs out of the box at a 64K-token...
arXiv cs.LG
·Woernle Frank, Fedosov Vladimir, Grinenko Artemiy
·
// relacionados
Leia também
Editorial
O modelo que continua aprendendo depois de entregue: dentro do Macaron-V1
Blog
Novo Nordisk e AWS levam IA agêntica à descoberta de medicamentos
Blog
Nvidia garante o valor de seus próprios chips para destravar US$ 500 bilhões em financiamento de infraestrutura de IA
Modelo