Blog
Dados & Embeddings
Fingerprint, Not Blueprint: How Positional Schemes Set the Default Spectral Algebra of Attention
arXiv:2607.06621v1 Announce Type: new Abstract: The pre-softmax score of an attention head is a bilinear form $score(i,j) = x_i^T M x_j$ in a learned operator $M = W_q^T W_k$. Because M is generally non-symmetric, hence non-normal, it has a complex eigenspectrum and non-orthogonal eigenvectors, the regime where non-Hermitian and random-matrix tools apply. We ask what this spectrum encodes, at three levels for previous-token and induction circuits. Statically, across seven pretrained models spann...
arXiv cs.LG
·Li Hengyu (Institute for Solid State Physics, The University of Tokyo)
·
// relacionados
Leia também
Blog
Soofi Consortium lança o Soofi S 30B-A3B: um modelo de fundação MoE híbrido Mamba-Transformer aberto para alemão e inglês
Blog
Construindo um pipeline PyTorch controlado pelo Gin Config com variantes de MLP configuráveis, agendamento por cosseno e substituições de parâmetros em tempo de execução
Blog
Bonsai 27B é um modelo de raciocínio totalmente aberto que cabe em um iPhone
Blog