VASAE: Naming SAE Dictionary Directions with Vocabulary-Aligned Anchoring
arXiv:2606.27941v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) provide useful decompositions of Transformer residual streams, but their learned features are usually named post hoc rather than directly connected to the Transformer's token vocabulary. We introduce Vocabulary-Aligned Sparse Autoencoder (VASAE), a method that trains SAE features under vocabulary-aligned anchoring and assigns each feature an intrinsic token name: the token string whose embedding is nearest to that feature...
arXiv cs.CL
·Kairui Zhang, Ziwen Yu, Zahraa S. Abdallah, Martha Lewis
·
// relacionados
Leia também
Editorial
Kimi K3: a China lança o maior modelo aberto do mundo — e ele não é o maior por acaso
Blog
Modelos de peso aberto agora igualam o desempenho cibernético de ponta de apenas quatro meses atrás por uma fração do custo
Blog
O novo manual de IA do Pentágono trata a adoção lenta como um risco maior do que o alinhamento imperfeito
Blog