Blog
Dados & Embeddings
Aurora: A Leverage-Aware Spectral Optimizer
arXiv:2606.27715v1 Announce Type: new Abstract: We show that for tall matrix parameters, like projection matrices in the MLP layers, the Muon update can have row norms that are arbitrarily non-uniform. This can lead to a self-reinforcing feedback loop whereby neurons receive persistently small updates and eventually do not contribute meaningfully to network outputs. This problem is effectively mitigated by an additional row normalization step, but current methods do this in a way that moves the ...
arXiv cs.LG
·Alec Dewulf, Dhruv Pai, Li Yang, Ashley Zhang, Ben Keigwin
·
// relacionados
Leia também
Editorial
O modelo que entra em espiral: a receita da Liquid AI contra os "doom loops"
Blog
Consórcio alemão de IA lança o Soofi S, um modelo aberto de 30B que lidera benchmarks tanto em inglês quanto em alemão
Blog
SensorFM, do Google, transforma dados desorganizados de sensores de wearables em uma camada de inteligência de saúde de uso geral
Blog