Blog
LLMs & Texto
How to Allocate Your Tokens? Scaling Laws with Training Steps and Batch Size
arXiv:2607.01487v1 Announce Type: new Abstract: We propose a scaling law that takes into account model size and training data while explicitly splitting the latter into training steps and batch size (called three-term law). Fitting the proposed law on a large set of training runs, we find that it correctly recovers the scaling of the optimal batch size. Moreover, because it makes use of training runs with suboptimal batch size, our proposed law can be robustly fit with a significantly smaller am...
arXiv cs.LG
·Fabian Schaipp
·
// relacionados
Leia também
Editorial
O modelo que continua aprendendo depois de entregue: dentro do Macaron-V1
Blog
Novo Nordisk e AWS levam IA agêntica à descoberta de medicamentos
Blog
Nvidia garante o valor de seus próprios chips para destravar US$ 500 bilhões em financiamento de infraestrutura de IA
Modelo