Blog
LLMs & Texto
What Tokens are Learned when Tokenization is Optimized Jointly with Language Modeling?
arXiv:2608.17325v1 Announce Type: new Abstract: Tokenization is a fundamental component of language modeling pipelines. Despite its importance, it is often fixed, even though it significantly impacts model performance across languages. In this work, we analyze what tokens are learned when tokenization is jointly optimized with language modeling. We compare tokenizer-free approaches such as SSLMs and H-Nets with fixed tokenizers across 18 typologically and script-diverse languages. Our results sh...
arXiv cs.CL
·Saketh Reddy Vemula, Parameswari Krishnamurthy
·