Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See

arXiv:2608.17744v1 Announce Type: new Abstract: Take three frontier mixture-of-experts models (Alibaba, OpenAI, NVIDIA; 3.6-4.0B active parameters each) and fine-tune them to reason in a low-resource language. On accuracy benchmarks almost nothing happens, and the benchmark itself is noise at this scale: changing only the random seed moves the score by 7.7 points, more than every data and recipe effect we measured. That null is our first result. The real changes live where accuracy cannot see. B...

arXiv cs.CL ·Ayoub Kirouane, Christos Petrocheilos ·
compartilhar: