SDO: Structure-Aware Data Organization for Efficient LLM Post-Training

arXiv:2607.27273v1 Announce Type: new Abstract: Post-training of large language models is expensive, and existing efficiency improvements mainly focus on selecting informative samples or designing training schedules. However, data organization itself is usually treated as a static preprocessing step: embedding-based grouping methods construct fixed partitions before training and cannot adapt to the evolving sample exposure during optimization. As a result, all samples receive similar exposure de...

arXiv cs.LG ·Jinliang Gao, Ning Yang, Hai Wang, Baili Xiao, Pin Lyu ·
compartilhar: