On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment
arXiv:2607.27081v1 Announce Type: new Abstract: Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers can embed harmful behaviors into downstream corpora, creating models that retain professional skills while violating human values on demand. Existing safety-realignment defenses often fail in practice due to three key limitations: they frequently cause catastrophic forgetting of specialized skills; th...
arXiv cs.AI
·Yongjian Guo, Wanlun Ma, Lingyu Shen, Xi Xiao, Sheng Wen
·
// relacionados