DUET: Dual-Teacher On-Policy Distillation via Same-Weight Disagreement for Prohibition Compliance

arXiv:2608.14644v1 Announce Type: new Abstract: Real-world LLM deployments increasingly rely on runtime-injected prohibitions--enterprise policies, PII redlines, tool boundaries--that vary per request and per tenant. Conventional post-training is structurally ill-suited: SFT hides the violation signal in compliant labels, and DPO's sequence-level preferences mismatch token-localized violations. We propose DUET, a token-selective on-policy distillation method for prohibition compliance. DUET pair...

arXiv cs.LG ·Zihan Li, Feifei Li, Wenhui Que ·
compartilhar: