ADS-C: Antidistillation Sampling for Classification

arXiv:2607.15467v1 Announce Type: new Abstract: Knowledge distillation enables an adversary to replicate a proprietary classifier by querying its prediction interface and training a surrogate on the returned probability vectors. Antidistillation sampling, proposed for large language models, counters this threat with an input-dependent, gradient-directed perturbation of the served distribution; its transfer to classification has not been studied. Adapting the defense to classification, we show it...

arXiv cs.LG ·Khawaja Abaid Ullah, Mohammad Javad Khojasteh ·
compartilhar: