Conformal Cascade: Distribution-Free Accuracy Guarantees for Multi-Tier LLM Inference

arXiv:2607.25018v1 Announce Type: new Abstract: Large language model (LLM) cascades reduce inference cost by routing easy queries to a small model and deferring hard queries to a larger one. Production cascades govern this deferral through a confidence threshold, but LLM confidence scores are miscalibrated, the threshold must be tuned per model pair and per domain, and no setting yields a formal bound on cascade accuracy. We introduce \textbf{Conformal Cascade} (CC), a multi-tier inference frame...

arXiv cs.LG ·Yifan Dou, Shikan Fang, Shibo Li ·
compartilhar: