Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents
arXiv:2607.27275v1 Announce Type: new Abstract: Post-training quantization to 4-bit weights is widely reported to be nearly lossless. We test this claim for multi-turn, tool-calling agents, where it now matters most. On $\tau^2$-bench, across two open-weight model families in dense and MoE variants and two domains (eight cells, 456 episodes each, at 16-, 8-, and 4-bit weights), quantization indeed looks free on the standard metric. No cell shows a score change that survives multiple-comparison c...
arXiv cs.LG
·Jiwon Jang, Kisu Yang, Heuiseok Lim, Hyunwoo Park
·
// relacionados