RoPoLL: Robust Panel of LLM Judges
arXiv:2606.30931v1 Announce Type: new Abstract: The LLM Jury, a Panel of LLM Evaluators (PoLL) reporting consensus scores, has become a practical alternative to single-judge LLM evaluation, yet its statistical behavior remains poorly understood. We formalize the LLM Jury under the Huber contamination model and show that PoLL incurs unbounded bias under any positive contamination, regardless of jury size, whenever a single judge fails in a biased, LLM-typical way (mode collapse, sycophancy, safet...
arXiv cs.AI
·Anish Acharya, Kris W Pan, Brian Verkhovsky
·
// relacionados
Leia também
Blog
Um Guia de Programação para a Programação de GPU Baseada em Tiles da NVIDIA: De cuTile e Kernels Triton até Flash Attention
Blog
OpenAI's GPT-5.6 Sol Ultra reportedly solves a 50-year-old math problem in under an hour
Blog
Grupos terroristas estão usando todos os principais chatbots de IA para planejamento de ataques e desenvolvimento de armas
Blog