Blog
LLMs & Texto
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning
arXiv:2607.15388v1 Announce Type: new Abstract: Many math- and science-oriented agent systems use hierarchical designs with specialized reviewer roles, assuming that a dedicated review stage should help turn wrong candidates into correct ones. We test this assumption on 4,181 verifier-grounded Omni-MATH problems using matched gpt-oss-120b actors. Collaboration adds little on the easiest tiers, but from tier 4 onward the gains open sharply; in this harder regime, broadcast-style peer discussion r...
arXiv cs.AI
·Chih-Hsuan Yang, Jingyan Jiang, Vikram Vasudevan, Cheng-Hau Yang, Huihuo Zheng, Le Chen, Eliu A. Huerta, Venkatram Vishwanath, Ian T. Foster, Rajeev Thakur
·
// relacionados
Leia também
Blog
Hugging Face says an AI agent hacked its infrastructure, and it used AI to fight back
Editorial
Cura 1T: um modelo que aprende medicina treinando a si mesmo, sob supervisão humana
Blog
Beyond grep: The case for a context-rich AI coding harness
Blog