Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning
arXiv:2607.16057v1 Announce Type: new Abstract: Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic tool-use. What remains poorly measured is AI progress on the analytical knowledge work white-collar professionals perform daily, including synthesizing complex information, exercising judgment under uncertainty an...
arXiv cs.CL
·Ajay Patel, Kartik Hosanagar, Ramayya Krishnan, Chris Callison-Burch, Karim Lakhani, Mitch Weiss
·
// relacionados
Leia também
Editorial
O modelo que continua aprendendo depois de entregue: dentro do Macaron-V1
Blog
Novo Nordisk e AWS levam IA agêntica à descoberta de medicamentos
Blog
Nvidia garante o valor de seus próprios chips para destravar US$ 500 bilhões em financiamento de infraestrutura de IA
Modelo