CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models

arXiv:2607.24999v1 Announce Type: new Abstract: LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively to matched interventions, and generalize beyond the models used to define them. We introduce CogArena, a procedurally generated 13-paradigm benchmark built around a multimethod framework for determining when cognitive-task scores warrant dimensional labels across five theory-motivated groupings. Across 55 open-...

arXiv cs.CL ·Dengzhe Hou, Lingyu Jiang, Fangzhou Lin, Kazunori D Yamada ·
compartilhar: