BehaviorBench: Benchmarking Foundation Models for Behavioral Science Tasks
arXiv:2606.24162v1 Announce Type: new Abstract: Foundation models have been increasingly applied to behavioral science domains such as psychology, sociology, and economics. While these models show promise in individual tasks such as survey response prediction and human-subject experiment simulation, there remains no systematic understanding of how well they perform across diverse behavioral science tasks, contexts, and populations. We introduce BehaviorBench, a comprehensive benchmark that evalu...
arXiv cs.CL
·Jin Huang, Yutong Xie, Wanli Song, Xingjian Zhang, Walter Yuan, Matthew O. Jackson, Qiaozhu Mei
·
// relacionados
Leia também
Blog
Unicorn, pelican, Middle-earth: OpenAI co-founder Karpathy is looking for the next AI vibe test
Editorial
CAPA: o benchmark que mede se o assistente de código aprende com você — ou repete a mesma pergunta
Blog
Why biological data matters more in AI drug discovery
Blog