StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
StartupBench evaluates end-to-end AI agents on real-world startup workflows and reveals that even top models complete only about 30% of tasks, highlighting gaps in instruction foll…
Hugging Face · Daily Papers
·Liya Zhu, Xin Ma
·
·▲ 8 upvotes
Este artigo está em destaque na seleção diária de papers do Hugging Face, curada pela comunidade de pesquisa em IA.
Autores: Liya Zhu, Xin Ma, Tao Liu, Haodong Wang, Ge Zhang, Jingzhe Ding
- 8 upvotes da comunidade
- Temas: Large Language Models, agents, StartupBench, E2E agent benchmark, agent harness, complex instruction following
Resumo
Resumo original (em inglês), extraído do paper:
StartupBench evaluates end-to-end AI agents on real-world startup workflows and reveals that even top models complete only about 30% of tasks, highlighting gaps in instruction following and domain expertise.