StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling

StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling

StateM is a runtime system that improves long-horizon agent execution through durable states, recoverable runbooks, and enforceable procedural controls without altering model weigh…

Hugging Face · Daily Papers ·Ziheng Qin, Yaxin Lu · ·▲ 346 upvotes

Este artigo está em destaque na seleção diária de papers do Hugging Face, curada pela comunidade de pesquisa em IA.

Autores: Ziheng Qin, Yaxin Lu, Zhangyang Atlas Wang, Kai Wang

  • 346 upvotes da comunidade
  • Temas: agent-native runtime, durable states, phase-local context, checked transitions, recoverable runbooks, versioned procedural practices

Resumo

Resumo original (em inglês), extraído do paper:

StateM is a runtime system that improves long-horizon agent execution through durable states, recoverable runbooks, and enforceable procedural controls without altering model weights.

Onde ler

compartilhar: