StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
StateM is a runtime system that improves long-horizon agent execution through durable states, recoverable runbooks, and enforceable procedural controls without altering model weigh…
Hugging Face · Daily Papers
·Ziheng Qin, Yaxin Lu
·
·▲ 346 upvotes
Este artigo está em destaque na seleção diária de papers do Hugging Face, curada pela comunidade de pesquisa em IA.
Autores: Ziheng Qin, Yaxin Lu, Zhangyang Atlas Wang, Kai Wang
- 346 upvotes da comunidade
- Temas: agent-native runtime, durable states, phase-local context, checked transitions, recoverable runbooks, versioned procedural practices
Resumo
Resumo original (em inglês), extraído do paper:
StateM is a runtime system that improves long-horizon agent execution through durable states, recoverable runbooks, and enforceable procedural controls without altering model weights.