AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation
We present AgentLens, a production-assessed benchmark for interactive code agents.
Hugging Face · Daily Papers
·Andrey Podivilov, Vadim Lomshakov
·
·▲ 7 upvotes
Este artigo está em destaque na seleção diária de papers do Hugging Face, curada pela comunidade de pesquisa em IA.
Autores: Andrey Podivilov, Vadim Lomshakov, Sergey Savin, Matvei Startsev, Roman Pozharskiy, Maksim Parshin
- 7 upvotes da comunidade
Resumo
Resumo original (em inglês), extraído do paper:
We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit -- did the task pass? -- but the people who actually use these agents experience the entire trajectory: how the agent follows instructions, uses its tools, verifies its own work, recovers from mistakes, and talks to them along the way. AgentLens evaluates that whole trajectory. It pairs formal verification, where an objective check exists, with LLM-written trajectory reviews and side-by-side comparisons, so that each run yields a readable explanation of why the score is what it is. This makes AgentLens useful for more than ranking models: we use it to diagnose model behavior, compare successive versions of our own agent, and catch product regressions in a nightly evaluation pipeline. We release the benchmark as open source at https://github.com/agent-lens/agent-lens-bench.Onde ler
// relacionados
Leia também
Blog
Orquestração agêntica: as organizações de IA corporativa têm um problema de implantação, não um problema de plataforma — e a maioria está chamando chatbots de agentes
Blog
Soofi Consortium lança o Soofi S 30B-A3B: um modelo de fundação MoE híbrido Mamba-Transformer aberto para alemão e inglês
Modelo
thinkingmachines/Inkling
Blog