When Retrieval Metrics Mislead: Measuring Policy Signal in Long-Horizon Tool-Use Agents
arXiv:2606.23937v1 Announce Type: new Abstract: Exact-match retrieval recall is often used as a proxy for whether a retriever supplies useful policy context to a downstream decision model. We test this proxy for pre-action policy classification in tau-bench using Qwen2.5-3B/7B classifiers. Under gold-policy conditioning, a compact structured state improves macro-F1 over raw trajectories by 0.13-0.17 after tuning. We then replace the benchmark-designated policy clause with the top-ranked clause r...
arXiv cs.CL
·Tianyu Ding, Juan Pablo De la Cruz Weinstein
·
// relacionados
Leia também
Blog
Unicorn, pelican, Middle-earth: OpenAI co-founder Karpathy is looking for the next AI vibe test
Editorial
CAPA: o benchmark que mede se o assistente de código aprende com você — ou repete a mesma pergunta
Blog
Why biological data matters more in AI drug discovery
Blog