LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation
arXiv:2607.27353v1 Announce Type: new Abstract: Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, or session-state layer. We introduce LayerRAG-Bench, a controlled cross-layer reliability benchmark with 8 enterprise domains, 240 tasks, 9 fault scenarios, 2 contract modes, and 38,880 live task-level records across nine models from OpenAI, Anthropic, and Gemini. Schema normalization raises schema-dri...
arXiv cs.CL
·Musa Shams (Independent Researcher)
·
// relacionados
Leia também
Blog
Claude published malicious code to the Internet and attacked 3 real companies
Blog
LingBot-Map Tutorial: GPU-Aware Inference and Point Cloud Export
Blog
Thinking Machines bets on efficiency over size with its second model, Inkling Small
Blog