Towards Safer RAG: Only Agents Capable of System 2 Thinking may Access Untrusted Documents

arXiv:2608.17153v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has significantly enhanced the performance of large language models (LLMs), yet these systems remain vulnerable to knowledge-poisoning attacks, in which misinformation in retrieved documents can influence the model's final outputs. Notably, an LLM may correctly detect that a document contains incorrect information while nevertheless being influenced by it. Prior work has addressed this vulnerability through the ...

arXiv cs.CL ·Mehrdad Ghassabi ·
compartilhar: