From Plausible to Actionable: A Position on LLM Self-Explanations

arXiv:2607.15957v1 Announce Type: new Abstract: Large Language Models (LLMs) can generate natural language explanations that rationalize their own decisions, a phenomenon commonly referred to as self-explanations.Such explanations have emerged as a promising direction for explainable artificial intelligence (XAI), particularly for interpreting LLM behavior.However, while self-explanations often appear plausible, whether they faithfully reflect a model's underlying reasoning process remains an op...

arXiv cs.CL ·Elize Herrewijnen, Benedetta Muscato, Gizem Gezici, Fosca Giannotti ·
compartilhar: