Token-Level Diagnosis of Sycophancy in LLMs with Attribution-Guided Steering

arXiv:2607.28906v1 Announce Type: new Abstract: Sycophancy refers to the tendency for large language models (LLMs) to match user beliefs at the cost of factual correctness, thereby undermining model reliability. Prior work on evaluating sycophancy in LLMs aims to assess whether a model's output matches an authority's claim, but cannot reveal which part of the prompt drives this sycophantic behavior. To bridge this gap, we investigate the relationship of sycophantic responses with an authority's ...

arXiv cs.CL ·Hieu Nguyen, Mahammed Kamruzzaman, Anshuman Chhabra, Gene Louis Kim ·
compartilhar: