On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness
arXiv:2607.29062v1 Announce Type: new Abstract: Model capabilities have improved in large part due to scaling chain of thought. This has been a promising development for AI safety--where models verbalize their reasoning, it is possible to monitor it. However, in some cases, models do not verbalize important steps in their reasoning process. For example, models prompted with a cue suggesting the incorrect answer may fail to acknowledge that cue, even when it appears instrumental to their conclusi...
arXiv cs.AI
·Matthew Nguyen, Kyle Cox, Austin Meek, Iv\'an Arcuschin
·