Inference-Time Attention Steering for Vision-Language-Action Driving Models

arXiv:2608.17095v1 Announce Type: new Abstract: Vision-language-action (VLA) driving models couple a reasoning stage with a diffusion-based trajectory decoder, but do not give a direct way to redirect attention toward safety-critical actors at inference time without retraining. We studied a bounded additive pre-softmax attention bias on the visual tokens of detector localized traffic actors on Alpamayo-R1's Qwen3-VL backbone. It is applied as a fail open forward pre-hook with no weight changes. ...

arXiv cs.CV ·Darshan Nagendra Prasad, Lars Ullrich, Knut Graichen ·
compartilhar: