Cross-Model Memory Transfer via Target-Side Reader Adaptation
arXiv:2608.17050v1 Announce Type: new Abstract: Methods for improving knowledge use in large language models typically fall into two regimes. Non-parametric retrieval offers flexible access to external knowledge, but adds retrieval latency, context overhead, and only shallow integration with the backbone. Parametric adaptation is efficient at inference time, but entangles knowledge with model weights and can be hard to update, audit, or transfer. Engram-style hashed memory occupies a middle regi...
arXiv cs.CL
·Mingyuan Li, Guangsheng Yu, Xu Wang, Shaoxiong Ji
·