Models for minimalist RAG: B1ade 335M Embedding and 1B Parameter Small Language Models
arXiv:2607.27506v1 Announce Type: new Abstract: Language and embedding models used in RAG systems are conventionally assumed to require large-scale pretraining and explicit grounding supervision. We present B1ade, an efficient RAG architecture comprising two purpose-built components: a compact embedding model and a purpose-built SLM. B1ade-embed, a 335M parameter retrieval model constructed via parameter-free fusion of five pretrained encoders achieves top MTEB scores among sub-500M models with ...
arXiv cs.CL
·Shreyas Subramanian, Mecit Gungor, Vikram Elango
·
// relacionados
Leia também
Blog
Claude published malicious code to the Internet and attacked 3 real companies
Blog
LingBot-Map Tutorial: GPU-Aware Inference and Point Cloud Export
Blog
Thinking Machines bets on efficiency over size with its second model, Inkling Small
Blog