Paper
LLMs & Texto
FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution
FreeToken is an edge-native Mixture-of-Experts serving system that dynamically maps computation and model state onto heterogeneous local hardware to run large open-weight models on…
Hugging Face · Daily Papers
·Shuo Yang, Xiaoze Fan
·
·▲ 43 upvotes
Este artigo está em destaque na seleção diária de papers do Hugging Face, curada pela comunidade de pesquisa em IA.
Autores: Shuo Yang, Xiaoze Fan, Melissa Pan, Haocheng Xi, Zhe Wang, Shanlin Sun
- 43 upvotes da comunidade
- Temas: MoE serving, expert residency, CPU-GPU execution, agentic state reuse, runtime memory management, offloading strategy
Resumo
Resumo original (em inglês), extraído do paper:
FreeToken is an edge-native Mixture-of-Experts serving system that dynamically maps computation and model state onto heterogeneous local hardware to run large open-weight models on personal machines.