An MLIR-Based Compilation Method for Large Language Models

arXiv:2607.15865v1 Announce Type: new Abstract: Large Language Models (LLMs) have become the dominant workload on modern AI accelerators, yet deploying them on specialized hardware still faces two core challenges: how to import a trained model into a compiler-friendly intermediate representation, and how to efficiently schedule the autoregressive inference loop under limited on-chip memory. This paper presents an MLIR (Multi-Level Intermediate Representation) based compilation method for large l...

arXiv cs.CL ·Pengchao Hu, Zhibin Xin, Yifan Chen, Yangyang Zhou, Liang Wang ·
compartilhar: