Skip to main navigation Skip to search Skip to main content

CIM for Transformer Models: Enhancing Large Language Model Inference Efficiency

  • Meng Syuan Li
  • , Jung Fang Ke
  • , En Ming Huang
  • , Zhi Wei Liu
  • , Yu Guang Chen
  • , Chun Yi Lee

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

In the field of large language model (LLM) inference, the high computational demand and extensive memory requirements for weights and key-value (KV) cache storage present significant challenges. This issue becomes especially problematic when relying exclusively on GPUs, as they often lack the capacity to accommodate the entire KV cache, particularly in larger LLMs. In the absence of direct communications like NVlink among multiple GPUs, LLMs typically require offloading the KV cache to the CPU for storage and computation, followed by transferring the multi-head attention results back to the GPU for subsequent transformer computations. Given that attention score computation is computationally demanding on the CPU and requires substantial data movement between KV caches and memory, the direct computation of attention scores and even the feedforward layers on Compute-in-Memory (CIM) systems emerges as a viable alternative. This paper is at the forefront of integrating CIM technology in LLM inference, and proposes an innovative architecture that leverages this emerging technology to enhance inference efficiency. Specifically, we present a tailored CIM-based dataflow and hierarchy design for optimize the computation of attention scores and feed-forward layers using CIMs. The results show improvements in performance, with 0.026 × inference latency and 1.199 × 10-3 × energy as compared to a CPU-based implementation.

Original languageEnglish
Title of host publicationIEEE Computer Society Annual Symposium on VLSI, ISVLSI 2025 - Conference Proceedings
PublisherIEEE Computer Society
ISBN (Electronic)9798331534776
DOIs
StatePublished - 2025
Event28th IEEE Computer Society Annual Symposium on VLSI, ISVLSI 2025 - Kalamata, Greece
Duration: 6 Jul 20259 Jul 2025

Publication series

NameProceedings of IEEE Computer Society Annual Symposium on VLSI, ISVLSI
ISSN (Print)2159-3469
ISSN (Electronic)2159-3477

Conference

Conference28th IEEE Computer Society Annual Symposium on VLSI, ISVLSI 2025
Country/TerritoryGreece
CityKalamata
Period6/07/259/07/25

Fingerprint

Dive into the research topics of 'CIM for Transformer Models: Enhancing Large Language Model Inference Efficiency'. Together they form a unique fingerprint.

Cite this