Heterogeneous Memory Chiplets Accelerate Multi-Request LLM Inference (NUS)

Researchers at the National University of Singapore published a technical paper titled “CHIPSMORE: Compute-in-Interconnect and -Memory Chiplets for Multi-Mode Multi-Request LLM Inference Acceleration.”
Abstract Excerpt:
“This paper presents CHIPSMORE, a multi-mode and multi-request LLM inference accelerator that integrates compute-in-interconnect and CIM to support both base-mode and low-rank adaptation (LoRA) inference under diverse workloads. CHIPSMORE employs heterogeneous processing elements consisting of resistive RAM analog compute-in-memory (RRAM-ACIM) and static RAM digital compute-in-memory (SRAM-DCIM) interconnected through a programmable Inter-PE computational network (IPCN). ”
Find the technical paper here. August 2026.
Chong, Yue Jiet, Yimin Wang, Zhen Wu, Zixuan Wang, Wei Zhang, and Xuanyao Fong. “CHIPSMORE: Compute-in-Interconnect and -Memory Chiplets for Multi-Mode Multi-Request LLM Inference Acceleration.” arXiv preprint arXiv:2608.30509 (August 2026). https://doi.org/10.48550/arXiv.2608.30509
The post Heterogeneous Memory Chiplets Accelerate Multi-Request LLM Inference (NUS) appeared first on Semiconductor Engineering.