Chip Architecture · Research
A Chiplet Expert-Reuse Architecture for MoE Inference
Mixture-of-experts models expand parameter capacity through sparse activation, but Top-k routing creates expert hotspots, random weight loading, inter-chiplet scatter-gather traffic, and synchronization tail latency in chiplet systems.
The Evolution of Memory-Centric Architectures for Large-Model Inference
As reasoning models, long contexts, and agent applications evolve, the main inference bottleneck is shifting from insufficient peak compute to frequent movement of model weights, KV caches, expert parameters, and intermediate state across memory hierarchies and interconnects.