Quantization & Compilation · Research
Compiler Optimization for Large Models on Edge NPUs
Edge neural processing units offer energy-efficient matrix arrays, tightly coupled on-chip memory, and asynchronous DMA. However, static graphs, limited dynamic shapes, restricted address spaces, and power and thermal constraints conflict structurally with autoregressive state, variable sequences, and mixed operators in large models.
Portable FP4 Quantization for Heterogeneous Accelerators
FP4 computing for large models is developing into a hardware ecosystem where MXFP4, NVFP4, and vendor-specific variants coexist.