Ashwin Kurella on MegaScale-Infer - Efficient Mixture-of-Experts Model Serving with Disaggregated Expert Parallelism
202509191108
Status: #idea
Tags: Microsoft Research
Ashwin Kurella on MegaScale-Infer - Efficient Mixture-of-Experts Model Serving with Disaggregated Expert Parallelism
References
- MegaScale-Infer: Efficient Mixture-of-Experts Model Serving with Disaggregated Expert Parallelism | Proceedings of the ACM SIGCOMM 2025 Conference
- MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism