Ashwin Kurella on MegaScale-Infer - Efficient Mixture-of-Experts Model Serving with Disaggregated Expert Parallelism

202509191108
Status: #idea
Tags: Microsoft Research

Ashwin Kurella on MegaScale-Infer - Efficient Mixture-of-Experts Model Serving with Disaggregated Expert Parallelism


References

  1. MegaScale-Infer: Efficient Mixture-of-Experts Model Serving with Disaggregated Expert Parallelism | Proceedings of the ACM SIGCOMM 2025 Conference
  2. MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism