How It Works
Mixture of Experts (MoE)
A model design where many specialized “expert” sub-networks exist, but only a small subset activates for any given input — letting a model have a huge total size while keeping the computation per query relatively cheap.
Origin
Foundational concept from Robert Jacobs, Michael Jordan, and Geoffrey Hinton (1991); revived for large-scale deep learning by Noam Shazeer and colleagues at Google in 2017, becoming mainstream in LLMs from 2020 onward.
