Mixture of Experts vs Dense Models: Key Differences
The short answer A dense model runs every parameter on every token. A mixture of experts instead routes each token through a few chosen experts. Mixtral’s router selects 2 of 8 experts. So it touches only 13B parameters per token,…