A learned router scores experts and selects one or more for each token or example. Sparse activation can increase parameter capacity without making the computation per token grow at the same rate, though routing and load balancing add complexity.
A mixture-of-experts model routes each input to a subset of specialized parameter blocks instead of activating every block.
A learned router scores experts and selects one or more for each token or example. Sparse activation can increase parameter capacity without making the computation per token grow at the same rate, though routing and load balancing add complexity.