{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-09-20-mixture-of-experts-explained/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"c2138c41-dd09-5194-ba8b-65db2d390778","excerpt":"The first time I looked at deploying a Mixture-of-Experts (MoE) model, the spec sheet confused me. Mixtral 8x7B is described as an 8×7B model that “only…","html":"<p>The first time I looked at deploying a Mixture-of-Experts (MoE) model, the spec sheet confused me. Mixtral 8x7B is described as an 8×7B model that “only activates about 13B parameters per token.” I read that as “runs like a 13B model” and sized a GPU for a 13B model. The weights would not fit.</p>\n<p>That gap, between what a MoE model computes and what it costs to hold in memory, is the whole story. It is also the part most summaries skip.</p>\n<p>This post is for engineers who are choosing a model to serve, or trying to understand why a MoE model is fast to run but expensive to host. I cover what an expert actually is, how the router picks experts, why compute and memory scale off completely different numbers, and the failure modes that show up once real traffic hits the model.</p>\n<h2>What an “expert” actually is</h2>\n<p>Start with an ordinary transformer block. Attention mixes information across tokens. Then a feed-forward network (FFN) processes each token position on its own. In a dense model, that FFN is a single pair of large matrices, and every token goes through all of it.</p>\n<p>A Mixture-of-Experts layer replaces that one FFN with several FFNs, called <strong>experts</strong>, plus a small <strong>router</strong> that decides which experts handle each token. Each expert has the same shape as the FFN it replaced; there are just more of them. The idea goes back to the <a href=\"https://arxiv.org/abs/1701.06538\">sparsely-gated MoE layer from Shazeer et al. (2017)</a>. Google’s <a href=\"https://arxiv.org/abs/2101.03961\">Switch Transformer</a> scaled it up and cut routing down to a single expert per token to keep things cheap.</p>\n<p>The word “expert” oversells it. Nobody assigns Expert 3 to Python and Expert 6 to French. Experts start random and specialize during training, in ways that mostly do not map to anything a human would name. What matters operationally is the count: more experts means more total parameters, and the router means only a few of them run for any given token.</p>\n<h2>How the router picks: a softmax and a top-k</h2>\n<p>For each token, the router:</p>\n<ol>\n<li>takes the token’s hidden state (the vector that represents the token at this layer),</li>\n<li>multiplies it by a small weight matrix to get one score per expert,</li>\n<li>applies a softmax and keeps the <code class=\"language-text\">k</code> highest-scoring experts.</li>\n</ol>\n<p><a href=\"https://arxiv.org/abs/2401.04088\">Mixtral 8x7B</a> uses eight experts per layer and <code class=\"language-text\">k=2</code>. Two experts run, the router’s weights combine their outputs, and the other six do nothing for that token.</p>\n<p>Here is the routing math with nothing else around it. <code class=\"language-text\">scores.topk</code> picks each token’s <code class=\"language-text\">k</code> experts, the softmax turns their scores into weights, and the loop adds up each chosen expert’s output scaled by its weight:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">import</span> torch\n<span class=\"token keyword\">import</span> torch<span class=\"token punctuation\">.</span>nn<span class=\"token punctuation\">.</span>functional <span class=\"token keyword\">as</span> F\n\n<span class=\"token keyword\">def</span> <span class=\"token function\">moe_layer</span><span class=\"token punctuation\">(</span>x<span class=\"token punctuation\">,</span> router<span class=\"token punctuation\">,</span> experts<span class=\"token punctuation\">,</span> k<span class=\"token operator\">=</span><span class=\"token number\">2</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    <span class=\"token comment\"># x: (num_tokens, d_model)</span>\n    scores <span class=\"token operator\">=</span> router<span class=\"token punctuation\">(</span>x<span class=\"token punctuation\">)</span>                      <span class=\"token comment\"># (num_tokens, num_experts)</span>\n    topk_w<span class=\"token punctuation\">,</span> topk_idx <span class=\"token operator\">=</span> scores<span class=\"token punctuation\">.</span>topk<span class=\"token punctuation\">(</span>k<span class=\"token punctuation\">,</span> dim<span class=\"token operator\">=</span><span class=\"token operator\">-</span><span class=\"token number\">1</span><span class=\"token punctuation\">)</span>\n    topk_w <span class=\"token operator\">=</span> F<span class=\"token punctuation\">.</span>softmax<span class=\"token punctuation\">(</span>topk_w<span class=\"token punctuation\">,</span> dim<span class=\"token operator\">=</span><span class=\"token operator\">-</span><span class=\"token number\">1</span><span class=\"token punctuation\">)</span>      <span class=\"token comment\"># normalize over the chosen experts</span>\n\n    out <span class=\"token operator\">=</span> torch<span class=\"token punctuation\">.</span>zeros_like<span class=\"token punctuation\">(</span>x<span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">for</span> slot <span class=\"token keyword\">in</span> <span class=\"token builtin\">range</span><span class=\"token punctuation\">(</span>k<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n        idx <span class=\"token operator\">=</span> topk_idx<span class=\"token punctuation\">[</span><span class=\"token punctuation\">:</span><span class=\"token punctuation\">,</span> slot<span class=\"token punctuation\">]</span>             <span class=\"token comment\"># which expert each token picked</span>\n        w <span class=\"token operator\">=</span> topk_w<span class=\"token punctuation\">[</span><span class=\"token punctuation\">:</span><span class=\"token punctuation\">,</span> slot<span class=\"token punctuation\">]</span><span class=\"token punctuation\">.</span>unsqueeze<span class=\"token punctuation\">(</span><span class=\"token operator\">-</span><span class=\"token number\">1</span><span class=\"token punctuation\">)</span>   <span class=\"token comment\"># its weight</span>\n        <span class=\"token keyword\">for</span> e<span class=\"token punctuation\">,</span> expert <span class=\"token keyword\">in</span> <span class=\"token builtin\">enumerate</span><span class=\"token punctuation\">(</span>experts<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n            mask <span class=\"token operator\">=</span> idx <span class=\"token operator\">==</span> e\n            <span class=\"token keyword\">if</span> mask<span class=\"token punctuation\">.</span><span class=\"token builtin\">any</span><span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n                out<span class=\"token punctuation\">[</span>mask<span class=\"token punctuation\">]</span> <span class=\"token operator\">+=</span> w<span class=\"token punctuation\">[</span>mask<span class=\"token punctuation\">]</span> <span class=\"token operator\">*</span> expert<span class=\"token punctuation\">(</span>x<span class=\"token punctuation\">[</span>mask<span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">return</span> out</code></pre></div>\n<p>The nested loop is written for clarity, not speed. A real kernel groups tokens by their chosen expert and runs each expert once over its batch, and that is where the engineering effort actually goes. But the shape is right: the softmax picks, the top-k prunes, and the outputs get summed with the router’s weights.</p>\n<p><img src=\"/54fe54d20975860a27a656abd38b92b8/moe-routing.svg\" alt=\"One token passing through a Mixture-of-Experts layer. The token&#x27;s hidden state goes into a router that runs a softmax over eight experts and keeps the top two. Two experts, drawn highlighted, are active with routing weights 0.7 and 0.3; the other six are greyed out and idle. The two active outputs feed a weighted-sum block that computes 0.7 times expert two plus 0.3 times expert five, which becomes the output for the next layer. A caption notes that compute paid is two of eight experts, twenty-five percent, while memory paid is all eight experts held in GPU VRAM, one hundred percent.\"></p>\n<h2>Why compute and memory scale differently</h2>\n<p>This is the part I got wrong. The two costs scale off different numbers.</p>\n<p><strong>Compute per token scales with the <em>active</em> parameters</strong>, because a token only touches <code class=\"language-text\">k</code> experts. Mixtral activates roughly 12.9B parameters per token even though the model holds 46.7B in total. A forward pass therefore costs about what a 13B dense model costs, not a 47B one. That is the selling point: you get the quality that comes with a large parameter count while paying a small model’s FLOPs per token.</p>\n<p><strong>Memory scales with the <em>total</em> parameters</strong>, because the router can send the next token to any expert. You do not know in advance which two of the eight a token will pick, and across a batch every expert gets used. So all of them have to be resident in GPU VRAM. There is no version of this where you keep only the active experts loaded.</p>\n<p>That means Mixtral 8x7B needs memory for its full ~47B parameters. At bf16 (two bytes per parameter) that is roughly 94GB of weights before you add anything else, which does not fit on a single 80GB card. You are in multi-GPU territory for a model whose per-token compute looks like a 13B.</p>\n<p>On top of the weights you still pay for the KV cache, and MoE does not shrink it at all. Attention is shared, so the <a href=\"/blog/2026-07-12-llm-kv-cache-gpu-memory/\">KV cache math</a> is the same as for a dense model of the same width and depth.</p>\n<p>In short, MoE trades memory for compute: you spend VRAM to buy cheaper tokens. Whether that is a good trade depends entirely on whether you have the VRAM to spend.</p>\n<h2>Load balancing: stopping one expert from eating everything</h2>\n<p>Left alone, routers collapse. Early in training a couple of experts get slightly better, so the router sends them more tokens. They get more gradient, so they get better still. After a while most tokens funnel into a small set of experts while the rest sit dead. You paid for eight experts and trained three.</p>\n<p>The standard fix is an auxiliary load-balancing loss: an extra training term that penalizes uneven routing and pushes the router to spread tokens across experts. The Switch Transformer paper describes the version most implementations follow.</p>\n<p>This mostly matters at training time, but it leaks into serving in one way worth knowing. Even a well-balanced router only balances <em>on average</em>. Any given batch can send an uneven share of tokens to one expert, and that expert becomes the batch’s bottleneck.</p>\n<p>Production MoE serving handles this with a <strong>capacity factor</strong>, a cap on how many tokens one expert will accept per batch. Tokens over the cap get dropped from that expert. (They still pass through via the residual connection, just without expert processing.) Setting the cap is a genuine tradeoff, with no default that is right for every workload:</p>\n<ul>\n<li><strong>Too low</strong>, and you drop real work and lose quality.</li>\n<li><strong>Too high</strong>, and you waste memory padding every expert’s buffer to the worst case.</li>\n</ul>\n<h2>What changes when you serve it</h2>\n<p>Three things follow once you actually put a MoE model behind traffic.</p>\n<p><strong>Sharding works differently.</strong> A dense 47B model splits across GPUs with tensor parallelism (each GPU holds a slice of every weight matrix). A MoE model can also use <strong>expert parallelism</strong>: put different experts on different GPUs, and route each token to whichever card holds its expert. That cuts per-GPU memory, but it adds a network hop in the middle of the layer and makes throughput sensitive to how evenly tokens land across cards. A skewed batch means one GPU works while the others wait.</p>\n<p><strong>Batching still helps, but the arithmetic shifts.</strong> With <a href=\"/blog/2026-08-31-continuous-batching-llm-serving/\">continuous batching</a> on a dense model, more concurrent requests keep the matrix multiplies full. On a MoE model those tokens scatter across experts, so a small batch can leave individual experts underfed even while the GPU as a whole looks busy. You often need more concurrency before a MoE model reaches the efficiency a dense model hits sooner.</p>\n<p><strong>The sizing rule flips from what the marketing implies.</strong> Budget memory for the total parameter count and compute for the active count. If you size the GPU off the active count, the weights will not load, which is exactly the mistake I opened with.</p>\n<h2>Tradeoffs and what I watch</h2>\n<p>MoE is a good deal when you are throughput-bound and have the VRAM: you serve big-model quality at small-model token cost. It is a bad deal when you are memory-bound, because you pay to hold parameters that any single token rarely runs through. For a lot of self-hosted setups on one or two GPUs, a dense model in the same memory budget is simpler, and its latency is easier to reason about.</p>\n<p>When I do run one, the metric I care about is routing balance under real traffic, not the average the training loss optimized. If one expert is consistently hot, latency gets spiky in a way that is hard to trace back. The model is “working”; a subset of it is just overloaded. Per-expert token counts, exported the same way I’d wire up any other <a href=\"/blog/2026-07-01-opensearch-workflow-monitoring/\">service metric into OpenSearch</a>, turn that mystery into a dashboard.</p>\n<p>Most of the LLM plumbing I built for <a href=\"/project/archi/\">Archi</a>, the RAG copilot for CMS operations, sits in front of whatever model is behind the endpoint. So for me the MoE-versus-dense choice is a hosting-cost and latency question rather than an application one. But it is the kind of question that decides your GPU bill, and “only activates 13B parameters” is the phrasing that gets people to size it wrong. Read it as: cheap to run, expensive to hold.</p>\n<h2>References</h2>\n<ul>\n<li>Shazeer et al., <a href=\"https://arxiv.org/abs/1701.06538\">Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer</a> (2017)</li>\n<li>Fedus et al., <a href=\"https://arxiv.org/abs/2101.03961\">Switch Transformers</a> (2021)</li>\n<li>Jiang et al., <a href=\"https://arxiv.org/abs/2401.04088\">Mixtral of Experts</a> (2024)</li>\n<li>NVIDIA, <a href=\"https://developer.nvidia.com/blog/applying-mixture-of-experts-in-llm-architectures/\">Applying Mixture of Experts in LLM Architectures</a></li>\n</ul>","frontmatter":{"title":"Mixture of Experts: Sparse Compute, Dense VRAM","date":"2026-09-20T00:00:00.000Z","description":"Mixture of Experts routes each token to a couple of expert layers, so the model runs cheaper per token yet still needs every expert sitting in GPU memory.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAAB4ElEQVQoz0WQ227TQBCGjYDE9q69u453fVh7N7ZzdtLEdYmiqDmgVAIJlbY3CHGBhAQS4im4g2t4H654MCZxaaVPvzwz/z+2x5ioKVCqGehQjs/79WF52FTby/lmNV3vzvf7+iUolFfLq6pfg+fBb7RI2iZpy00wL1g8Fno2qffDapuX66xcDRfbUbUbzDfZZDVabIWauUEf+Tn4IWVgphrQEW3TtIVj0FOpbJLaVEEJmK6EjuN10f+IgcAHM5L6Qa/fq7JsVuRzx9OIJs2oATpeUHiiYKKgPAc/NA2bJA0s6As58o46dv3soQ9YroSAHl1EvcVgspTdmenE0IewRCR57oRa5H/eff37/tuP3Y1hcUwhJhsQkW2alvrsslc/9TLiw5shLA3bjQHTifLu2ffXH36+/fRlf2f7XQuHlhO1ndByIxuHLS+7vnj1+83HzfzA1dx2QkgZMGuAZU+QeOZpWxRcjvx4CAzUlPEcszTOqrDc36yvf91+ZlmNThHDgt0nIOxQid0IBsfPcSLEklhXTPTgSHCOWJVhWppU0o5CbgwRCAcNpKOJrwnvUp5hevxVFo19/YIEQ0Tijpd0OqnvK8E155owaTqBYWLxCHrU5iQ2idBxUdzG4h50/wCef14NSLeaKCNzAAAAAElFTkSuQmCC","aspectRatio":1.899441340782123,"src":"/static/af4d473859fb3f274c785c8efe369f46/40a76/hero.png","srcSet":"/static/af4d473859fb3f274c785c8efe369f46/c972b/hero.png 340w,\n/static/af4d473859fb3f274c785c8efe369f46/27625/hero.png 680w,\n/static/af4d473859fb3f274c785c8efe369f46/40a76/hero.png 1360w,\n/static/af4d473859fb3f274c785c8efe369f46/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-09-20-mixture-of-experts-explained/","previous":"blog/2026-09-21-binary-quantization-vector-search/","next":"blog/2026-09-19-minhash-lsh-near-duplicate-detection-rag/"}},"staticQueryHashes":["32046230"]}