Sovereignty · 2026-09-16 · 1:42

Why a sparse model can outrun a bigger one on a small GPU

On integrated GPUs, generation speed is limited by memory bandwidth: every token streams the active weights past the chip. That is why a mixture-of-experts model that activates only a few billion parameters per token can feel faster than a smaller dense one.

In favour
  • Sparse (mixture-of-experts) models with few active parameters suit bandwidth-bound hardware well.
  • It gives budget-friendly local setups a path to larger models without a discrete GPU.
Worth watching
  • Total model size still has to fit in memory, even if only part is active per token.
  • Benchmark claims move fast; verify a specific model on your own workload before trusting the numbers.
Our takeTwo honest paths for a modest local box: a dense model with an external GPU, or a sparse mixture-of-experts that stays fast on the integrated one. Match the architecture to the bottleneck.

Source: On mixture-of-experts and memory-bound hardware

← All Focus posts