Sovereignty · 2026-09-16 · 1:40

Running a 27B model locally: it's the memory pipe, not the chip

A hands-on test of a 27B model on integrated graphics shows the real bottleneck: on shared memory, every token streams all the weights past the chip, so speed is bound by memory bandwidth, not raw compute.

In favour
  • Free wins that actually help: a tighter quantisation that fits in VRAM, plus speculative decoding, roughly double throughput.
  • An external GPU over USB4/Thunderbolt is enough for inference — text generation barely uses the link once weights are resident.
Worth watching
  • Integrated GPUs stay slow for large dense models however you tune them.
  • Blindly offloading all layers when the model does not fit makes it dramatically slower, not faster.
Our takeFor local, sovereign LLMs the rule is blunt: buy VRAM, not bandwidth. Last year's model that you own at zero cost per token often beats this year's that you rent.

Source: «I ran a 27B model on a hand-sized PC» (hardware review)

← All Focus posts