Running a 27B model locally: it's the memory pipe, not the chip
A hands-on test of a 27B model on integrated graphics shows the real bottleneck: on shared memory, every token streams all the weights past the chip, so speed is bound by memory bandwidth, not raw compute.
In favour
Free wins that actually help: a tighter quantisation that fits in VRAM, plus speculative decoding, roughly double throughput.
An external GPU over USB4/Thunderbolt is enough for inference — text generation barely uses the link once weights are resident.
Worth watching
Integrated GPUs stay slow for large dense models however you tune them.
Blindly offloading all layers when the model does not fit makes it dramatically slower, not faster.
Our takeFor local, sovereign LLMs the rule is blunt: buy VRAM, not bandwidth. Last year's model that you own at zero cost per token often beats this year's that you rent.