Sovereignty · 2026-09-18 · 1:43

What a regular computer can actually run — and the number that decides it

A benchmark on a small mini PC (unified memory, video RAM set in the BIOS, Ollama over SSH) puts real numbers on a question worth asking before renting more cloud: what can you run at home? The answer turns on one metric — how many parameters are active per token, not how many the model has in total.

In favour
  • A 35B mixture-of-experts model reached ~38 tokens/second, while a smaller 12B dense model managed only ~10. Sparse-but-large beats small-but-dense on this class of hardware.
  • Several capable models (a new 27B, a 26B, a 20B) ran at readable speed on a machine that costs a fraction of a workstation.
Worth watching
  • "A regular computer" is generous: this is a mini PC bought with local models in mind, not a typical home machine with no dedicated graphics.
  • Numbers are single-run, one quantisation, default settings. Treat them as a direction, not a spec — measure your own model on your own task.
Our takeThe lesson is not a hardware recommendation, it is a selection rule: for a modest box, pick models by active parameters per token. It is the cheapest path to keeping capable inference on hardware you own.

Source: Local LLM speed tests on a mini PC

← All Focus posts