What a regular computer can actually run — and the number that decides it
A benchmark on a small mini PC (unified memory, video RAM set in the BIOS, Ollama over SSH) puts real numbers on a question worth asking before renting more cloud: what can you run at home? The answer turns on one metric — how many parameters are active per token, not how many the model has in total.
In favour- A 35B mixture-of-experts model reached ~38 tokens/second, while a smaller 12B dense model managed only ~10. Sparse-but-large beats small-but-dense on this class of hardware.
- Several capable models (a new 27B, a 26B, a 20B) ran at readable speed on a machine that costs a fraction of a workstation.
- "A regular computer" is generous: this is a mini PC bought with local models in mind, not a typical home machine with no dedicated graphics.
- Numbers are single-run, one quantisation, default settings. Treat them as a direction, not a spec — measure your own model on your own task.
Our takeThe lesson is not a hardware recommendation, it is a selection rule: for a modest box, pick models by active parameters per token. It is the cheapest path to keeping capable inference on hardware you own.
Source: Local LLM speed tests on a mini PC
← All Focus posts