There is a moment in almost every compliance engagement when the conversation stops being about paperwork. Someone at the table says: fine, but we cannot put this in a chatbot, so is there a version of this that stays in the building? The honest answer, in 2026, is yes. It costs less than people assume, it works better than sceptics assume, and the hard part is none of the things they worry about.

What we are actually talking about

A private assistant is an open-weight model running on hardware you control, wrapped in an interface your people use, with access to your documents. Nothing leaves. There is no per-token bill, no vendor terms of service governing your prompts, no international transfer to justify, and no third party in the chain to add to a supplier register.

It is not a frontier model. It will not match the best commercial system on the hardest reasoning tasks, and anyone who tells you otherwise is selling something. The right frame is the one we keep coming back to: last year's frontier, which you own and run at zero cost per token. For a growing list of jobs, that trade is now obviously worth making.

The one number that decides everything

Almost every mistake we see in local deployments comes from choosing the model by its size. The number that matters on modest hardware is how many parameters are active per token.

Here is why. On a machine with shared or integrated memory, generation speed is bound by memory bandwidth rather than raw compute: every token produced streams the active weights past the chip. A dense model activates all of its parameters for every token. A mixture-of-experts model activates a small fraction. The consequence is counter-intuitive and reliable.

In one benchmark on a small mini PC with unified memory that we covered in our Focus notes, a 35-billion-parameter mixture-of-experts model reached roughly 38 tokens per second, while a smaller 12-billion-parameter dense model managed about 10. Sparse-but-large beat small-but-dense by nearly four to one on the same box. Those are single-run numbers at one quantisation with default settings, so treat them as a direction rather than a specification, and measure your own model on your own task. The direction, though, has held every time we have checked it.

Two corollaries follow. The whole model still has to fit in memory even if only part of it is active per token, so total size sets what you can load and active size sets how fast it runs. And if you have a discrete card, the rule inverts towards dense models: buy VRAM, not bandwidth.

Three tiers that actually exist

A mini PC with unified memory, for a small team. Set the video memory allocation in the BIOS, serve with Ollama, and you have a machine that runs a capable mixture-of-experts model at readable speed for a handful of concurrent users. This is the tier that surprises people. It is also the tier where the word "regular computer" gets stretched: this is a small machine bought deliberately with local models in mind, not the laptop somebody already has.

A workstation with a 24GB card, for a department. A dense 27B model with a vision head quantises to roughly 13GB and sits comfortably in 24GB with room for context. Vision matters more than it sounds: a model that can read a screenshot, a scanned invoice or a diagram covers a surprising share of real office requests. Two free improvements roughly double throughput on constrained setups: a tighter quantisation that actually fits in VRAM, and speculative decoding.

A two-card server, for a company. At this point you are running a service, with queuing, concurrency and uptime expectations, and you should serve with vLLM rather than a desktop tool. You are also, whether you admit it or not, running a small piece of production infrastructure, with everything that implies about backups, monitoring and somebody being on call.

One trap worth naming: if the model does not fit, do not blindly offload every layer to the GPU. It makes things dramatically slower, not faster. And an external GPU over USB4 or Thunderbolt is perfectly adequate for inference, because once the weights are resident the link is barely used.

The rest of the stack

The model is the part everyone talks about and the smallest part of the work.

  • Serving. Ollama for a team, vLLM when you need throughput and concurrency. Both run entirely on your hardware and expose an OpenAI-compatible interface, which means most tooling written for commercial APIs points at your box with a changed base URL.
  • The application layer. A self-hosted platform for building the actual assistants, so the people who understand the work can assemble a workflow without waiting for a developer. This is where adoption is won or lost.
  • Retrieval over your documents. The value is rarely the model's general knowledge. It is that it can answer from your contracts, your procedures and your last three years of correspondence.
  • Redaction at the boundary. Even in a local deployment, mask personal data before it goes into indexes and logs. On-device redaction tools now cover Greek reasonably well, though coverage varies by field type and should be tested on your own formats.

The compliance dividend

This is the part that justifies the project to a board, and it is not hand-waving.

With a local assistant there is no processor to add to your NIS2 supplier inventory and no supplier contract to renegotiate for incident notification. There is no international transfer to assess and no sub-processor list to monitor under GDPR. Under the AI Act you remain a deployer with literacy and transparency duties, but the questions about where prompts go, which the omnibus did not touch, answer themselves. And the most common shadow-AI failure mode, staff pasting client material into a free consumer tool because the sanctioned one is unusable, loses its motive when the sanctioned one is fast, unmetered and does not refuse.

The four costs nobody puts in the proposal

There is no one to call. When a commercial API degrades, you open a ticket. When your box degrades, you are the ticket. For a great many organisations this single sentence is the correct reason not to do it.

Evaluation becomes your job. Nobody is silently improving the model behind your back, which is both the point and the problem. You need a small set of real tasks with known-good answers, run after every change, or you will discover a regression through a colleague's bad afternoon.

Model refresh becomes a decision. Open weights move fast. Someone has to notice that a better model now fits in the same memory, test it, and switch. Left alone, a local deployment quietly ages into a worse product than the cloud option it replaced.

One named person has to own it. Hardware, updates, the eval set, the index of documents, the access rules. Half a day a month when it is healthy, a bad week when it is not. Projects that skip this step do not fail loudly; they just stop being used.

What we would actually recommend

Hybrid, almost always. Put the work that is sensitive, repetitive and well-defined on your own hardware: summarising internal documents, drafting from templates, searching your own corpus, classification, transcription, first-pass translation. Keep the frontier subscription for the hard, occasional, non-sensitive reasoning that justifies it. Route deliberately rather than by habit.

Start with one workload, one machine, one owner and a written measurement of whether it worked. If the assistant is faster than the alternative and the answers hold up, the second workload will arrive on its own. If it is not, you will have spent a mid-range workstation finding that out, which is a cheap answer to an expensive question.


The observations on active parameters, memory bandwidth and quantisation come from hands-on benchmarks we summarised in our Focus notes during September 2026; the figures are single-run and hardware-specific by nature. Related reading: Shadow AI and Before You Adopt AI.