Why Kep Runs Entirely on Ollama: The Architecture Behind Local-Only AI for Regulated Industries
Kep is a private AI assistant for law firms and accountants that runs on a dedicated machine in the client's own office. Every architectural decision behind it follows from one rule: a client install only ever talks to a model running in the room it's sitting in. Here's how that rule shapes the code.
The Core Design Constraint
This isn't a philosophical stance against cloud AI. It's a response to the position regulated firms are in:
- They answer to clients, insurers and auditors who increasingly ask one question about any AI tool: where do our documents go?
- "To a third-party model provider, under a data-processing agreement" is a defensible answer — but it is still a promise, not an architecture.
- Accounting work needs exact figures. "Approximately" isn't an acceptable answer when the numbers end up in a tax return.
So the rule is enforced in code, not in a policy document. In llm/__init__.py:
class CloudBackendNotAllowed(RuntimeError):
pass
def _effective_provider() -> str:
provider = settings.llm_provider
if provider in ("claude", "openrouter") and settings.node_type != "internal_dev":
raise CloudBackendNotAllowed(
f"node_type={settings.node_type!r} may not use provider={provider!r}. "
"Only node_type='internal_dev' may use a cloud LLM backend."
)
return provider
The default node_type is kep_node — anything installed at a client. A client machine that is misconfigured to point at a cloud provider doesn't quietly fall back to it; it refuses to start generating.
The one place a cloud model is allowed is an internal instance — which is exactly what our public online demo is. It runs a fictional law firm on a cloud model so anyone can try Kep from a browser, and it says so in a banner on every page. A real Kep never does.
The Hardware Reality
One dedicated machine per client solves three problems at once:
| Problem | Cloud Approach | Kep Approach |
|---|---|---|
| Data residency | Multi-region deployment, data-processing agreements | Physical hardware inside the client's own office |
| Audit trail | Logs scattered across providers | A single SQLite file the client can back up themselves |
| Performance ceiling | Per-seat API quotas | One client = one entire machine's resources |
The Ollama-only rule closes two doors: documents can't leak through a provider's SDK, and a model can't change behaviour mid-contract because a provider updated it. New models ship as a signed package installed on-site; the machine verifies the signature before applying it. The only network traffic is a small outbound health-and-licence check — never documents, questions or answers.
How Retrieval Actually Works
Most RAG systems trust the model to decide which documents it should even look at. For legal work that's not good enough: retrieval itself has to be scoped by ownership, before the model ever gets involved.
- The query is embedded locally.
sqlite-vecsearches only documents the requester owns or has been explicitly given access to (visible_document_ids()inbrain/retrieve.py).- Only that already-filtered result set is handed to the model.
The model can't leak a colleague's private document because it never receives it. test_documents.py covers this from the outside — for example, a document one user uploads is invisible to another user unless it's explicitly shared.
Ledger's Dirty Secret
The accounting module deliberately doesn't trust the model with arithmetic. When a document is ingested:
- A conservative regex extracts amounts (
£1,234.56,GBP 1,234.56,1,234.56 GBP) with their line items. Anything it can't match with confidence is simply not extracted, rather than guessed. - Amounts are stored as integer pence in their own SQLite table — no floating-point drift.
- A chat query gets the semantic RAG results and a deterministic block summed in SQL, scoped by the same ownership rule as everything else:
Computed totals (from 14 structured record(s), computed exactly — not by you):
- Sum: £7,407.36 GBP
- ...
The module's system prompt tells the model to treat that block as authoritative over any figure it reads out of retrieved text.
The Single-Process Bottleneck
Running Flask in production with a single worker sounds unusual until you consider the actual shape of the load: one firm, one machine, not a thousand paralegals hammering it at once. The semaphore that queues LLM generation in api/chat.py is process-local, so it only means something in a single process — which is why the service is configured as gunicorn with --workers 1 --threads N, on purpose. Chat streaming reuses the SSE pattern we'd already proven in Sentinel, which is part of why the stack is Flask rather than FastAPI.
What We Had to Sacrifice
- Instant model updates — a signed on-site package means models change on a maintenance schedule, not the moment a new one drops. For a firm that needs the tool to behave the same way next month, that's a feature.
- Frontier-level raw capability — local open models are very good, but the biggest cloud models still lead. In our own small test (one sample contract with 12 planted problems) the model Kep runs found 11 and the cloud model we compared it with found 10 — one test, not a benchmark.
- Scale illusions — this doesn't scale horizontally by design. Each new client is another machine.
- Interface polish — the web UI and the Word/Excel panel are the product for now; a native menu-bar app and voice input aren't built yet.
None of that is an accident. It's the direct cost of the one rule everything else follows.
Kep is built by FixFlex Ltd, London. Read the non-technical story in Why we built an AI that never touches the cloud, or try the online demo — a fictional law firm, no sign-up — at kep.fixflex.co.uk.