How Quorra works

This page explains Quorra’s design in plain terms. For the full engineering detail, see the technical documentation.

The shape of the system

Everything runs inside one home server. At the center is QuorraAPI, the service that does the thinking. Around it:

  • Clients — the ways you talk to Quorra. Today that’s a chat interface (Matrix / Element) and a web chat (Open WebUI); a dedicated app is planned.
  • Local services — your file storage, photo library, finance app, media tools, and so on. Quorra doesn’t replace these — she talks to them for you.
  • The model — a large language model running locally on the server’s GPUs. No request for inference ever leaves the machine.

You send a message; QuorraAPI gathers what it needs, decides what to do, possibly takes an action, and replies.

The local services keep their own interfaces — you can still open your finance app or photo library directly. But the design intent is that you rarely need to. Quorra is the front door; the service UIs are there for visual review or the occasional task that conversation handles awkwardly.

The conversation loop

When you send a message, Quorra:

  1. Works out who is asking — each household member has their own identity and private context.
  2. Loads recent conversation history and the core facts she knows about you and the household.
  3. Passes everything to the local model, along with the tools she’s allowed to use right now.
  4. The model may call tools — search your documents, look up a memory, log a transaction — and then writes a reply.

Tools and permission tiers

Every action Quorra can take is a tool, and every tool has a permission tier fixed when it’s built — it can never be escalated at runtime:

  • Tier 0 — read-only. Looking things up. Always allowed, no confirmation.
  • Tier 1 — low-risk changes. Logging a transaction, creating a reminder. Done, then reported.
  • Tier 2 — destructive or outward-facing. Deleting things, sending a message, making a purchase. Quorra describes exactly what she’s about to do and waits for you to confirm.

This is built in from the start, not bolted on — it’s the difference between an assistant you can trust with real tasks and one you can’t.

Memory

Quorra’s memory is layered, and the layers are kept honest:

  • Operational memory — specific facts (“the wifi password is…”, “the car takes 5W-30”). She saves these as you talk, and you can view, edit, or delete any of them.
  • Knowledge retrieval — your documents and notes, searchable so she can answer from them and cite where the answer came from.
  • Episodic memory — short summaries of significant interactions, distilled overnight.
  • Style — the household’s tone and phrasing, learned separately. This layer captures how to say things, never facts.

The rule underneath all of it: facts live in storage you can inspect and delete. Quorra should never “know” something about you that you can’t find and remove.

Running on local hardware

The language model runs on the server’s GPUs through a local inference engine. This is what makes the privacy real: the model is on the machine, so a conversation with Quorra is just two programs talking on the same computer. It also means Quorra keeps working when the internet doesn’t.

Scopes — the right tools in the right place

Different conversations need different abilities. A finance chat should have finance tools; a general chat shouldn’t. Quorra uses scopes to match capability to context: each conversation is locked to a scope, and only the relevant tools are even reachable. It keeps her focused, and it’s a strong safety boundary — an out-of-scope action isn’t just hidden, it’s structurally unavailable.

Two agents, kept separate

Quorra the conversational assistant is one program. A separate diagnostic agent watches the health of the server itself — disks, temperatures, services. They’re deliberately kept apart: a confident-sounding wrong answer is annoying in conversation and dangerous in diagnostics, so the health monitor is built to be deterministic and careful, not chatty.

Want the detail?

The technical documentation has the full architecture, the decision record explains why each major choice was made, and the benchmarks back the performance claims with real measurements.