Quorra — Architecture

Last updated: 2026-07-26 (concierge-extraction design — MCP-app target in §4.1d/§6.4, personas→roles naming note in §4.1c; earlier same day: workspace creation — registry-driven wizard + integration_optin opt-in table + POST /workspaces//preview-notes + GET /integrations//budgets; finance offered bind-only, in §4.1c) Status: M3 complete; M4 in progress — QuorraAPI core, finance scope, modes, and observability done; M2 (RAG) shipped (in-process rag/, per-owner Qdrant collections); planning re-platforming onto Radicale CalDAV; workspaces/personas/guest-concierge designed 2026-07-09 (phased build — see plans/workspaces-personas-concierge.md), not yet implemented; remaining service integrations (files, Jellyfin, email, PMS) pending


1. System boundaries

What Quorra owns

  • All AI and ML logic (inference, embeddings, RAG, memory management)
  • The Quorra agent runtime and the diagnostic agent runtime
  • Quorra API — the HTTP service layer connecting clients to Quorra’s intelligence
  • User-facing interfaces (API, PWA client)
  • Orchestration of local service integrations (storage, photos, calendar, media, home automation)
  • The overnight consolidation job
  • Audit logging of agent actions

What Quorra integrates with (but does not own)

The following services are already running on JUNC1. Quorra calls their local APIs; it does not manage or replace them.

ServicePurposeLocal address
NextcloudFile storage, calendar, contactscloud.juncyard.com
ImmichPhoto backup and ML (face recognition, object detection)photos.juncyard.com
JellyfinMedia streamingwatch.juncyard.com
Home AssistantHome automation:8123
Actual BudgetPersonal financebudget.juncyard.com
GiteaCode hostinggit.juncyard.com
Matrix/SynapseMessaging — primary user-facing interface for prototypematrix.juncyard.com
AuthentikSSO/identity — OIDC provider; Authentik UUID is canonical user IDauth.juncyard.com

What Quorra explicitly does not do

  • Manage JUNC1 OS, Proxmox, or Docker infrastructure
  • Replace local services — it wraps and queries them via their local APIs
  • Store user data that a dedicated service already manages (files in Nextcloud, photos in Immich, etc.)

2. Physical and virtual architecture

JUNC1 runs Proxmox as the hypervisor. The junc stack lives entirely inside Ubuntu Server VM 100.

┌──────────────────────── Physical machine ─────────────────────────┐
│  CPU: Ryzen 7 3700X  |  RAM: 32GB DDR4  |  PSU: 1000W            │
│  Storage: 2×1TB NVMe (1 Ubuntu VM / 1 Windows VM), 2×4TB HDD, 1×8TB HDD │
│  GPU 1: RTX 5080 (16GB GDDR7)  |  GPU 2: RTX 3070 (8GB GDDR6)   │
│                                                                     │
│  ┌──────────────────── Proxmox (192.168.0.100) ─────────────────┐ │
│  │                                                                │ │
│  │  ┌─────── Ubuntu VM 100 (192.168.0.49) ──────────────────┐   │ │
│  │  │  Both GPUs passed through to this VM                  │   │ │
│  │  │  All Docker Compose stacks run here                   │   │ │
│  │  │  Domain: juncyard.com                                 │   │ │
│  │  └───────────────────────────────────────────────────────┘   │ │
│  │                                                                │ │
│  │  ┌─────── Windows VM 101 ────────────────────────────────┐   │ │
│  │  │  Gaming only — unrelated to this project              │   │ │
│  │  └───────────────────────────────────────────────────────┘   │ │
│  └────────────────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────────┘

3. High-level system diagram

┌──────────────── Ubuntu VM 100 (junc stack) ─────────────────────────┐
│                                                                       │
│  ┌──────────────────────────────────────────────────────────────┐   │
│  │                    Infrastructure layer                       │   │
│  │  Nextcloud · Immich · Jellyfin · Home Assistant              │   │
│  │  Actual Budget · Gitea                                        │   │
│  │  Authentik (SSO) · Matrix/Synapse · Element Web              │   │
│  └──────────────────────────┬───────────────────────────────────┘   │
│                              │ local API calls                       │
│  ┌──────────────────────────▼───────────────────────────────────┐   │
│  │                      QuorraAPI                                │   │
│  │  /chat  /memory  /context  /events  /health                  │   │
│  │  Tool system with permission tiers                            │   │
│  │  Multi-user session management (Authentik identity)           │   │
│  └────┬──────────┬──────────┬──────────┬────────────────────────┘   │
│       │          │          │          │                             │
│  ┌────▼───┐ ┌───▼────┐ ┌───▼────┐ ┌──▼─────────────────────────┐  │
│  │Memory  │ │History │ │ RAG    │ │  Inference layer             │  │
│  │layer   │ │SQLite  │ │service │ │  llama.cpp · RTX5080+3070   │  │
│  │SQLite+ │ │        │ │        │ │  OpenAI-compatible API       │  │
│  │files   │ │        │ │        │ │                              │  │
│  └────────┘ └────────┘ └────────┘ └────────────────────────────┘  │
│                                                                       │
│  ┌──────────────────────────────────────────────────────────────┐   │
│  │           Diagnostic agent (separate process)                 │   │
│  │           Smaller model · deterministic playbooks             │   │
│  └──────────────────────────────────────────────────────────────┘   │
│                                                                       │
│  ┌──────────────────────────────────────────────────────────────┐   │
│  │                   ~/data/knowledge/                           │   │
│  │    household/ · members/kurt/ · guests/                      │   │
│  │    Live markdown vault — RAG service indexes this            │   │
│  └──────────────────────────────────────────────────────────────┘   │
└──────────────────────────────┬────────────────────────────────────────┘
                               │ relay only (remote access)
                    ┌──────────▼──────────┐
                    │   Clients           │
                    │   Element Web       │
                    │   Web client (M7)   │
                    └─────────────────────┘

4. Core components

4.1 Quorra agent (QuorraAPI)

The user-facing conversational layer, exposed as an HTTP service.

  • Receives messages from clients (Matrix bot, future PWA, direct HTTP)
  • Assembles context from the memory layer and RAG retrieval
  • Calls tools via the permission-tiered action layer
  • Manages per-user session history and identity (Authentik UUID)
  • Does not directly manipulate storage — goes through local service APIs

API surface (implemented in repos/quorra-api/):

  • POST /chat — send message, get reply; accepts origin field (standard, matrix, open_webui, voice) to tailor response style per surface; origin is persisted to the session on first message
  • POST /chat/stream — SSE variant of /chat. Tool calls run synchronously and emit status events; the final LLM reply streams as token events carrying real llama.cpp deltas as they generate, then a terminal done event carries session_id and any pending_confirmation. A failure mid-generation ends the stream with a terminal error event instead of done. Same auth and body as /chat. Separate endpoint because the return type (StreamingResponse) differs fundamentally.
    • The loop assembles deltas itself (chat/streaming.py): content fragments pass through an incremental cleaner and out to the client immediately, while tool-call fragments accumulate silently, so a turn is identified as tool-calling-or-final only once the upstream stream ends — without ever buffering the reply. Model reasoning (reasoning_content, and any inline <think> span) is dropped before it reaches a client.
    • Closing the client connection closes the upstream llama.cpp stream, which aborts the generation and frees the server’s single slot. Whatever text had already been sent is persisted with an in-band interrupted marker, so an interrupted reply survives in session history rather than vanishing. See decisions.md “Real token streaming”.
  • GET|PUT|DELETE /sessions/{session_id}/purpose — manage a per-session “purpose” directive (stored on sessions.purpose, ownership-checked). Injected into the system prompt between the base template and the origin style by _build_session_context(); additive, never replaces Quorra’s core identity. Session state endpoints live under /sessions (not /chat) to separate state management from chat operations.
  • GET|PUT|DELETE /sessions/{session_id}/mode — manage the active mode for a session (see §4.1b Modes below). Validated against the session’s committed scope; returns 400 for unknown modes, 409 for pre-commit sessions.
  • GET|POST|PUT|DELETE /memory — CRUD for user-scoped memories (core/contextual, tag-filtered)
  • GET|POST|PUT|DELETE /household/memory — CRUD for household-scoped memories (any authenticated user can edit/delete)
  • POST /events — inbound webhook from local services (stub — dispatching TBD)
  • GET /health — checks LLM primary and actual-bridge reachability

System prompt architecture: the system prompt is assembled additively from fragments. Base identity (src/quorra_api/prompts/base.md) always; one src/quorra_api/prompts/scopes/<scope>.md fragment per active session scope; one src/quorra_api/prompts/modes/<scope>/<mode>.md fragment if a mode is active (see §4.1a); core memories when any exist; the session purpose only when scope is exactly ["general"]; per-origin style directives (src/quorra_api/prompts/origin/*.md) last. All .md files load once at import. Adding a new origin, scope, or mode is a file + registry entry — no code branches. Session history (load_history) injects time-gap markers ([N hours passed], [next day — …]) between consecutive messages more than 4 hours apart so the LLM can treat pre-gap context as stale; a midnight crossing alone does not trigger one. Time gaps also auto-clear the active mode (if any) so Quorra returns to default scope behavior after inactivity.

Chat scope: every session locks to a scope (a list of capability surfaces, e.g. ["general"], ["finance"]) on first commit. Scope is sent by the client as scope: list[str] on /chat//chat/stream, persisted on sessions.scope_json, and immutable for the session’s lifetime — re-scoping means starting a new session. While scope is NULL (pre-commit) the active tool list is empty and the model can chat but cannot fire actions; this is the hook the future LLM-guided setup-conversation feature will plug into. The tool registry exposes for_scope(scope)/get_for_scope(name, scope) so the active tool list per turn is the intersection of the session’s scope and each tool’s declared scopes frozenset — out-of-scope tools are not just hidden from the LLM, they’re structurally unreachable (defense in depth against prompt injection or name-guessing). Known scopes live in src/quorra_api/chat/scopes.py; the minimal rollout ships general and finance. Mismatched scope on an existing session returns 400.

4.1a Modes

Modes are dynamic behavioral overlays within a scope. Where scopes gate which tools are reachable (structural, registry-level, immutable per session), modes guide how Quorra uses those tools (prompt-level, mutable within a session). A mode never expands the tool surface beyond the session’s scope — worst-case “leak” is irrelevant prompt guidance, not tool access.

Why modes exist: On a 14B model with limited prompt attention, loading every workflow’s instructions into every session wastes context budget. Modes load workflow-specific prompt fragments and pre-fetched context only when actively needed, keeping the effective prompt lean and focused. This is a capability multiplier at small model sizes.

Activation: Externally-driven — clients set the mode via PUT /sessions/{id}/mode. The LLM does not decide to enter a mode on its own (LLM-driven activation can be added later without breaking changes). This keeps scope prompts lean — no mode descriptions are injected until a mode is active.

Storage: active_mode: str | None column on Session (mutable, nullable). Set/cleared via the mode endpoint. NULL means no mode active — the session uses only its scope fragment (default behavior).

Auto-expiry: Piggybacks on the existing 4+ hour time-gap detection in the chat loop. When a gap is detected and a mode is active, active_mode is auto-cleared to NULL before building the system prompt. No new timestamps or columns needed. Matches user intuition: “I walked away and came back, start fresh.”

Prompt composition order (with mode):

base identity → scope fragment(s) → MODE FRAGMENT (if active) → core memories → session purpose → response style

Mode fragments live at src/quorra_api/prompts/modes/{scope}/{mode_name}.md, loaded at import like scope fragments. Entry context (if defined for the mode) is injected as an additional section below the mode fragment on the first turn after mode activation.

Mode registry: ModeDefinition dataclass (name, scope, description, optional entry_context async callable) stored in a ModeRegistry keyed by (scope, name). ModeRegistry.for_scope(scope) returns available modes; validate(mode_name, scope) raises on invalid mode/scope combinations.

Planned modes:

ScopeModeWorkflow
financebudgetPlan next month’s budget — proactive spending analysis, category comparison, allocation setting
financereconcileRapid-fire transaction review and categorization correction
financereviewAnalytical deep-dive into spending trends and anomalies
planningdaily-briefingStructured morning overview — pending items, reminders, schedule (deferred; see plans/daily-briefing-mode.md)
generalresearchExtended topic exploration with active knowledge-base searching
generalhousehold-planningStructured planning for trips, events, projects — checklists and timelines
mediamovie-nightInteractive recommendation narrowing for a group (future — ships with Jellyfin)
planningweekly-reviewWeek-ahead calendar layout with conflict detection (future — ships with Radicale CalDAV)
planningschedulingFind-time-for-X with availability checking (future)
mailinbox-reviewGuided inbox triage — prioritize, summarize, act (future — ships with Protonmail)
mailcomposeContext-aware email drafting with Tier 2 send (future)
photosmemory-laneConversational photo browsing by person/date/event (future — ships with Immich)

4.1b Matrix bot adapter (repos/quorra-matrix/)

A standalone matrix-nio service — the Matrix “mouth and ears” for Quorra, with no LLM or tool logic of its own.

  • DMs (2-member rooms) respond to every message; group rooms respond only when the bot is @mentioned (mention stripped before forwarding, sender display name prepended as [Name]).
  • Per-message identity: static MATRIX_USER_MAP (matrix-id→email) + trusted-service auth (X-Quorra-User-Email + shared secret) — the same auth path as the Open WebUI Pipe. Unmapped senders get a polite refusal.
  • Room→scope mapping via the MATRIX_ROOM_SCOPE_MAP env (JSON {room_id: ["finance", …]}), wrapped in resolve_scope(room_id) -> list[str] so a future swap to a user-defined DB-backed lookup is a one-function change. DMs and unmapped rooms resolve to ["general"]. The bot includes the resolved scope on the first /chat/stream request of a session; the API commits and locks it from there.
  • Room→session mapping and pending Tier 2 confirmations persisted in bot-local SQLite, so a bot restart preserves continuity.
  • Responses use POST /chat (non-streaming, single-message reply). A progressive-edit streaming path using POST /chat/stream + m.replace (batched ~400 ms) was implemented but reverted 2026-05-22 — edit ghosting in Matrix clients made the experience worse than a single reply. Streaming code is preserved as a dormant fallback for future re-evaluation.
  • !quorra purpose|forget|status|mode handled entirely bot-side via the session state endpoints (/sessions/{id}/purpose, /sessions/{id}/mode); never forwarded to the LLM. !quorra mode budget activates a mode; !quorra mode with no argument clears it.

4.1c Workspaces, service bindings & personas

A Workspace is a first-class “brief” Quorra operates within — the generalization of scope along a principal/trust axis. Where scope alone models one user wearing hats, a workspace bundles what Quorra can touch, who it is being and for whom, how much rope it has, and how it carries itself — the way a human executive assistant context-switches between the household, a rental business, and a job search: one assistant, different information boundaries. The motivating case is helping run an STR/MTR rental (including interacting with guests on the owner’s behalf), but the rental is just the first instance.

The four composable axes of a session:

Session ─┬─ workspace_slug  → the brief: a set of service bindings + personas + standing policies
         ├─ scope           → capability domains active this session (⊆ the workspace's bound services)
         ├─ persona         → who Quorra is being: tool allow-list (subtractive), retrieval binding, prompt family, principal kind
         └─ active_mode     → behavioral overlay within a scope (§4.1a — unchanged)

Owner sessions resolve to workspace = <personal/household>, persona = owner (null allow-list) → behavior identical to today. A guest session (§4.1d) resolves to workspace = str, persona = guest-concierge, principal = a reservation.

Service bindings — the uniform capability unit: a workspace’s capabilities are a set of (service, data_window, tier_policy) rows. service names a tool family (≈ a scope): finance, calendar, knowledge, notes, pms, messaging. data_window identifies the slice — which budget, which collection, which PMS account. This is the whole generalization: PMS-on-account-W and finance-on-budget-X are the same kind of row, so nothing is service-specific in the model and the user’s “which services this brief uses” multi-select maps 1:1 to binding rows. A session’s scope is validated to be ⊆ the workspace’s bound services.

Personas — the trust/actor profile (subtractive): a code-level registry entry (like ModeRegistry) carrying tool_allowlist (None = unrestricted), retrieval_binding (None = user-derived; else a fixed knowledge window), prompt_family, and principal_kind. A workspace declares which personas it offers. The allow-list is subtractive — it intersects the scope-reachable tool set — which the additive scope registry does not otherwise express. The owner persona has an all-null profile, so existing sessions are unaffected byte-for-byte; guest-concierge is the restricted one. Personas gate capability (registry layer); modes overlay prompt behavior (prompt layer) — different layers, so modes are untouched. (Naming note, 2026-07-26: “persona” is transitional vocabulary — the target model re-keys this axis as workspace roles resolved from the counterparty’s principal (definitions stay in code; assignments become DB membership rows), landing with multi-user-per-workspace. See decisions.md “Personas dissolve into workspace roles”.)

Principal kinds: the actor generalizes from AuthentikUser to a Principal protocol — AuthenticatedPrincipal (existing), ReservationPrincipal (a guest; a reservation-scoped identity with no owner identity, expiring at checkout + grace), and a documented AnonymousPrincipal seam (front-desk gatekeeper; not built).

Storage: a workspaces table + a workspace_service_bindings table; new nullable workspace_slug and persona columns on Session (backfilled to owner defaults, immutable-after-commit like scope). Personas and the principal protocol are code, not tables (mirroring modes/scopes). New-path-only: the Workspace is the authority for the new (STR/guest) path; the live owner/household retrieval derivation (collections_for_user) is left working as-is — no regression to shipped M2 RAG.

Generic vs. instance: the model (Workspace, ServiceBinding, Persona, Principal, policy) is generic; the str workspace, the guest-concierge persona, this PMS account, and this reservation are instances. Seams the model accommodates but this design does not build: the anonymous front-desk principal, expires_at on a workspace (event/travel briefs), reduced-trust authenticated personas (a kid’s or an elderly parent’s brief), and a sanctioned cross-workspace read-grant (an “accountant” join at tax time) — each a persona/principal/binding variation, not a schema change.

Owner-in-workspace (the owner face; built 2026-07-11). When the owner inhabits a workspace (a session with workspace_slug set, persona = owner), the context is sealed both ways while the identity layer stays universal:

  • RAG: the workspace’s data_dir is indexed into its own collection kb_ws_<slug> and excluded from kb_member_<slug>; in-workspace search_knowledge hardwires to kb_ws_<slug> via search_window. So workspace knowledge is reachable only in-workspace, and never pollutes personal chat. collections_for_user never returns kb_ws_*.
  • Memory: memories.workspace_slug (NULL = personal/household). A workspace turn reads only workspace_slug == X; a personal turn excludes workspace rows (workspace_slug IS NULL). Sealed symmetrically, like RAG.
  • Identity stays universal: Quorra’s self, the user’s name/locale (from user_preferences, not the memory corpus), her capabilities, and the awareness that she’s in this workspace all persist — only the retrievable data is workspace-local. A prompts/workspaces/base.md block frames the sealed context; in personal context, an awareness block lists the user’s workspaces (name + description) so Quorra can offer to switch without seeing specifics.
  • Scope derivation: the owner’s in-workspace scope = owner_workspace_scope(bindings) = (bound ∩ known scopes) − guest-only (hosting) − deferred (finance, until its per-workspace budget binding is wired), at least ["general"]. Owner and guest thus see different tool surfaces of the same workspace.
  • Entry: one Matrix room per workspace (Workspace.matrix_room_id, 1-1). The bot resolves room → workspace via GET /workspaces/by-room and sends the generic workspace_slug on /chat (the chat contract never speaks Matrix room IDs). General/personal is the default home; entering a workspace is the deliberate, mutually-exclusive switch. Threading is small: the loop forwards session.workspace_slug into the prompt builder and the tool executor’s user-scoped arg injection, and the memory/knowledge tools consume it.

Workspace authoring (the write face; built 2026-07-23). Quorra authors markdown documents into her workspace via the owner-only document tool — the workspace is now a living project space, not a read-only shelf. Every write is a validator-gated round-trip: kb.serialize assembles a schema-conformant file (canonical field order, flow-style tags, bare ISO dates, auto-filled quorra provenance) → kb.parse + kb.validate → block on any error-severity issue (write nothing, return the issues so the model self-corrects) → write into the session workspace’s data_dir → rag.indexer.index_file reindexes just that file into kb_ws_<slug> (non-destructive ensure_collection + delete_by_path-before-upsert, since chunk IDs are positional) so it’s searchable the same session. WARNING/INFO issues (relative-date, no-summary-lead) are surfaced but don’t block. create is Tier 1 (write + audit); update is Tier 2 (overwrite → confirmation). Workspace-only for this cut: guarded on the session’s committed workspace_slug, unavailable in personal/general chat. Guest-unreachable by construction — document isn’t in the guest-concierge allow-list, so the persona gate fails it closed. The whole-vault mount is now :rw; the tool’s data_dir containment (kebab_cased filename, no separators) + the persona gate are the real boundary. See plans/vivid-roaming-popcorn.md.

Workspace reorganization (built 2026-07-25). Beyond authoring single files, Quorra restructures the workspace directory via the owner-only organize tool — create_folder, move, rename, promote_to_folder (foo.md → foo/overview.md), delete (dead/empty files only). The link-and-index-preserving engine from the one-time corpus migration is lifted into a pure, workspace-scoped kernel (kb/reorg.py): a single move-map drives every rename, each internal link is re-expressed against its target’s new path (never re-derived from link text) — including the entities: frontmatter cross-refs the migration engine skipped — and a simulate pass asserts zero-new-dangling before anything is written. The RAG index is kept in sync by move-safe helpers (reindex_move/reindex_delete) that clear the old path from both kb_ws_<slug> and its guest window before indexing the new path (positional chunk IDs would otherwise orphan). Like document, organize never writes on call: it validates + simulates then enqueues a reorg kb_suggestion (new target_kind; the ops plan lives in proposed_content), reviewed as a before/after tree; on approval a two-phase journaled apply runs — disk is all-or-nothing (reverse-on-failure), the reindex is best-effort — and it re-validates against current disk so a stale plan fails cleanly instead of half-applying. Tier 1 (the inbox is the confirmation); guest-unreachable (not in the guest-concierge allow-list; the egress suite proves it). The in-workspace prompt block shows Quorra the current file tree so she targets ops at real paths. See plans/quorra-needs-the-ability-delegated-wand.md.

Workspace creation (built 2026-07-26). A personal workspace is now created from the web app via a guided wizard, replacing seed/DB surgery. POST /workspaces (owner-gated) derives a unique kebab slug (collision-suffixed), provisions the vault directory members/<owner>/knowledge/Workspaces/<slug> (a plain mkdir — quorra:vault + setgid + umask 002 inherit ownership/mode; never chown as non-root), writes the Workspace row (kind="personal", persona=["owner"], matrix_room_id NULL — room binding deferred) plus its service bindings, and optionally stores an LLM-drafted, owner-approved starter ## Workspace notes as a workspace-scoped core memory — all before commit, so a provisioning failure rolls back with no orphan row. The offered services are not hardcoded: an integration registry (integrations/registry.py, a frozen-dataclass catalog mirroring ModeRegistry/personas) declares each built-in integration’s metadata (scope, view_key, always_on, owner_creatable, requires_config, an is_configured(settings) probe, and a plugin seam), and a per-instance integration_optin table (seeded idempotently from the configured, non-plugin catalog at startup) records what this junc has enabled. GET /integrations?creatable=true = the wizard’s checkbox feed (enabled ∩ owner-creatable — hosting/pms are owner_creatable=False, never offered). The create path validates requested services against that source of truth and against the owner scope ceiling (on the & KNOWN_SCOPES subset, since knowledge is a binding, not a scope). Finance is offered bind-only: selecting it requires a budget pick (GET /budgets over list_accessible) captured into the binding’s data_window ({"budget": alias}, read by workspace_budget()), but threading that pinned budget through the live finance resolver + lifting the finance-scope deferral is a follow-up — the Finances tab stays present-but-disabled. Starter notes are generated server-side (POST /workspaces/preview-notes, reusing the reflection LLM transport — the llama-server isn’t browser-reachable) and shown editable before Create. See plans/let-s-get-to-work-velvet-quiche.md.

Deferred (next): the finance live-wiring (pinned budget → finance_dispatch, lift the finance deferral, enable the Finances view), plugin/third-party integrations (the registry plugin seam) + a Settings UI for the opt-in table beyond the PATCH /integrations/{service} toggle, workspace-creation follow-ons (Matrix-room auto-provisioning, hosting/guest workspaces, starter-paperwork upload, members & roles), personal-vault authoring (kb_member_<slug> in personal context), guest-content authoring (guest_visible house-guides into the guest window), reflection-driven reorganization proposals (on-demand only for this cut), the PWA’s before/after-tree renderer, and a general inotify watcher for direct human vault edits (this phase reindexes only Quorra’s own writes).

4.1d Guest concierge

The guest concierge is the first externally-facing workspace instance: Quorra fielding guest messages for a rental. It is not bolted-on code — it is the generic chat loop run with persona = guest-concierge and principal = ReservationPrincipal. Swap the persona and principal and the same handler serves any external-facing brief.

Flow:

guest (Airbnb/VRBO/direct) ─msg─▶ PMS ─webhook─▶ /integrations/pms/webhook
                                                      │ normalize → upsert reservation
                                                      ▼
                                             concierge handler (constrained loop as ReservationPrincipal)
                                                      │ tools ∈ {search_knowledge(guest window), check_this_booking, escalate_to_owner}
                                                      ▼
                                             {draft, category, confidence, escalation_reason}
                                                      │ autonomous_ok(...) → False (always, for now)
                                                      ▼
                                             guest_message review row (draft_ready | needs_you)
                                                      │
                          Matrix DM to owner ◀─alert─▶ owner: "send" / "edit: …" / "escalate"
                                                      │ approve
                                                      ▼
                                             PMS send-message API → sent

Egress boundary — three structural controls, never the prompt (a guest may be adversarial; prompt rules don’t hold against injection, so enforcement is at code choke points):

  • Capability: the persona tool_allowlist (exactly search_knowledge, check_this_booking, escalate_to_owner) intersected with scope-reachable tools at the registry + chat-loop gates. Memory tools are excluded entirely — a guest could otherwise poison or read cognition.
  • Retrieval: the persona’s retrieval_binding hardwires the knowledge tool to a dedicated guest-window collection (kb_str_guest); the guest has no user_uuid, so the owner-collection resolver is unreachable.
  • Content: a default-deny guest_visible frontmatter gate — only opted-in vault content is indexed into the guest window.

Because the ReservationPrincipal carries no owner identity, even a bug in the above cannot surface owner data: there is no owner collection to derive. This is Design principle 2 (privacy-by-architecture) applied to an untrusted counterparty.

Draft-and-approve: every reply is drafted for owner approval; nothing sends autonomously today. The concierge emits {draft, category, confidence, escalation_reason} each turn and calls autonomous_ok(workspace, category, confidence), which returns False for all inputs (empty allow-list). Whitelisting categories later (wifi, checkout) lets those auto-send while everything else still routes to review — one policy function plus an already-captured signal. Human-in-loop is also the injection defence during trust-building: the owner sees any injected draft before it leaves.

Review surface: the owner’s existing Matrix/Quorra chat. The concierge posts the draft (or an escalation) into the owner’s room via the bot; the owner replies send / edit: … / escalate. One guest_message review queue with two states (draft_ready, needs_you), reusing the Notification status-state-machine pattern (§4.1b delivery). No new push channel; high-urgency items fire an immediate Matrix DM, otherwise they queue.

Runtime: in-process (the generic loop with a guest persona + reservation principal). The extraction target is now decided (2026-07-26): not a standalone cognitive service but an installable MCP app — the future concierge app carries only PMS/channel I/O behind MCP tools (get_inbound_messages/get_reservation/get_thread/send_reply, the PMSClient surface), while cognition, the egress controls, and the review queue stay in quorra-api (“apps expose capabilities; Quorra thinks” — see §6.4 and decisions.md “Concierge extraction target”). send_reply is an ordinary Tier-2 tool, so draft-and-approve is the standard autonomy-tier system and the egress boundary never moves. Appification waits until Phase 3 has validated live guest behavior. The PMS is a vendor-abstracted service adapter (§6.3/§6.4; Hostaway or Lodgify) — see service-integrations.md — whose PMSClient methods are written as candidate MCP tools for that lift.

4.1e Web client (repos/quorra-web/) — M7

The first-party household client: a mobile-first responsive SPA that becomes an installable PWA as its final, thin layer. “PWA vs web app” is a false dichotomy — a PWA is a web app plus a manifest and a service worker — so the build order is web app first, installability last (matching the M7 checklist). The DoD device is a phone; the design scales up to desktop, where Open WebUI remains the power-user surface until the web client supersedes it.

Stack: React + Vite + TypeScript, Tailwind CSS v4 with components from the neobrutalism.com registry (Base UI variant — the registry’s Radix variant is a broken port, see decisions.md “Web client styling”), lucide-react icons, TanStack Query (server state), react-router, oidc-client-ts (OIDC). SSE consumption is a small hand-written fetch-streaming reader (native EventSource is GET-only; /chat/stream is POST). Static build output — no SSR (authenticated app, no SEO surface).

Identity & auth: the first client to exercise QuorraAPI’s existing Authentik JWT path (auth/middleware.py:get_current_user — JWKS, RS256, audience/issuer checks). quorra-web registers as a public OIDC client in Authentik; the SPA runs the authorization-code + PKCE redirect flow (never popups — redirect survives installed/standalone PWA mode unchanged) and sends the access token as Bearer on every API call. Refresh-token rotation + offline_access give persistent phone login. No new auth code in QuorraAPI; authentik_audience must be configured to accept the new client. participants = [self] for MVP; multi-user sessions are a later UX problem, not an API gap.

Serving (same-origin, no CORS): the container’s nginx serves the static build and proxies /api/* → quorra-api:8000 (SSE: proxy_buffering off on the stream route); host nginx terminates TLS at app.juncyard.com. QuorraAPI deliberately has no CORS middleware — same-origin keeps it that way (no preflights on the SSE POST, no token-bearing cross-origin surface). Deployed as a compose service alongside the other adapters (~/projects/junc1/compose/quorra).

Feature phases:

  • A — shell + auth: scaffold, OIDC login round-trip, authenticated /health call, compose + TLS serving pipeline live.
  • B — chat MVP: new-session flow with scope picker (the session commits scope on first message — the same flow OWUI hardwires to ["general"]), /chat/stream rendering (status events as tool-activity chips, token accumulation, markdown rendering), session sidebar + history. Needs new QuorraAPI surface: GET /sessions (user-scoped list; the /sessions namespace was shaped for this) and a full-history variant of /sessions/{id}/messages/recent.
  • C — structured UI: Tier 2 confirmation cards (render pending_confirmation; approve = next /chat call carrying confirmation_id, decline = let it expire — an explicit cancel endpoint is an open item), notifications outbox surface, purpose/mode controls.
  • D — PWA layer: manifest, icons, minimal app-shell service worker (vite-plugin-pwa), install tested on iOS Safari + Android Chrome.

Deliberately deferred: offline data/sync (meaningless while inference lives on JUNC1 and the client is LAN-only until M8 — the service worker precaches the app shell, nothing more); Web Push (backlog — transits vendor push services, a local-first tension needing its own decision; note iOS delivers Web Push only to installed PWAs, one reason the install layer exists even while thin); inline photo results (blocked on the M4 Immich integration — Phase B builds a generic rich tool-result rendering seam that photo results plug into); a dedicated web origin prompt fragment (MVP sends origin: "standard"; add the origin + its regression tests when style tuning warrants).

4.1e Notepad

The daily-workflow emulation of a paper notepad: frictionless capture all day, one batch of judgment overnight, one review in the morning. Lives in the personal (null-workspace) context — deliberately not a workspace, because workspaces are sealed data boundaries and the notepad is an unsealed intake funnel whose job is routing outward (calendar, tasks, finance, memory, workspace KBs).

Capture (POST /notepad/items, quorra-web Notepad page): zero-inference — the raw jot text plus a timestamp lands in notepad_items, instantly acknowledged. The web client writes every jot to a localStorage queue first and syncs opportunistically (reconnect/focus/interval); a client-generated client_id with a unique (user_uuid, client_id) index makes blind replay idempotent. No service worker in v1 — the queue survives reload and capture only happens with the page open.

Triage (notepad/triage.py, nightly worker + POST /notepad/triage-now): one LLM pass over the day’s captured items (14B, plain chat-completion, prompt-demanded JSON, tolerant parse — the reflection/reflect.py idiom). Three-way outcome per item, flag-don’t-guess:

  • propose → one or more proposed_actions rows (an item can yield several — “dinner w/ Sam Fri, I owe him $20” → event + transaction);
  • ask → an ask-kind queue row carrying one clarifying question;
  • no_action → recorded on the item itself (triage_note), not the queue.

Full accounting: items the model fails to account for stay captured and roll into the next batch — the item’s state transition is the watermark, so re-running is harmless. The prompt injects the newest notepad-correction memories (“Sam always means Sam Reilly”) and the live workspace slugs for routing.

Review (Notepad page, /notepad/actions/*): per-kind cards with structured field editing (no JSON blobs). Approve applies through the existing direct apply functions — RadicaleStore (event/task/reminder), budget resolve + _log_transaction (transaction), create_memory (memory), and enqueue_suggestion (doc — a handoff to the KB inbox, which owns validation and final review). The review click is the Tier-2 confirmation; on apply failure the row reverts to draft_ready with the error surfaced (the kb_suggestions semantics). Answering an ask runs an instant single-item re-triage with the Q/A appended — follow-ups appear in the same review session. Defer rolls a jot to tomorrow’s batch. Any resolve verb can carry a remember note → a notepad-correction memory, closing the learning loop.

proposed_actions is deliberately producer-agnostic (source_kind/source_id, typed kind + validated JSON payload) — it is the intended generalized approvals queue that kb_suggestions and the concierge review may migrate onto later (see decisions.md “Notepad & the approvals queue”).

Config: QUORRA_NOTEPAD_ENABLED (nightly worker only — endpoints are always on), QUORRA_NOTEPAD_TRIAGE_TIME (default 03:30, household_tz).

4.2 Diagnostic agent

A separate process from the Quorra agent — this separation is deliberate and must be maintained.

  • Runs a smaller, more deterministic model (exact model TBD — see open questions)
  • Structured playbooks: SMART data, thermal trends, GPU health, disk space, service health, network
  • Communicates with Quorra agent over an internal IPC interface (format TBD — see open questions)
  • Self-healing for common issues: disk full → suggest cleanup, service crash → restart, DB corruption → restore snapshot
  • Escalates to Quorra agent only when a human-readable explanation or nuanced judgment is needed
  • Must be independently restartable without affecting Quorra agent

Rationale for separation: hallucination is bad in conversation; it is catastrophic in diagnostics.

4.3 Memory layer

A recurring distinction governs where data lives:

  • Operational data — structured, stateful, acted upon by the system. Memories, tasks, reminders, budget state. Lives in the database (SQLite memories table today; additional tables as domains grow). The system reads it to decide what to do.
  • Reference data — natural-language, knowledge-oriented, referred to by the system. Notes, documents, household knowledge. Lives in the markdown vault (~/data/knowledge/) and is retrieved via RAG. The system reads it to inform answers.

The test: does the system act on this data, or refer to it? Act on it → DB. Refer to it → vault. This already holds in practice (memories live in SQLite, not markdown) and should hold as new domains are added. Don’t conflate the two just because they’re both “about the user.”

Five components, in ascending order of write cost:

ComponentTechnologyRoleUpdate frequency
Operational memorySQLite memories tableStructured facts, preferences, household info — scoped by user or household, tiered by importance (core/contextual)Real-time (LLM memory_save tool call)
Vector / RAGQdrant (decided)Fast document retrieval from ~/data/knowledge/Continuous (on ingest)
Knowledge graphTBDRelationships, curated structured factsOn significant events; overnight consolidation
Episodic memoryFile-basedRecent interaction context; significant event summariesOvernight consolidation
LoRA weightsFine-tune checkpointsHousehold tone and style — NOT factsOvernight (rolling window)

Design constraint: facts must live in retrievable, deletable, inspectable storage. LoRA weights capture style only.

Operational memory scoping:

Memories have two dimensions:

  • Scope: user_uuid is set for user-specific memories, NULL for household-wide memories
  • Importance: core memories are injected into every system prompt automatically; contextual memories are retrieved on demand via memory_recall tool when the LLM determines they’re topically relevant

Context injection model — core memories always present, contextual via tools:

  • Core memories (both user and household) are loaded from the DB and injected into the system prompt on every request, between the base identity template and session purpose
  • Contextual memories are retrieved via memory_recall (Tier 0 tool) when the LLM needs domain-specific context — the LLM provides topic tags and the DB returns matching memories
  • Knowledge base is searched via search_knowledge_base (Tier 0 tool, stub until M2 RAG)
  • Conversation history (last 5–8 messages) is always present
  • The LLM chains memory and RAG when needed (two-hop retrieval: memory gives personal facts, RAG gives reference material)

Memory creation: The LLM creates memories at runtime via memory_save (Tier 1 tool call) during normal conversation. No separate extraction model or async post-inference pass. The system prompt includes specific guidance on what to save vs. not save, and when to ask before superseding ambiguous facts. Behavioral regression tests enforce prompt calibration over time.

Memory replacement: memory_save accepts an optional supersedes parameter — a keyword from the old memory to replace. The service layer searches and removes matches, then creates the new memory. The LLM must confirm with the user before superseding when old and new facts could coexist (e.g., a person can own two cars).

Information type taxonomy:

All information Quorra works with falls into one of four categories. The category determines the retrieval mechanism and storage target:

TypeExampleRetrieval patternInterim homeTarget home
Factual”I prefer the long route to avoid traffic”Tag / key-value lookupSQLite memoriesSQLite memories
Entity knowledge”Cara is Gary’s wife”, “Audi takes premium fuel”Graph traversalSQLite memories (interim)Knowledge graph
Episodic”I had a difficult conversation with my brother today”Time + context indexOvernight summaries (partial)Episodic tier
Reference / DocumentsCar manual, saved recipes, personal notesSemantic search~/data/knowledge/ vaultVault + RAG (Qdrant)

Factual memories have the user as their subject — preferences, personal state, biographical facts. Entity knowledge covers properties of and relationships between named entities the user interacts with (people, objects, places); this merges what might colloquially be called “relational” and “domain knowledge” since both require entity-identity lookup rather than semantic similarity, and both target the knowledge graph. The routing rule for ambiguous cases: if the subject is the user themselves → Factual; if the fact creates or models a named entity → Entity knowledge. Reference/Documents are not memory rows — they are vault files indexed by the RAG pipeline.

The memories table carries a memory_type enum column (factual | entity | episodic) so that entity-knowledge rows can be migrated to the knowledge graph when that layer is ready, without having to re-infer type from content.

Expiring and consume-once memories:

Some memories are inherently transient. “I’m not feeling well today” should surface once in the next morning brief and then be discarded — persisting it as a normal contextual memory would cause it to recur indefinitely.

Two additional columns on the memories table handle this:

  • expires_at — nullable timestamp; a memory past this time is treated as deleted.
  • surface_once — nullable boolean; when true, the memory is queued for the next relevant surface opportunity (typically the daily briefing mode), then deleted after it is surfaced.

The daily-briefing mode’s entry context callable queries for pending briefing-queue items: memories where surface_once = true that have not yet been consumed, or where expires_at is imminent. After the briefing runs, consumed surface_once memories are deleted.

Memory legibility surface (built 2026-07-28): principle 7 promises the user can see, edit, and delete anything Quorra knows about them; until now nothing rendered the memories table. Settings → Memory (quorra-web, /settings/memory) lists every row the owner holds, grouped by scope, with create / edit / archive / restore / permanent-delete. Three things make it a legibility surface rather than a table viewer: provenance is rendered in plain language (the six source values become “Quorra inferred this”, “From your notepad”, …), core vs contextual is stated as what it actually means (in every prompt vs looked up by topic), and a tag-less contextual memory is flagged — tag matching is the only content path into contextual recall, so an untagged row is stored but unreachable.

It reads across workspaces. The seal is a cognition boundary — it stops the model seeing across contexts — not an audit boundary against the owner, who owns all of these workspaces. list_memories_for_owner implements that separately from _scope_filter, which stays untouched on the cognition path; the read is owner-scoped, guest-unreachable, and never feeds a prompt. See decisions.md.

Soft delete (2026-07-28): deleting a memory sets memories.archived_at — invisible to every read at once, hard-purged by a nightly worker (memory/purge.py) after memory_archive_retention_days (30). supersede_and_create archives too, which is the change that mattered most: it matches by naive substring and the 14B is known to over-supersede, so it had been silent unrecoverable loss. The LLM tool path can only archive; permanent deletion is reachable only by the owner through the UI (principle 4, expressed in the data layer).

Memory reconciliation (overnight, M5): Deduplication, contradiction detection, stale cleanup, importance promotion (contextual → core for frequently-referenced facts), and synthesis of observations from daily patterns. Passive behavioral pattern synthesis — detecting “Kurt frequently attends his nephew’s baseball games” across many sessions — is handled here. A distinction applies: patterns from stated behavior (explicit mentions across conversations) can be synthesized from session history alone. Patterns from observed behavior (calendar attendance, location data, activity streams) depend on M4+ service integrations and are not possible before those exist. Synthesized patterns are promoted into the memories table as core importance facts.

Context-triggered reminders:

Context-triggered reminders are a distinct primitive from time-based reminders. A time-based reminder fires at a scheduled moment; a context-triggered reminder fires when a specific situation is recognized in conversation.

Example: “Next time I’m getting gas, remind me to use premium in the Audi.” There is no schedule — the trigger condition is semantic, not temporal.

These are stored in a separate context_triggers table (not the memories table):

FieldTypeDescription
idinteger PK
user_uuidtextOwner
descriptiontextWhat to surface (“use premium fuel in the Audi”)
trigger_hinttextNatural language trigger condition (“when Kurt mentions getting gas or stopping for fuel”)
created_attimestamp
session_idtextSession in which the trigger was created

At the start of each planning-scope session turn, the LLM is given the list of active context triggers for the user. When a turn’s content matches a trigger — recognized semantically by the LLM, not by string matching — Quorra surfaces the reminder and marks the trigger consumed (deleted) or recurring per user preference. This belongs to the planning scope’s tool surface.

4.4 Action layer (permission tiers)

Every tool registered with Quorra declares a tier. Tier cannot be promoted at runtime.

TierExamplesConfirmation
0 — Read-onlyQuery files, search photos, read calendar, check system healthNone
1 — Low-risk writesCreate reminder, add calendar event, request media download, restart non-critical serviceNotification after
2 — Destructive or externalDelete files, send messages, make purchases, modify system configExplicit user confirmation before

For Tier 2 actions: draft the action and ask for approval. Never execute silently.

Tool naming convention: Tools use underscore-prefixed namespaces matching their integration (finance_log_transaction, calendar_query_events, system_get_health). Cross-cutting tools have no prefix (search_knowledge_base, get_my_context). The OpenAI tool-calling format only allows [a-zA-Z0-9_-] in tool names — slashes are not valid.

4.5 Inference layer

Local LLM serving for the Quorra agent.

  • Framework: llama.cpp (native — replacing Ollama)
  • Hardware: RTX 5080 (16GB) + RTX 3070 (8GB) — both passed through to Ubuntu VM, each running an independent llama.cpp server process (CUDA_VISIBLE_DEVICES pinned)
  • Target API: OpenAI-compatible local endpoint per GPU (allows swapping models without changing QuorraAPI)
  • Dual-GPU strategy: Separate models per GPU — enables true parallel inference between Quorra agent and diagnostic agent, lower perceived latency per model, and tractable LoRA fine-tuning on smaller models

GPU assignments:

GPUVRAMRoleModel
RTX 508016GB GDDR7Quorra agent (user-facing)Qwen 3 14B Q4_K_M — confirmed (10.7 GB VRAM, ctx 8192, --parallel 1; 78.8 tok/s P50; benchmarked 2026-05-21)
RTX 30708GB GDDR6Diagnostic agent (M6)Qwen 3 8B Q4 (~4–5GB) — validated but not yet deployed; reserved for M6

LoRA fine-tuning: QLoRA on Qwen 3 14B via Unsloth typically requires 8–12GB VRAM at Q4 quantization — tractable overnight on the RTX 5080. Fine-tuning is scoped to smaller models only (not a 32B split model).

Embedding model: nomic-embed-text or similar (~137M params, <1GB) — can run on either GPU without meaningful VRAM impact.

4.6 The knowledge base (~/data/knowledge/)

The primary fact layer. Quorra is the primary author; users have direct edit access.

~/data/knowledge/
├── household/           shared household resources
│   ├── calendar/
│   ├── finances/
│   ├── home/
│   ├── pets/
│   └── vehicles/
├── members/
│   ├── kurt/
│   │   ├── knowledge/   personal vault (675+ markdown files, indexed by RAG)
│   │   └── shared/
│   └── example/         template for new members
├── guests/
└── templates/

The vault holds reference knowledge only. Agent cognition (learned preferences, observations) and per-user settings live in the quorra-api DB (memories, user_preferences) — not the vault. The pre-DB agent-memory/ and household-agent/ directories were retired 2026-06-15; the access model they documented moved to authorization.md. See decisions.md.

User access: files are readable and editable directly via Nextcloud (OnlyOffice for rich editing). Direct edits are detected by the inotify watcher, logged, and audited during overnight consolidation. The long-term UX goal is all edits flowing through Quorra — view the file in OnlyOffice, ask Quorra to make the change — so the write schema is enforced at authoring time rather than corrected overnight. Direct edit access is preserved for small changes that don’t warrant an inference call.

Design constraint: the RAG corpus should be predominantly prose. Embedding models produce low-discrimination vectors for terse, structured content — checkbox task lines, bullet fragments, key-value pairs — leading to noisy retrieval. Embeddings also cannot capture state (checked vs. unchecked, active vs. resolved), so structured task data belongs in the database, not the vault. Inline structured content dilutes the surrounding prose chunks it sits within, degrading retrieval quality for the prose too. When structured data exists in the vault for human readability (e.g., a markdown file summarizing preferences), the authoritative copy for Quorra’s actions is the DB row — the vault copy is a reference artifact. This constraint shapes what the M2 embedding pipeline should index, what it should skip, and how chunk boundaries are drawn.

4.7 Embedding update strategy

Embedding updates are tiered by latency urgency — heavier processing is deferred to overnight where possible.

TriggerWhatWhen
Quorra writes a knowledge fileRe-embed the affected file immediatelySynchronous — Quorra authored it
User edits a file directlyLog the change, re-embed the fileNear-real-time: inotify watcher on knowledge/ detects changes, batches every ~5 minutes; change logged for overnight schema audit
User hands Quorra a document to ingestAsk: “Learn this now or later tonight?”User-directed; respects hardware constraints honestly
Full consistency passRe-index everything not yet embedded or with stale embeddingsOvernight only

The “learn now or later?” prompt is a deliberate UX choice — it sets honest expectations about the cost of immediate indexing rather than hiding latency behind a spinner.

4.8 Overnight consolidation job

Scheduled window: 2–4 AM via systemd timer on JUNC1.

Tasks (in order):

  1. Full consistency pass — re-index any files not yet embedded or with stale embeddings
  2. Direct-edit audit — review the inotify change log for files edited directly by users; check each against the vault schema; normalize structural drift; re-embed affected files
  3. Refresh stale embeddings in the vector store
  4. Generate episodic summaries of significant interactions from the day
  5. Update structured fact files (relationships, preferences inferred from recent events)
  6. Memory reconciliation — runs against the memories table:
    • Deduplicate redundant memories (same fact saved in different sessions)
    • Detect and resolve contradictions (e.g., two different cars listed as “Kurt’s car”)
    • Remove stale memories contradicted by newer ones that the LLM missed at runtime
    • Promote importance: contextual → core for facts referenced repeatedly
    • Synthesize observations from daily interaction patterns
    • Clean up orphan memories (no tags, empty content)
  7. Generate LoRA fine-tune training candidates from the recent interaction window
  8. Run LoRA fine-tune on RTX 5080 if candidates exceed threshold (Qwen 3 14B, QLoRA via Unsloth)
  9. Write consolidation report to logs

Constraints:

  • Must not run during active sessions
  • Must be independently restartable (each task idempotent)
  • Produces a report readable by the user

4.9 Observability

Two views over the same metric pipeline — a live operational snapshot and a historical comparison view.

QuorraAPI                llama-server (systemd, host)
  └─ /metrics                └─ /metrics  (--metrics flag)
        │                         │
        └──────────┬──────────────┘
                   ▼
              Prometheus  (admin compose stack, 15s scrape, 90d retention)
                   │
       ┌───────────┼────────────────────────┐
       ▼                                    ▼
  quorra-dashboards                      Grafana
  (live, 5s refresh,                     (historical + week-over-week,
  in-memory app counters)                annotated with deploys + model switches)

Metric sources:

  • QuorraAPI — application-level metrics (LLM call duration / tokens, chat-loop iterations, tool calls, HTTP latency from prometheus-fastapi-instrumentator) defined in repos/quorra-api/src/quorra_api/metrics.py. Exposed at quorra-api:8000/metrics.
  • llama-server — inference internals (llamacpp:tokens_predicted_total, prompt_tokens_total, kv_cache_usage_ratio, requests_processing, …) when started with --metrics. Exposed at 172.17.0.1:11435/metrics (docker0). Prometheus reaches it via host.docker.internal (extra_hosts: host-gateway).

Surfaces:

  • quorra-dashboards (http://192.168.0.49:8081/) — single static HTML page, vanilla JS, parses QuorraAPI’s /metrics directly every 5s. Optimized for “what’s it doing right now”; in-memory counters reset on quorra-api restart.
  • Grafana Quorra Inference History (http://192.168.0.49:3100/d/quorra-inference-history) — historical view, default range 7d. Each panel renders the live series alongside the same query with offset $comparison_offset so week-over-week regressions and improvements are visible at a glance. Survives container restarts (Prometheus TSDB on host bind mount).

Annotations for change correlation:

  • QuorraAPI deploys — quorra_build_info{git_commit, git_branch, build_time} gauge set to 1 at startup with deploy-identity labels (populated by ~/projects/junc1/compose/quorra/scripts/deploy.sh). Grafana queries changes(quorra_build_info[1m]) > 0 and marks the timeline with the new commit on every redeploy.
  • Model switches — scripts/switch-model.sh POSTs a tagged annotation to Grafana’s /api/annotations after the new model passes its health check, using a service-account token at /etc/quorra/grafana-annotations.token. Missing token → silent no-op, the switch itself never fails on annotation push.

4.9a LLM turn inspector (inspect/)

Prometheus answers how fast and how often. The inspector answers what was actually sent — the assembled system prompt, the tool schemas, the history window, the retrieved chunks, and the model’s reasoning. Those are the inputs that actually determine a reply’s quality, and without them a wrong answer can only be debugged by re-reading source.

_call_llm / _call_llm_stream ──┐
notepad.triage                 │
reflection                     ├─► inspect.recorder (in-memory ring, N=50)
workspaces.init                │        │
                               │        ├─► GET /inspect/turns        (owner + developer_mode)
concierge  ── EXCLUDED ────────┘        ├─► DELETE /inspect/turns     (principle 7)
                                        └─► POST /bugs (opt-in attach) ──► bugs/*.md + *-payload.json

Storage is memory-only, by design. A collections.deque(maxlen=…) of TurnRecords, appended at turn start and mutated in place (so an in-flight or crashed turn is still inspectable). There is no table and no migration. The assembled prompt inlines core memories, budget balances, today’s CalDAV tasks, and the workspace file tree — persisting it would create a durable shadow copy outside the memory-deletion surfaces that principle 7 guarantees. Nothing reaches disk unless the owner explicitly attaches a turn to a bug report.

maxlen bounds record count, not bytes, so the recorder also enforces per-message and per-tool-result truncation plus a per-record byte budget, and exposes truncated so the UI never implies it is showing more than it kept.

Capture is explicit, never ambient. Each call site passes a TurnRecord and calls begin_call / finish_call. An httpx event hook on the shared client would have been fewer lines, but it would capture Radicale, Qdrant, Actual, the embedding server, and the guest concierge, and would then need a maintained deny-list to exclude the egress boundary — policy where principle 2 requires architecture. The concierge path has no recorder call, and an import-assertion test in the guest egress suite is the structural guard on that.

The payload is serialized eagerly, before the request goes out: the chat loop appends to its messages list across tool iterations, and build_chat_payload appends the configured prefill to the payload’s own list — so the payload and the loop’s messages are not the same object and diverge whenever a prefill is set. Capturing before the POST also means a failed call still has its payload, which is the case worth having.

Access is gated three ways: normal auth, per-record owner scoping (mismatch → 404, no existence leak), and developer_mode — the first server-side enforcement of that flag, which until now gated client UX only. A ReservationPrincipal is not an AuthentikUser, so /inspect is unreachable from the guest path by construction.

Reasoning (delta.reasoning_content, or an inline <think> span depending on the llama.cpp build) travels a channel of its own: a separate SSE frame type, never mixed into reply text, never persisted, and never fed back into history. StreamCleaner’s guarantee that emitted text is byte-identical to clean_reply under any chunking is unchanged — reasoning capture is a sink attached to the existing discard path, not a relaxation of it. Live display is a normal user preference (show_thinking, default on); only the inspector is developer-gated.


5. Data flows

5.1 User query → response

user message (Matrix DM or PWA)
  → client adapter → POST /chat
  → identity resolution (Authentik UUID)
  → session history load (last 5–8 messages, SQLite)
  → core memories loaded (user + household, from memories table)
  → LLM call (system prompt [with core memories] + history + tool list)
  → tool call loop:
      LLM calls memory_recall if topic-specific context needed (Tier 0)
      LLM calls memory_save if user shared a memorable fact (Tier 1)
      LLM calls search_knowledge_base if document search needed (Tier 0)
      LLM calls domain tools as appropriate (finance, calendar, etc.)
  → response → client
  → history persisted (SQLite)

5.2 Photo ingestion (existing pipeline — Immich)

photo uploaded to Immich
  → CSAM hash check (required on all ingestion)
  → Immich CUDA ML: face clustering, object detection
  → Quorra integration (future): semantic embedding → vector store
  → knowledge graph update: face cluster → household member

5.3 Overnight consolidation

systemd timer (2 AM)
  → check active sessions → abort if any
  → embed new documents/photos added since last run
  → distill episodic summaries from day's interactions
  → update knowledge graph
  → memory reconciliation (deduplicate, resolve contradictions, promote, synthesize)
  → generate LoRA training candidates
  → [optional] run LoRA fine-tune
  → write report to the quorra-api data volume (/data/, alongside the audit log; final path an M5 decision)

6. Interface contracts (stubs)

6.1 QuorraAPI

  • Transport: HTTP (exact: TBD — HTTP/WebSocket for streaming)
  • Primary external interface: POST /chat (custom JSON) — all clients connect here
  • Auth: two paths, both resolve to Authentik UUID:
    • Authentik OIDC JWT (Bearer token) — for direct API callers with OIDC tokens
    • Trusted-service auth (Bearer service-secret + X-Quorra-User-Email header) — for internal adapters (Open WebUI Pipe, Matrix bot) that identify users on behalf of the chat interface; email → UUID resolved via config-based user map
  • Format: JSON (QuorraAPI’s native format, not OpenAI-compatible — adapters handle per-client translation)
  • The /v1/chat/completions OpenAI-compat endpoint was considered and deliberately skipped; revisit only if a third-party client that exclusively speaks OpenAI format emerges

6.2 Quorra ↔ Diagnostic agent IPC

  • Internal only; no external exposure
  • Format: TBD (Unix socket JSON messages or lightweight message queue)
  • Quorra can request: health summary, specific diagnostic check results
  • Diagnostic agent can push: critical alerts, threshold breaches

6.3 External service integrations

ServiceInterfaceDirection
NextcloudLocal WebDAV / REST APIQuorra reads/writes
ImmichLocal REST APIQuorra reads; ingest events via webhook
JellyfinLocal REST APIQuorra reads (library, watch history)
Home AssistantLocal REST API + WebSocketQuorra reads/writes (device states, automations)
Actual Budget@actual-app/api or RESTQuorra reads/writes (transaction logging)
MatrixClient-Server APIQuorra sends notifications; bot adapter handles inbound
GiteaREST APIQuorra reads (repos, issues)

6.4 Capability interface principle

Quorra depends on capabilities (“log a transaction,” “search photos,” “get open tasks”), never on implementations (Actual Budget, Immich, a specific task app). Each integration domain exposes a lean interface shaped by what Quorra actually needs to do — not a mirror of the backend’s full API surface. Swapping a backend means writing a new adapter behind the same capability interface, not re-architecting the tool surface or prompt fragments.

The existing tool naming convention already supports this: finance_log_transaction is the capability; the actual-bridge sidecar is the current adapter behind it. Today’s service choices are provisional — they prove the concept now, but may be replaced later with native implementations or different third-party services. Routing through thin adapters behind small capability interfaces makes “swap the backend” mean “write a new adapter,” not “re-architect Quorra.”

Keep interfaces lean, not baroque. Cover only the operations Quorra actually uses. Avoid both extremes: too implementation-specific (leaks backend details, defeats the purpose) and too elaborate (over-engineering for swaps that may never happen). Draw the interface at the level of what Quorra needs to do — nothing more.

Installable integrations = MCP apps (target, decided 2026-07-26). The wire format for installable third-party integrations is MCP: an integration is an app packaged as an MCP server exposing tools/resources/prompts; quorra-api is the sole MCP host and the only inference loop (“apps expose capabilities; Quorra thinks”). A thin Quorra manifest (declared tools, requested tiers — platform-clamped, external send ≥ Tier 2, fixed at install; emitted events) plus a workspace ServiceBinding row forms the install/opt-in grant, registered through the integration registry’s plugin seam; a single authenticated, content-free event-poke endpoint lets an app announce inbound work, which Quorra then pulls through the app’s own tools. Apps shipping their own cognition are a documented non-goal of this contract. First planned instance: the concierge’s PMS/channel layer (§4.1d). See decisions.md “Concierge extraction target — an installable MCP app”.


7. Multi-user model

Every data structure must accommodate multiple users from the start.

  • Household-level state: facts shared across all users (household/ in the knowledge base)
  • User-level state: per-user agent memory, episodic history, private notes, preferences
  • Identity: Authentik UUID is the canonical user identifier — consistent across Matrix, Nextcloud, Jellyfin, and any future interface
  • Memory namespacing: vector store and knowledge graph entries tagged with user scope (household or per-user)
  • Quorra agent context: knows which user is speaking and scopes retrieval accordingly
  • Workspaces: household and each member are the first instances of a first-class Workspace (§4.1c) — a data-partition + egress-policy boundary. A rental (or side business, project, delegated brief) is another. Sessions resolve into a workspace; the workspace holds its bindings, personas, and policies.
  • Principal kinds: the actor is a Principal protocol, not only an Authentik user — AuthenticatedPrincipal (household members), ReservationPrincipal (a guest, with no owner identity, expiring at checkout), and a documented AnonymousPrincipal seam. Authentik UUID remains canonical for authenticated users; non-authenticated principals never carry one.

8. Privacy and security model

  • No data leaves JUNC1 without explicit per-request opt-in
  • Cloud dependencies enumerated and opt-out-able individually (initially: relay for remote access, OTA updates)
  • Agent action audit log: all Tier 1+ actions logged to an append-only file at /data/audit.log (config audit_log_path) and dual-written to the SQLite audit_log table
  • Memory inspection: users can view and delete any stored fact about themselves via the /memory API
  • CSAM detection: hash-based scan on all photo ingestion using established libraries
  • Encryption at rest: TBD
  • Guest / external-party egress boundary (§4.1d): when Quorra acts toward an untrusted external party (e.g. a rental guest), the data boundary is enforced structurally, never by prompt. Three controls — a subtractive persona tool allow-list, a retrieval binding hardwired to a guest-only collection, and a default-deny guest_visible content gate — plus a principal that carries no owner identity, so owner data is unreachable by construction even under prompt injection. The external-facing agent runs draft-and-approve (the owner reviews every outbound message) until autonomy is explicitly granted per category.

9. The Grid (long-term vision)

The Grid is a federated network of personally-owned junc nodes. Each piece of junc is independently functional — self-contained, locally intelligent, privately owned — but capable of communicating with others directly, with no platform in the middle.

Design for it; don’t build it yet. Architectural choices that affect Grid readiness:

  • Use open standards for data formats (CalDAV, CardDAV, ActivityPub-compatible schemas) where practical
  • Authentik identity maps cleanly to a future portable, self-sovereign identity layer
  • Per-user memory namespacing is compatible with federated access (explicit trust grants per sharing request)

Current focus: make a single piece of junc work exceptionally well. The Grid becomes possible once the single-node product is validated.


10. Open questions

Track unresolved design decisions here. When resolved, move the answer into the relevant section and note the worklog date.

QuestionStatusCandidates / Notes
LLM modelsDecidedRTX 5080: Qwen 3 14B Q4_K_M — P50 TTFT 44.5ms, 10,719 MB VRAM, 5,122 MB free. RTX 3070: Qwen 3 8B Q4 — P50 TTFT 50.5ms, 5,717 MB VRAM, 2,124 MB free. Both live as systemd services.
Dual-GPU inference strategyDecidedSeparate llama.cpp processes per GPU (see section 4.5); tensor split approach set aside
Vector databaseDecidedQdrant — better standalone performance; ChromaDB set aside
Agent framework architectureDecidedCustom tool-calling in Python; permission tiers declared at registration, not enforced by framework
Implementation languageDecidedPython 3.12; FastAPI for QuorraAPI; uv for dependency management
Tier 2 confirmation UXDecidedLLM generates a human-readable description of the pending action; client re-submits with confirmation_id. Natural-language confirmation (“yes”, “go ahead”) works. Implemented in M3.
Context injection: system vs. user messageDecidedRAG and episodic memory exposed as Tier 0 tools the LLM calls selectively — not pre-assembled. Only conversation history always injected. Eliminates context waste on pure tool calls.
Quorra ↔ diagnostic IPCOpenUnix socket JSON vs. lightweight message queue; deferred to M6
Diagnostic agent modelOpenQwen 3 8B Q4 (small quantized LLM) vs. rule-based without LLM; deferred to M6
Encryption at restOpenProxmox/LUKS level vs. ZFS native encryption vs. application-level
LoRA training pipelineOpenScoped to Qwen 3 14B on RTX 5080; tooling: Unsloth (QLoRA); frequency and candidate selection TBD
Knowledge graph technologyDeferredTier 2 memory will use structured markdown files initially. Graph layer (Neo4j, SQLite + custom, or DuckDB) deferred to a later milestone. Design context assembly to abstract “retrieve facts” from storage format so graph can be added without rewriting the pipeline