Quorra — Architecture
Last updated: 2026-07-26 (concierge-extraction design — MCP-app target in §4.1d/§6.4, personas→roles naming note in §4.1c; earlier same day: workspace creation — registry-driven wizard + integration_optin opt-in table + POST /workspaces//preview-notes + GET /integrations//budgets; finance offered bind-only, in §4.1c)
Status: M3 complete; M4 in progress — QuorraAPI core, finance scope, modes, and observability done; M2 (RAG) shipped (in-process rag/, per-owner Qdrant collections); planning re-platforming onto Radicale CalDAV; workspaces/personas/guest-concierge designed 2026-07-09 (phased build — see plans/workspaces-personas-concierge.md), not yet implemented; remaining service integrations (files, Jellyfin, email, PMS) pending
1. System boundaries
What Quorra owns
- All AI and ML logic (inference, embeddings, RAG, memory management)
- The Quorra agent runtime and the diagnostic agent runtime
- Quorra API — the HTTP service layer connecting clients to Quorra’s intelligence
- User-facing interfaces (API, PWA client)
- Orchestration of local service integrations (storage, photos, calendar, media, home automation)
- The overnight consolidation job
- Audit logging of agent actions
What Quorra integrates with (but does not own)
The following services are already running on JUNC1. Quorra calls their local APIs; it does not manage or replace them.
| Service | Purpose | Local address |
|---|---|---|
| Nextcloud | File storage, calendar, contacts | cloud.juncyard.com |
| Immich | Photo backup and ML (face recognition, object detection) | photos.juncyard.com |
| Jellyfin | Media streaming | watch.juncyard.com |
| Home Assistant | Home automation | :8123 |
| Actual Budget | Personal finance | budget.juncyard.com |
| Gitea | Code hosting | git.juncyard.com |
| Matrix/Synapse | Messaging — primary user-facing interface for prototype | matrix.juncyard.com |
| Authentik | SSO/identity — OIDC provider; Authentik UUID is canonical user ID | auth.juncyard.com |
What Quorra explicitly does not do
- Manage JUNC1 OS, Proxmox, or Docker infrastructure
- Replace local services — it wraps and queries them via their local APIs
- Store user data that a dedicated service already manages (files in Nextcloud, photos in Immich, etc.)
2. Physical and virtual architecture
JUNC1 runs Proxmox as the hypervisor. The junc stack lives entirely inside Ubuntu Server VM 100.
┌──────────────────────── Physical machine ─────────────────────────┐
│ CPU: Ryzen 7 3700X | RAM: 32GB DDR4 | PSU: 1000W │
│ Storage: 2×1TB NVMe (1 Ubuntu VM / 1 Windows VM), 2×4TB HDD, 1×8TB HDD │
│ GPU 1: RTX 5080 (16GB GDDR7) | GPU 2: RTX 3070 (8GB GDDR6) │
│ │
│ ┌──────────────────── Proxmox (192.168.0.100) ─────────────────┐ │
│ │ │ │
│ │ ┌─────── Ubuntu VM 100 (192.168.0.49) ──────────────────┐ │ │
│ │ │ Both GPUs passed through to this VM │ │ │
│ │ │ All Docker Compose stacks run here │ │ │
│ │ │ Domain: juncyard.com │ │ │
│ │ └───────────────────────────────────────────────────────┘ │ │
│ │ │ │
│ │ ┌─────── Windows VM 101 ────────────────────────────────┐ │ │
│ │ │ Gaming only — unrelated to this project │ │ │
│ │ └───────────────────────────────────────────────────────┘ │ │
│ └────────────────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────────┘
3. High-level system diagram
┌──────────────── Ubuntu VM 100 (junc stack) ─────────────────────────┐
│ │
│ ┌──────────────────────────────────────────────────────────────┐ │
│ │ Infrastructure layer │ │
│ │ Nextcloud · Immich · Jellyfin · Home Assistant │ │
│ │ Actual Budget · Gitea │ │
│ │ Authentik (SSO) · Matrix/Synapse · Element Web │ │
│ └──────────────────────────┬───────────────────────────────────┘ │
│ │ local API calls │
│ ┌──────────────────────────▼───────────────────────────────────┐ │
│ │ QuorraAPI │ │
│ │ /chat /memory /context /events /health │ │
│ │ Tool system with permission tiers │ │
│ │ Multi-user session management (Authentik identity) │ │
│ └────┬──────────┬──────────┬──────────┬────────────────────────┘ │
│ │ │ │ │ │
│ ┌────▼───┐ ┌───▼────┐ ┌───▼────┐ ┌──▼─────────────────────────┐ │
│ │Memory │ │History │ │ RAG │ │ Inference layer │ │
│ │layer │ │SQLite │ │service │ │ llama.cpp · RTX5080+3070 │ │
│ │SQLite+ │ │ │ │ │ │ OpenAI-compatible API │ │
│ │files │ │ │ │ │ │ │ │
│ └────────┘ └────────┘ └────────┘ └────────────────────────────┘ │
│ │
│ ┌──────────────────────────────────────────────────────────────┐ │
│ │ Diagnostic agent (separate process) │ │
│ │ Smaller model · deterministic playbooks │ │
│ └──────────────────────────────────────────────────────────────┘ │
│ │
│ ┌──────────────────────────────────────────────────────────────┐ │
│ │ ~/data/knowledge/ │ │
│ │ household/ · members/kurt/ · guests/ │ │
│ │ Live markdown vault — RAG service indexes this │ │
│ └──────────────────────────────────────────────────────────────┘ │
└──────────────────────────────┬────────────────────────────────────────┘
│ relay only (remote access)
┌──────────▼──────────┐
│ Clients │
│ Element Web │
│ Web client (M7) │
└─────────────────────┘
4. Core components
4.1 Quorra agent (QuorraAPI)
The user-facing conversational layer, exposed as an HTTP service.
- Receives messages from clients (Matrix bot, future PWA, direct HTTP)
- Assembles context from the memory layer and RAG retrieval
- Calls tools via the permission-tiered action layer
- Manages per-user session history and identity (Authentik UUID)
- Does not directly manipulate storage — goes through local service APIs
API surface (implemented in repos/quorra-api/):
POST /chat— send message, get reply; acceptsoriginfield (standard,matrix,open_webui,voice) to tailor response style per surface; origin is persisted to the session on first messagePOST /chat/stream— SSE variant of/chat. Tool calls run synchronously and emitstatusevents; the final LLM reply streams astokenevents carrying real llama.cpp deltas as they generate, then a terminaldoneevent carriessession_idand anypending_confirmation. A failure mid-generation ends the stream with a terminalerrorevent instead ofdone. Same auth and body as/chat. Separate endpoint because the return type (StreamingResponse) differs fundamentally.- The loop assembles deltas itself (
chat/streaming.py): content fragments pass through an incremental cleaner and out to the client immediately, while tool-call fragments accumulate silently, so a turn is identified as tool-calling-or-final only once the upstream stream ends — without ever buffering the reply. Model reasoning (reasoning_content, and any inline<think>span) is dropped before it reaches a client. - Closing the client connection closes the upstream llama.cpp stream, which aborts the generation and frees the server’s single slot. Whatever text had already been sent is persisted with an in-band interrupted marker, so an interrupted reply survives in session history rather than vanishing. See decisions.md “Real token streaming”.
- The loop assembles deltas itself (
GET|PUT|DELETE /sessions/{session_id}/purpose— manage a per-session “purpose” directive (stored onsessions.purpose, ownership-checked). Injected into the system prompt between the base template and the origin style by_build_session_context(); additive, never replaces Quorra’s core identity. Session state endpoints live under/sessions(not/chat) to separate state management from chat operations.GET|PUT|DELETE /sessions/{session_id}/mode— manage the active mode for a session (see §4.1b Modes below). Validated against the session’s committed scope; returns400for unknown modes,409for pre-commit sessions.GET|POST|PUT|DELETE /memory— CRUD for user-scoped memories (core/contextual, tag-filtered)GET|POST|PUT|DELETE /household/memory— CRUD for household-scoped memories (any authenticated user can edit/delete)POST /events— inbound webhook from local services (stub — dispatching TBD)GET /health— checks LLM primary and actual-bridge reachability
System prompt architecture: the system prompt is assembled additively from fragments. Base identity (src/quorra_api/prompts/base.md) always; one src/quorra_api/prompts/scopes/<scope>.md fragment per active session scope; one src/quorra_api/prompts/modes/<scope>/<mode>.md fragment if a mode is active (see §4.1a); core memories when any exist; the session purpose only when scope is exactly ["general"]; per-origin style directives (src/quorra_api/prompts/origin/*.md) last. All .md files load once at import. Adding a new origin, scope, or mode is a file + registry entry — no code branches. Session history (load_history) injects time-gap markers ([N hours passed], [next day — …]) between consecutive messages more than 4 hours apart so the LLM can treat pre-gap context as stale; a midnight crossing alone does not trigger one. Time gaps also auto-clear the active mode (if any) so Quorra returns to default scope behavior after inactivity.
Chat scope: every session locks to a scope (a list of capability surfaces, e.g. ["general"], ["finance"]) on first commit. Scope is sent by the client as scope: list[str] on /chat//chat/stream, persisted on sessions.scope_json, and immutable for the session’s lifetime — re-scoping means starting a new session. While scope is NULL (pre-commit) the active tool list is empty and the model can chat but cannot fire actions; this is the hook the future LLM-guided setup-conversation feature will plug into. The tool registry exposes for_scope(scope)/get_for_scope(name, scope) so the active tool list per turn is the intersection of the session’s scope and each tool’s declared scopes frozenset — out-of-scope tools are not just hidden from the LLM, they’re structurally unreachable (defense in depth against prompt injection or name-guessing). Known scopes live in src/quorra_api/chat/scopes.py; the minimal rollout ships general and finance. Mismatched scope on an existing session returns 400.
4.1a Modes
Modes are dynamic behavioral overlays within a scope. Where scopes gate which tools are reachable (structural, registry-level, immutable per session), modes guide how Quorra uses those tools (prompt-level, mutable within a session). A mode never expands the tool surface beyond the session’s scope — worst-case “leak” is irrelevant prompt guidance, not tool access.
Why modes exist: On a 14B model with limited prompt attention, loading every workflow’s instructions into every session wastes context budget. Modes load workflow-specific prompt fragments and pre-fetched context only when actively needed, keeping the effective prompt lean and focused. This is a capability multiplier at small model sizes.
Activation: Externally-driven — clients set the mode via PUT /sessions/{id}/mode. The LLM does not decide to enter a mode on its own (LLM-driven activation can be added later without breaking changes). This keeps scope prompts lean — no mode descriptions are injected until a mode is active.
Storage: active_mode: str | None column on Session (mutable, nullable). Set/cleared via the mode endpoint. NULL means no mode active — the session uses only its scope fragment (default behavior).
Auto-expiry: Piggybacks on the existing 4+ hour time-gap detection in the chat loop. When a gap is detected and a mode is active, active_mode is auto-cleared to NULL before building the system prompt. No new timestamps or columns needed. Matches user intuition: “I walked away and came back, start fresh.”
Prompt composition order (with mode):
base identity → scope fragment(s) → MODE FRAGMENT (if active) → core memories → session purpose → response style
Mode fragments live at src/quorra_api/prompts/modes/{scope}/{mode_name}.md, loaded at import like scope fragments. Entry context (if defined for the mode) is injected as an additional section below the mode fragment on the first turn after mode activation.
Mode registry: ModeDefinition dataclass (name, scope, description, optional entry_context async callable) stored in a ModeRegistry keyed by (scope, name). ModeRegistry.for_scope(scope) returns available modes; validate(mode_name, scope) raises on invalid mode/scope combinations.
Planned modes:
| Scope | Mode | Workflow |
|---|---|---|
finance | budget | Plan next month’s budget — proactive spending analysis, category comparison, allocation setting |
finance | reconcile | Rapid-fire transaction review and categorization correction |
finance | review | Analytical deep-dive into spending trends and anomalies |
planning | daily-briefing | Structured morning overview — pending items, reminders, schedule (deferred; see plans/daily-briefing-mode.md) |
general | research | Extended topic exploration with active knowledge-base searching |
general | household-planning | Structured planning for trips, events, projects — checklists and timelines |
media | movie-night | Interactive recommendation narrowing for a group (future — ships with Jellyfin) |
planning | weekly-review | Week-ahead calendar layout with conflict detection (future — ships with Radicale CalDAV) |
planning | scheduling | Find-time-for-X with availability checking (future) |
mail | inbox-review | Guided inbox triage — prioritize, summarize, act (future — ships with Protonmail) |
mail | compose | Context-aware email drafting with Tier 2 send (future) |
photos | memory-lane | Conversational photo browsing by person/date/event (future — ships with Immich) |
4.1b Matrix bot adapter (repos/quorra-matrix/)
A standalone matrix-nio service — the Matrix “mouth and ears” for Quorra, with no LLM or tool logic of its own.
- DMs (2-member rooms) respond to every message; group rooms respond only when the bot is @mentioned (mention stripped before forwarding, sender display name prepended as
[Name]). - Per-message identity: static
MATRIX_USER_MAP(matrix-id→email) + trusted-service auth (X-Quorra-User-Email+ shared secret) — the same auth path as the Open WebUI Pipe. Unmapped senders get a polite refusal. - Room→scope mapping via the
MATRIX_ROOM_SCOPE_MAPenv (JSON{room_id: ["finance", …]}), wrapped inresolve_scope(room_id) -> list[str]so a future swap to a user-defined DB-backed lookup is a one-function change. DMs and unmapped rooms resolve to["general"]. The bot includes the resolved scope on the first/chat/streamrequest of a session; the API commits and locks it from there. - Room→session mapping and pending Tier 2 confirmations persisted in bot-local SQLite, so a bot restart preserves continuity.
- Responses use
POST /chat(non-streaming, single-message reply). A progressive-edit streaming path usingPOST /chat/stream+m.replace(batched ~400 ms) was implemented but reverted 2026-05-22 — edit ghosting in Matrix clients made the experience worse than a single reply. Streaming code is preserved as a dormant fallback for future re-evaluation. !quorra purpose|forget|status|modehandled entirely bot-side via the session state endpoints (/sessions/{id}/purpose,/sessions/{id}/mode); never forwarded to the LLM.!quorra mode budgetactivates a mode;!quorra modewith no argument clears it.
4.1c Workspaces, service bindings & personas
A Workspace is a first-class “brief” Quorra operates within — the generalization of scope along a principal/trust axis. Where scope alone models one user wearing hats, a workspace bundles what Quorra can touch, who it is being and for whom, how much rope it has, and how it carries itself — the way a human executive assistant context-switches between the household, a rental business, and a job search: one assistant, different information boundaries. The motivating case is helping run an STR/MTR rental (including interacting with guests on the owner’s behalf), but the rental is just the first instance.
The four composable axes of a session:
Session ─┬─ workspace_slug → the brief: a set of service bindings + personas + standing policies
├─ scope → capability domains active this session (⊆ the workspace's bound services)
├─ persona → who Quorra is being: tool allow-list (subtractive), retrieval binding, prompt family, principal kind
└─ active_mode → behavioral overlay within a scope (§4.1a — unchanged)
Owner sessions resolve to workspace = <personal/household>, persona = owner (null allow-list) → behavior identical to today. A guest session (§4.1d) resolves to workspace = str, persona = guest-concierge, principal = a reservation.
Service bindings — the uniform capability unit: a workspace’s capabilities are a set of (service, data_window, tier_policy) rows. service names a tool family (≈ a scope): finance, calendar, knowledge, notes, pms, messaging. data_window identifies the slice — which budget, which collection, which PMS account. This is the whole generalization: PMS-on-account-W and finance-on-budget-X are the same kind of row, so nothing is service-specific in the model and the user’s “which services this brief uses” multi-select maps 1:1 to binding rows. A session’s scope is validated to be ⊆ the workspace’s bound services.
Personas — the trust/actor profile (subtractive): a code-level registry entry (like ModeRegistry) carrying tool_allowlist (None = unrestricted), retrieval_binding (None = user-derived; else a fixed knowledge window), prompt_family, and principal_kind. A workspace declares which personas it offers. The allow-list is subtractive — it intersects the scope-reachable tool set — which the additive scope registry does not otherwise express. The owner persona has an all-null profile, so existing sessions are unaffected byte-for-byte; guest-concierge is the restricted one. Personas gate capability (registry layer); modes overlay prompt behavior (prompt layer) — different layers, so modes are untouched. (Naming note, 2026-07-26: “persona” is transitional vocabulary — the target model re-keys this axis as workspace roles resolved from the counterparty’s principal (definitions stay in code; assignments become DB membership rows), landing with multi-user-per-workspace. See decisions.md “Personas dissolve into workspace roles”.)
Principal kinds: the actor generalizes from AuthentikUser to a Principal protocol — AuthenticatedPrincipal (existing), ReservationPrincipal (a guest; a reservation-scoped identity with no owner identity, expiring at checkout + grace), and a documented AnonymousPrincipal seam (front-desk gatekeeper; not built).
Storage: a workspaces table + a workspace_service_bindings table; new nullable workspace_slug and persona columns on Session (backfilled to owner defaults, immutable-after-commit like scope). Personas and the principal protocol are code, not tables (mirroring modes/scopes). New-path-only: the Workspace is the authority for the new (STR/guest) path; the live owner/household retrieval derivation (collections_for_user) is left working as-is — no regression to shipped M2 RAG.
Generic vs. instance: the model (Workspace, ServiceBinding, Persona, Principal, policy) is generic; the str workspace, the guest-concierge persona, this PMS account, and this reservation are instances. Seams the model accommodates but this design does not build: the anonymous front-desk principal, expires_at on a workspace (event/travel briefs), reduced-trust authenticated personas (a kid’s or an elderly parent’s brief), and a sanctioned cross-workspace read-grant (an “accountant” join at tax time) — each a persona/principal/binding variation, not a schema change.
Owner-in-workspace (the owner face; built 2026-07-11). When the owner inhabits a workspace (a session with workspace_slug set, persona = owner), the context is sealed both ways while the identity layer stays universal:
- RAG: the workspace’s
data_diris indexed into its own collectionkb_ws_<slug>and excluded fromkb_member_<slug>; in-workspacesearch_knowledgehardwires tokb_ws_<slug>viasearch_window. So workspace knowledge is reachable only in-workspace, and never pollutes personal chat.collections_for_usernever returnskb_ws_*. - Memory:
memories.workspace_slug(NULL = personal/household). A workspace turn reads onlyworkspace_slug == X; a personal turn excludes workspace rows (workspace_slug IS NULL). Sealed symmetrically, like RAG. - Identity stays universal: Quorra’s self, the user’s name/locale (from
user_preferences, not the memory corpus), her capabilities, and the awareness that she’s in this workspace all persist — only the retrievable data is workspace-local. Aprompts/workspaces/base.mdblock frames the sealed context; in personal context, an awareness block lists the user’s workspaces (name +description) so Quorra can offer to switch without seeing specifics. - Scope derivation: the owner’s in-workspace scope =
owner_workspace_scope(bindings)=(bound ∩ known scopes) − guest-only (hosting) − deferred (finance, until its per-workspace budget binding is wired), at least["general"]. Owner and guest thus see different tool surfaces of the same workspace. - Entry: one Matrix room per workspace (
Workspace.matrix_room_id, 1-1). The bot resolves room → workspace viaGET /workspaces/by-roomand sends the genericworkspace_slugon/chat(the chat contract never speaks Matrix room IDs). General/personal is the default home; entering a workspace is the deliberate, mutually-exclusive switch. Threading is small: the loop forwardssession.workspace_sluginto the prompt builder and the tool executor’s user-scoped arg injection, and the memory/knowledge tools consume it.
Workspace authoring (the write face; built 2026-07-23). Quorra authors markdown documents into her workspace via the owner-only document tool — the workspace is now a living project space, not a read-only shelf. Every write is a validator-gated round-trip: kb.serialize assembles a schema-conformant file (canonical field order, flow-style tags, bare ISO dates, auto-filled quorra provenance) → kb.parse + kb.validate → block on any error-severity issue (write nothing, return the issues so the model self-corrects) → write into the session workspace’s data_dir → rag.indexer.index_file reindexes just that file into kb_ws_<slug> (non-destructive ensure_collection + delete_by_path-before-upsert, since chunk IDs are positional) so it’s searchable the same session. WARNING/INFO issues (relative-date, no-summary-lead) are surfaced but don’t block. create is Tier 1 (write + audit); update is Tier 2 (overwrite → confirmation). Workspace-only for this cut: guarded on the session’s committed workspace_slug, unavailable in personal/general chat. Guest-unreachable by construction — document isn’t in the guest-concierge allow-list, so the persona gate fails it closed. The whole-vault mount is now :rw; the tool’s data_dir containment (kebab_cased filename, no separators) + the persona gate are the real boundary. See plans/vivid-roaming-popcorn.md.
Workspace reorganization (built 2026-07-25). Beyond authoring single files, Quorra restructures the workspace directory via the owner-only organize tool — create_folder, move, rename, promote_to_folder (foo.md → foo/overview.md), delete (dead/empty files only). The link-and-index-preserving engine from the one-time corpus migration is lifted into a pure, workspace-scoped kernel (kb/reorg.py): a single move-map drives every rename, each internal link is re-expressed against its target’s new path (never re-derived from link text) — including the entities: frontmatter cross-refs the migration engine skipped — and a simulate pass asserts zero-new-dangling before anything is written. The RAG index is kept in sync by move-safe helpers (reindex_move/reindex_delete) that clear the old path from both kb_ws_<slug> and its guest window before indexing the new path (positional chunk IDs would otherwise orphan). Like document, organize never writes on call: it validates + simulates then enqueues a reorg kb_suggestion (new target_kind; the ops plan lives in proposed_content), reviewed as a before/after tree; on approval a two-phase journaled apply runs — disk is all-or-nothing (reverse-on-failure), the reindex is best-effort — and it re-validates against current disk so a stale plan fails cleanly instead of half-applying. Tier 1 (the inbox is the confirmation); guest-unreachable (not in the guest-concierge allow-list; the egress suite proves it). The in-workspace prompt block shows Quorra the current file tree so she targets ops at real paths. See plans/quorra-needs-the-ability-delegated-wand.md.
Workspace creation (built 2026-07-26). A personal workspace is now created from the web app via a guided wizard, replacing seed/DB surgery. POST /workspaces (owner-gated) derives a unique kebab slug (collision-suffixed), provisions the vault directory members/<owner>/knowledge/Workspaces/<slug> (a plain mkdir — quorra:vault + setgid + umask 002 inherit ownership/mode; never chown as non-root), writes the Workspace row (kind="personal", persona=["owner"], matrix_room_id NULL — room binding deferred) plus its service bindings, and optionally stores an LLM-drafted, owner-approved starter ## Workspace notes as a workspace-scoped core memory — all before commit, so a provisioning failure rolls back with no orphan row. The offered services are not hardcoded: an integration registry (integrations/registry.py, a frozen-dataclass catalog mirroring ModeRegistry/personas) declares each built-in integration’s metadata (scope, view_key, always_on, owner_creatable, requires_config, an is_configured(settings) probe, and a plugin seam), and a per-instance integration_optin table (seeded idempotently from the configured, non-plugin catalog at startup) records what this junc has enabled. GET /integrations?creatable=true = the wizard’s checkbox feed (enabled ∩ owner-creatable — hosting/pms are owner_creatable=False, never offered). The create path validates requested services against that source of truth and against the owner scope ceiling (on the & KNOWN_SCOPES subset, since knowledge is a binding, not a scope). Finance is offered bind-only: selecting it requires a budget pick (GET /budgets over list_accessible) captured into the binding’s data_window ({"budget": alias}, read by workspace_budget()), but threading that pinned budget through the live finance resolver + lifting the finance-scope deferral is a follow-up — the Finances tab stays present-but-disabled. Starter notes are generated server-side (POST /workspaces/preview-notes, reusing the reflection LLM transport — the llama-server isn’t browser-reachable) and shown editable before Create. See plans/let-s-get-to-work-velvet-quiche.md.
Deferred (next): the finance live-wiring (pinned budget → finance_dispatch, lift the finance deferral, enable the Finances view), plugin/third-party integrations (the registry plugin seam) + a Settings UI for the opt-in table beyond the PATCH /integrations/{service} toggle, workspace-creation follow-ons (Matrix-room auto-provisioning, hosting/guest workspaces, starter-paperwork upload, members & roles), personal-vault authoring (kb_member_<slug> in personal context), guest-content authoring (guest_visible house-guides into the guest window), reflection-driven reorganization proposals (on-demand only for this cut), the PWA’s before/after-tree renderer, and a general inotify watcher for direct human vault edits (this phase reindexes only Quorra’s own writes).
4.1d Guest concierge
The guest concierge is the first externally-facing workspace instance: Quorra fielding guest messages for a rental. It is not bolted-on code — it is the generic chat loop run with persona = guest-concierge and principal = ReservationPrincipal. Swap the persona and principal and the same handler serves any external-facing brief.
Flow:
guest (Airbnb/VRBO/direct) ─msg─▶ PMS ─webhook─▶ /integrations/pms/webhook
│ normalize → upsert reservation
▼
concierge handler (constrained loop as ReservationPrincipal)
│ tools ∈ {search_knowledge(guest window), check_this_booking, escalate_to_owner}
▼
{draft, category, confidence, escalation_reason}
│ autonomous_ok(...) → False (always, for now)
▼
guest_message review row (draft_ready | needs_you)
│
Matrix DM to owner ◀─alert─▶ owner: "send" / "edit: …" / "escalate"
│ approve
▼
PMS send-message API → sent
Egress boundary — three structural controls, never the prompt (a guest may be adversarial; prompt rules don’t hold against injection, so enforcement is at code choke points):
- Capability: the persona
tool_allowlist(exactlysearch_knowledge,check_this_booking,escalate_to_owner) intersected with scope-reachable tools at the registry + chat-loop gates. Memory tools are excluded entirely — a guest could otherwise poison or read cognition. - Retrieval: the persona’s
retrieval_bindinghardwires the knowledge tool to a dedicated guest-window collection (kb_str_guest); the guest has nouser_uuid, so the owner-collection resolver is unreachable. - Content: a default-deny
guest_visiblefrontmatter gate — only opted-in vault content is indexed into the guest window.
Because the ReservationPrincipal carries no owner identity, even a bug in the above cannot surface owner data: there is no owner collection to derive. This is Design principle 2 (privacy-by-architecture) applied to an untrusted counterparty.
Draft-and-approve: every reply is drafted for owner approval; nothing sends autonomously today. The concierge emits {draft, category, confidence, escalation_reason} each turn and calls autonomous_ok(workspace, category, confidence), which returns False for all inputs (empty allow-list). Whitelisting categories later (wifi, checkout) lets those auto-send while everything else still routes to review — one policy function plus an already-captured signal. Human-in-loop is also the injection defence during trust-building: the owner sees any injected draft before it leaves.
Review surface: the owner’s existing Matrix/Quorra chat. The concierge posts the draft (or an escalation) into the owner’s room via the bot; the owner replies send / edit: … / escalate. One guest_message review queue with two states (draft_ready, needs_you), reusing the Notification status-state-machine pattern (§4.1b delivery). No new push channel; high-urgency items fire an immediate Matrix DM, otherwise they queue.
Runtime: in-process (the generic loop with a guest persona + reservation principal). The extraction target is now decided (2026-07-26): not a standalone cognitive service but an installable MCP app — the future concierge app carries only PMS/channel I/O behind MCP tools (get_inbound_messages/get_reservation/get_thread/send_reply, the PMSClient surface), while cognition, the egress controls, and the review queue stay in quorra-api (“apps expose capabilities; Quorra thinks” — see §6.4 and decisions.md “Concierge extraction target”). send_reply is an ordinary Tier-2 tool, so draft-and-approve is the standard autonomy-tier system and the egress boundary never moves. Appification waits until Phase 3 has validated live guest behavior. The PMS is a vendor-abstracted service adapter (§6.3/§6.4; Hostaway or Lodgify) — see service-integrations.md — whose PMSClient methods are written as candidate MCP tools for that lift.
4.1e Web client (repos/quorra-web/) — M7
The first-party household client: a mobile-first responsive SPA that becomes an installable PWA as its final, thin layer. “PWA vs web app” is a false dichotomy — a PWA is a web app plus a manifest and a service worker — so the build order is web app first, installability last (matching the M7 checklist). The DoD device is a phone; the design scales up to desktop, where Open WebUI remains the power-user surface until the web client supersedes it.
Stack: React + Vite + TypeScript, Tailwind CSS v4 with components from the neobrutalism.com registry (Base UI variant — the registry’s Radix variant is a broken port, see decisions.md “Web client styling”), lucide-react icons, TanStack Query (server state), react-router, oidc-client-ts (OIDC). SSE consumption is a small hand-written fetch-streaming reader (native EventSource is GET-only; /chat/stream is POST). Static build output — no SSR (authenticated app, no SEO surface).
Identity & auth: the first client to exercise QuorraAPI’s existing Authentik JWT path (auth/middleware.py:get_current_user — JWKS, RS256, audience/issuer checks). quorra-web registers as a public OIDC client in Authentik; the SPA runs the authorization-code + PKCE redirect flow (never popups — redirect survives installed/standalone PWA mode unchanged) and sends the access token as Bearer on every API call. Refresh-token rotation + offline_access give persistent phone login. No new auth code in QuorraAPI; authentik_audience must be configured to accept the new client. participants = [self] for MVP; multi-user sessions are a later UX problem, not an API gap.
Serving (same-origin, no CORS): the container’s nginx serves the static build and proxies /api/* → quorra-api:8000 (SSE: proxy_buffering off on the stream route); host nginx terminates TLS at app.juncyard.com. QuorraAPI deliberately has no CORS middleware — same-origin keeps it that way (no preflights on the SSE POST, no token-bearing cross-origin surface). Deployed as a compose service alongside the other adapters (~/projects/junc1/compose/quorra).
Feature phases:
- A — shell + auth: scaffold, OIDC login round-trip, authenticated
/healthcall, compose + TLS serving pipeline live. - B — chat MVP: new-session flow with scope picker (the session commits scope on first message — the same flow OWUI hardwires to
["general"]),/chat/streamrendering (statusevents as tool-activity chips,tokenaccumulation, markdown rendering), session sidebar + history. Needs new QuorraAPI surface:GET /sessions(user-scoped list; the/sessionsnamespace was shaped for this) and a full-history variant of/sessions/{id}/messages/recent. - C — structured UI: Tier 2 confirmation cards (render
pending_confirmation; approve = next/chatcall carryingconfirmation_id, decline = let it expire — an explicit cancel endpoint is an open item), notifications outbox surface, purpose/mode controls. - D — PWA layer: manifest, icons, minimal app-shell service worker (
vite-plugin-pwa), install tested on iOS Safari + Android Chrome.
Deliberately deferred: offline data/sync (meaningless while inference lives on JUNC1 and the client is LAN-only until M8 — the service worker precaches the app shell, nothing more); Web Push (backlog — transits vendor push services, a local-first tension needing its own decision; note iOS delivers Web Push only to installed PWAs, one reason the install layer exists even while thin); inline photo results (blocked on the M4 Immich integration — Phase B builds a generic rich tool-result rendering seam that photo results plug into); a dedicated web origin prompt fragment (MVP sends origin: "standard"; add the origin + its regression tests when style tuning warrants).
4.1e Notepad
The daily-workflow emulation of a paper notepad: frictionless capture all day, one batch of judgment overnight, one review in the morning. Lives in the personal (null-workspace) context — deliberately not a workspace, because workspaces are sealed data boundaries and the notepad is an unsealed intake funnel whose job is routing outward (calendar, tasks, finance, memory, workspace KBs).
Capture (POST /notepad/items, quorra-web Notepad page): zero-inference — the raw jot text plus a timestamp lands in notepad_items, instantly acknowledged. The web client writes every jot to a localStorage queue first and syncs opportunistically (reconnect/focus/interval); a client-generated client_id with a unique (user_uuid, client_id) index makes blind replay idempotent. No service worker in v1 — the queue survives reload and capture only happens with the page open.
Triage (notepad/triage.py, nightly worker + POST /notepad/triage-now): one LLM pass over the day’s captured items (14B, plain chat-completion, prompt-demanded JSON, tolerant parse — the reflection/reflect.py idiom). Three-way outcome per item, flag-don’t-guess:
- propose → one or more
proposed_actionsrows (an item can yield several — “dinner w/ Sam Fri, I owe him $20” → event + transaction); - ask → an
ask-kind queue row carrying one clarifying question; - no_action → recorded on the item itself (
triage_note), not the queue.
Full accounting: items the model fails to account for stay captured and roll into the next batch — the item’s state transition is the watermark, so re-running is harmless. The prompt injects the newest notepad-correction memories (“Sam always means Sam Reilly”) and the live workspace slugs for routing.
Review (Notepad page, /notepad/actions/*): per-kind cards with structured field editing (no JSON blobs). Approve applies through the existing direct apply functions — RadicaleStore (event/task/reminder), budget resolve + _log_transaction (transaction), create_memory (memory), and enqueue_suggestion (doc — a handoff to the KB inbox, which owns validation and final review). The review click is the Tier-2 confirmation; on apply failure the row reverts to draft_ready with the error surfaced (the kb_suggestions semantics). Answering an ask runs an instant single-item re-triage with the Q/A appended — follow-ups appear in the same review session. Defer rolls a jot to tomorrow’s batch. Any resolve verb can carry a remember note → a notepad-correction memory, closing the learning loop.
proposed_actions is deliberately producer-agnostic (source_kind/source_id, typed kind + validated JSON payload) — it is the intended generalized approvals queue that kb_suggestions and the concierge review may migrate onto later (see decisions.md “Notepad & the approvals queue”).
Config: QUORRA_NOTEPAD_ENABLED (nightly worker only — endpoints are always on), QUORRA_NOTEPAD_TRIAGE_TIME (default 03:30, household_tz).
4.2 Diagnostic agent
A separate process from the Quorra agent — this separation is deliberate and must be maintained.
- Runs a smaller, more deterministic model (exact model TBD — see open questions)
- Structured playbooks: SMART data, thermal trends, GPU health, disk space, service health, network
- Communicates with Quorra agent over an internal IPC interface (format TBD — see open questions)
- Self-healing for common issues: disk full → suggest cleanup, service crash → restart, DB corruption → restore snapshot
- Escalates to Quorra agent only when a human-readable explanation or nuanced judgment is needed
- Must be independently restartable without affecting Quorra agent
Rationale for separation: hallucination is bad in conversation; it is catastrophic in diagnostics.
4.3 Memory layer
A recurring distinction governs where data lives:
- Operational data — structured, stateful, acted upon by the system. Memories, tasks, reminders, budget state. Lives in the database (SQLite
memoriestable today; additional tables as domains grow). The system reads it to decide what to do. - Reference data — natural-language, knowledge-oriented, referred to by the system. Notes, documents, household knowledge. Lives in the markdown vault (
~/data/knowledge/) and is retrieved via RAG. The system reads it to inform answers.
The test: does the system act on this data, or refer to it? Act on it → DB. Refer to it → vault. This already holds in practice (memories live in SQLite, not markdown) and should hold as new domains are added. Don’t conflate the two just because they’re both “about the user.”
Five components, in ascending order of write cost:
| Component | Technology | Role | Update frequency |
|---|---|---|---|
| Operational memory | SQLite memories table | Structured facts, preferences, household info — scoped by user or household, tiered by importance (core/contextual) | Real-time (LLM memory_save tool call) |
| Vector / RAG | Qdrant (decided) | Fast document retrieval from ~/data/knowledge/ | Continuous (on ingest) |
| Knowledge graph | TBD | Relationships, curated structured facts | On significant events; overnight consolidation |
| Episodic memory | File-based | Recent interaction context; significant event summaries | Overnight consolidation |
| LoRA weights | Fine-tune checkpoints | Household tone and style — NOT facts | Overnight (rolling window) |
Design constraint: facts must live in retrievable, deletable, inspectable storage. LoRA weights capture style only.
Operational memory scoping:
Memories have two dimensions:
- Scope:
user_uuidis set for user-specific memories, NULL for household-wide memories - Importance:
corememories are injected into every system prompt automatically;contextualmemories are retrieved on demand viamemory_recalltool when the LLM determines they’re topically relevant
Context injection model — core memories always present, contextual via tools:
- Core memories (both user and household) are loaded from the DB and injected into the system prompt on every request, between the base identity template and session purpose
- Contextual memories are retrieved via
memory_recall(Tier 0 tool) when the LLM needs domain-specific context — the LLM provides topic tags and the DB returns matching memories - Knowledge base is searched via
search_knowledge_base(Tier 0 tool, stub until M2 RAG) - Conversation history (last 5–8 messages) is always present
- The LLM chains memory and RAG when needed (two-hop retrieval: memory gives personal facts, RAG gives reference material)
Memory creation: The LLM creates memories at runtime via memory_save (Tier 1 tool call) during normal conversation. No separate extraction model or async post-inference pass. The system prompt includes specific guidance on what to save vs. not save, and when to ask before superseding ambiguous facts. Behavioral regression tests enforce prompt calibration over time.
Memory replacement: memory_save accepts an optional supersedes parameter — a keyword from the old memory to replace. The service layer searches and removes matches, then creates the new memory. The LLM must confirm with the user before superseding when old and new facts could coexist (e.g., a person can own two cars).
Information type taxonomy:
All information Quorra works with falls into one of four categories. The category determines the retrieval mechanism and storage target:
| Type | Example | Retrieval pattern | Interim home | Target home |
|---|---|---|---|---|
| Factual | ”I prefer the long route to avoid traffic” | Tag / key-value lookup | SQLite memories | SQLite memories |
| Entity knowledge | ”Cara is Gary’s wife”, “Audi takes premium fuel” | Graph traversal | SQLite memories (interim) | Knowledge graph |
| Episodic | ”I had a difficult conversation with my brother today” | Time + context index | Overnight summaries (partial) | Episodic tier |
| Reference / Documents | Car manual, saved recipes, personal notes | Semantic search | ~/data/knowledge/ vault | Vault + RAG (Qdrant) |
Factual memories have the user as their subject — preferences, personal state, biographical facts. Entity knowledge covers properties of and relationships between named entities the user interacts with (people, objects, places); this merges what might colloquially be called “relational” and “domain knowledge” since both require entity-identity lookup rather than semantic similarity, and both target the knowledge graph. The routing rule for ambiguous cases: if the subject is the user themselves → Factual; if the fact creates or models a named entity → Entity knowledge. Reference/Documents are not memory rows — they are vault files indexed by the RAG pipeline.
The memories table carries a memory_type enum column (factual | entity | episodic) so that entity-knowledge rows can be migrated to the knowledge graph when that layer is ready, without having to re-infer type from content.
Expiring and consume-once memories:
Some memories are inherently transient. “I’m not feeling well today” should surface once in the next morning brief and then be discarded — persisting it as a normal contextual memory would cause it to recur indefinitely.
Two additional columns on the memories table handle this:
expires_at— nullable timestamp; a memory past this time is treated as deleted.surface_once— nullable boolean; whentrue, the memory is queued for the next relevant surface opportunity (typically the daily briefing mode), then deleted after it is surfaced.
The daily-briefing mode’s entry context callable queries for pending briefing-queue items: memories where surface_once = true that have not yet been consumed, or where expires_at is imminent. After the briefing runs, consumed surface_once memories are deleted.
Memory legibility surface (built 2026-07-28): principle 7 promises the user can see, edit, and delete anything Quorra knows about them; until now nothing rendered the memories table. Settings → Memory (quorra-web, /settings/memory) lists every row the owner holds, grouped by scope, with create / edit / archive / restore / permanent-delete. Three things make it a legibility surface rather than a table viewer: provenance is rendered in plain language (the six source values become “Quorra inferred this”, “From your notepad”, …), core vs contextual is stated as what it actually means (in every prompt vs looked up by topic), and a tag-less contextual memory is flagged — tag matching is the only content path into contextual recall, so an untagged row is stored but unreachable.
It reads across workspaces. The seal is a cognition boundary — it stops the model seeing across contexts — not an audit boundary against the owner, who owns all of these workspaces. list_memories_for_owner implements that separately from _scope_filter, which stays untouched on the cognition path; the read is owner-scoped, guest-unreachable, and never feeds a prompt. See decisions.md.
Soft delete (2026-07-28): deleting a memory sets memories.archived_at — invisible to every read at once, hard-purged by a nightly worker (memory/purge.py) after memory_archive_retention_days (30). supersede_and_create archives too, which is the change that mattered most: it matches by naive substring and the 14B is known to over-supersede, so it had been silent unrecoverable loss. The LLM tool path can only archive; permanent deletion is reachable only by the owner through the UI (principle 4, expressed in the data layer).
Memory reconciliation (overnight, M5): Deduplication, contradiction detection, stale cleanup, importance promotion (contextual → core for frequently-referenced facts), and synthesis of observations from daily patterns. Passive behavioral pattern synthesis — detecting “Kurt frequently attends his nephew’s baseball games” across many sessions — is handled here. A distinction applies: patterns from stated behavior (explicit mentions across conversations) can be synthesized from session history alone. Patterns from observed behavior (calendar attendance, location data, activity streams) depend on M4+ service integrations and are not possible before those exist. Synthesized patterns are promoted into the memories table as core importance facts.
Context-triggered reminders:
Context-triggered reminders are a distinct primitive from time-based reminders. A time-based reminder fires at a scheduled moment; a context-triggered reminder fires when a specific situation is recognized in conversation.
Example: “Next time I’m getting gas, remind me to use premium in the Audi.” There is no schedule — the trigger condition is semantic, not temporal.
These are stored in a separate context_triggers table (not the memories table):
| Field | Type | Description |
|---|---|---|
id | integer PK | |
user_uuid | text | Owner |
description | text | What to surface (“use premium fuel in the Audi”) |
trigger_hint | text | Natural language trigger condition (“when Kurt mentions getting gas or stopping for fuel”) |
created_at | timestamp | |
session_id | text | Session in which the trigger was created |
At the start of each planning-scope session turn, the LLM is given the list of active context triggers for the user. When a turn’s content matches a trigger — recognized semantically by the LLM, not by string matching — Quorra surfaces the reminder and marks the trigger consumed (deleted) or recurring per user preference. This belongs to the planning scope’s tool surface.
4.4 Action layer (permission tiers)
Every tool registered with Quorra declares a tier. Tier cannot be promoted at runtime.
| Tier | Examples | Confirmation |
|---|---|---|
| 0 — Read-only | Query files, search photos, read calendar, check system health | None |
| 1 — Low-risk writes | Create reminder, add calendar event, request media download, restart non-critical service | Notification after |
| 2 — Destructive or external | Delete files, send messages, make purchases, modify system config | Explicit user confirmation before |
For Tier 2 actions: draft the action and ask for approval. Never execute silently.
Tool naming convention: Tools use underscore-prefixed namespaces matching their integration (finance_log_transaction, calendar_query_events, system_get_health). Cross-cutting tools have no prefix (search_knowledge_base, get_my_context). The OpenAI tool-calling format only allows [a-zA-Z0-9_-] in tool names — slashes are not valid.
4.5 Inference layer
Local LLM serving for the Quorra agent.
- Framework: llama.cpp (native — replacing Ollama)
- Hardware: RTX 5080 (16GB) + RTX 3070 (8GB) — both passed through to Ubuntu VM, each running an independent llama.cpp server process (
CUDA_VISIBLE_DEVICESpinned) - Target API: OpenAI-compatible local endpoint per GPU (allows swapping models without changing QuorraAPI)
- Dual-GPU strategy: Separate models per GPU — enables true parallel inference between Quorra agent and diagnostic agent, lower perceived latency per model, and tractable LoRA fine-tuning on smaller models
GPU assignments:
| GPU | VRAM | Role | Model |
|---|---|---|---|
| RTX 5080 | 16GB GDDR7 | Quorra agent (user-facing) | Qwen 3 14B Q4_K_M — confirmed (10.7 GB VRAM, ctx 8192, --parallel 1; 78.8 tok/s P50; benchmarked 2026-05-21) |
| RTX 3070 | 8GB GDDR6 | Diagnostic agent (M6) | Qwen 3 8B Q4 (~4–5GB) — validated but not yet deployed; reserved for M6 |
LoRA fine-tuning: QLoRA on Qwen 3 14B via Unsloth typically requires 8–12GB VRAM at Q4 quantization — tractable overnight on the RTX 5080. Fine-tuning is scoped to smaller models only (not a 32B split model).
Embedding model: nomic-embed-text or similar (~137M params, <1GB) — can run on either GPU without meaningful VRAM impact.
4.6 The knowledge base (~/data/knowledge/)
The primary fact layer. Quorra is the primary author; users have direct edit access.
~/data/knowledge/
├── household/ shared household resources
│ ├── calendar/
│ ├── finances/
│ ├── home/
│ ├── pets/
│ └── vehicles/
├── members/
│ ├── kurt/
│ │ ├── knowledge/ personal vault (675+ markdown files, indexed by RAG)
│ │ └── shared/
│ └── example/ template for new members
├── guests/
└── templates/
The vault holds reference knowledge only. Agent cognition (learned preferences, observations) and per-user settings live in the quorra-api DB (memories, user_preferences) — not the vault. The pre-DB agent-memory/ and household-agent/ directories were retired 2026-06-15; the access model they documented moved to authorization.md. See decisions.md.
User access: files are readable and editable directly via Nextcloud (OnlyOffice for rich editing). Direct edits are detected by the inotify watcher, logged, and audited during overnight consolidation. The long-term UX goal is all edits flowing through Quorra — view the file in OnlyOffice, ask Quorra to make the change — so the write schema is enforced at authoring time rather than corrected overnight. Direct edit access is preserved for small changes that don’t warrant an inference call.
Design constraint: the RAG corpus should be predominantly prose. Embedding models produce low-discrimination vectors for terse, structured content — checkbox task lines, bullet fragments, key-value pairs — leading to noisy retrieval. Embeddings also cannot capture state (checked vs. unchecked, active vs. resolved), so structured task data belongs in the database, not the vault. Inline structured content dilutes the surrounding prose chunks it sits within, degrading retrieval quality for the prose too. When structured data exists in the vault for human readability (e.g., a markdown file summarizing preferences), the authoritative copy for Quorra’s actions is the DB row — the vault copy is a reference artifact. This constraint shapes what the M2 embedding pipeline should index, what it should skip, and how chunk boundaries are drawn.
4.7 Embedding update strategy
Embedding updates are tiered by latency urgency — heavier processing is deferred to overnight where possible.
| Trigger | What | When |
|---|---|---|
| Quorra writes a knowledge file | Re-embed the affected file immediately | Synchronous — Quorra authored it |
| User edits a file directly | Log the change, re-embed the file | Near-real-time: inotify watcher on knowledge/ detects changes, batches every ~5 minutes; change logged for overnight schema audit |
| User hands Quorra a document to ingest | Ask: “Learn this now or later tonight?” | User-directed; respects hardware constraints honestly |
| Full consistency pass | Re-index everything not yet embedded or with stale embeddings | Overnight only |
The “learn now or later?” prompt is a deliberate UX choice — it sets honest expectations about the cost of immediate indexing rather than hiding latency behind a spinner.
4.8 Overnight consolidation job
Scheduled window: 2–4 AM via systemd timer on JUNC1.
Tasks (in order):
- Full consistency pass — re-index any files not yet embedded or with stale embeddings
- Direct-edit audit — review the inotify change log for files edited directly by users; check each against the vault schema; normalize structural drift; re-embed affected files
- Refresh stale embeddings in the vector store
- Generate episodic summaries of significant interactions from the day
- Update structured fact files (relationships, preferences inferred from recent events)
- Memory reconciliation — runs against the
memoriestable:- Deduplicate redundant memories (same fact saved in different sessions)
- Detect and resolve contradictions (e.g., two different cars listed as “Kurt’s car”)
- Remove stale memories contradicted by newer ones that the LLM missed at runtime
- Promote importance: contextual → core for facts referenced repeatedly
- Synthesize observations from daily interaction patterns
- Clean up orphan memories (no tags, empty content)
- Generate LoRA fine-tune training candidates from the recent interaction window
- Run LoRA fine-tune on RTX 5080 if candidates exceed threshold (Qwen 3 14B, QLoRA via Unsloth)
- Write consolidation report to logs
Constraints:
- Must not run during active sessions
- Must be independently restartable (each task idempotent)
- Produces a report readable by the user
4.9 Observability
Two views over the same metric pipeline — a live operational snapshot and a historical comparison view.
QuorraAPI llama-server (systemd, host)
└─ /metrics └─ /metrics (--metrics flag)
│ │
└──────────┬──────────────┘
▼
Prometheus (admin compose stack, 15s scrape, 90d retention)
│
┌───────────┼────────────────────────┐
▼ ▼
quorra-dashboards Grafana
(live, 5s refresh, (historical + week-over-week,
in-memory app counters) annotated with deploys + model switches)
Metric sources:
- QuorraAPI — application-level metrics (LLM call duration / tokens, chat-loop iterations, tool calls, HTTP latency from
prometheus-fastapi-instrumentator) defined inrepos/quorra-api/src/quorra_api/metrics.py. Exposed atquorra-api:8000/metrics. - llama-server — inference internals (
llamacpp:tokens_predicted_total,prompt_tokens_total,kv_cache_usage_ratio,requests_processing, …) when started with--metrics. Exposed at172.17.0.1:11435/metrics(docker0). Prometheus reaches it viahost.docker.internal(extra_hosts: host-gateway).
Surfaces:
quorra-dashboards(http://192.168.0.49:8081/) — single static HTML page, vanilla JS, parses QuorraAPI’s/metricsdirectly every 5s. Optimized for “what’s it doing right now”; in-memory counters reset on quorra-api restart.- Grafana
Quorra Inference History(http://192.168.0.49:3100/d/quorra-inference-history) — historical view, default range 7d. Each panel renders the live series alongside the same query withoffset $comparison_offsetso week-over-week regressions and improvements are visible at a glance. Survives container restarts (Prometheus TSDB on host bind mount).
Annotations for change correlation:
- QuorraAPI deploys —
quorra_build_info{git_commit, git_branch, build_time}gauge set to 1 at startup with deploy-identity labels (populated by~/projects/junc1/compose/quorra/scripts/deploy.sh). Grafana querieschanges(quorra_build_info[1m]) > 0and marks the timeline with the new commit on every redeploy. - Model switches —
scripts/switch-model.shPOSTs a tagged annotation to Grafana’s/api/annotationsafter the new model passes its health check, using a service-account token at/etc/quorra/grafana-annotations.token. Missing token → silent no-op, the switch itself never fails on annotation push.
4.9a LLM turn inspector (inspect/)
Prometheus answers how fast and how often. The inspector answers what was actually sent — the assembled system prompt, the tool schemas, the history window, the retrieved chunks, and the model’s reasoning. Those are the inputs that actually determine a reply’s quality, and without them a wrong answer can only be debugged by re-reading source.
_call_llm / _call_llm_stream ──┐
notepad.triage │
reflection ├─► inspect.recorder (in-memory ring, N=50)
workspaces.init │ │
│ ├─► GET /inspect/turns (owner + developer_mode)
concierge ── EXCLUDED ────────┘ ├─► DELETE /inspect/turns (principle 7)
└─► POST /bugs (opt-in attach) ──► bugs/*.md + *-payload.json
Storage is memory-only, by design. A collections.deque(maxlen=…) of TurnRecords, appended at turn start and mutated in place (so an in-flight or crashed turn is still inspectable). There is no table and no migration. The assembled prompt inlines core memories, budget balances, today’s CalDAV tasks, and the workspace file tree — persisting it would create a durable shadow copy outside the memory-deletion surfaces that principle 7 guarantees. Nothing reaches disk unless the owner explicitly attaches a turn to a bug report.
maxlen bounds record count, not bytes, so the recorder also enforces per-message and per-tool-result truncation plus a per-record byte budget, and exposes truncated so the UI never implies it is showing more than it kept.
Capture is explicit, never ambient. Each call site passes a TurnRecord and calls begin_call / finish_call. An httpx event hook on the shared client would have been fewer lines, but it would capture Radicale, Qdrant, Actual, the embedding server, and the guest concierge, and would then need a maintained deny-list to exclude the egress boundary — policy where principle 2 requires architecture. The concierge path has no recorder call, and an import-assertion test in the guest egress suite is the structural guard on that.
The payload is serialized eagerly, before the request goes out: the chat loop appends to its messages list across tool iterations, and build_chat_payload appends the configured prefill to the payload’s own list — so the payload and the loop’s messages are not the same object and diverge whenever a prefill is set. Capturing before the POST also means a failed call still has its payload, which is the case worth having.
Access is gated three ways: normal auth, per-record owner scoping (mismatch → 404, no existence leak), and developer_mode — the first server-side enforcement of that flag, which until now gated client UX only. A ReservationPrincipal is not an AuthentikUser, so /inspect is unreachable from the guest path by construction.
Reasoning (delta.reasoning_content, or an inline <think> span depending on the llama.cpp build) travels a channel of its own: a separate SSE frame type, never mixed into reply text, never persisted, and never fed back into history. StreamCleaner’s guarantee that emitted text is byte-identical to clean_reply under any chunking is unchanged — reasoning capture is a sink attached to the existing discard path, not a relaxation of it. Live display is a normal user preference (show_thinking, default on); only the inspector is developer-gated.
5. Data flows
5.1 User query → response
user message (Matrix DM or PWA)
→ client adapter → POST /chat
→ identity resolution (Authentik UUID)
→ session history load (last 5–8 messages, SQLite)
→ core memories loaded (user + household, from memories table)
→ LLM call (system prompt [with core memories] + history + tool list)
→ tool call loop:
LLM calls memory_recall if topic-specific context needed (Tier 0)
LLM calls memory_save if user shared a memorable fact (Tier 1)
LLM calls search_knowledge_base if document search needed (Tier 0)
LLM calls domain tools as appropriate (finance, calendar, etc.)
→ response → client
→ history persisted (SQLite)
5.2 Photo ingestion (existing pipeline — Immich)
photo uploaded to Immich
→ CSAM hash check (required on all ingestion)
→ Immich CUDA ML: face clustering, object detection
→ Quorra integration (future): semantic embedding → vector store
→ knowledge graph update: face cluster → household member
5.3 Overnight consolidation
systemd timer (2 AM)
→ check active sessions → abort if any
→ embed new documents/photos added since last run
→ distill episodic summaries from day's interactions
→ update knowledge graph
→ memory reconciliation (deduplicate, resolve contradictions, promote, synthesize)
→ generate LoRA training candidates
→ [optional] run LoRA fine-tune
→ write report to the quorra-api data volume (/data/, alongside the audit log; final path an M5 decision)
6. Interface contracts (stubs)
6.1 QuorraAPI
- Transport: HTTP (exact: TBD — HTTP/WebSocket for streaming)
- Primary external interface:
POST /chat(custom JSON) — all clients connect here - Auth: two paths, both resolve to Authentik UUID:
- Authentik OIDC JWT (Bearer token) — for direct API callers with OIDC tokens
- Trusted-service auth (Bearer service-secret +
X-Quorra-User-Emailheader) — for internal adapters (Open WebUI Pipe, Matrix bot) that identify users on behalf of the chat interface; email → UUID resolved via config-based user map
- Format: JSON (QuorraAPI’s native format, not OpenAI-compatible — adapters handle per-client translation)
- The
/v1/chat/completionsOpenAI-compat endpoint was considered and deliberately skipped; revisit only if a third-party client that exclusively speaks OpenAI format emerges
6.2 Quorra ↔ Diagnostic agent IPC
- Internal only; no external exposure
- Format: TBD (Unix socket JSON messages or lightweight message queue)
- Quorra can request: health summary, specific diagnostic check results
- Diagnostic agent can push: critical alerts, threshold breaches
6.3 External service integrations
| Service | Interface | Direction |
|---|---|---|
| Nextcloud | Local WebDAV / REST API | Quorra reads/writes |
| Immich | Local REST API | Quorra reads; ingest events via webhook |
| Jellyfin | Local REST API | Quorra reads (library, watch history) |
| Home Assistant | Local REST API + WebSocket | Quorra reads/writes (device states, automations) |
| Actual Budget | @actual-app/api or REST | Quorra reads/writes (transaction logging) |
| Matrix | Client-Server API | Quorra sends notifications; bot adapter handles inbound |
| Gitea | REST API | Quorra reads (repos, issues) |
6.4 Capability interface principle
Quorra depends on capabilities (“log a transaction,” “search photos,” “get open tasks”), never on implementations (Actual Budget, Immich, a specific task app). Each integration domain exposes a lean interface shaped by what Quorra actually needs to do — not a mirror of the backend’s full API surface. Swapping a backend means writing a new adapter behind the same capability interface, not re-architecting the tool surface or prompt fragments.
The existing tool naming convention already supports this: finance_log_transaction is the capability; the actual-bridge sidecar is the current adapter behind it. Today’s service choices are provisional — they prove the concept now, but may be replaced later with native implementations or different third-party services. Routing through thin adapters behind small capability interfaces makes “swap the backend” mean “write a new adapter,” not “re-architect Quorra.”
Keep interfaces lean, not baroque. Cover only the operations Quorra actually uses. Avoid both extremes: too implementation-specific (leaks backend details, defeats the purpose) and too elaborate (over-engineering for swaps that may never happen). Draw the interface at the level of what Quorra needs to do — nothing more.
Installable integrations = MCP apps (target, decided 2026-07-26). The wire format for installable third-party integrations is MCP: an integration is an app packaged as an MCP server exposing tools/resources/prompts; quorra-api is the sole MCP host and the only inference loop (“apps expose capabilities; Quorra thinks”). A thin Quorra manifest (declared tools, requested tiers — platform-clamped, external send ≥ Tier 2, fixed at install; emitted events) plus a workspace ServiceBinding row forms the install/opt-in grant, registered through the integration registry’s plugin seam; a single authenticated, content-free event-poke endpoint lets an app announce inbound work, which Quorra then pulls through the app’s own tools. Apps shipping their own cognition are a documented non-goal of this contract. First planned instance: the concierge’s PMS/channel layer (§4.1d). See decisions.md “Concierge extraction target — an installable MCP app”.
7. Multi-user model
Every data structure must accommodate multiple users from the start.
- Household-level state: facts shared across all users (
household/in the knowledge base) - User-level state: per-user agent memory, episodic history, private notes, preferences
- Identity: Authentik UUID is the canonical user identifier — consistent across Matrix, Nextcloud, Jellyfin, and any future interface
- Memory namespacing: vector store and knowledge graph entries tagged with user scope (household or per-user)
- Quorra agent context: knows which user is speaking and scopes retrieval accordingly
- Workspaces: household and each member are the first instances of a first-class Workspace (§4.1c) — a data-partition + egress-policy boundary. A rental (or side business, project, delegated brief) is another. Sessions resolve into a workspace; the workspace holds its bindings, personas, and policies.
- Principal kinds: the actor is a
Principalprotocol, not only an Authentik user —AuthenticatedPrincipal(household members),ReservationPrincipal(a guest, with no owner identity, expiring at checkout), and a documentedAnonymousPrincipalseam. Authentik UUID remains canonical for authenticated users; non-authenticated principals never carry one.
8. Privacy and security model
- No data leaves JUNC1 without explicit per-request opt-in
- Cloud dependencies enumerated and opt-out-able individually (initially: relay for remote access, OTA updates)
- Agent action audit log: all Tier 1+ actions logged to an append-only file at
/data/audit.log(configaudit_log_path) and dual-written to the SQLiteaudit_logtable - Memory inspection: users can view and delete any stored fact about themselves via the
/memoryAPI - CSAM detection: hash-based scan on all photo ingestion using established libraries
- Encryption at rest: TBD
- Guest / external-party egress boundary (§4.1d): when Quorra acts toward an untrusted external party (e.g. a rental guest), the data boundary is enforced structurally, never by prompt. Three controls — a subtractive persona tool allow-list, a retrieval binding hardwired to a guest-only collection, and a default-deny
guest_visiblecontent gate — plus a principal that carries no owner identity, so owner data is unreachable by construction even under prompt injection. The external-facing agent runs draft-and-approve (the owner reviews every outbound message) until autonomy is explicitly granted per category.
9. The Grid (long-term vision)
The Grid is a federated network of personally-owned junc nodes. Each piece of junc is independently functional — self-contained, locally intelligent, privately owned — but capable of communicating with others directly, with no platform in the middle.
Design for it; don’t build it yet. Architectural choices that affect Grid readiness:
- Use open standards for data formats (CalDAV, CardDAV, ActivityPub-compatible schemas) where practical
- Authentik identity maps cleanly to a future portable, self-sovereign identity layer
- Per-user memory namespacing is compatible with federated access (explicit trust grants per sharing request)
Current focus: make a single piece of junc work exceptionally well. The Grid becomes possible once the single-node product is validated.
10. Open questions
Track unresolved design decisions here. When resolved, move the answer into the relevant section and note the worklog date.
| Question | Status | Candidates / Notes |
|---|---|---|
| LLM models | Decided | RTX 5080: Qwen 3 14B Q4_K_M — P50 TTFT 44.5ms, 10,719 MB VRAM, 5,122 MB free. RTX 3070: Qwen 3 8B Q4 — P50 TTFT 50.5ms, 5,717 MB VRAM, 2,124 MB free. Both live as systemd services. |
| Dual-GPU inference strategy | Decided | Separate llama.cpp processes per GPU (see section 4.5); tensor split approach set aside |
| Vector database | Decided | Qdrant — better standalone performance; ChromaDB set aside |
| Agent framework architecture | Decided | Custom tool-calling in Python; permission tiers declared at registration, not enforced by framework |
| Implementation language | Decided | Python 3.12; FastAPI for QuorraAPI; uv for dependency management |
| Tier 2 confirmation UX | Decided | LLM generates a human-readable description of the pending action; client re-submits with confirmation_id. Natural-language confirmation (“yes”, “go ahead”) works. Implemented in M3. |
| Context injection: system vs. user message | Decided | RAG and episodic memory exposed as Tier 0 tools the LLM calls selectively — not pre-assembled. Only conversation history always injected. Eliminates context waste on pure tool calls. |
| Quorra ↔ diagnostic IPC | Open | Unix socket JSON vs. lightweight message queue; deferred to M6 |
| Diagnostic agent model | Open | Qwen 3 8B Q4 (small quantized LLM) vs. rule-based without LLM; deferred to M6 |
| Encryption at rest | Open | Proxmox/LUKS level vs. ZFS native encryption vs. application-level |
| LoRA training pipeline | Open | Scoped to Qwen 3 14B on RTX 5080; tooling: Unsloth (QLoRA); frequency and candidate selection TBD |
| Knowledge graph technology | Deferred | Tier 2 memory will use structured markdown files initially. Graph layer (Neo4j, SQLite + custom, or DuckDB) deferred to a later milestone. Design context assembly to abstract “retrieve facts” from storage format so graph can be added without rewriting the pipeline |