Quorra — Milestones
Updated: 2026-07-27 (audit — cross-referenced every marker against shipped code + the worklog; Workspaces Phases 1–2 marked done/deployed, post-Phase-2 increments listed, CalDAV planning marked complete, M7 Phase A marked done)
Status markers: [ ] not started · [~] in progress · [x] complete · [?] blocked
Infrastructure baseline (pre-existing — not our work to implement)
The following was built and is operational before this design-first restart. Quorra builds on top of it.
Core infrastructure ✅
- Proxmox hypervisor on physical machine; Ubuntu Server VM (ID 100) hosts all services
- Both GPUs (RTX 5080 + RTX 3070) passed through to Ubuntu VM
- Docker and Docker Compose; all services containerized
- nginx reverse proxy + Let’s Encrypt SSL; domain juncyard.com
Identity and access ✅
- Authentik SSO — central OIDC provider; all services authenticated through it
- Authentik UUID established as canonical per-user identity
Personal data services ✅
- Nextcloud (file storage and sync; Collabora for document editing)
- Immich (photo backup; CUDA-accelerated face recognition and ML)
- Obsidian notes sync: Nextcloud sync →
~/data/knowledge/members/kurt/knowledge/(vault opens directly from Nextcloud directory on all devices) - Household knowledge base structured at
~/data/knowledge/(675+ files) - Contacts migration from cloud: not yet complete (CardDAV deferred; Radicale is the provider). Calendar consolidation planned 2026-07-09:
plans/calendar-consolidation.md
Communication ✅
- Matrix/Synapse + Element Web — primary chat interface
- Authentik SSO wired to all services
Media ✅
- Jellyfin (SSO wired; Quorra integration not yet built)
Finance ✅
- Actual Budget (SSO wired; prior Quorra integration exists in old codebase)
Current LLM runtime (to be replaced)
- Ollama running Gemma 4B on RTX 5080 via GPU passthrough
- Prior QuorraAPI (TypeScript/Fastify) + quorra-rag (Python/ChromaDB) archived at
~/projects/junc1/_archive/quorra-2026-05-18/— intentionally set aside
M0 — Foundation: design and decisions
Goal: All major architectural questions surfaced. Project structure established. Developer can sit down and immediately know what the next task is.
- Project folder created at
~/projects/quorra - CLAUDE.md written (AI assistant context, infrastructure baseline documented)
- docs/technical/architecture.md first draft (components, memory tiers, permission tiers, open questions)
- docs/technical/milestones.md created
- logs/ worklog initialized
- Dual-GPU inference strategy decided — separate llama.cpp processes per GPU
- Vector database decided — Qdrant
- Agent framework architecture decided — custom tool-calling (Python)
- Development toolchain defined — Python 3.12 (system), FastAPI, uv
- LLM models validated on JUNC1 hardware — Qwen 3 14B Q4_K_M (RTX 5080), Qwen 3 8B Q4 (RTX 3070); document P50 TTFT and VRAM headroom
- JUNC1 OS environment documented (Ubuntu version, CUDA version, available mounts)
Definition of done: A new session starts from CLAUDE.md and knows exactly what to build first. The inference framework and base LLM are validated on JUNC1 hardware. All open questions in architecture.md have at least a tentative decision.
M1 — LLM serving (llama.cpp) ✅
Goal: Local LLM accessible via OpenAI-compatible API using llama.cpp. Both GPUs utilized via separate server processes. Latency acceptable for interactive use. Ollama is replaced.
- CUDA toolkit 12.8 installed on JUNC1 (nvcc at /usr/local/cuda-12.8/bin/nvcc)
- llama.cpp b9159 built from source with CUDA (sm_120 + sm_86 architectures)
- Qwen 3 14B Q4_K_M extracted from Ollama → /home/kurt/data/models/ (SSD, 8.7 GB)
- Qwen 3 8B Q4 extracted from Ollama → /home/kurt/data/models/ (SSD, 4.9 GB)
- quorra-llm-primary.service: RTX 5080 (CUDA 0), port 11435, always-loaded
- quorra-llm-secondary.service: RTX 3070 (CUDA 1), port 11436, always-loaded
- Both services enabled and starting on boot (systemd)
- OpenAI-compatible API operational on both ports (validated with curl)
- P50 / P95 TTFT benchmarked: primary 44.5ms / 65.5ms; secondary 50.5ms / 81.6ms
- VRAM headroom documented: RTX 5080 5,122 MB free; RTX 3070 2,124 MB free
- Ollama container stopped (compose stack at ~/projects/junc1/compose/llm/)
- notes-exporter and couchdb containers removed (compose stack at ~/projects/junc1/compose/notes/)
- Model keep-alive: always-loaded (—keep -1)
- repos/llama-serving/ git repo created with systemd units, scripts, README
Definition of done: curl to the local llama.cpp API returns a coherent response. Both GPUs active under load. P50 TTFT documented and acceptable for conversation (target: <3s warm). Ollama is no longer the runtime.
M2 — RAG layer ✅ (retrieval-first cut, 2026-06-17)
Goal: Vector database operational. Embedding pipeline feeding it from ~/data/knowledge/. Retrieval-augmented generation working end to end.
Shipped as an in-process rag/ subsystem in quorra-api (not a sidecar — cf. decisions “RAG implementation”). DoD verified live: /chat “what fluids does my Audi take?” → accurate cited answer from vehicles/2011-audi-a3/overview.md.
- Qdrant deployed and operational (
qdrant/qdrant:v1.12.4, compose,127.0.0.1:6333) - Embedding model selected and running — nomic-embed-text-v1.5 on RTX 5080:11437 (
--embeddings, systemd, co-resident with the 14B) - Indexing pipeline operational: markdown → embeddings → vector store (
scripts/rag_indexing/backfill.py; 449/450 chunks, 1 pathological URL-dump skip) - [~] Incremental reindex — backfill is re-runnable/idempotent (manual reindex); Quorra’s own writes reindex-on-write since 2026-07-23 (
document/organize→rag.indexer), so the deferred inotify watcher (below) now only covers direct human edits - Basic RAG query: question → retrieved context → LLM → answer with source citation (verified live via
/chat+search_knowledge) - Source deduplication (best chunk per source file)
- [~] Markdown score boost — N/A for now (vault is markdown-only; no raw PDFs until the doc-ingestion/priming pipeline)
- Memory namespace isolation — per-owner Qdrant collections (
kb_household+kb_member_<slug>); a query searches household ∪ the acting user’s own collection only - User can inspect retrieved chunks for a given query —
GET /rag/search(legibility) - Embedding pipeline implements vault schema spec (§7): H2/H3 chunking (+ word/char split caps for nomic’s 2048-tok ctx), all frontmatter + path-derived
owneras Qdrant payload for pre-filtering - inotify watcher on ~/data/knowledge/ — near-real-time re-embed + overnight-audit log — deferred fast-follow (
rag/watcher.py; shares the edit-detection seam with session 3’s write-gate) - Review v1 RAG tuning notes (
~/data/archive/quorra-v1/docs/WORKLOG.md, 2026-05-12) — optional, low priority; informs relevance tuning
Knowledge graph infrastructure is deferred. Entity-knowledge facts live in the memories table (memory_type=entity) as an interim representation; overnight consolidation will migrate them to the graph when that milestone arrives.
M2 follow-ups (backlog): the inotify watcher (above); retrieval relevance tuning (single-word queries rank loosely — score thresholds / hybrid / rerank); UUID→slug DB table to replace the interim MEMBER_SLUG_MAP_JSON config var (rag/service.py:member_slug_for is the single swap point); pin Qdrant client(1.18)/server(1.12.4) versions; the doc-ingestion/priming pipeline (plans/document-ingestion-priming.md, gated on the session-3 write-gate).
Definition of done: Asking a question about a document in ~/data/knowledge/ returns an accurate, cited answer. A per-user query does not surface another user’s private documents. ✅
M3 — Core Quorra agent (QuorraAPI) ✅
Goal: QuorraAPI redesigned from scratch. Chat loop with permission-tiered tools and memory integration working end to end. Multi-user sessions working via Authentik.
- QuorraAPI project scaffolded in
repos/quorra-api/— FastAPI, uv, Python 3.12, Alembic/SQLite -
POST /chat— session-based conversation with history (SQLite, rolling 5–8 message window) - Identity resolution — Authentik OIDC JWT → UUID as session identity (JWKS-cached middleware)
- Context assembly — RAG and episodic memory exposed as Tier 0 tools (
search_knowledge_base,get_my_context); conversation history always injected; no unconditional pre-assembly - Tool registration system — each tool declares its tier at registration; tier cannot be promoted at runtime
- [~] Tier 0 tools:
search_knowledge_base(stub — M2 pending) andget_my_context(reads core + contextual memories from the DBmemoriestable) remain cross-cutting (all scopes). The calendar/contacts/media/jellyfin/email/photos/system-health stubs were unregistered in the room-scoping rollout (2026-05-21) — they returned placeholder data and were misleading; each will reappear when its real integration and dedicated scope land. (unregistered by design — will reappear under dedicated scopes in M4) - [~] Tier 1 tools:
calendar_create_event,contacts_create,email_sendlikewise unregistered until their real CalDAV/Proton wiring (M4) ships under their respective scopes (unregistered by design — will reappear under dedicated scopes in M4) - Session purpose — per-session directive stored on
sessions.purpose;GET|PUT|DELETE /sessions/{id}/purpose; injected into the system prompt between base and origin style - Time-gap markers —
load_historyinjects[N hours passed]/[next day — …]separators when consecutive messages are >4h apart (midnight crossing alone does not trigger one) - Streaming chat —
POST /chat/stream(SSE) emitsstatus/token/doneevents; tool calls run synchronously, the final reply streams - Tier 2 confirmation flow: LLM-generated description + re-submit with
confirmation_id; never executes silently -
GET|POST|PUT|DELETE /memory— full CRUD for user-scoped memories (core/contextual, tag-filtered, DB-backed) -
GET|POST|PUT|DELETE /household/memory— full CRUD for household-scoped memories -
memory_save(Tier 1) andmemory_recall(Tier 0) tools — LLM creates/retrieves memories during conversation - Core memories injected into system prompt via async
build_system_prompt_async -
memory_savesupportssupersedesparameter for replacing obsolete memories - System prompt updated with memory save/recall guidance (what to save, when to ask before superseding)
- Behavioral regression tests for memory tool calibration (save/don’t-save, supersede, coreference)
-
POST /events— inbound webhook stub; accepts and acknowledges; dispatching TBD -
GET /health— checks LLM primary and actual-bridge reachability - Agent action audit log — dual-write to SQLite
audit_logtable and append-only flat file - Review v1 system prompt — closed as superseded; v2 built from scratch with scope/origin/mode fragments, memory guidance, and vault authoring conventions; v1 was Gemma 4B era and no longer applicable
- Matrix bot adapter (thin) —
repos/quorra-matrix/; matrix-nio bot, DM/@mention gating, per-user service auth, room→session mapping, progressive streaming viam.replace,!quorracommands; no LLM logic in the bot - Multi-user session isolation — sessions scoped to Authentik UUID; cross-user session re-use rejected
- Origin-aware system prompt adaptation —
ChatOriginenum (matrix, open_webui, voice, standard); per-origin style directives inprompts/origin/*.md; origin persisted to session metadata - Room/session scoping —
scoperequest param on/chat//chat/stream; nullablesessions.scope_json(locked once committed); scope-aware tool registry (for_scope+get_for_scopedefense-in-depth); pre-commit tool gate; additive per-scope system-prompt fragments underprompts/scopes/; Matrix bot resolves room→scope viaMATRIX_ROOM_SCOPE_MAP+resolve_scope(); Open WebUI Pipe commits new sessions to["general"]. Shipsgeneral+financescopes; tool stubs for unimplemented integrations unregistered. - Direct-response short-circuit for finance tool actions —
direct_responseflag onToolDefinition;action_direct_responsefor consolidated tools; skips second LLM inference when tool returns user-ready text (finance query latency: 10.2s → 2.4s avg) - Account names injected into finance scope prompt context —
_fetch_account_namesincontext.pycalls actual-bridge/balances; result injected intorender_budgets_block;accounttool parameter references the list explicitly; graceful degradation on bridge unavailability - Observability —
quorra-dashboardsstatic live metrics page (:8081); Prometheus scraping QuorraAPI + llama-server--metricsendpoint viahost.docker.internal; GrafanaQuorra Inference Historydashboard with week-over-week offset comparison; deploy annotations viaquorra_build_infogauge; model-switch annotations fromswitch-model.shAPI push - memory_delete tool (Tier 1) — LLM can delete memories during conversation by keyword match
- Modes designed (§4.1a architecture, decisions.md 2026-05-25) — implementation deferred to M4
Definition of done: Full conversation loop — user sends message via Matrix, Quorra retrieves context, calls a tool, returns answer — working end to end. Tier 1 actions log a notification. Tier 2 actions require explicit confirmation. Audit log is populated.
Room-scoping follow-up items (LLM-guided scope-setup, user-defined room scope, multi-scope sessions, Open WebUI scope selector, Matrix no-markdown fix) moved to backlog below.
M4 — Service integrations
Goal: Quorra connected to the running services on JUNC1. Conversation with Quorra is the primary interface for all integrated services; native UIs are visual aids and backups. See service-integrations.md for full detail on each integration.
Modes (§4.1a) — shipped 2026-05-25 (code in modes/, not chat/modes.py)
-
active_modenullable text column on sessions table (Alembic migrationg1b2c3d4e5f6) - ModeRegistry + ModeDefinition dataclass (
src/quorra_api/modes/registry.py) -
GET|PUT|DELETE /sessions/{id}/modeendpoint wired in sessions router; validates against session scope; returns 409 for pre-commit sessions - [~] Mode prompt files — finance done (
prompts/modes/finance/{budget,reconcile,review}.md); general/planning modes pending (daily-briefingis now a planning-scope mode tracked inplans/daily-briefing-mode.md;researchnot started) - Mode-aware prompt composition: base → scope → mode fragment (if active) → core memories → purpose → style
- Auto-expire on 4+ hour time-gap: clear
active_modebefore building system prompt when gap detected - Entry context callable for
daily-briefingmode — deferred toplans/daily-briefing-mode.md(the briefing composer) - Matrix bot
!quorra mode [name]wired to mode endpoint;!quorra modewith no argument clears mode - [~] Behavioral regression tests — finance modes covered; general/planning modes pending
Memory schema extensions (designed 2026-06-03)
- Alembic migration: add
memory_typeenum column to memories table (factual|entity|episodic) - Alembic migration: add
expires_at(nullable timestamp) andsurface_once(nullable boolean) to memories table - Update
memory_savetool schema: optionalexpires_atandsurface_onceparameters with schema guidance - Alembic migration: create
context_triggerstable (id,user_uuid,description,trigger_hint,created_at,session_id) - Regression tests for expiry and surface_once behavior
Planning scope — shipped; CalDAV re-platform complete 2026-06-10 (plans/caldav-planning-replatform.md)
A task/reminder/scheduling planning consolidated tool shipped on SQLite 2026-05-26 and was fully re-platformed onto Radicale CalDAV (Phases 1/2a/2b/2c complete 2026-06-10; SQLite planning tables dropped, migration k5e6f7a8b9c0). Next evolution: multi-list read + multi-user household store — plans/calendar-consolidation.md (planned 2026-07-09).
-
planningregistered inscopes.py;prompts/scopes/planning.mdfragment - Tasks + time-based reminders + scheduling —
planningtool actions read/write CalDAV viaRadicaleStore; calendar events (create/list/update/delete_event) +complete_taskadded; deletes promoted to Tier 2;reminder_loopreads CalDAV withReminderDispatchLogdedup;{planning_block}prompt injection -
planning_set_trigger/planning_check_triggers(context_triggers) — not started; separate from the CalDAV work, still future - Matrix room mapping for a dedicated planning room in
MATRIX_ROOM_SCOPE_MAP - Behavioral regression tests for context trigger creation and matching
Workspaces, personas & guest concierge — designed 2026-07-09 (plans/workspaces-personas-concierge.md)
Introduces the Workspace abstraction (a “brief” bundling generic service bindings + personas + standing policies), a subtractive persona axis, a Principal protocol, and the first external-facing instance — a guest concierge for an STR/MTR rental (guest messages via a PMS, draft-and-approve replies surfaced in the owner’s Matrix chat). Generic core; the rental is an instance. Four phases; hard dependency chain 0→1→2→3.
Phase 0 — design docs (design-first gate)
- decisions.md — new “Workspaces & guest access” group (D1–D9)
- architecture.md — §4.1c Workspaces/personas + §4.1d Guest concierge; §7 principal model; §8 egress boundary
- service-integrations.md — PMS section + overview-table row
- milestones.md — this subsection
- CLAUDE.md Current State +
logs/2026-07.mdworklog entry
Phase 1 — Workspace + persona core — implemented + deployed 2026-07-09 (quorra-api d07974d; live DB o9c0d1e2f3a4); zero owner-behavior change held
-
workspaces+workspace_service_bindingstables/models + migration; seeded (later refined:household= ownerless shared root,kurtretired — personal context is the null-workspace default,strowned +hosting-bound) -
personas/registry (owner+guest-concierge);auth/principal.pyprotocol +AuthenticatedPrincipal -
Session.workspace_slug+Session.personacolumns + migration (backfill owner defaults; immutable-after-commit) - Subtractive persona filter at
tools/registry.pychoke points +chat/loop.pygates; scope ⊆ bindings validation -
prompts/personas/loader inchat/context.py; register personas inmain.py - Exit: existing suite green (owner unchanged); egress suite proves the guest persona is denied every non-allow-listed tool at every choke point
Phase 2 — egress-hard concierge, offline (no real PMS) — implemented 2026-07-09; deployed 2026-07-09 (quorra-api 7e1008c; live DB r2f3a4b5c6d7; /concierge live + auth-gated)
-
guest_visiblefrontmatter field (kb/schema.py+kb/validator.py,invalid-boolcheck);kb_str_guestindex split (collections_for_chunkinchunker.py, routed inbackfill.py; guest-visible chunks stay in their owner collection too) - Retrieval-binding path —
rag/service.pysearch_window()(one named collection, nevercollections_for_user) +search_guest_knowledge()hardwired tokb_str_guest;search_knowledgesplit into a standalone guest-safe tool (the consolidatedmemorytool is untouched for owners — a name allow-list couldn’t admit itssearch_knowledgeaction without also admittingmemory_save) -
concierge/package:ReservationPrincipal(owner_uuid=None, checkout+grace expiry), constrained-loophandler.py(own mini-loop, NOT the chat loop; same registry choke point;reservation_idpinned on every call; Tier 1+ audited underreservation:<id>), structured{reply_draft,category,confidence,urgency,escalation_reason}(unparseable → escalate),policy.pyautonomous_ok()(frozen empty allow-list) -
reservations+guest_messagestables + migrationp0d1e2f3a4b5; review service (get_pending/approve/reject, edit-on-approve, escalation round-trip);search_knowledge/check_this_booking/escalate_to_ownerunder the newstrscope (+prompts/scopes/str.md) - Matrix review surface — review cards via the existing
Notificationoutbox the bot already polls (zero bot changes);/conciergerouter (inboundseam the Phase 3 webhook will call,messages/pending,approve,reject); send stubbed (send_replylogs, Phase 3 = PMS adapter) - Exit (offline): adversarial egress suite green (
tests/test_egress_guest.py, 63 tests): guest surface = exactly 3 tools, every other tool unreachable in every scope, retrieval queries onlykb_str_guest(owner-resolver tripwired), owner secret never surfaces, injection corpus → escalation not leak, memory-poisoning attempt writes nothing, cross-booking probe pinned,autonomous_okfalse on full grid, expired reservation never reaches the LLM - [~] Exit (live): manual E2E on the deployed stack — hand-crafted
/concierge/inbound→ review card in the real Matrix room → approve →sent. Deploy done; owner routing resolved (review cards go toWorkspace.owner_uuid;QUORRA_BOOTSTRAP_OWNER_UUIDset); vault location resolved (workspacedata_dir). Still pending: STR guest content authored under thedata_dirwithguest_visible: true+ a backfill to populatekb_str_guest, and the bot joined to a Matrix room bound to thestrworkspace (matrix_room_idset)
Phase 3 — PMS adapter (external I/O; vendor chosen here)
-
integrations/pms/PMSClientinterface + Hostaway/Lodgify impl; per-workspace encrypted credentials -
/integrations/pms/webhookreceiver (signature-verified) → normalize → upsert reservation → concierge handler - Real send-on-approval; reservation sync (webhook + backfill); OTA content-rule enforcement; urgency → immediate Matrix DM
- Extraction-readiness constraints (decided 2026-07-26 — see decisions.md “Concierge extraction target”): (a) all PMS send stays on the trusted approve path — move the
autonomous_okauto-send out ofconcierge/handler.py::_persistinto the approve flow inconcierge/service.py; (b)PMSClientmethods are shaped as candidate MCP tools (get_inbound_messages/get_reservation/get_thread/send_reply) - Exit: against the chosen PMS sandbox — inbound guest message → draft → approved reply delivered through the original channel
Gate: Phase 3 (real channel) does not start until the Phase 2 egress suite is green — the boundary is proven offline before a real guest can reach it. (The egress suite is green; Phase 3 is unblocked.)
Post-Phase-2 workspace evolution — shipped increments 2026-07-11 → 2026-07-26 (details: decisions.md “Workspaces & guest access” + the worklog; added in the 2026-07-27 audit)
- Owner-in-workspace (the sealed inhabited context) — deployed 2026-07-11 (
s3a4b5c6d7e8):Session.workspace_slugseals RAG (kb_ws_<slug>, excluded from personal) and memory (workspace_slugdimension) both ways; Matrix room-per-workspace entry (GET /workspaces/by-room); owner scope derived from bindings (owner_workspace_scope); live seal matrix verified - Workspace authoring (
documentwrite-gate tool) — deployed 2026-07-23: kernelserialize→parse/validate→ block-on-error → write → reindex-on-write; owner-only, guest-unreachable; vault mount:ro→:rw - Rootless quorra-api (
quorra:vault) — deployed 2026-07-24: uid 1001 + sharedvaultgroup (Nextcloudwww-datajoined via entrypoint wrapper) so authored files landquorra:vault 664and stay OnlyOffice-editable. (Host-sidequorrauser/vaultgroup creation still pending — needs sudo, Kurt) - Passive reflection + KB-suggestion review inbox — deployed 2026-07-25 (
t4b5c6d7e8f9): reflect-on-idle worker →kb_suggestionsinbox → approval applies through the shared write-gate; review-time unified diffs +staleflag; enabled with test timings (2min idle/60s poll — raise to 15/300 later) - Explicit chat authoring routes to the inbox — 2026-07-25: chat
documentrequests validate then enqueue a suggestion (Tier 1 — the inbox review is the confirmation); the web editor’s direct Save stays a write-gate write - Workspace reorganization (
organizetool) — 2026-07-25: link-safekb/reorg.pykernel (move-map, zero-new-dangling simulate,entities:cross-refs), move-safe reindex,reorgsuggestion type, two-phase journaled apply - Workspace web app M1 (
repos/quorra-web/) — deployed 2026-07-25 at app.juncyard.com: chat SSE + files tree + editor + approval modal; Authentik PKCE (first client on the JWT path); shared markdown renderer; activity indicator; per-workspace session picker 2026-07-26 (see the future-phases entry below) - Workspace Tasks tab — 2026-07-25 (
ed7a8fa): project-scoped CalDAV task read/create surfaced as a workspace tab in the web app - Pinned-artifact chat scoping — 2026-07-25 (
73337e1; web pin icons077b207): optionalartifact_refon the chat request injects a scoping directive and pins the referenced task/file for theplanning/documenttools - Workspace creation wizard — deployed 2026-07-26 (
a0b9069; web2a38c5d+ login-stall fixddf5885; migrationu5c6d7e8f9a0, live DB stamped 2026-07-27): registry-drivenPOST /workspacesbacked by anintegrations/registry.pycatalog (withpluginseam) + idempotently-seededintegration_optintable;GET /integrations//budgets; finance offered bind-only (budget captured into the binding’sdata_window); LLM-drafted starter notes via/preview-notes. See decisions.md “Workspace creation is a registry-driven wizard”
Future phases — to design later (captured 2026-07-25; not yet scheduled)
- [~] Workspace creation workflow — steps (1)+(2) shipped 2026-07-26 as the registry-driven wizard (see the shipped-increments list above); remaining: (3) starter paperwork and (4) members & roles, plus Matrix-room auto-provisioning and hosting/guest-workspace creation. Original scope: a new workspace starts from a blank template, created through a guided setup flow rather than seed/DB surgery. The flow bundles: (1) service selection — manually pick which service bindings this workspace needs (budget, tasks/calendar, hosting, files, etc.); (2) prompt initialization — the user describes in detail what the workspace is for, and Quorra generates a starter workspace prompt/
## Workspace notesblock from that description; (3) starter paperwork — the user uploads relevant initial documents that seed the workspace’s KB as founding context (routes through the authoring/indexer path); (4) members & roles (looks ahead to multi-user-per-workspace) — add other users to the workspace and set their roles/permissions as part of setup. Design must reconcile with the existingWorkspace.owner_uuidsingle-owner model and the workspace roles model — the “members & roles” step is where the personas→roles unification lands (see decisions.md “Personas dissolve into workspace roles”). - Multiple chat sessions per workspace (web app) — done 2026-07-26.
GET /workspaces/{slug}/sessions(owner-gated, workspace-scoped) + a header-dropdownSessionPickerin the web app: list/switch/create/delete sessions, per-workspace active-session pointer in localStorage coexisting with N server-side sessions. Every session still seals to its workspace viaworkspace_slug(proven by a backend test excluding personal/other-workspace sessions). Matrix stays one-room-per-workspace (out of scope); thematrix_room_id1-1→1-many relaxation is a noted follow-up, not built. See decisions.md “Multiple chat sessions per workspace”. (Still overlaps the M7 global session-sidebar work — a globalGET /sessionsremains separate.) - Quorra as MCP host + app platform (direction decided 2026-07-26 — see decisions.md “Concierge extraction target — an installable MCP app”): the platform stance is apps expose capabilities; Quorra thinks — a 3p integration is an installable app packaged as an MCP server (tools/resources/suggested prompt); quorra-api is the sole brain and sole MCP host. Pieces: an MCP client in quorra-api; a per-app manifest (declared tools, requested tiers, emitted events) with platform-clamped tiers (external send ≥ Tier 2; fixed at install — no runtime promotion); install = manifest registration via the integration registry’s
pluginseam, workspace opt-in = a ServiceBinding row; one authenticated content-free event-poke endpoint (app announces inbound work; Quorra pulls via the app’s MCP tools). Documented seam, not built: apps shipping their own cognition (would require capability-scoped service tokens). Build when the host has a second customer (e.g. Home Assistant’s MCP server) — not concierge-first. - Concierge appification (supersedes “concierge as a standalone external service”; same 2026-07-26 decision): after Phase 3 has run against live guest traffic, lift the PMS/channel I/O into the first installed MCP app —
PMSClientmethods become the app’s tools;send_replyis an ordinary Tier-2 tool, so draft-and-approve is the standard tier system (review queue = the confirmation surface) andautonomous_okis a per-category waiver on the approve path. Cognition, the egress controls, and the review queue never move — the adversarial egress suite stays in-process and needs no re-derivation. - Personal + household vault authoring (deferred; scoped 2026-07-28 — see decisions.md “Vault authoring extends beyond workspaces”): let Quorra author into
household/andmembers/<slug>/knowledge/, not only workspacedata_dirs, by extendingapply_documentwith a destination rather than adding a second write path. This is the standing “Knowledge base authorship model” commitment (2026-06-03), still unbuilt;base.mdwas trimmed to today’s reach so the prompt stops claiming it. Seams: target resolution for the null-workspace case (wants theMEMBER_SLUG_MAP_JSON→DB-table TODO inrag/service.py:member_slug_forfirst); household-vs-personal routing (open question — does Quorra pick the destination or the user?); reindex intokb_member_<slug>/kb_household;document_dispatchaccepting a destination instead of refusing, staying guest-unreachable (extend the egress suite); a personal case for the workspace-keyedkb_suggestionsrows; widening thevaultgroup pastWorkspaces/; and sealing direction — a workspace session must not author into the personal tree (needs its own egress test). - Workspace roles (personas → roles) (vocabulary decided 2026-07-26 — see decisions.md “Personas dissolve into workspace roles”): re-key the persona axis as audience-relative roles — definitions in a frozen code registry (allow-list, data-window grants ⊆ workspace bindings, prompt block, tier ceiling), assignments in a DB membership mapping (principal → workspace → role), guests as implicit ephemeral memberships via
ReservationPrincipal. Lands as the opening move of multi-user-per-workspace, not as a standalone rename; later unification:Workspace.owner_uuid→ “the member holding roleowner”.
Service integrations
Priority order based on daily value:
-
Actual Budget — 9 finance actions (Tier 0/1/2) including
switch_budgetandlist_budgets; multi-budget access viabudget_accessDB table (many-to-many user↔budget, per-user aliases, per-session active budget with natural-language per-call override);actual-bridgeNode.js sidecar keys workers by(userUuid, syncId)so one user can hold multiple budgets in flight; per-requestparticipants-intersection enforces privacy in shared rooms; fully wired into QuorraAPI tool registry -
Calendar / Tasks / Reminders (Radicale CalDAV) — live since 2026-06-10: standalone Radicale store (internal
radicale:5232, external dav.juncyard.com) decoupled from Nextcloud, folded into theplanningscope viaRadicaleStore; iPhone CalDAV interop verified. Remaining evolution tracked inplans/calendar-consolidation.md(multi-list adapter read — the 20 per-project VTODO lists from the Stage-4 import; multi-user household store + reminder fan-out; Nextcloud Calendar retirement). Contacts/CardDAV deferred. -
Nextcloud — files / knowledge base — Quorra as primary file maintainer (M4 concern: active file ops — organize, ingest, move, delete via WebDAV); document ingestion flow (drop + describe → Quorra files it); inotify watcher for proactive organization (note: the re-embed trigger on user edits is the M2 inotify watcher; the M4 watcher drives filing decisions)
-
Jellyfin — library search, watchlist/playlist management (Tier 0/1); no playback control yet
-
Protonmail (email) — read inbox, extract action items to knowledge base (Tier 0/1); send email with confirmation (Tier 2); depends on Proton Mail Bridge Docker setup
-
System / service health — CPU/RAM/GPU/disk, temperatures, Docker container statuses (Tier 0, reactive); Diagnostic Agent (M6) slots in behind same tool interface
-
Open WebUI — Pipe function (
repos/quorra-openwebui-pipe/) calls QuorraAPI/chatvia trusted-service auth; appears as “Quorra” in model dropdown; identity via Authentik SSO → Pipe__user__→ email header -
Immich — photo search by person/date/location (Tier 0); lower priority
-
Home Assistant — deferred; minimal HA devices; revisit when PoE cameras are added
-
Review v1 service docs (
~/data/archive/quorra-v1/services/) for integration details, port mappings, and capability notes relevant to v2 service-integrations.md
Each integration: read interface first, write interface second. Each tool registers its tier at definition.
Definition of done: Can log a transaction, query calendar events, and check system health entirely through conversation. Actual Budget, calendar/tasks/reminders (Radicale CalDAV), and system health integrated. Planning scope operational with context triggers and reminders. Modes implemented and tested.
M5 — Overnight consolidation
Goal: Nightly consolidation job runs reliably and produces updated memory state and a readable report.
- systemd timer configured on JUNC1 (target: 2–4 AM)
- Active session check — job aborts cleanly if sessions are active
- Embedding refresh: new files in
~/data/knowledge/indexed since last run - Episodic summary generation from the day’s significant interactions
- Knowledge graph update pass (relationships, preferences inferred from recent events)
- Memory reconciliation — runs against the
memoriestable:- Deduplication (merge redundant memories saved in different sessions)
- Contradiction detection (flag or resolve conflicting memories)
- Stale memory cleanup (remove memories superseded by newer ones the LLM missed)
- Importance promotion (contextual → core for frequently-referenced facts)
- Memory synthesis (derive observations from daily interaction patterns)
- Reconciliation report included in nightly consolidation output
- LoRA training candidate generation from recent interaction window
- [Optional] LoRA fine-tune run if candidates exceed threshold
- Consolidation report written to the quorra-api data volume (
/data/, alongside the audit log; final path a M5 decision) - Each task confirmed idempotent (safe to re-run if interrupted)
- Job confirmed non-interfering with active sessions (tested overnight)
Definition of done: Job runs at 2 AM, produces a log entry, and the morning session shows updated memory state. Tested through at least one full overnight run.
M6 — Diagnostic agent
Goal: Separate diagnostic agent process running independently, monitoring JUNC1 health, surfacing summaries through Quorra.
- Diagnostic agent process defined and deployable independently of QuorraAPI
- Model selected (small quantized model on RTX 3070 or rule-based — see open questions)
- Quorra ↔ diagnostic IPC interface defined and implemented
- SMART data monitoring
- Thermal monitoring (GPU and CPU temperatures)
- Disk space monitoring
- Service health checks (all running Docker containers)
- At least one self-healing action implemented (e.g., restart crashed service)
- Health summaries surfaced to user via Quorra on request
- Diagnostic agent restarts independently without affecting QuorraAPI
Definition of done: Quorra can answer “is everything okay?” with a real system health summary. Diagnostic agent survives an independent restart. At least one self-healing action has been triggered and logged.
M7 — Web client (PWA)
Goal: Mobile-usable web interface as an alternative to the Matrix/Element chat interface. Started 2026-07-09 — design in architecture.md §4.1e (mobile-first web app first; PWA installability is the final, thin layer; offline/push deferred). Repo: repos/quorra-web/.
The workspace-face M1 vertical slice shipped from this repo 2026-07-25 (deployed at app.juncyard.com — chat + files + editor + approval inbox; see the Workspaces shipped-increments list in M4). M7’s household chat MVP continues in the same codebase; markers below reflect what that slice already delivered.
Phase A — shell + auth: ✅ (2026-07-25)
- Web app scaffolded (React + Vite + TS, oidc-client-ts; plain-CSS tokens rather than Tailwind — superseded 2026-08-19: migrating to Tailwind v4 + the neobrutalism.com registry, see decisions.md “Web client styling”)
- Authentik OIDC client registered (app slug
quorra-web-app, client_idquorra-web, public, code+PKCE,sub_mode=user_uuid); quorra-api JWT config repointed (AUTHENTIK_ISSUER/JWKS_URL/AUDIENCE) - OIDC login round-trip working; authenticated API calls from the browser (first client to exercise the JWT path)
- Served from JUNC1: compose service, container nginx static +
/apiproxy (SSE unbuffered), TLS at app.juncyard.com
Phase B — chat MVP:
- [~] QuorraAPI:
GET /sessions— the personal (null-workspace) cut shipped 2026-08-19 for Home, recency-capped, sharing one query with the workspace variant (sessions/summaries.py); a global cross-context list is still separate and unbuilt - Home — personal chat (2026-08-19):
/homeis a full-width conversation outside any workspace.ChatPanetakes aChatContextrather than aWorkspace, so one chat component serves both faces; personal sessions commit["general"]. Desktop lands here; the phone still lands on the Notepad. See decisions.md “Home is the personal chat surface” - [~] QuorraAPI: history loading — recent-history load shipped in the M1 slice; full-history variant of
/sessions/{id}/messages/recentstill pending - [~] New-session flow with scope picker (commits scope on first message) — both faces mint sessions from the SessionPicker; no picker exists because neither face lets the user choose (a workspace derives its scope, personal is a constant). A picker only becomes meaningful if personal capability stays client-chosen, which the containment-axes resolver removes
-
/chat/streamSSE rendering:statuschips (activity indicator),tokenaccumulation, markdown (react-markdown+remark-gfm, nodangerouslySetInnerHTML) - [~] Session sidebar + history view — per-workspace and personal header-dropdown pickers shipped (2026-07-26, 2026-08-19); global cross-context sidebar pending
- Generic rich tool-result rendering seam (photo results plug in when M4 Immich lands)
- Works on mobile browser (tested: iOS Safari, Android Chrome)
Phase C — structured UI:
- Tier 1/2 confirmation flows work on mobile (confirmation card; approve carries
confirmation_id) - Notifications outbox surface
- Purpose/mode controls
- Memory legibility UI (principle 7) — built 2026-07-28. Settings → Memory lists every memory the owner holds across personal, household and all workspaces (the seal is a cognition boundary, not an audit boundary), with plain-language provenance, a core-vs-contextual explainer, a tag-less-recall warning, and create/edit/archive/restore/permanent-delete. Backend: the M3 router modernized (
workspace/archived/q/limitparams,workspace_slug+archived_aton the wire, service-or-JWT auth) plus soft delete (memories.archived_at, migrationz0b1c2d3e4f5, nightlymemory/purge.py). First HTTP test coverage for these ten endpoints — which surfaced a live bug:MemoryUpdate.tags’ Ellipsis “sentinel” madetagsrequired in Pydantic v2, so every PUT that omitted it had been 422ing since M3.
Phase D — PWA layer:
- Installable as a PWA (manifest, icons, minimal app-shell service worker)
- Install tested on iOS Safari + Android Chrome
De-scoped from M7: photo search results inline (blocked on M4 Immich integration — the Phase B rendering seam is the M7 deliverable); Web Push (backlog; local-first tension — needs its own decision); offline data/sync (pointless until M8 remote access).
Definition of done: A household member can open the web app on a phone browser, authenticate, and have a full conversation with Quorra including calendar queries. Works without the Matrix/Element app. (Photo search joins the DoD when the M4 Immich integration exists.)
M8 — Remote access
Goal: Quorra accessible outside the home network. No unencrypted traffic. LAN access continues if relay goes down.
- Relay solution chosen (Tailscale, custom WireGuard, or equivalent)
- Remote access functional from outside home network (tested on cellular)
- TLS on all traffic confirmed
- LAN access continues if relay is unavailable (tested by disabling relay)
- Privacy audit: no user data transits relay unencrypted
- Power user self-hosting path documented
Definition of done: Can access Quorra from a phone on cellular with the same experience as on LAN. Relay outage does not break LAN access.
Backlog — unscheduled
Items not yet assigned to a milestone. Will be triaged as earlier milestones complete.
- Contacts migration from cloud services (CardDAV on Radicale — provider decided 2026-06-09, deferred; calendar migration is no longer backlog: planned in
plans/calendar-consolidation.md) - PDF-to-markdown conversion pipeline (OCR for scanned PDFs; improves RAG quality)
- Retrieval feedback loop — Quorra flags low-confidence retrievals, proposes knowledge file improvements
- LoRA fine-tuning pipeline (full training run tooling, not just candidate generation)
- Voice interface (Whisper + Piper on Raspberry Pi or similar)
- CSAM detection pipeline on Immich ingestion (hash-based; required before any photo sharing features)
- Security camera integration (Home Assistant + PoE cameras)
- The Grid architecture (peer-to-peer between junc nodes; design before build)
- Second JUNC node deployment (proof of concept for product readiness)
- Notifications: push to PWA or mobile when external events trigger Quorra
- CardDAV contacts on Radicale (it already speaks CardDAV — zero-cost when needed; Quorra reads via standard protocol)
- Review v1 PRFAQ (
~/data/archive/quorra-v1/docs/PRFAQ.md) — product vision doc with no v2 equivalent; consider porting to~/projects/quorra/docs/ - Review v1 milestones (
~/data/archive/quorra-v1/milestones.md) for any tasks or context worth carrying forward - LLM-guided scope-setup conversation: new session starts with scope NULL; model interviews user to determine scope before committing; pre-commit tool gate is the hook (no API contract changes needed)
- User-defined room scope: replace env-based
MATRIX_ROOM_SCOPE_MAPwith bot-DB-backed lookup so room admins can set/change scope through chat commands; already isolated behindresolve_scope() - Multi-scope sessions: users opt into combined scopes (e.g.
["finance", "planning"]); schema already supports list; work is UX and prompt tuning - Open WebUI scope selector: surface scope choice in the Pipe UI so OWUI sessions can start outside
general - Matrix no-markdown compliance (
test_no_bullet_lists_in_matrix): flaky on 14B — stronger formatting directive needed inprompts/origin/matrix.mdand/or a few-shot example - Test debt:
sync_idparam mismatch intests/finance/test_actual_tools.py;participantsparam mismatch intests/test_executor.py - Session-active account default (
switch_account): if user’s dominant case is always the same account, a session-active default follows theswitch_budgetpattern