Quorra — Knowledge Base Schema
Last updated: 2026-06-16 (kernel redesign — provenance quad, ownership-by-location, confidence moved to the DB, preference dropped / pet added)
This document is the canonical schema reference for ~/data/knowledge/. It governs how Quorra writes vault files and how the M2 embedding pipeline indexes them. Since Quorra is the primary author of the vault, the write format and retrieval format are co-designed here.
This describes the target. The live vault has not been migrated to this schema yet — it is the spec the one-time migration (and the production reconciliation tool) will enforce. See kb-kernel.md for the validator/parser that checks conformance, and decisions.md for the rationale behind the kernel decisions.
1. Frontmatter standard
Every vault file Quorra creates or edits must include the following YAML frontmatter block at the top of the file.
Required fields
The author (Quorra) really only chooses type and tags — the provenance quad is auto-filled at write time.
| Field | Type | Valid values | Notes |
|---|---|---|---|
type | string | see §2 taxonomy | Document class — determines routing, chunking, and consolidation behaviour. A class, never a topic. |
tags | list | lowercase strings | At least 1 required; hyphens for multi-word (home-maintenance). Topics live here (health, finance, travel), never in type. |
created_by | string | quorra | user | Who authored the file |
created_at | date | ISO 8601 YYYY-MM-DD | Creation date |
updated_by | string | quorra | user | Who last significantly edited it |
updated_at | date | ISO 8601 YYYY-MM-DD | Bumped on every write, even minor edits |
Optional fields
| Field | Type | Notes |
|---|---|---|
status | string | active (default) | draft | archived — indexing lifecycle of the file (see §1.1); archived is excluded from RAG, draft is indexed but flagged |
sensitivity | string | routine (default) | private | restricted — privacy gradation within an owner |
entities | list | Cross-references to other vault files by relative path; used by overnight consolidation to maintain the knowledge graph |
source | string | Provenance of the facts — where the information came from ("car manual", "conversation 2026-06-03") |
expires_at | date | ISO 8601 — hint to overnight consolidation that this content may be stale after this date |
There is no scope field (ownership is by location — see §3) and no confidence field
(epistemic stance lives in the DB — see §1.2).
Minimal example
---
type: vehicle
tags: [vehicles, subaru, maintenance]
created_by: quorra
created_at: 2026-06-16
updated_by: quorra
updated_at: 2026-06-16
---1.1 status means indexing lifecycle, not subject state
status answers “should this file be indexed?” — not “how is the project/thing going?“. A finished
garage renovation keeps status: active (it is current reference knowledge) until someone archives it;
its completion is expressed in prose and in its task state, not in status. Don’t conflate the two.
1.2 No confidence field — epistemic stance is DB-only
stated/observed/inferred describe Quorra’s cognition about the user, and that lives in the
DB memories table (the source column: user-explicit | llm-inferred | imported). The vault holds
asserted or sourced knowledge; a tentative inference about the user is a memory, not a vault file.
Provenance on a vault file is carried by created_by (+ source for an external artifact).
2. Document type taxonomy
type is a document class — what the file is. What the file is about (its topic) goes in tags.
There are 11 classes:
| Category | Types |
|---|---|
| Entity Knowledge | person, pet, vehicle, property, project, business, account |
| Episodic | event, log |
| Reference | reference, note |
type | Subject | Example |
|---|---|---|
person | A named person | Profile for a family member, friend, or contact |
pet | A named animal | The household dog’s profile, vet/care notes |
vehicle | A vehicle | Ownership, specs, maintenance history |
property | A home or property | Address, features, maintenance records |
project | An initiative | The reference record of a project (see §6) |
business | A company or service | Account details, notes, contact info |
account | A financial or service account | Account notes, credentials context |
event | A significant event | A trip itinerary, a milestone, a planned gathering |
log | A time-indexed record | Trip log, maintenance log, journal entry, progress log |
reference | A document or resource | Guide, manual, recipe, research, saved article |
note | Free-form content | Anything that genuinely doesn’t fit the above |
preference is not a type. Preferences are DB-primary (memories table), not vault files — see §8.
Topics are not types. health, finance, travel, legal are tags; the class comes from what
the file is (a reference doc, a log, an account), not what it’s about.
3. Ownership is by location
A file’s owner — who it is visible to — is determined by where it lives, not by a frontmatter field. The vault is partitioned by directory, and the path is the single source of truth for ownership:
household/…→household(visible to all household members)members/<slug>/knowledge/…→member:<slug>(private to that member)
The M2 indexer derives an owner value from the path into the retrieval index payload (call it
owner/visibility in the index — deliberately not “scope”, which is the unrelated session/tool
feature). Household-vs-member is a storage partition; the retrieval overlap a member experiences
(household ∪ member:<self>) is a query-time filter over one index, not a per-file field.
Why location and not a field: a scope field would be a second source of truth for something the path
already encodes (drift risk), and it collides with the existing session-scope feature. Location-as-truth
also aligns the ownership boundary with the future per-user encryption boundary and the index
namespace — a member’s files and their (content-leaking) vectors sit inside the same sealed boundary,
while household/ is the shared domain. To change a file’s ownership, you move it.
Finer-grained privacy uses sensitivity (within an owner) and shared/with-<slug>/ paths (explicit
cross-member sharing).
4. Authoring conventions
These rules govern how Quorra writes vault file content. They exist to make every chunk semantically dense and self-contained when retrieved in isolation — the retriever never has the surrounding file for context.
1. Self-contained sections
Every H2 and H3 section must make sense read alone. Use full entity names (“Kurt’s 2019 Subaru Outback”, not “it”); repeat key context rather than relying on an earlier section; never use a pronoun as the primary subject.
2. One topic per section
Each H2/H3 section covers exactly one fact cluster. If a section starts covering a second topic, split it.
3. Absolute dates
Never write relative time (“recently”, “currently”, “last year”). Always use explicit dates or year-month qualifiers (“As of 2026-05…”, “Purchased 2024-03”).
4. Prose over structure — and never task content
Section bodies are prose paragraphs. Bullet lists are acceptable only for genuinely enumerable items (contact details, step-by-step instructions, short unordered lists). Do not use bullets for narrative facts. Never use checkbox/task lists in the vault — operational tasks live in the planning DB, not the vault. Checkbox content in a vault file is a hard validator error (it gets evicted to the DB).
5. Lead with the summary sentence
The first sentence of every section states the key fact; the rest elaborates. This makes the first sentence alone a useful retrieval chunk.
6. Keep tentative inferences out of the vault
The vault holds asserted or sourced knowledge. If something is a tentative inference about the user, it belongs in the DB memories table (with source: llm-inferred), not a vault file. When a vault file does include a qualified statement, say so in prose (“As of 2026-05, based on the owner’s manual…”) rather than asserting uncertainty as fact.
7. Section length
Target 50–200 words per H3 section. If a section exceeds ~200 words, split it into sub-sections. Files longer than ~500 lines should be split using the entity-folder pattern (§5).
5. File and folder naming conventions
- Filenames: kebab-case, all lowercase.
subaru-outback.md,mom.md,europe-trip-2026.md - No spaces or uppercase in any filename or directory name (exception:
README.md) - Entity folders: when a single entity accumulates 2 or more document types, promote it to a folder:
- Before:
subaru-outback.md - After:
subaru-outback/overview.md,subaru-outback/maintenance-log.md
- Before:
- Date-prefixed logs:
YYYY-MM-DD-description.mdfor time-indexed records - README.md files are directory metadata only — not indexed by the RAG pipeline
6. The composition model — projects and rich entities
A rich entity (a project, a vehicle, a person) is not one monolithic file. It is a composition across stores, joined by a stable identifier. A garage-renovation project, for example:
- Vault record (
type: project) — the durable reference: what the project is, goals, key decisions, and a prose “current state” section Quorra refreshes. Prose only — no checklists, no shopping lists. - Tasks and purchase lists → the planning DB (CalDAV), tagged to the project (
X-QUORRA-PROJECT). These have real due dates, completion, and reminders — state that embeddings cannot represent. - Progress updates →
type: logentries (retrieval-worthy milestones only). - The unified “everything about the project” view is composed on demand, not stored — Quorra joins the vault record + the project’s open tasks + recent log entries. A future UI renders the same join as a project page.
On disk, entity-folder promotion applies:
members/kurt/knowledge/projects/garage-renovation/
├── overview.md # type: project — the reference record
└── progress-log.md # type: log — milestone updates
The cross-store join key is the project identifier (e.g. the slug garage-renovation, shared by the
vault folder and the tasks’ X-QUORRA-PROJECT). The composition is the same whether Quorra assembles it
conversationally or a UI renders it — build the join once, as a shared service.
7. What the embedding pipeline indexes
Indexing is a deterministic rule. There is no “content-class” field — operational/task content doesn’t belong in the vault by design (it lives in the planning DB), so it never reaches the indexer.
Indexed =
status∈ {active,draft} and ≥ ~100 words of prose and not aREADMEortemplates/file. Excluded =status: archived, sub-threshold files, and directory/template metadata.
draft files are indexed but flagged (the status payload lets Quorra surface them as working drafts).
Chunking strategy (for M2 implementation)
Chunk at H2/H3 section boundaries. Each chunk’s index payload should include:
file_path— relative path from vault roottitle— H1 heading of the filesection— the H2/H3 heading for this chunkparent_section— the H2 heading when the chunk is H3-level (null otherwise)owner— derived from the file path (household|member:<slug>), the retrieval privacy filter- All frontmatter fields:
type,tags,status,sensitivity,created_at,updated_at,source
This payload enables pre-filtering before semantic search — e.g. “find type: vehicle files tagged
maintenance owned by member:kurt” runs as an index filter, not a semantic query.
8. Vault vs. memory decision rule
The vault and the DB memory system are complementary. The test:
Does the system act on this data, or refer to it?
- Act on it → DB (
memoriestable)- Refer to it → vault (indexed by RAG)
Vault: reference documents, entity records, itineraries, guides, logs. Things Quorra and users read to get information.
Memory: preferences, operational facts, observations, session context. Things the system uses to
decide what to do or say — including the epistemic stance (stated/observed/inferred) the vault no
longer carries.
Preferences are DB-primary: a type: preference vault file is not allowed (it would re-create the
dual-source-of-truth the agent-memory retirement eliminated).
9. Example compliant file
---
type: vehicle
tags: [vehicles, subaru, maintenance, fuel, insurance]
status: active
sensitivity: routine
created_by: quorra
created_at: 2026-06-16
updated_by: quorra
updated_at: 2026-06-16
entities:
- household/vehicles/insurance.md
source: conversation 2026-06-16
---
# Kurt's 2019 Subaru Outback
## Overview
Kurt's daily driver is a 2019 Subaru Outback in Magnetite Gray Metallic. It was purchased new in March 2019 from Subaru of Portland. As of 2026-06, the odometer reads approximately 74,000 miles.
## Fuel
Kurt's 2019 Subaru Outback takes regular 87 octane fuel. The owner's manual confirms that premium is not required and provides no measurable benefit for this engine. Kurt fills up roughly once per week, typically at Costco.
## Insurance
Kurt's Subaru Outback is insured through Progressive under a policy held jointly with the household. The policy renews annually in September. Full details are in household/vehicles/insurance.md.
## Maintenance notes
The Subaru Outback is on a synthetic oil change interval of 6,000 miles or 6 months, whichever comes first. As of 2026-04, the last oil change was performed at 71,200 miles. The next service is due around 77,200 miles or October 2026.This file is owned by member:kurt because of its location (members/kurt/knowledge/…), not a field.