Quorra — Knowledge Base Kernel

Last updated: 2026-06-16 (initial design)

The kernel is the durable, reusable core that turns raw vault markdown into a validated, schema-conformant representation. It is intentionally small and dependency-light, because it is reused by three different consumers. This document is the developer-facing design; the user/Quorra-facing schema it enforces is knowledge-base-schema.md.

Why a kernel: shared core, divergent shells

The knowledge-base reconciliation effort splits into three sessions — (1) design the kernel, (2) migrate the legacy corpus once, (3) build the production reconciliation feature. A key design decision: the kernel is shared across those jobs (and M2), while the shells around it diverge and are deliberately not merged.

Shared kernelDivergent shell
Parser (markdown → structure)✅ all consumers—
Deterministic validator (conformance → issues)✅ all consumers—
Schema spec (the target)✅ all consumers—
Atomic LLM fixes (infer tags, write lead)✅ largely shared—
Link-graph + bulk rename/move + 1000-link rewrite—migration-only
Bulk orchestration, legacy-type remap, mass eviction—migration-only
Overnight loop, write-gate, inotify watcher—production-only

The overlap that’s worth sharing (parser + validator + schema) costs ~nothing extra to share — you’d build it for the migration anyway, and M2 needs it too. The migration’s heavy structural machinery is one-time and should be a pragmatic script that consumes the kernel, not generalized into it. See decisions.md.

The three consumers

  • Migration (session 2): runs the kernel over the whole corpus once; auto-applies deterministic fixes, routes judgment issues to the LLM, queues human issues for review. Supervised, reversible (git), bulk.
  • Production reconciliation (session 3): the overnight direct-edit audit and the authoring write-gate. Trickle, unattended, conservative. The write-gate blocks a commit on any error.
  • M2 RAG indexer: reuses the parser (for chunking) and the indexing rule (for inclusion/exclusion).

Parser contract

Pure, no I/O — the caller reads the file and passes text; the parser returns structure. This is the one representation every consumer (validator, migration, M2 chunker) operates on.

parse(path, text) -> VaultDoc {
  path:          str
  frontmatter:   dict          # parsed YAML, or {} if absent
  body:          str           # markdown after the frontmatter
  sections:      [Section{ level: 2|3, heading: str, text: str, word_count: int }]   # split at H2/H3
  links:         [Link{ kind: wikilink|embed|markdown, target: str, alias: str|None }]
  word_count:    int           # prose only (excludes frontmatter and headings)
  has_frontmatter: bool
}

New dependencies (none exist in the codebase today): python-frontmatter + pyyaml, plus a light markdown parser (mistune) or regex for section/link extraction. Mirror the self-contained subsystem layout of integrations/caldav/ (pure mapping-style modules separated from I/O).

Validator rule set

Pure function validate(doc: VaultDoc) -> [Issue]. The validator only reports; it never mutates. Each issue is tagged so a driver knows who fixes it.

Issue {
  code:          str            # stable identifier, e.g. "missing-required-field"
  field:         str | None     # the frontmatter field or structural element
  severity:      error | warning
  kind:          deterministic | judgment | human
  message:       str
  suggested_fix: str | None
}
kindMeaningHandled by
deterministicmechanically fixable, no inferencemigration auto-fix
judgmentneeds inferenceLLM remediation
humanmust not be guessedflagged for review

Representative rules:

RuleSeverityKind
missing required fielderrorprovenance (created_*/updated_*) = deterministic; tags/type = judgment
field value not in valid set (type not in taxonomy, malformed dates, tags not lowercase-hyphenated)errorformat = deterministic; reclassification = judgment
filename not kebab-caseerrordeterministic
checkbox / task content in bodyerrorhuman (evict → planning DB)
section > ~200 wordswarningjudgment (split)
relative date expressions (“recently”, “currently”)warningjudgment
no summary-lead / pronoun-subject sectionwarningjudgment

Drivers set policy on top of the same issue list:

  • write-gate → block on any error.
  • migration → apply deterministic, send judgment to the LLM, queue human.
  • M2 → consumes only the indexing rule (status + word-floor + not-metadata), not the full set.

Convergence/idempotency is a driver guarantee, not a kernel one: LLM remediation fires only on a specific validator failure, targets that defect, and stops when the file passes — no open-ended “improve this.” That makes repeated runs settle.

Forward constraints to honour

  • Encryption = ownership = index boundary. Ownership is by directory location (schema §3). A member’s vectors can leak content (embedding inversion), so M2 must keep member content in per-member index namespaces inside the same boundary as the files; household/ is the shared domain. Don’t design M2 as one commingled collection.
  • Stable cross-store join id. A project/entity view joins the vault record with planning-DB tasks (X-QUORRA-PROJECT), logs, etc. The join needs a stable identifier (stable id + slug alias, or slug-as-key with reconciliation-tooling rename-cascade) so renames don’t break references — the recurring UUID-vs-slug pattern in this project.
  • Compose once. The “project view” join should be a single shared aggregation service consumed by both Quorra (as a tool) and a future UI (as a data source) — not implemented twice.

Resolved

  • Journaling is composition, not a third store. Journals are type: log (no type: journal): task dailies → CalDAV, reflective prose → vault type: log, distilled significance → episodic memory; the “journal for date X” view composes the three on demand. The validator suppresses the reference-prose warnings for journal logs (first-person by nature), and personal dailies are distill-but-don’t-raw-index (≥100-word indexing floor still applies). See decisions.md → “Journaling is composition, not a third store”.