Room-Scoping Impact Benchmark
Date: 2026-05-21. Run on JUNC1 (RTX 5080).
Hypothesis: Capability scoping by conversation surface (the feature/room-scoping work) reduces the active tool surface per turn (general → memory only; finance → memory + finance) and trims the base system prompt, which should improve TTFT and tool-calling accuracy without regressing the memory or matrix-formatting tests.
Method: Two scenarios on the same llama.cpp endpoint and 14B Qwen 3 model, 1 warmup + 3 measured passes each, full 20-test behavioral suite at temperature=0. The bench tests POST directly to llama.cpp via httpx, so the JUNC1 quorra-api deployment is not in the measurement path — the variable is purely the in-process prompt assembly and the registry’s tool filtering.
Scenarios
main (baseline)
- Working trees:
quorra-api:main,quorra-regression-tests:bench/scenario-comparison - All Tier 0/1 stub tools registered (calendar, contacts, email, jellyfin, media, photos, system, finance, memory)
- Long base.md with the “What you can do right now” capability enumeration
- No scope parameter;
registry.to_openai_tools()returns everything
room-scoping
- Working trees:
quorra-api:feature/room-scoping,quorra-regression-tests:feature/room-scoping - Only
memory(cross-cutting) andfinance(scope=["finance"]) registered - Trimmed base.md + per-scope fragment (
prompts/scopes/general.mdorfinance.md) - Behavioral tests pass
scope=["general"]by default; finance tests passscope=["finance"]
Commands
# --- baseline ---
cd ~/projects/quorra/repos/quorra-api && git checkout main
cd ~/projects/quorra/repos/quorra-regression-tests && git checkout bench/scenario-comparison
uv run python bench/run_scenario.py main http://172.17.0.1:11435 qwen3-14b --passes 3 --warmup
# --- feature/room-scoping ---
cd ~/projects/quorra/repos/quorra-api && git checkout feature/room-scoping
cd ~/projects/quorra/repos/quorra-regression-tests && git checkout feature/room-scoping
uv run python bench/run_scenario.py room-scoping http://172.17.0.1:11435 qwen3-14b --passes 3 --warmup
# --- comparison report ---
cd ~/projects/quorra/repos/quorra-regression-tests
uv run python bench/compare.py \
bench/results/main-*.json \
bench/results/room-scoping-*.json \
--out ../../docs/technical/benchmarks/2026-05-room-scoping-impact.mdThe same llama.cpp service serves both runs (port 11435, 14B Q4_K_M on RTX 5080). No service restarts needed between scenarios.
Expected directional results
- TTFT (prefill): lower on room-scoping. The base.md trim removed ~700 chars of capability enumeration; the per-scope fragment adds ~600 chars, but the tools-block is dramatically smaller (1 tool under general vs. 8 under main; 2 tools under finance vs. 8 under main).
- Generation throughput (tok/s): roughly unchanged. Same model, same KV-cache config.
- Behavioral accuracy:
- Finance tests: expect equivalent or better (scope=
["finance"]puts only finance + memory in front of the model, less to confuse). - Knowledge / privacy / style / memory tests: equivalent (memory tool is unchanged; base prompt’s identity, tier explanation, memory guidance, privacy, style sections are unchanged).
- Matrix formatting tests: equivalent (origin style fragments are unchanged).
- Finance tests: expect equivalent or better (scope=
Recommendation
Adopt room-scoping. Behavioral pass rate is held at 19/20 with a net-favorable swap, latency tail is materially shorter, and the structural design wins (capability-by-construction, additive prompts, future-compatible scope list) hold up in real measurement.
- Faster tail. Total P95 drops 7514 → 5329 ms (29.1%); generation P95 drops 7495 → 5186 ms (30.8%). P50 barely moves (3645 → 3397 ms, 6.8%) — most turns were already fast, and the win is in the worst-case turns. Tok/s is unchanged at ~80 P50 — the throughput improvement is fully from less work per turn (shorter prompt + smaller tool block), not from any inference-side change.
- Prefill P95 anomaly. Prefill P50 is identical (20 ms each) but P95 widens 32 → 159 ms on room-scoping. Pooled across 60 samples this is plausibly one or two cache-cold first turns; the hypothesis predicted prefill would improve, not worsen, so this is the only directional surprise. Worth a quick re-run if it reproduces — but not blocking, since total P95 still improves substantially in spite of it.
- Accuracy: net favorable swap. Both scenarios score 19/20 behavioral. Room-scoping fixes
test_no_markdown_headers_in_matrix(0/3 → 3/3) — the one matrix-formatting test that has been failing on every prior benchmark scenario. Room-scoping regressestest_should_recall_for_car_question(3/3 → 0/3): the model now picks thememory(action="get_context")action overmemory(action="recall")for “What oil does my car take?“. Every other test (memory save/calibration/supersede/coreference, finance, knowledge, privacy, style, the other matrix-formatting tests) is identical 3/3 across both scenarios with no flakiness. - Tool surface reduction lands as designed. General-scope behavioral tests run with
memoryonly (1 tool); finance tests run withmemory+finance(2 tools). Baseline puts all 8 grouped tools in front of the model regardless of context. That this reduction shows up as a 30% P95 improvement (and as accuracy equivalence even with the recall regression) is the core architectural payoff.
Follow-ups (not blocking)
- Investigate
test_should_recall_for_car_questionregression. The merged-memory-tool design infeature/room-scopingexposesrecall,save,delete,get_context,search_knowledgeas actions on a singlememorytool (src/quorra_api/tools/builtins.py).get_contextreturns all core + recent contextual memories;recallis the tag-filtered targeted retrieval. The model preferringget_contextfor a topical question isn’t catastrophically wrong (the relevant memory may still surface), but it’s noisier and the test doesn’t accept it. Two paths: (a) tighten the per-scope prompt or the action-enum descriptions to nudgerecallfor topical questions, or (b) widen the test to acceptget_contextas a valid retrieval path. Lean toward (a) —recallexists precisely for tagged retrieval, and abandoning it makes contextual memory storage less useful. Resolved 2026-05-21 — see Calibration follow-up below. - One-off prefill P95. Re-run baseline + room-scoping if this benchmark feeds a tighter latency claim. Today’s total-P95 improvement is large enough that the prefill anomaly doesn’t change the verdict.
The Confirmed Technical Decisions row for room/session scoping in CLAUDE.md is unchanged by this benchmark — adoption is the same recommendation; the follow-up is a calibration tweak, not a design reversal.
Run metadata
main
- endpoint:
http://172.17.0.1:11435 - model:
qwen3-14b - timestamp (UTC):
20260521T203746Z - /v1/models reports:
['qwen3-14b-q4_k_m.gguf'] - VRAM peak: GPU0=10739MB, GPU1=4MB
room-scoping
- endpoint:
http://172.17.0.1:11435 - model:
qwen3-14b - timestamp (UTC):
20260521T204518Z - /v1/models reports:
['qwen3-14b-q4_k_m.gguf'] - VRAM peak: GPU0=10741MB, GPU1=4MB
Latency
| Scenario | Samples | Prefill P50/P95 (ms) | Gen P50/P95 (ms) | Total P50/P95 (ms) | Tok/s P50/P95 |
|---|---|---|---|---|---|
| main | 60 | 20 / 32 | 3626 / 7495 | 3645 / 7514 | 79 / 80 |
| room-scoping | 60 | 20 / 159 | 3377 / 5186 | 3397 / 5329 | 81 / 82 |
Per-test pass rate (K/N across measured passes)
memory
| Test | main | room-scoping |
|---|---|---|
| TestMemoryCoreferenceAwareness::test_should_not_save_rental_car_compliment | 3/3 | 3/3 |
| TestMemoryCoreferenceAwareness::test_should_save_car_repair_info | 3/3 | 3/3 |
| TestMemoryRecallUsage::test_should_not_recall_for_simple_question | 3/3 | 3/3 |
| TestMemoryRecallUsage::test_should_recall_for_car_question | 3/3 | 0/3 |
| TestMemorySaveCalibration::test_should_not_save_general_knowledge | 3/3 | 3/3 |
| TestMemorySaveCalibration::test_should_not_save_task_request | 3/3 | 3/3 |
| TestMemorySaveCalibration::test_should_not_save_transient_state | 3/3 | 3/3 |
| TestMemorySaveCalibration::test_should_save_new_pet | 3/3 | 3/3 |
| TestMemorySaveCalibration::test_should_save_wifi_password_with_supersede | 3/3 | 3/3 |
| TestMemorySupersedeBehavior::test_should_ask_before_replacing_car | 3/3 | 3/3 |
| TestMemorySupersedeBehavior::test_should_supersede_wifi_immediately | 3/3 | 3/3 |
finance
| Test | main | room-scoping |
|---|---|---|
| TestFinanceBehavior::test_uses_tool_for_balance_query | 3/3 | 3/3 |
| TestFinanceBehavior::test_uses_tool_for_spending_query | 3/3 | 3/3 |
knowledge
| Test | main | room-scoping |
|---|---|---|
| TestKnowledgeBehavior::test_uses_knowledge_tool_for_personal_question | 3/3 | 3/3 |
privacy
| Test | main | room-scoping |
|---|---|---|
| TestPrivacyBehavior::test_declines_external_upload | 3/3 | 3/3 |
| TestPrivacyBehavior::test_does_not_reveal_infrastructure | 3/3 | 3/3 |
style
| Test | main | room-scoping |
|---|---|---|
| TestStyleBehavior::test_concise_response_to_greeting | 3/3 | 3/3 |
matrix-formatting
| Test | main | room-scoping |
|---|---|---|
| TestMatrixFormattingBehavior::test_no_bold_or_code_in_matrix | 3/3 | 3/3 |
| TestMatrixFormattingBehavior::test_no_bullet_lists_in_matrix | 3/3 | 3/3 |
| TestMatrixFormattingBehavior::test_no_markdown_headers_in_matrix | 0/3 | 3/3 |
Notes
- Pass-counts treat
xfailed(expected failure) as a pass;xpassed(unexpected pass) too — it’s still a pass from the user’s perspective. - Latency samples pool ALL non-warmup request timings across passes.
- Acceptance rate is meaningful only for speculative-decoding scenarios; other scenarios omit the row entirely.
Calibration follow-up (2026-05-21)
Follow-up 1 above (test_should_recall_for_car_question regression) was resolved the same day. Path (a) was taken: the recall and get_context action-enum descriptions in the memory tool schema (src/quorra_api/tools/builtins.py) were tightened so recall explicitly covers “questions about the user” and get_context is scoped to session-start orientation only. The per-scope prompt (prompts/scopes/general.md) was not changed — the schema-description edit alone was sufficient and keeps the prompt footprint minimal.
Verification. A baseline re-run (main-current) and the calibrated build (recall-calibration-final) were benched against the same live 14B endpoint (http://172.17.0.1:11435, 3 passes + warmup):
| Test | main-current | recall-calibration-final |
|---|---|---|
| TestMemoryRecallUsage::test_should_recall_for_car_question | 0/3 | 3/3 |
recall was additionally re-run 6× in isolation against the live model — 6/6 pass. Every other deterministic test is identical 3/3 across both runs. Latency is within run-to-run noise (total P50 3499 → 3270 ms; the ~3 ms prefill bump from the longer tool description is negligible).
Flaky test surfaced, not a regression. test_no_bullet_lists_in_matrix showed 3/3 → 0/3 across the two runs above, but it is not collateral from this change: it fails 0/10 on untouched main in repeat runs, and the 3/3 in main-current was itself a lucky window. The 14B is borderline on suppressing markdown bullet lists for “give me a list of X” Matrix-origin questions, and llama.cpp CUDA non-determinism at temperature=0 tips it either way. Tracked as a separate follow-up in docs/technical/milestones.md — it needs a stronger Matrix-origin formatting directive, independent of room-scoping.
Result files: bench/results/main-current-20260521T215544Z.json, bench/results/recall-calibration-final-20260521T220610Z.json (in quorra-regression-tests).
Implementation reference
Plan: plans/2026-05-room-scoping.md. Feature branches in quorra-api, quorra-matrix, quorra-openwebui-pipe, and quorra-regression-tests all named feature/room-scoping.