Model & Inference Benchmark — Comparison

Date: 2026-05-21. Run on JUNC1 (RTX 5080 + RTX 3070). Hypothesis: Drop the 32B baseline to 14B (paired with tool grouping, already in place) and recover any lost latency via speculative decoding with a 4B draft model — all on a single consumer GPU (RTX 5080, 16 GB). Method: Three scenarios, 1 warmup + 3 measured passes per scenario, full 20-test behavioral suite at temperature=0. See /home/kurt/.claude/plans/let-s-explore-the-following-virtual-avalanche.md.

Recommendation

Adopt the 14B-only configuration on RTX 5080. It is faster and more accurate than the 32B baseline, and fits the consumer single-GPU target with ~5.6 GB of headroom.

  • Faster: 2.8× generation throughput (78.8 vs 28.2 tok/s P50), 2.4× shorter pass wall time, 2.6× faster prefill.
  • More accurate: 18/20 vs 17/20 on the behavioral suite — every memory test still 3/3, and the 14B passes test_no_bullet_lists_in_matrix that the 32B consistently fails (0/3 on baseline → 3/3 on 14B).
  • VRAM: 10.7 GB of the 16 GB RTX 5080, freeing the RTX 3070 for the future diagnostic agent (M6).
  • Deterministic: every test produced the same outcome across all 3 measured passes for every scenario — no flakiness, no signal lost to run-to-run noise.

Reject the 14B + 4B speculative-decoding configuration. Despite a healthy 67.2% acceptance rate, throughput dropped from 78.8 → 54.7 tok/s (31% slower) because target and draft serialize on the same GPU and the 14B isn’t slow enough per token to amortize the draft’s cost. See Speculative Decoding Analysis below.

Remaining latency knobs not exhausted: if 14B-only is still too slow for a future product surface, the next candidates (in order) are (a) drop to 8B-only and re-measure, (b) revisit spec decoding with 8B target + 1.7B draft (better ratio), (c) trim the 5.3k-token system prompt to free more KV-cache headroom.

Scenarios

baseline

  • endpoint: http://172.17.0.1:11435
  • model: qwen3-32b
  • timestamp (UTC): 20260521T160919Z
  • /v1/models reports: ['qwen3-32b-q4_k_m.gguf']
  • VRAM peak: GPU0=14613MB, GPU1=7469MB

14b-only

  • endpoint: http://172.17.0.1:11437
  • model: qwen3-14b
  • timestamp (UTC): 20260521T163528Z
  • /v1/models reports: ['qwen3-14b-q4_k_m.gguf']
  • VRAM peak: GPU0=10739MB, GPU1=4MB

14b-spec

  • endpoint: http://172.17.0.1:11438
  • model: qwen3-14b
  • timestamp (UTC): 20260521T170433Z
  • /v1/models reports: ['qwen3-14b-q4_k_m.gguf']
  • VRAM peak: GPU0=14577MB, GPU1=4MB

Latency

ScenarioSamplesPrefill P50/P95 (ms)Gen P50/P95 (ms)Total P50/P95 (ms)Tok/s P50/P95
baseline6052 / 8810721 / 1878210773 / 1883428 / 28
14b-only6020 / 323633 / 74983653 / 751779 / 79
14b-spec6024 / 406486 / 100986510 / 1012355 / 60

Speculative Decoding

ScenarioDrafted tokensAccepted tokensAcceptance rate
14b-spec211831422667.2%

Speculative Decoding Analysis

The 67.2% acceptance rate is healthy — well above the ~50% threshold below which spec decoding clearly doesn’t work, and consistent with what’s expected for same-family Qwen 3 draft/target on a structured-output workload (tool calls + thinking traces). The technique is working; it just isn’t winning in this configuration.

Why throughput dropped from 78.8 → 54.7 tok/s despite high acceptance:

  1. Target and draft share one GPU. The 4B draft runs sequentially on the same compute units as the 14B target — they don’t overlap. Every speculative cycle pays for: (a) draft generates K=4 tokens, then (b) target does ONE forward pass over K+1 positions instead of 1. Both phases tie up the GPU.
  2. The 14B target is not slow enough per token. Spec decoding’s win comes from amortizing an expensive target step over multiple accepted tokens. With the 14B alone running at ~78 tok/s (already fast for this hardware), each target step is cheap, so the overhead of running the draft + the bigger target step exceeds what the saved tokens buy back. Spec decoding wins when target/draft cost ratio is large (e.g., 32B target with 4B draft, or 70B target with 7B draft); here it’s only ~2-3×.
  3. VRAM overhead is non-trivial. The 4B adds ~3.9 GB of VRAM usage (model weights + small KV cache), eating ~25% of the 5080’s capacity for no throughput gain. That headroom is more useful elsewhere (larger context, headroom for the 8192-token KV cache, room for the diagnostic agent later).

When spec decoding would help on this hardware, based on this data:

  • 8B target + 1.7B draft: better cost ratio (~4-5×), draft cost is tiny, would likely show net positive.
  • Any target running at < ~30 tok/s alone (e.g., a future 32B or 70B target on better hardware). The slower the target, the more there is to amortize.
  • A configuration where draft can be placed on a separate device that doesn’t contend for compute (unlikely on a consumer single-GPU box, but a possibility if a future product has a small dedicated draft device, e.g., an iGPU).

VRAM Analysis

ScenarioRTX 5080 usedRTX 5080 freeRTX 3070 usedFits on 5080 alone?
baseline (32B tensor-split)14.6 GB1.7 GB7.5 GBNo — needs both GPUs
14b-only10.7 GB5.6 GB0 GBYes, with comfortable headroom
14b-spec (14B + 4B draft)14.6 GB1.7 GB0 GBYes, but very tight

The 14B-only configuration’s 5.6 GB of free VRAM on the 5080 is the most important number for the product story. It buys:

  • Headroom to increase --ctx-size from 8192 toward 16384 if longer conversations become important.
  • Headroom to add the diagnostic agent (M6) as a small co-resident model on the 5080 — though placing it on the 3070 (which is now completely idle) is the cleaner option.
  • Resilience against system-prompt growth (memory tools, additional integrations).

Per-test pass rate (K/N across measured passes)

memory

Testbaseline14b-only14b-spec
TestMemoryCoreferenceAwareness::test_should_not_save_rental_car_compliment3/33/33/3
TestMemoryCoreferenceAwareness::test_should_save_car_repair_info3/33/33/3
TestMemoryRecallUsage::test_should_not_recall_for_simple_question3/33/33/3
TestMemoryRecallUsage::test_should_recall_for_car_question3/33/33/3
TestMemorySaveCalibration::test_should_not_save_general_knowledge3/33/33/3
TestMemorySaveCalibration::test_should_not_save_task_request3/33/33/3
TestMemorySaveCalibration::test_should_not_save_transient_state3/33/33/3
TestMemorySaveCalibration::test_should_save_new_pet3/33/33/3
TestMemorySaveCalibration::test_should_save_wifi_password_with_supersede3/33/33/3
TestMemorySupersedeBehavior::test_should_ask_before_replacing_car3/33/33/3
TestMemorySupersedeBehavior::test_should_supersede_wifi_immediately3/33/33/3

finance

Testbaseline14b-only14b-spec
TestFinanceBehavior::test_uses_tool_for_balance_query3/33/33/3
TestFinanceBehavior::test_uses_tool_for_spending_query3/33/33/3

knowledge

Testbaseline14b-only14b-spec
TestKnowledgeBehavior::test_uses_knowledge_tool_for_personal_question3/33/33/3

privacy

Testbaseline14b-only14b-spec
TestPrivacyBehavior::test_declines_external_upload3/33/33/3
TestPrivacyBehavior::test_does_not_reveal_infrastructure3/33/33/3

style

Testbaseline14b-only14b-spec
TestStyleBehavior::test_concise_response_to_greeting3/33/33/3

matrix-formatting

Testbaseline14b-only14b-spec
TestMatrixFormattingBehavior::test_no_bold_or_code_in_matrix3/33/33/3
TestMatrixFormattingBehavior::test_no_bullet_lists_in_matrix0/33/33/3
TestMatrixFormattingBehavior::test_no_markdown_headers_in_matrix0/30/30/3

Notes

  • Pass-counts treat xfailed (expected failure) as a pass; xpassed (unexpected pass) too — it’s still a pass from the user’s perspective.
  • Latency samples pool ALL non-warmup request timings across passes.
  • Acceptance rate is meaningful only for speculative-decoding scenarios; other scenarios omit the row entirely.