Date: 2026-05-21. Run on JUNC1 (RTX 5080 + RTX 3070).
Hypothesis: Drop the 32B baseline to 14B (paired with tool grouping, already in place) and recover any lost latency via speculative decoding with a 4B draft model — all on a single consumer GPU (RTX 5080, 16 GB).
Method: Three scenarios, 1 warmup + 3 measured passes per scenario, full 20-test behavioral suite at temperature=0. See /home/kurt/.claude/plans/let-s-explore-the-following-virtual-avalanche.md.
Recommendation
Adopt the 14B-only configuration on RTX 5080. It is faster and more accurate than the 32B baseline, and fits the consumer single-GPU target with ~5.6 GB of headroom.
Faster: 2.8× generation throughput (78.8 vs 28.2 tok/s P50), 2.4× shorter pass wall time, 2.6× faster prefill.
More accurate: 18/20 vs 17/20 on the behavioral suite — every memory test still 3/3, and the 14B passes test_no_bullet_lists_in_matrix that the 32B consistently fails (0/3 on baseline → 3/3 on 14B).
VRAM: 10.7 GB of the 16 GB RTX 5080, freeing the RTX 3070 for the future diagnostic agent (M6).
Deterministic: every test produced the same outcome across all 3 measured passes for every scenario — no flakiness, no signal lost to run-to-run noise.
Reject the 14B + 4B speculative-decoding configuration. Despite a healthy 67.2% acceptance rate, throughput dropped from 78.8 → 54.7 tok/s (31% slower) because target and draft serialize on the same GPU and the 14B isn’t slow enough per token to amortize the draft’s cost. See Speculative Decoding Analysis below.
Remaining latency knobs not exhausted: if 14B-only is still too slow for a future product surface, the next candidates (in order) are (a) drop to 8B-only and re-measure, (b) revisit spec decoding with 8B target + 1.7B draft (better ratio), (c) trim the 5.3k-token system prompt to free more KV-cache headroom.
Scenarios
baseline
endpoint: http://172.17.0.1:11435
model: qwen3-32b
timestamp (UTC): 20260521T160919Z
/v1/models reports: ['qwen3-32b-q4_k_m.gguf']
VRAM peak: GPU0=14613MB, GPU1=7469MB
14b-only
endpoint: http://172.17.0.1:11437
model: qwen3-14b
timestamp (UTC): 20260521T163528Z
/v1/models reports: ['qwen3-14b-q4_k_m.gguf']
VRAM peak: GPU0=10739MB, GPU1=4MB
14b-spec
endpoint: http://172.17.0.1:11438
model: qwen3-14b
timestamp (UTC): 20260521T170433Z
/v1/models reports: ['qwen3-14b-q4_k_m.gguf']
VRAM peak: GPU0=14577MB, GPU1=4MB
Latency
Scenario
Samples
Prefill P50/P95 (ms)
Gen P50/P95 (ms)
Total P50/P95 (ms)
Tok/s P50/P95
baseline
60
52 / 88
10721 / 18782
10773 / 18834
28 / 28
14b-only
60
20 / 32
3633 / 7498
3653 / 7517
79 / 79
14b-spec
60
24 / 40
6486 / 10098
6510 / 10123
55 / 60
Speculative Decoding
Scenario
Drafted tokens
Accepted tokens
Acceptance rate
14b-spec
21183
14226
67.2%
Speculative Decoding Analysis
The 67.2% acceptance rate is healthy — well above the ~50% threshold below which spec decoding clearly doesn’t work, and consistent with what’s expected for same-family Qwen 3 draft/target on a structured-output workload (tool calls + thinking traces). The technique is working; it just isn’t winning in this configuration.
Why throughput dropped from 78.8 → 54.7 tok/s despite high acceptance:
Target and draft share one GPU. The 4B draft runs sequentially on the same compute units as the 14B target — they don’t overlap. Every speculative cycle pays for: (a) draft generates K=4 tokens, then (b) target does ONE forward pass over K+1 positions instead of 1. Both phases tie up the GPU.
The 14B target is not slow enough per token. Spec decoding’s win comes from amortizing an expensive target step over multiple accepted tokens. With the 14B alone running at ~78 tok/s (already fast for this hardware), each target step is cheap, so the overhead of running the draft + the bigger target step exceeds what the saved tokens buy back. Spec decoding wins when target/draft cost ratio is large (e.g., 32B target with 4B draft, or 70B target with 7B draft); here it’s only ~2-3×.
VRAM overhead is non-trivial. The 4B adds ~3.9 GB of VRAM usage (model weights + small KV cache), eating ~25% of the 5080’s capacity for no throughput gain. That headroom is more useful elsewhere (larger context, headroom for the 8192-token KV cache, room for the diagnostic agent later).
When spec decoding would help on this hardware, based on this data:
8B target + 1.7B draft: better cost ratio (~4-5×), draft cost is tiny, would likely show net positive.
Any target running at < ~30 tok/s alone (e.g., a future 32B or 70B target on better hardware). The slower the target, the more there is to amortize.
A configuration where draft can be placed on a separate device that doesn’t contend for compute (unlikely on a consumer single-GPU box, but a possibility if a future product has a small dedicated draft device, e.g., an iGPU).
VRAM Analysis
Scenario
RTX 5080 used
RTX 5080 free
RTX 3070 used
Fits on 5080 alone?
baseline (32B tensor-split)
14.6 GB
1.7 GB
7.5 GB
No — needs both GPUs
14b-only
10.7 GB
5.6 GB
0 GB
Yes, with comfortable headroom
14b-spec (14B + 4B draft)
14.6 GB
1.7 GB
0 GB
Yes, but very tight
The 14B-only configuration’s 5.6 GB of free VRAM on the 5080 is the most important number for the product story. It buys:
Headroom to increase --ctx-size from 8192 toward 16384 if longer conversations become important.
Headroom to add the diagnostic agent (M6) as a small co-resident model on the 5080 — though placing it on the 3070 (which is now completely idle) is the cleaner option.
Resilience against system-prompt growth (memory tools, additional integrations).