Direct-Response Optimization — Latency Comparison
Date: 2026-05-24. Run on JUNC1, Qwen 3 14B Q4_K_M on RTX 5080.
Change: Skip the second LLM inference for single-tool calls when the tool output is already user-ready (direct_response flag on ToolDefinition). Applied to all finance tool actions.
Method: 3 runs per query type, wall-clock timed via curl through the nginx proxy to QuorraAPI /chat (non-streaming). Each run creates a new session (no session_id reuse).
Results
| Query | Before (ms) | After (ms) | Reduction |
|---|---|---|---|
| Account balances | 8850, 9617, 13908 | 2581, 2016, 1933 | 80% |
| Recent transactions | 9392, 8531, 8345 | 2220, 2442, 2105 | 74% |
| Monthly spending | 11505, 9489, 12314 | 3759, 2646, 2295 | 74% |
| Control (no tools) | 1628, 1019, 1017 | 1586, 1081, 1017 | No change |
Averages
| Query | Before avg (ms) | After avg (ms) | Speedup |
|---|---|---|---|
| Account balances | 10,792 | 2,177 | 5.0x |
| Recent transactions | 8,756 | 2,256 | 3.9x |
| Monthly spending | 11,103 | 2,900 | 3.8x |
| Control (no tools) | 1,221 | 1,228 | 1.0x (unchanged) |
Finance query average: 10,217ms → 2,444ms (4.2x faster, 76% reduction)
Why it works
Before this change, every finance query required two full LLM inferences:
- LLM decides which tool to call (~2-4s generation)
- Tool executes and returns formatted text (~100-500ms)
- Tool result appended to messages, sent back to LLM for a second inference (~3-8s generation) just to rephrase what the tool already returned
The second inference is redundant for tools that already return user-ready text. The direct_response flag lets tools opt in to returning their output directly, skipping step 3.
The control query (no tools involved) is unaffected, confirming the optimization only applies to tool-call paths.
Scope
Applied to all finance tool actions (query_balances, query_transactions, query_categories, query_monthly_spending, list_budgets, switch_budget, log_transaction, modify_transaction, delete_transaction). Memory tools remain unaffected — their output is internal-format data that the LLM must interpret.
The short-circuit fires only when:
- Exactly 1 tool call in the iteration
- Tool execution succeeded
- Tool definition has
direct_response=True(or per-action override)
Multi-tool calls, failed executions, and tools without the flag always get the second LLM pass.