Direct-Response Optimization — Latency Comparison

Date: 2026-05-24. Run on JUNC1, Qwen 3 14B Q4_K_M on RTX 5080. Change: Skip the second LLM inference for single-tool calls when the tool output is already user-ready (direct_response flag on ToolDefinition). Applied to all finance tool actions. Method: 3 runs per query type, wall-clock timed via curl through the nginx proxy to QuorraAPI /chat (non-streaming). Each run creates a new session (no session_id reuse).

Results

QueryBefore (ms)After (ms)Reduction
Account balances8850, 9617, 139082581, 2016, 193380%
Recent transactions9392, 8531, 83452220, 2442, 210574%
Monthly spending11505, 9489, 123143759, 2646, 229574%
Control (no tools)1628, 1019, 10171586, 1081, 1017No change

Averages

QueryBefore avg (ms)After avg (ms)Speedup
Account balances10,7922,1775.0x
Recent transactions8,7562,2563.9x
Monthly spending11,1032,9003.8x
Control (no tools)1,2211,2281.0x (unchanged)

Finance query average: 10,217ms → 2,444ms (4.2x faster, 76% reduction)

Why it works

Before this change, every finance query required two full LLM inferences:

  1. LLM decides which tool to call (~2-4s generation)
  2. Tool executes and returns formatted text (~100-500ms)
  3. Tool result appended to messages, sent back to LLM for a second inference (~3-8s generation) just to rephrase what the tool already returned

The second inference is redundant for tools that already return user-ready text. The direct_response flag lets tools opt in to returning their output directly, skipping step 3.

The control query (no tools involved) is unaffected, confirming the optimization only applies to tool-call paths.

Scope

Applied to all finance tool actions (query_balances, query_transactions, query_categories, query_monthly_spending, list_budgets, switch_budget, log_transaction, modify_transaction, delete_transaction). Memory tools remain unaffected — their output is internal-format data that the LLM must interpret.

The short-circuit fires only when:

  • Exactly 1 tool call in the iteration
  • Tool execution succeeded
  • Tool definition has direct_response=True (or per-action override)

Multi-tool calls, failed executions, and tools without the flag always get the second LLM pass.