Service — llama.cpp (LLM serving)

What it is

llama.cpp is the local LLM runtime on JUNC1. It serves open-weight language models over an OpenAI-compatible HTTP API, running inference directly on the machine’s GPUs. It is what powers Quorra at inference time — no request for generation ever leaves the junc.

llama.cpp replaced Ollama as the runtime. Running llama.cpp natively gives direct control over context length, parallelism, and GPU placement that Ollama wrapped but did not fully expose. Ollama is stopped.

Serving setup

Two independent llama-server processes are configured, one pinned to each GPU via CUDA_VISIBLE_DEVICES, each managed as a systemd service:

ServiceGPUPortModelStatus
quorra-llm-primaryRTX 5080 (16 GB)11435Qwen 3 14B Q4_K_MRunning — always loaded
quorra-llm-secondaryRTX 3070 (8 GB)11436—Disabled — reserved for the diagnostic agent (M6)

The primary service runs Qwen 3 14B Q4_K_M fully in VRAM on the RTX 5080 (~10.7 GB of 16 GB used), leaving headroom for the KV cache. This model was chosen after a benchmark against a 32B tensor-split configuration — see Quorra’s model latency benchmark and the decision record.

API

Each process exposes an OpenAI-compatible HTTP API:

  • Primary: http://<host>:11435/v1/...
  • Endpoints: /v1/chat/completions, /v1/completions, /v1/models

Because the API is OpenAI-compatible, the model can be swapped without changing QuorraAPI. The model in use can be switched via scripts/switch-model.sh in the llama-serving repo.

Deployment

The systemd units and helper scripts live in the llama-serving repo (git.juncyard.com/kurt/llama-serving). Both GPUs are passed through to the Ubuntu VM at the Proxmox level — see Hardware.

Notes

  • The model is kept always-loaded (--keep -1) so there is no cold-start delay on the first message of the day.
  • The RTX 3070 sits idle today; it is deliberately reserved for the separate diagnostic agent rather than being used to split a larger model across both GPUs.