Service — llama.cpp (LLM serving)
What it is
llama.cpp is the local LLM runtime on JUNC1. It serves open-weight language models over an OpenAI-compatible HTTP API, running inference directly on the machine’s GPUs. It is what powers Quorra at inference time — no request for generation ever leaves the junc.
llama.cpp replaced Ollama as the runtime. Running llama.cpp natively gives direct control over context length, parallelism, and GPU placement that Ollama wrapped but did not fully expose. Ollama is stopped.
Serving setup
Two independent llama-server processes are configured, one pinned to each GPU
via CUDA_VISIBLE_DEVICES, each managed as a systemd service:
| Service | GPU | Port | Model | Status |
|---|---|---|---|---|
quorra-llm-primary | RTX 5080 (16 GB) | 11435 | Qwen 3 14B Q4_K_M | Running — always loaded |
quorra-llm-secondary | RTX 3070 (8 GB) | 11436 | — | Disabled — reserved for the diagnostic agent (M6) |
The primary service runs Qwen 3 14B Q4_K_M fully in VRAM on the RTX 5080 (~10.7 GB of 16 GB used), leaving headroom for the KV cache. This model was chosen after a benchmark against a 32B tensor-split configuration — see Quorra’s model latency benchmark and the decision record.
API
Each process exposes an OpenAI-compatible HTTP API:
- Primary:
http://<host>:11435/v1/... - Endpoints:
/v1/chat/completions,/v1/completions,/v1/models
Because the API is OpenAI-compatible, the model can be swapped without changing
QuorraAPI. The model in use can be switched via scripts/switch-model.sh in the
llama-serving repo.
Deployment
The systemd units and helper scripts live in the llama-serving repo
(git.juncyard.com/kurt/llama-serving). Both GPUs are passed through to the
Ubuntu VM at the Proxmox level — see Hardware.
Notes
- The model is kept always-loaded (
--keep -1) so there is no cold-start delay on the first message of the day. - The RTX 3070 sits idle today; it is deliberately reserved for the separate diagnostic agent rather than being used to split a larger model across both GPUs.