fix(llm): enable flash-attn on qwen-classifier, real VRAM cause found

Batch/ubatch reduction barely moved measured VRAM (~768MB free, same as
before) — wrong lever. llama-server runs with --flash-attn on; this
service didn't. Without it, the unfused attention compute buffer at
65536 ctx is far larger than flash-attn's fused workspace, which is
what the naive weights+KV estimate missed. Matches llama-server's flag.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
2026-09-09 17:04:05 +02:00
co-authored by Claude-Bot
parent ba9ace6f71
commit 1102273384
+12 -9
View File
@@ -68,15 +68,17 @@ services:
#
# VRAM: weights+KV math (~4.9GiB) predicted comfortable headroom in the
# ~6.1GiB free on the R9700, but measured live it actually used ~5.85GiB —
# left only ~700MB free, too tight (compute-buffer/graph overhead the
# naive weights+KV estimate didn't account for). --batch-size/--ubatch-size
# dropped well below llama-server's defaults (2048/512) to shrink that
# buffer — this classifier processes one request at a time, no batching
# throughput to lose. --n-gpu-layers 999 (full offload) kept since the
# ~1.5GB freed here should cover it; re-verify with rocm-smi after deploy
# before trusting this number again. GPU_MAX_HW_QUEUES=1 carried over from
# llama-server's comment above — same ROCm/ROCm#5706 clock-pinning bug
# applies now that two HIP contexts (this + llama-server) share the card.
# left only ~700MB free, too tight. Dropping --batch-size/--ubatch-size
# barely moved it (~768MB free) — wrong lever. Actual cause: llama-server
# runs with --flash-attn on but this service was missing it — without
# flash attention the unfused attention compute buffer at 65536 ctx is
# much larger (roughly O(n^2) intermediate buffers vs flash-attn's fused,
# near-linear workspace), dwarfing the naive weights+KV estimate. Added
# --flash-attn on to match llama-server; re-verify with rocm-smi after
# deploy before trusting any of these numbers again. GPU_MAX_HW_QUEUES=1
# carried over from llama-server's comment above — same ROCm/ROCm#5706
# clock-pinning bug applies now that two HIP contexts (this +
# llama-server) share the card.
qwen-classifier:
image: ghcr.io/ggml-org/llama.cpp:server-rocm
container_name: qwen-classifier
@@ -102,6 +104,7 @@ services:
--parallel 1
--batch-size 512
--ubatch-size 128
--flash-attn on
--cache-type-k q4_0
--cache-type-v q4_0
--jinja