Files
qwentts.cpp/docs/DOCKER.md
T
978fc08f5a Add Dockerfile (cpu/cuda/vulkan), entrypoint, docs, and a GHCR build/release workflow (#15)
* Add Dockerfile (cpu/cuda), entrypoint, docs, and a GHCR release workflow

Multi-stage Dockerfile with cpu and cuda targets, built from the
existing build scripts' cmake invocations. The CUDA target documents
and applies the two docker-build-specific gotchas we hit running this
in production: CMAKE_CUDA_ARCHITECTURES needs an explicit override for
GPUs older than this project's own default arch list when building
without GPU device access, and the CUDA driver stub library needs an
explicit -L/-lcuda at link time since ggml's VMM pool allocator needs
driver-API symbols that aren't present without a real driver. Also
installs make/pkg-config/libopenblas-dev (CPU build) and libgomp1
(CUDA runtime), and copies the ggml shared libraries alongside the
binaries with LD_LIBRARY_PATH set, since the binaries' baked-in RPATH
points at the build-tree location that doesn't exist in the final
stage.

Adds a GitHub Actions workflow that builds both variants on every
push to master and version tag, publishing to
ghcr.io/serveurpersocom/qwentts.cpp, and validates the build (without
pushing) on PRs that touch the Docker files.

Verified end-to-end on a real GTX 1070 (Pascal): both cpu and cuda
targets build clean, run, and produce valid synthesized WAV output
through tts-server's HTTP API.

* Dockerfile: add a vulkan build target (AMD/Intel GPUs)

Uses the LunarG Vulkan SDK apt repo for glslc (ggml-vulkan's shader
compiler), which Ubuntu 22.04's own repos don't package. Runtime image
ships Mesa's Vulkan drivers; NVIDIA users should prefer :cuda instead,
since NVIDIA's Vulkan ICD isn't bundled.

* docker workflow: bump actions to latest majors (checkout v7, buildx v4, login v4, build-push v7)

* Support Pascal (sm_61) in the default CUDA architecture list

Adds 61-real to CMAKE_CUDA_ARCHITECTURES' default so Pascal cards
(GTX 10-series etc.) work without an explicit override -- including
in the Docker cuda target, where docker build has no GPU device to
autodetect against in the first place. Real-only (no virtual/PTX):
Pascal is now the oldest supported card, so it doesn't need to seed
forward JIT compatibility for anything older the way the 75-virtual
entry does for 7.5+.

Verified end-to-end on a real GTX 1070 (sm_61): built the cuda Docker
target with no --build-arg override, ran it with --gpus all, and
confirmed via container logs that the GPU loaded the model and served
a real synthesis request producing valid WAV output.

Adjusts the Dockerfile comments and docs/DOCKER.md accordingly.

* docker: pass through --max-prefill-tokens, add ref_text voice cloning

- entrypoint.sh: forward MAX_PREFILL_TOKENS to --max-prefill-tokens,
  matching the existing optional-flag passthrough pattern
- entrypoint.sh: a same-stem .txt next to a voice's .wav now supplies
  ref_text, enabling ICL clone mode instead of the x_vector_only
  fallback used when no transcript is given
- switch voice registration's JSON construction from raw printf to jq
  for safe escaping of arbitrary transcript text; base64 payloads go
  through jq's --rawfile (not --arg) since large files blow past
  ARG_MAX as a command-line argument
- add jq to all three runtime image stages (cpu/cuda/vulkan) for the
  above
- docs/DOCKER.md: document both additions

Verified end-to-end on a CPU build: both the ref_text and no-ref_text
registration paths log correctly (ref_text=yes / ref_text=no) and
/health responds after voice registration completes.

* docker: opt-in fatal worst-case warmup synthesis

ggml_backend_sched grows its compute buffer to fit the largest graph
it has ever built and never shrinks it back, so VRAM use can ratchet
up the first time a long reply arrives on live traffic. This adds an
opt-in startup warmup: when WARMUP_VOICE is set, entrypoint.sh runs
one real synthesis capped at WARMUP_MAX_NEW_TOKENS (default 750)
frames after voice registration, forcing that worst-case decode/
codec-decode buffer growth to happen at startup instead of mid-request.
Complements --max-prefill-tokens, which only covers the input side.

Off by default (WARMUP_VOICE unset skips it entirely, matching every
other optional flag in this entrypoint), and fatal on failure: if the
warmup synthesis fails (most likely OOM), the container kills the
server and exits non-zero rather than come up healthy and fail
unpredictably later -- a deployment that can't afford its own
configured worst case should know that at startup, not on a live
request.

Verified on a CPU build: warmup runs after voice registration, hits
the configured frame cap exactly, and the container stays healthy.

---------

Co-authored-by: Gary <gitea@gerasch.dev>
2026-08-07 07:52:22 +02:00

5.4 KiB

Docker

Pre-built images: ghcr.io/serveurpersocom/qwentts.cpp:cpu, :cuda and :vulkan (also tagged per release, e.g. :cuda-v1.2.3). All three run tts-server; qwen-tts and qwen-codec are included in the same image at /app/.

docker run --rm -p 8080:8080 \
    -v /path/to/models:/models:ro \
    -e MODEL_PATH=/models/qwen-talker-1.7b-base-Q8_0.gguf \
    -e CODEC_PATH=/models/qwen-tokenizer-12hz-Q8_0.gguf \
    ghcr.io/serveurpersocom/qwentts.cpp:cpu

CUDA image, with GPU access and a directory of reference WAVs to auto-register as cloned voices on startup:

docker run --rm --gpus all -p 8080:8080 \
    -v /path/to/models:/models:ro \
    -v /path/to/voices:/voices:ro \
    -e MODEL_PATH=/models/qwen-talker-1.7b-base-Q8_0.gguf \
    -e CODEC_PATH=/models/qwen-tokenizer-12hz-Q8_0.gguf \
    ghcr.io/serveurpersocom/qwentts.cpp:cuda

Vulkan image (AMD/Intel GPUs), passing through the DRI device node:

docker run --rm --device /dev/dri -p 8080:8080 \
    -v /path/to/models:/models:ro \
    -e MODEL_PATH=/models/qwen-talker-1.7b-base-Q8_0.gguf \
    -e CODEC_PATH=/models/qwen-tokenizer-12hz-Q8_0.gguf \
    ghcr.io/serveurpersocom/qwentts.cpp:vulkan

The :vulkan image bundles Mesa's Vulkan drivers (AMD/Intel). On an NVIDIA GPU, prefer :cuda; running :vulkan there would additionally need the host's proprietary NVIDIA Vulkan ICD mounted in, which the image does not provide.

Entrypoint environment variables

Variable Default
MODEL_PATH /models/qwen-talker-1.7b-base-Q8_0.gguf
CODEC_PATH /models/qwen-tokenizer-12hz-Q8_0.gguf
TTS_LANG auto
HOST 0.0.0.0
PORT 8080
MODEL_ALIAS unset (reports the GGUF file name)
CODEC_CHUNK_DUR unset (server default: 24.0)
CODEC_LEFT_DUR unset (server default: 2.0)
MAX_BATCH unset (server default: 1)
MAX_PREFILL_TOKENS unset (server default: 0, disabled)
NO_FA unset; set to 1 to disable flash attention
CLAMP_FP16 unset; set to 1 to clamp hidden states
WARMUP_VOICE unset; set to a registered voice name to enable the startup warmup below
WARMUP_MAX_NEW_TOKENS 750 (only used when WARMUP_VOICE is set)
WARMUP_TEXT a generic filler sentence (only used when WARMUP_VOICE is set)

Every *.wav placed in /voices is registered as a cloned voice under its filename stem (e.g. /voices/freeman.wav -> voice freeman) once /health responds. A same-stem .txt file (e.g. /voices/freeman.txt) supplies that voice's ref_text -- the transcript of the reference clip -- which enables higher-fidelity ICL clone mode instead of the x_vector_only fallback used when no transcript is given.

Worst-case VRAM warmup

ggml_backend_sched grows its compute buffer to fit the largest graph it has ever built and never shrinks it back, so VRAM use can ratchet up the first time a long prompt or a long reply arrives on live traffic. Setting WARMUP_VOICE runs one real synthesis at container startup, capped at WARMUP_MAX_NEW_TOKENS frames, to force that worst-case buffer growth to happen up front instead of on a live request. Combine with --max-prefill-tokens (MAX_PREFILL_TOKENS above) for the input side of the same problem. If the warmup synthesis fails (most likely out of VRAM), the container exits non-zero rather than come up healthy and fail unpredictably later -- by design, since a deployment that can't afford its own configured worst case should know that at startup.

Building locally

git clone --recurse-submodules https://github.com/ServeurpersoCom/qwentts.cpp.git
cd qwentts.cpp
docker build --target cpu    -t qwentts.cpp:cpu    .
docker build --target cuda   -t qwentts.cpp:cuda   .
docker build --target vulkan -t qwentts.cpp:vulkan .

--target is required to pick a variant; without it, docker build uses the last stage in the Dockerfile (cuda).

Older GPUs (pre-Pascal)

docker build never has GPU device access (unlike docker run --gpus), so CMake's CUDA-architecture autodetection has nothing to detect against. This project's own CMakeLists.txt already handles that by defaulting CMAKE_CUDA_ARCHITECTURES to a fixed Pascal-and-newer list (61-real;75-virtual;80-virtual;86-real;89-real, plus Blackwell with CUDA 12.8+) when the variable isn't set, so Pascal cards (sm_61, e.g. the GTX 10-series) work out of the box with no override. GPUs older than Pascal (Maxwell and earlier) still need the architecture passed explicitly:

docker build --target cuda -t qwentts.cpp:cuda \
    --build-arg CMAKE_CUDA_ARCHITECTURES=50 .   # Maxwell

Find your GPU's compute capability at https://developer.nvidia.com/cuda-gpus.

The CUDA build links against libcuda.so (the driver API, used by ggml's VMM pool allocator) at build time even though no driver is present. The Dockerfile already points the linker at the devel image's lib64/stubs/libcuda.so for this; it's mentioned here only in case you customize CUDA_BUILD_IMAGE to a base that ships that stub somewhere else.