Files
978fc08f5a Add Dockerfile (cpu/cuda/vulkan), entrypoint, docs, and a GHCR build/release workflow (#15)
* Add Dockerfile (cpu/cuda), entrypoint, docs, and a GHCR release workflow

Multi-stage Dockerfile with cpu and cuda targets, built from the
existing build scripts' cmake invocations. The CUDA target documents
and applies the two docker-build-specific gotchas we hit running this
in production: CMAKE_CUDA_ARCHITECTURES needs an explicit override for
GPUs older than this project's own default arch list when building
without GPU device access, and the CUDA driver stub library needs an
explicit -L/-lcuda at link time since ggml's VMM pool allocator needs
driver-API symbols that aren't present without a real driver. Also
installs make/pkg-config/libopenblas-dev (CPU build) and libgomp1
(CUDA runtime), and copies the ggml shared libraries alongside the
binaries with LD_LIBRARY_PATH set, since the binaries' baked-in RPATH
points at the build-tree location that doesn't exist in the final
stage.

Adds a GitHub Actions workflow that builds both variants on every
push to master and version tag, publishing to
ghcr.io/serveurpersocom/qwentts.cpp, and validates the build (without
pushing) on PRs that touch the Docker files.

Verified end-to-end on a real GTX 1070 (Pascal): both cpu and cuda
targets build clean, run, and produce valid synthesized WAV output
through tts-server's HTTP API.

* Dockerfile: add a vulkan build target (AMD/Intel GPUs)

Uses the LunarG Vulkan SDK apt repo for glslc (ggml-vulkan's shader
compiler), which Ubuntu 22.04's own repos don't package. Runtime image
ships Mesa's Vulkan drivers; NVIDIA users should prefer :cuda instead,
since NVIDIA's Vulkan ICD isn't bundled.

* docker workflow: bump actions to latest majors (checkout v7, buildx v4, login v4, build-push v7)

* Support Pascal (sm_61) in the default CUDA architecture list

Adds 61-real to CMAKE_CUDA_ARCHITECTURES' default so Pascal cards
(GTX 10-series etc.) work without an explicit override -- including
in the Docker cuda target, where docker build has no GPU device to
autodetect against in the first place. Real-only (no virtual/PTX):
Pascal is now the oldest supported card, so it doesn't need to seed
forward JIT compatibility for anything older the way the 75-virtual
entry does for 7.5+.

Verified end-to-end on a real GTX 1070 (sm_61): built the cuda Docker
target with no --build-arg override, ran it with --gpus all, and
confirmed via container logs that the GPU loaded the model and served
a real synthesis request producing valid WAV output.

Adjusts the Dockerfile comments and docs/DOCKER.md accordingly.

* docker: pass through --max-prefill-tokens, add ref_text voice cloning

- entrypoint.sh: forward MAX_PREFILL_TOKENS to --max-prefill-tokens,
  matching the existing optional-flag passthrough pattern
- entrypoint.sh: a same-stem .txt next to a voice's .wav now supplies
  ref_text, enabling ICL clone mode instead of the x_vector_only
  fallback used when no transcript is given
- switch voice registration's JSON construction from raw printf to jq
  for safe escaping of arbitrary transcript text; base64 payloads go
  through jq's --rawfile (not --arg) since large files blow past
  ARG_MAX as a command-line argument
- add jq to all three runtime image stages (cpu/cuda/vulkan) for the
  above
- docs/DOCKER.md: document both additions

Verified end-to-end on a CPU build: both the ref_text and no-ref_text
registration paths log correctly (ref_text=yes / ref_text=no) and
/health responds after voice registration completes.

* docker: opt-in fatal worst-case warmup synthesis

ggml_backend_sched grows its compute buffer to fit the largest graph
it has ever built and never shrinks it back, so VRAM use can ratchet
up the first time a long reply arrives on live traffic. This adds an
opt-in startup warmup: when WARMUP_VOICE is set, entrypoint.sh runs
one real synthesis capped at WARMUP_MAX_NEW_TOKENS (default 750)
frames after voice registration, forcing that worst-case decode/
codec-decode buffer growth to happen at startup instead of mid-request.
Complements --max-prefill-tokens, which only covers the input side.

Off by default (WARMUP_VOICE unset skips it entirely, matching every
other optional flag in this entrypoint), and fatal on failure: if the
warmup synthesis fails (most likely OOM), the container kills the
server and exits non-zero rather than come up healthy and fail
unpredictably later -- a deployment that can't afford its own
configured worst case should know that at startup, not on a live
request.

Verified on a CPU build: warmup runs after voice registration, hits
the configured frame cap exactly, and the container stays healthy.

---------

Co-authored-by: Gary <gitea@gerasch.dev>
2026-08-07 07:52:22 +02:00

99 lines
4.7 KiB
Bash

#!/bin/bash
# tts-server entrypoint: starts the server, then registers every reference
# voice found in /voices (each *.wav registers under its filename stem,
# optionally paired with a same-stem .txt for ref_text ICL cloning) once
# the server is ready to accept requests.
set -e
MODEL=${MODEL_PATH:-/models/qwen-talker-1.7b-base-Q8_0.gguf}
CODEC=${CODEC_PATH:-/models/qwen-tokenizer-12hz-Q8_0.gguf}
LANG=${TTS_LANG:-auto}
HOST=${HOST:-0.0.0.0}
PORT=${PORT:-8080}
ALIAS=${MODEL_ALIAS:-}
extra_args=()
[ -n "$ALIAS" ] && extra_args+=(--alias "$ALIAS")
[ -n "$CODEC_CHUNK_DUR" ] && extra_args+=(--codec-chunk-dur "$CODEC_CHUNK_DUR")
[ -n "$CODEC_LEFT_DUR" ] && extra_args+=(--codec-left-dur "$CODEC_LEFT_DUR")
[ -n "$MAX_BATCH" ] && extra_args+=(--max-batch "$MAX_BATCH")
[ -n "$MAX_PREFILL_TOKENS" ] && extra_args+=(--max-prefill-tokens "$MAX_PREFILL_TOKENS")
[ "$NO_FA" = "1" ] && extra_args+=(--no-fa)
[ "$CLAMP_FP16" = "1" ] && extra_args+=(--clamp-fp16)
/app/tts-server \
--model "$MODEL" \
--codec "$CODEC" \
--lang "$LANG" \
--host "$HOST" \
--port "$PORT" \
"${extra_args[@]}" &
SERVER_PID=$!
until curl -sf "http://localhost:${PORT}/health" > /dev/null 2>&1; do
kill -0 "$SERVER_PID" 2>/dev/null || { echo "tts-server exited before becoming healthy" >&2; wait "$SERVER_PID"; }
sleep 1
done
for wav in /voices/*.wav; do
[ -f "$wav" ] || continue
name=$(basename "$wav" .wav)
# base64 payload can be multiple MB -- too large for an argv string
# (ARG_MAX), so it's written to a temp file and read in via jq's
# --rawfile rather than passed as a --arg.
b64_file=$(mktemp)
base64 -w0 "$wav" > "$b64_file"
txt="${wav%.wav}.txt"
# A same-stem .txt supplies the reference transcript, enabling ICL
# clone mode (higher fidelity) instead of the x_vector_only fallback
# used when no ref_text is sent. jq handles JSON-escaping arbitrary
# transcript text safely (quotes, backslashes, ...).
if [ -f "$txt" ]; then
echo "Registering voice '$name' from $wav (with ref_text from $txt)"
jq -n --arg name "$name" --rawfile wav_b64 "$b64_file" --rawfile ref_text "$txt" \
'{name: $name, wav_b64: $wav_b64, ref_text: ($ref_text | sub("\n+$"; ""))}' \
> /tmp/voice_payload.json
else
echo "Registering voice '$name' from $wav"
jq -n --arg name "$name" --rawfile wav_b64 "$b64_file" \
'{name: $name, wav_b64: $wav_b64}' \
> /tmp/voice_payload.json
fi
curl -sf -X POST "http://localhost:${PORT}/v1/audio/voices" \
-H "Content-Type: application/json" \
-d @/tmp/voice_payload.json \
&& echo " -> ok" || echo " -> FAILED"
rm -f /tmp/voice_payload.json "$b64_file"
done
# Optional fatal worst-case warmup: exercises one real end-to-end synthesis
# capped at WARMUP_MAX_NEW_TOKENS frames, so the decode and codec-decode
# compute buffers (the part that scales with output AUDIO length) get
# reserved up front -- complementing --max-prefill-tokens above, which
# only reserves for input TEXT length. Off by default: only runs when
# WARMUP_VOICE names an already-registered voice. Failure is fatal (kills
# the server and exits non-zero) by design: a deployment that can't
# afford its own configured worst case should refuse to come up healthy,
# not fail unpredictably on a live request later. The exact wording of
# WARMUP_TEXT doesn't matter -- only that it's long enough the model
# doesn't hit EOS on its own before WARMUP_MAX_NEW_TOKENS does, so the
# warmup actually exercises the same truncation path a real over-length
# response would hit in production.
if [ -n "$WARMUP_VOICE" ]; then
WARMUP_MAX_NEW_TOKENS=${WARMUP_MAX_NEW_TOKENS:-750}
WARMUP_TEXT=${WARMUP_TEXT:-"This is a warmup sentence used to reserve worst case memory usage before accepting real requests. This is a warmup sentence used to reserve worst case memory usage before accepting real requests. This is a warmup sentence used to reserve worst case memory usage before accepting real requests."}
echo "Warming up with a ${WARMUP_MAX_NEW_TOKENS}-frame-capped synthesis to reserve worst-case VRAM..."
if ! curl -sf -X POST "http://localhost:${PORT}/v1/audio/speech" \
-H "Content-Type: application/json" \
-d "$(jq -n --arg voice "$WARMUP_VOICE" --arg input "$WARMUP_TEXT" --argjson max_new_tokens "$WARMUP_MAX_NEW_TOKENS" \
'{voice: $voice, input: $input, response_format: "pcm", max_new_tokens: $max_new_tokens}')" \
-o /dev/null; then
echo " -> FATAL: worst-case warmup synthesis failed (likely out of VRAM) -- refusing to come up healthy"
kill "$SERVER_PID" 2>/dev/null
exit 1
fi
echo " -> ok"
fi
wait "$SERVER_PID"