The CMakeLists.txt >=12.8 CUDA-architecture branch claimed 121a-real
(Blackwell Ultra) alongside 120a-real, but nvcc from CUDA 12.8.1
rejects compute_121 ("nvcc fatal: Unsupported gpu architecture
'compute_121'"). Confirmed against a real build: 12.8.1 compiles
cleanly with 120a-real but not 121a-real; 12.9.2 compiles both. Split
the >=12.8 branch into >=12.9 (full Blackwell + Ultra) and >=12.8
(Blackwell only, no Ultra).
Since no single CUDA toolkit spans the full arch range -- 12.9.x is
the newest 12.x that still emits Pascal (61-real) SASS, 13.x drops
Pascal entirely but is otherwise more current -- the single :cuda
image can no longer serve both audiences. Split it into two CI
variants, :cuda12 (12.9.2, Pascal through Blackwell Ultra) and
:cuda13 (13.3.1, Turing and newer). The Dockerfile's own --target
cuda stage is unchanged; the CUDA_BUILD_IMAGE/CUDA_RUNTIME_IMAGE ARGs
now default to 12.9.2 (was 12.4.1) and the CI matrix overrides them
per variant.
Both variants were build-tested against nvidia/cuda:12.8.1 and 12.9.2
locally and exercised with real end-to-end synthesis requests on an
actual sm_61 card (GTX 1070 Max-Q) -- RTF ~0.4 on both, no kernel
image / arch mismatch errors.
Co-authored-by: Gary <gitea@gerasch.dev>
* Add Dockerfile (cpu/cuda), entrypoint, docs, and a GHCR release workflow
Multi-stage Dockerfile with cpu and cuda targets, built from the
existing build scripts' cmake invocations. The CUDA target documents
and applies the two docker-build-specific gotchas we hit running this
in production: CMAKE_CUDA_ARCHITECTURES needs an explicit override for
GPUs older than this project's own default arch list when building
without GPU device access, and the CUDA driver stub library needs an
explicit -L/-lcuda at link time since ggml's VMM pool allocator needs
driver-API symbols that aren't present without a real driver. Also
installs make/pkg-config/libopenblas-dev (CPU build) and libgomp1
(CUDA runtime), and copies the ggml shared libraries alongside the
binaries with LD_LIBRARY_PATH set, since the binaries' baked-in RPATH
points at the build-tree location that doesn't exist in the final
stage.
Adds a GitHub Actions workflow that builds both variants on every
push to master and version tag, publishing to
ghcr.io/serveurpersocom/qwentts.cpp, and validates the build (without
pushing) on PRs that touch the Docker files.
Verified end-to-end on a real GTX 1070 (Pascal): both cpu and cuda
targets build clean, run, and produce valid synthesized WAV output
through tts-server's HTTP API.
* Dockerfile: add a vulkan build target (AMD/Intel GPUs)
Uses the LunarG Vulkan SDK apt repo for glslc (ggml-vulkan's shader
compiler), which Ubuntu 22.04's own repos don't package. Runtime image
ships Mesa's Vulkan drivers; NVIDIA users should prefer :cuda instead,
since NVIDIA's Vulkan ICD isn't bundled.
* docker workflow: bump actions to latest majors (checkout v7, buildx v4, login v4, build-push v7)
* Support Pascal (sm_61) in the default CUDA architecture list
Adds 61-real to CMAKE_CUDA_ARCHITECTURES' default so Pascal cards
(GTX 10-series etc.) work without an explicit override -- including
in the Docker cuda target, where docker build has no GPU device to
autodetect against in the first place. Real-only (no virtual/PTX):
Pascal is now the oldest supported card, so it doesn't need to seed
forward JIT compatibility for anything older the way the 75-virtual
entry does for 7.5+.
Verified end-to-end on a real GTX 1070 (sm_61): built the cuda Docker
target with no --build-arg override, ran it with --gpus all, and
confirmed via container logs that the GPU loaded the model and served
a real synthesis request producing valid WAV output.
Adjusts the Dockerfile comments and docs/DOCKER.md accordingly.
* docker: pass through --max-prefill-tokens, add ref_text voice cloning
- entrypoint.sh: forward MAX_PREFILL_TOKENS to --max-prefill-tokens,
matching the existing optional-flag passthrough pattern
- entrypoint.sh: a same-stem .txt next to a voice's .wav now supplies
ref_text, enabling ICL clone mode instead of the x_vector_only
fallback used when no transcript is given
- switch voice registration's JSON construction from raw printf to jq
for safe escaping of arbitrary transcript text; base64 payloads go
through jq's --rawfile (not --arg) since large files blow past
ARG_MAX as a command-line argument
- add jq to all three runtime image stages (cpu/cuda/vulkan) for the
above
- docs/DOCKER.md: document both additions
Verified end-to-end on a CPU build: both the ref_text and no-ref_text
registration paths log correctly (ref_text=yes / ref_text=no) and
/health responds after voice registration completes.
* docker: opt-in fatal worst-case warmup synthesis
ggml_backend_sched grows its compute buffer to fit the largest graph
it has ever built and never shrinks it back, so VRAM use can ratchet
up the first time a long reply arrives on live traffic. This adds an
opt-in startup warmup: when WARMUP_VOICE is set, entrypoint.sh runs
one real synthesis capped at WARMUP_MAX_NEW_TOKENS (default 750)
frames after voice registration, forcing that worst-case decode/
codec-decode buffer growth to happen at startup instead of mid-request.
Complements --max-prefill-tokens, which only covers the input side.
Off by default (WARMUP_VOICE unset skips it entirely, matching every
other optional flag in this entrypoint), and fatal on failure: if the
warmup synthesis fails (most likely OOM), the container kills the
server and exits non-zero rather than come up healthy and fail
unpredictably later -- a deployment that can't afford its own
configured worst case should know that at startup, not on a live
request.
Verified on a CPU build: warmup runs after voice registration, hits
the configured frame cap exactly, and the container stays healthy.
---------
Co-authored-by: Gary <gitea@gerasch.dev>
The left context of the buffered chunked decode is no longer a caller
knob: it derives from the codec's own sliding window (2x144 frames),
placing the default decode at the residual floor of the split.
codec_chunk_sec moves from qt_tts_params to qt_init_params, resolved
once to frames at load. The mid-struct removal bumps the ABI to a
closed range [QT_ABI_MIN_VERSION, QT_ABI_VERSION] = [4, 4]; the probe
asserts both bounds reject through the range check.
The 15 predictor flavors build and allocate once at load, positions,
kv rows, and mask baked as never freed graph outputs, and replay
directly on the backend. The prefill slices the last position before
lm_head so every flavor reads one logits row at offset zero. Replaces
the per step graph rebuild, sched allocation, and debug prints.
The speech body accepts seed, max_new_tokens, temperature, top_k,
top_p, and repetition_penalty. Unset fields keep the engine defaults,
a temperature of zero selects greedy decoding, and the subtalker
mirrors the talker knobs. A fixed seed makes a request reproducible.
seed_reference hashes the ICL reference codes and restores the conv
contexts, KV ring, and position from a per reference snapshot slot on
a repeat, saving the primed state device to device after a fresh
prime. The reference priming cost amortizes across repeated cloned
voice requests.
POST /v1/voices registers a voice from a WAV extracted server side
through qt_extract_voice_ref or from pre extracted .spk and .rvq
latents taken verbatim, DELETE drops it and GET lists it alongside the
model speakers. A registered voice wins over a speaker of the same
name and injects the reference latents into qt_tts_params, ref_text
present selects ICL clone mode. The registry lives in process RAM
under the synthesis mutex, so registration and lookups never race a
running synthesis. The audio and rvq readers gain buffer variants
factored from the file paths. The README and the architecture
document catch up on the streaming decode, the hidden bridge, and the
server endpoints.
With -o '-', stdin is read line by line and every line synthesises
immediately as its own utterance, model and speaker staying resident
across lines. Each utterance after the first opens with a fresh RIFF
header, armed at end of line and consumed lazily at the next audio,
so a client can split the stream into standalone WAV clips on the
RIFF magic. Port of the feature contributed to omnivoice.cpp in
ServeurpersoCom/omnivoice.cpp#11.
Co-authored-by: Jeffrey van Binsbergen <comgenie@comgenie.com>