This commit is contained in:
Pascal
2026-05-31 17:49:31 +02:00
parent 08b79d1209
commit 8aba0f012a
32 changed files with 1010 additions and 1010 deletions
+28 -28
View File
@@ -1,4 +1,4 @@
[Qwen] qwentts.cpp 5442c2f (2026-05-31)
[Qwen] qwentts.cpp 199a658 (2026-05-31)
ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97247 MiB):
Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97247 MiB
load_backend: loaded CUDA backend from /mnt/workspace/qwentts.cpp/build/libggml-cuda.so
@@ -39,15 +39,15 @@ load_backend: loaded CPU backend from /mnt/workspace/qwentts.cpp/build/libggml-c
[Prompt] Built: 38 ids, N_text=30, N_instruct=0, T_ctx=41, hidden=2048, lang=english (id=2050), speaker=vivian (id=3065) ref_spk_emb=no icl=no
[Debug] prompt-ids: [38] first4: 151644.000000 77091.000000 198.000000 80.000000
[Debug] talker-input-embed: [41, 2048] first4: 0.021718 -0.009369 0.007129 -0.019678
[Debug] trailing-text-hidden: [1, 2048] first4: -0.003786 0.008339 -0.005023 -0.000287
[Debug] tts-pad-embed: [2048] first4: -0.003786 0.008339 -0.005023 -0.000287
[Debug] trailing-text-hidden: [1, 2048] first4: -0.003815 0.008345 -0.005047 -0.000283
[Debug] tts-pad-embed: [2048] first4: -0.003815 0.008345 -0.005047 -0.000283
[Debug] talker-hidden-prefill-l0: [41, 2048] first4: -0.380321 -0.062937 0.001099 -0.016046
[Debug] talker-hidden-prefill-l7: [41, 2048] first4: -2.109848 0.164757 2.020530 1.130142
[Debug] talker-hidden-prefill-l14: [41, 2048] first4: -1.859514 0.329631 2.120106 1.083235
[Debug] talker-hidden-prefill-l21: [41, 2048] first4: -1.446755 0.266681 2.494964 0.577370
[Debug] talker-hidden-prefill-l27: [41, 2048] first4: -1.404861 -24.704107 -1.774197 24.936502
[Debug] talker-hidden-prefill-final: [41, 2048] first4: -0.043405 -0.797034 -0.050935 0.702261
[Debug] talker-logits-prefill: [3072] first4: -0.811035 -8.015625 -2.671875 -4.843750
[Debug] talker-logits-prefill: [3072] first4: -0.797852 -8.000000 -2.703125 -4.839844
[Sample] step=0 c0=1995 u=-1.0000000000 subseq=0
[Sample-CP] g=0 c=1159 u=-1.0000000000 subseq=1
[Sample-CP] g=1 c=355 u=-1.0000000000 subseq=2
@@ -65,8 +65,8 @@ load_backend: loaded CPU backend from /mnt/workspace/qwentts.cpp/build/libggml-c
[Sample-CP] g=13 c=812 u=-1.0000000000 subseq=14
[Sample-CP] g=14 c=901 u=-1.0000000000 subseq=15
[Debug] codes-step0: [16] first4: 1995.000000 1159.000000 355.000000 22.000000
[Debug] next-emb-step0: [2048] first4: -0.039680 -0.032027 -0.001058 0.180065
[Debug] talker-hidden-step1: [2048] first4: 0.583368 -1.895007 3.011910 -1.211491
[Debug] next-emb-step0: [2048] first4: -0.039709 -0.032021 -0.001082 0.180069
[Debug] talker-hidden-step1: [2048] first4: 0.584582 -1.898573 3.029192 -1.221801
[Sample] step=1 c0=259 u=-1.0000000000 subseq=16
[Sample-CP] g=0 c=259 u=-1.0000000000 subseq=17
[Sample-CP] g=1 c=1160 u=-1.0000000000 subseq=18
@@ -94,14 +94,14 @@ load_backend: loaded CPU backend from /mnt/workspace/qwentts.cpp/build/libggml-c
[Pipeline] Generation done : 64 frames
[Debug] codes-full: [64, 16] first4: 1995.000000 1159.000000 355.000000 22.000000
[Debug] output-audio: [122880] first4: 0.000015 0.000017 0.000015 0.000017
[Perf] PromptBuild 4.3 ms
[Perf] Prefill 10.6 ms (T_ctx prefill)
[Perf] TTFA 21.2 ms (first frame codes)
[Perf] TalkerDecode 193.9 ms (64 frames, 3.03 ms/frame)
[Perf] CodePredictor 350.7 ms (5.48 ms/frame)
[Perf] PromptBuild 2.8 ms
[Perf] Prefill 11.0 ms (T_ctx prefill)
[Perf] TTFA 20.2 ms (first frame codes)
[Perf] TalkerDecode 199.2 ms (64 frames, 3.11 ms/frame)
[Perf] CodePredictor 351.9 ms (5.50 ms/frame)
[Perf] HostCompose 1.0 ms (c0 sample + next emb)
[Perf] CodecDecode 21.8 ms
[Perf] Total 583.1 ms (64 frames, 8.53 ms/frame AR, audio 5.12 s, RTF 0.114)
[Perf] CodecDecode 23.2 ms
[Perf] Total 589.8 ms (64 frames, 8.63 ms/frame AR, audio 5.12 s, RTF 0.115)
[WAV] Wrote cpp/customvoice/customvoice-cpp.wav: 122880 samples, 24000 Hz, mono S16
[Pipeline] Wrote 122880 samples (5.12 s) -> cpp/customvoice/customvoice-cpp.wav
[Input] Prompt: 100 chars: qwentts.cpp is a minimal C++17 and GGML port of Qwen3-TTS fo...
@@ -116,18 +116,18 @@ load_backend: loaded CPU backend from /mnt/workspace/qwentts.cpp/build/libggml-c
[GGML] Cmd: ../build/qwen-tts --model ../models/qwen-talker-1.7b-customvoice-Q8_0.gguf --codec ../models/qwen-tokenizer-12hz-Q8_0.gguf --seed 42 --speaker vivian --lang english --max-new 64 --dump cpp/customvoice -o cpp/customvoice/customvoice-cpp.wav --greedy
[GGML] Audio: 122880 samples 24000 Hz 5.12s -> cpp/customvoice/customvoice-cpp.wav
[Cossim] PromptIDs exact: 100.00% (38 values)
[Cossim] Embed cos: 0.999962 max: 5.1095e-02 mean: 2.0833e-04
[Cossim] TrailingText cos: 0.999999 max: 1.8678e-04 mean: 4.3506e-05
[Cossim] TTSPadEmbed cos: 0.999999 max: 1.8678e-04 mean: 4.3506e-05
[Cossim] L0 cos: 0.999990 max: 5.8628e-02 mean: 6.5103e-04
[Cossim] L7 cos: 0.999982 max: 2.1607e+01 mean: 4.4860e-02
[Cossim] L14 cos: 0.999982 max: 2.1537e+01 mean: 5.5432e-02
[Cossim] L21 cos: 0.999979 max: 2.1858e+01 mean: 1.0403e-01
[Cossim] L27 cos: 0.999877 max: 7.6894e+01 mean: 4.5622e-01
[Cossim] Final cos: 0.998980 max: 1.0111e+01 mean: 3.4557e-02
[Cossim] Logits cos: 0.999883 max: 3.6396e-01 mean: 5.8559e-02
[Cossim] NextEmbStep0 cos: 0.999992 max: 1.1319e-03 mean: 2.3562e-04
[Cossim] TalkerHiddenStep1 cos: 0.999782 max: 3.5262e-01 mean: 3.5847e-02
[Cossim] CodesFull exact: 7.34% (1008 values)
[Cossim] Audio cos: 0.074544
[Cossim] WAV stft_cos: 0.345793 samples: 120960
[Cossim] Embed cos: 0.999962 max: 5.1083e-02 mean: 2.0858e-04
[Cossim] TrailingText cos: 0.999999 max: 2.2450e-04 mean: 4.4874e-05
[Cossim] TTSPadEmbed cos: 0.999999 max: 2.2450e-04 mean: 4.4874e-05
[Cossim] L0 cos: 0.999990 max: 5.8628e-02 mean: 6.5153e-04
[Cossim] L7 cos: 0.999982 max: 2.1607e+01 mean: 4.4834e-02
[Cossim] L14 cos: 0.999982 max: 2.1537e+01 mean: 5.5475e-02
[Cossim] L21 cos: 0.999979 max: 2.1858e+01 mean: 1.0392e-01
[Cossim] L27 cos: 0.999883 max: 6.6968e+01 mean: 4.5412e-01
[Cossim] Final cos: 0.998996 max: 9.8778e+00 mean: 3.4441e-02
[Cossim] Logits cos: 0.999884 max: 3.5614e-01 mean: 5.9515e-02
[Cossim] NextEmbStep0 cos: 0.999992 max: 1.1177e-03 mean: 2.3569e-04
[Cossim] TalkerHiddenStep1 cos: 0.999753 max: 3.8159e-01 mean: 3.7885e-02
[Cossim] CodesFull exact: 5.95% (1008 values)
[Cossim] Audio cos: 0.078188
[Cossim] WAV stft_cos: 0.223316 samples: 120960