test logs

This commit is contained in:
Pascal
2026-05-31 02:39:03 +02:00
parent a62fde62e6
commit 29a7fa0f97
48 changed files with 1475 additions and 1463 deletions
+24 -23
View File
@@ -1,3 +1,4 @@
[Qwen] qwentts.cpp a62fde6 (2026-05-30)
ggml_cuda_init: found 1 CUDA devices (Total VRAM: 97247 MiB):
Device 0: NVIDIA RTX PRO 6000 Blackwell Workstation Edition, compute capability 12.0, VMM: yes, VRAM: 97247 MiB
load_backend: loaded CUDA backend from /mnt/workspace/qwentts.cpp/build/libggml-cuda.so
@@ -32,21 +33,21 @@ load_backend: loaded CPU backend from /mnt/workspace/qwentts.cpp/build/libggml-c
[Pipeline] Ready: hop 1920 samples @ 24000 Hz mono, 16 codebooks @ 12.5 Hz
[KVCache] Allocated: 28 layers, 8 KV heads, head_dim 128, max_seq_len 4096 -> 896 MB
[KVCache] Allocated: 5 layers, 8 KV heads, head_dim 128, max_seq_len 16 -> 0 MB
[Pipeline] Loaded: arch=1b7 variant=custom_voice tokenizer=qwen3_tts_tokenizer_12hz codebooks=16 speaker_encoder=absent speakers=9
[Pipeline] Loaded: arch=1b7 variant=custom_voice tokenizer=qwen3_tts_tokenizer_12hz codebooks=16 speaker_encoder=absent speakers=9 fa=off clamp_fp16=off
[BPE] Loaded from GGUF: 151676 vocab, 151291 merges, eos_id=151643
[BPE] Registered 5 arch special tokens (total specials=6)
[Prompt] Built: 38 ids, N_text=30, N_instruct=0, T_ctx=41, hidden=2048, lang=english (id=2050), speaker=vivian (id=3065) ref_spk_emb=no icl=no
[Debug] prompt-ids: [38] first4: 151644.000000 77091.000000 198.000000 80.000000
[Debug] talker-input-embed: [41, 2048] first4: 0.021601 -0.009355 0.007161 -0.019744
[Debug] talker-input-embed: [41, 2048] first4: 0.021718 -0.009369 0.007129 -0.019678
[Debug] trailing-text-hidden: [1, 2048] first4: -0.003786 0.008339 -0.005023 -0.000287
[Debug] tts-pad-embed: [2048] first4: -0.003786 0.008339 -0.005023 -0.000287
[Debug] talker-hidden-prefill-l0: [41, 2048] first4: -0.385550 -0.059227 0.003153 -0.015722
[Debug] talker-hidden-prefill-l7: [41, 2048] first4: -2.096292 0.136028 2.073626 1.093323
[Debug] talker-hidden-prefill-l14: [41, 2048] first4: -1.841977 0.296974 2.177680 1.048685
[Debug] talker-hidden-prefill-l21: [41, 2048] first4: -1.422673 0.233760 2.557228 0.556667
[Debug] talker-hidden-prefill-l27: [41, 2048] first4: -1.257986 -24.830193 -1.723763 25.083599
[Debug] talker-hidden-prefill-final: [41, 2048] first4: -0.038877 -0.801303 -0.049500 0.706581
[Debug] talker-logits-prefill: [3072] first4: -0.937246 -7.968508 -2.695894 -4.725146
[Debug] talker-hidden-prefill-l0: [41, 2048] first4: -0.384422 -0.059876 0.002925 -0.015458
[Debug] talker-hidden-prefill-l7: [41, 2048] first4: -2.092169 0.135218 2.071625 1.087890
[Debug] talker-hidden-prefill-l14: [41, 2048] first4: -1.837527 0.295807 2.176088 1.042910
[Debug] talker-hidden-prefill-l21: [41, 2048] first4: -1.418551 0.228465 2.559490 0.549581
[Debug] talker-hidden-prefill-l27: [41, 2048] first4: -1.250818 -24.837687 -1.736372 25.114700
[Debug] talker-hidden-prefill-final: [41, 2048] first4: -0.038676 -0.801975 -0.049888 0.707837
[Debug] talker-logits-prefill: [3072] first4: -1.227702 -7.923600 -2.594268 -4.722540
[Sample] step=0 c0=1995 u=-1.0000000000 subseq=0
[Sample-CP] g=0 c=1159 u=-1.0000000000 subseq=1
[Sample-CP] g=1 c=355 u=-1.0000000000 subseq=2
@@ -65,7 +66,7 @@ load_backend: loaded CPU backend from /mnt/workspace/qwentts.cpp/build/libggml-c
[Sample-CP] g=14 c=901 u=-1.0000000000 subseq=15
[Debug] codes-step0: [16] first4: 1995.000000 1159.000000 355.000000 22.000000
[Debug] next-emb-step0: [2048] first4: -0.039680 -0.032027 -0.001058 0.180065
[Debug] talker-hidden-step1: [2048] first4: 0.458236 -1.977294 3.021161 -1.306485
[Debug] talker-hidden-step1: [2048] first4: 0.473627 -1.851452 2.993278 -1.176888
[Sample] step=1 c0=259 u=-1.0000000000 subseq=16
[Sample-CP] g=0 c=259 u=-1.0000000000 subseq=17
[Sample-CP] g=1 c=1160 u=-1.0000000000 subseq=18
@@ -104,21 +105,21 @@ load_backend: loaded CPU backend from /mnt/workspace/qwentts.cpp/build/libggml-c
[Python] Codes shape: (63, 16) (T_frames, num_code_groups)
[Python] Audio: 120960 samples 24000 Hz 5.04s -> python/customvoice/customvoice-python.wav
[Quant] Q8_0 -> ../models/qwen-talker-1.7b-customvoice-Q8_0.gguf + ../models/qwen-tokenizer-12hz-Q8_0.gguf
[GGML] Cmd: ../build/qwen-tts --model ../models/qwen-talker-1.7b-customvoice-Q8_0.gguf --codec ../models/qwen-tokenizer-12hz-Q8_0.gguf --seed 42 --text qwentts.cpp is a minimal C++17 and GGML port of Qwen3-TTS for zero shot multilingual text to speech. --speaker vivian --lang english --max-new 64 --dump cpp/customvoice -o cpp/customvoice/customvoice-cpp.wav --greedy
[GGML] Cmd: ../build/qwen-tts --model ../models/qwen-talker-1.7b-customvoice-Q8_0.gguf --codec ../models/qwen-tokenizer-12hz-Q8_0.gguf --seed 42 --speaker vivian --lang english --max-new 64 --dump cpp/customvoice -o cpp/customvoice/customvoice-cpp.wav --greedy
[GGML] Audio: 122880 samples 24000 Hz 5.12s -> cpp/customvoice/customvoice-cpp.wav
[Cossim] PromptIDs exact: 100.00% (38 values)
[Cossim] Embed cos: 0.999962 max: 5.1095e-02 mean: 2.0739e-04
[Cossim] Embed cos: 0.999961 max: 5.1095e-02 mean: 2.1507e-04
[Cossim] TrailingText cos: 0.999999 max: 1.8678e-04 mean: 4.3506e-05
[Cossim] TTSPadEmbed cos: 0.999999 max: 1.8678e-04 mean: 4.3506e-05
[Cossim] L0 cos: 0.999985 max: 5.0579e-02 mean: 8.8635e-04
[Cossim] L7 cos: 0.999985 max: 1.0991e+01 mean: 4.6925e-02
[Cossim] L14 cos: 0.999985 max: 1.1005e+01 mean: 6.0868e-02
[Cossim] L21 cos: 0.999980 max: 1.1358e+01 mean: 1.2149e-01
[Cossim] L27 cos: 0.999866 max: 5.1759e+01 mean: 5.5204e-01
[Cossim] Final cos: 0.998199 max: 1.2530e+01 mean: 4.4798e-02
[Cossim] Logits cos: 0.999809 max: 3.7629e-01 mean: 7.2518e-02
[Cossim] L0 cos: 0.999985 max: 5.6349e-02 mean: 8.8962e-04
[Cossim] L7 cos: 0.999986 max: 1.3070e+01 mean: 4.6807e-02
[Cossim] L14 cos: 0.999985 max: 1.3085e+01 mean: 6.0997e-02
[Cossim] L21 cos: 0.999980 max: 1.3418e+01 mean: 1.2221e-01
[Cossim] L27 cos: 0.999842 max: 6.6824e+01 mean: 5.7243e-01
[Cossim] Final cos: 0.998055 max: 1.2943e+01 mean: 4.5826e-02
[Cossim] Logits cos: 0.999601 max: 5.1484e-01 mean: 9.3705e-02
[Cossim] NextEmbStep0 cos: 0.999992 max: 1.1319e-03 mean: 2.3562e-04
[Cossim] TalkerHiddenStep1 cos: 0.999182 max: 7.1951e-01 mean: 6.7702e-02
[Cossim] CodesFull exact: 5.06% (1008 values)
[Cossim] Audio cos: 0.080945
[Cossim] WAV stft_cos: 0.162718 samples: 120960
[Cossim] TalkerHiddenStep1 cos: 0.999478 max: 7.5009e-01 mean: 5.4098e-02
[Cossim] CodesFull exact: 8.04% (1008 values)
[Cossim] Audio cos: 0.138139
[Cossim] WAV stft_cos: 0.223565 samples: 120960