Commit Graph
10 Commits
Author SHA1 Message Date
Pascal b1e864cb8d cli, facade: rename --ref-audio to --ref-wav, align with omnivoice convention 2026-05-14 15:46:46 +02:00
Pascal 6ad6a7db7a facade, abi: thin CLI on top of qwen_* public ABI, lock contract with C99 probe 2026-05-14 15:31:47 +02:00
Pascal eefb13043e tests logs 2026-05-11 16:27:12 +02:00
Pascal 42510fff59 tests logs 2026-05-11 13:38:21 +02:00
Pascal 59fda26827 tests logs 2026-05-11 07:23:36 +02:00
Pascal 7b9435c886 clone: mode B fix, librosa to torchaudio resample, plus SEANet bisection tooling 2026-05-10 22:11:45 +02:00
Pascal 7e89929a70 fix speaker encoder ECAPA forward cossim 0.86 -> 0.996
mel-spk and mel-mag dumps in speaker-encoder-extract.h applied an extra
ggml_transpose plus cont before write. Raw ggml ne=(C, T) already
streams as numpy [T, C], so the transpose was inverting axes vs the
python upstream. Removed it.
MelMag 0.04 -> 0.999, MelSpk 0.92 -> 0.998

spk_conv1d_same passed ggml_im2col a kernel ne=(K, 1, IC, 1) and an
input ne=(T_pad, 1, IC, 1) with IC in ne[2]. But the im2col impl reads
IC = b->ne[1] when is_2D=false, so it saw IC=1, wrote OW*K floats into
a buffer declared for OW*IC*K floats, and mul_mat consumed 99% garbage.
Moved IC into ne[1] for both kernel and input, which makes the impl
read the real IC and writes a buffer coherent with the declared ne. The
permute and the retranspose after pad become unnecessary, dropped both.
SpkFrontend 0.74 -> 0.994, SpeakerEmb 0.86 -> 0.996

Adds ECAPA bisection infrastructure : 4 stage out params in
speaker_encoder_forward (frontend, block3, mfa, asp), codec encoder
intermediate dumps in pipeline-codec.cpp (seanet-out, enc-transformer
out, codec-pre-fsq), matching pytorch hooks in debug-clone-cossim.py.
2026-05-10 20:56:14 +02:00
Pascal acb75fca36 tests 2026-05-10 17:24:33 +02:00
Pascal 4feb286f04 tests 2026-05-10 16:43:58 +02:00
Pascal add3f940a0 Initial release 2026-05-10 15:57:15 +02:00