Pascal 530eed6ac5 loader, quantize: align GGUF on llama.cpp, conv kernels widened to F16 at load
GGUF norm matches llama.cpp policy: F32 master stays F32, BF16
variant keeps source BF16, K-quants fall back to F16 when kernel
rows do not align. No conv override in pick_type.

Conv kernels widen to F16 at load through gf_load_conv (12 sites).
qwen_load_ctw_f32 accepts BF16 source.

TODO upstream GGML: ggml_conv_1d and ggml_conv_1d_dw force F16 on
their im2col output, while conv_2d picks the kernel dtype. This
crashes F32 and BF16 kernels on CPU (im2col only handles F16) and
BF16 on Vulkan (mul_mat refuses BF16 on the operand the kernel
ends up on). Aligning conv_1d on conv_2d removes the workaround.
2026-05-11 12:42:54 +02:00
2026-05-10 17:20:15 +02:00
2026-05-10 20:56:28 +02:00
2026-05-11 07:23:36 +02:00
2026-05-10 15:57:15 +02:00
2026-05-10 15:57:15 +02:00
2026-05-10 15:57:15 +02:00
2026-05-10 15:57:15 +02:00
2026-05-10 15:57:15 +02:00
2026-05-10 15:57:15 +02:00
2026-05-10 15:57:15 +02:00
2026-05-10 15:57:15 +02:00
2026-05-10 15:57:15 +02:00
2026-05-10 15:57:15 +02:00
2026-05-10 15:57:15 +02:00
2026-05-10 15:57:15 +02:00
2026-05-10 15:57:15 +02:00
2026-05-10 15:57:15 +02:00
2026-05-10 15:57:15 +02:00
2026-05-10 15:57:15 +02:00
2026-05-10 15:57:15 +02:00
2026-05-10 15:57:15 +02:00
2026-05-10 15:57:15 +02:00
2026-05-10 15:57:15 +02:00
S
Description
qwentts.cpp fork: tts-server Qwen3-TTS con containerizzazione Vulkan
MIT
2.9 MiB
Languages
C++ 81.6%
Python 13.1%
C 1.6%
Shell 1.4%
CMake 1.2%
Other 1.1%