530eed6ac5b72a76469034e85b168f9ba73a7c86
GGUF norm matches llama.cpp policy: F32 master stays F32, BF16 variant keeps source BF16, K-quants fall back to F16 when kernel rows do not align. No conv override in pick_type. Conv kernels widen to F16 at load through gf_load_conv (12 sites). qwen_load_ctw_f32 accepts BF16 source. TODO upstream GGML: ggml_conv_1d and ggml_conv_1d_dw force F16 on their im2col output, while conv_2d picks the kernel dtype. This crashes F32 and BF16 kernels on CPU (im2col only handles F16) and BF16 on Vulkan (mul_mat refuses BF16 on the operand the kernel ends up on). Aligning conv_1d on conv_2d removes the workaround.
Languages
C++
81.6%
Python
13.1%
C
1.6%
Shell
1.4%
CMake
1.2%
Other
1.1%