ggml: bump submodule for CUDA k-quant GET_ROWS support
Device-side embedding and codebook lookups now cover q2_K to q6_K, so a quantized token_embd no longer drops the graph out of the direct device path. i-quants are left as a TODO.
This commit is contained in: