The road to Ollama and LM Studio. Needs a new GGML tensor type for the NWC layout (8 x 512 blocks, 32 lane streams, mantissa plane), the fused CUDA mmv kernel ported to ggml-cuda, a CPU dequantization fallback, and a converter from the HF checkpoint. The format is specified in docs/format.md and reproduced in tests/test_format_cpu.py.
The road to Ollama and LM Studio. Needs a new GGML tensor type for the NWC layout (8 x 512 blocks, 32 lane streams, mantissa plane), the fused CUDA mmv kernel ported to ggml-cuda, a CPU dequantization fallback, and a converter from the HF checkpoint. The format is specified in docs/format.md and reproduced in tests/test_format_cpu.py.