Skip to content

llama.cpp port: GGML tensor type, CUDA mmv kernel, converter #6

Description

@parda21

The road to Ollama and LM Studio. Needs a new GGML tensor type for the NWC layout (8 x 512 blocks, 32 lane streams, mantissa plane), the fused CUDA mmv kernel ported to ggml-cuda, a CPU dequantization fallback, and a converter from the HF checkpoint. The format is specified in docs/format.md and reproduced in tests/test_format_cpu.py.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    help wantedwe cannot do this one aloneroadmapplanned work, shown in the README roadmap

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions