Skip to content

Tensor-core accumulation path (mma.m16n8k16) for A100 / H100 SXM #3

Description

@parda21

The v9 decoder is issue-bound at 11-13 SM clocks per 64 weights; on the A100 and H100 SXM the bandwidth is high enough that this loses. Plan (docs/format.md, section 4): one prmt per pair produces the bf16x2 A-fragment of an mma.m16n8k16 directly, activations as the B-fragment, accumulation in the D-fragment; the shuffle reduction and the 16 activation registers disappear. First step: measure the Hopper pipe model with experiments/pipe_bench.cu.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    roadmapplanned work, shown in the README roadmap

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions