Masked MHA Runtime

Allocation-free FP16/BF16 attention that masks padded logits inside softmax, removing the per-call -inf pre-fill. BF16 accepts fused-QKV token strides.

API

  • forward(q, k, v, *, scale=None)
  • forward_static(q, k, v, *, logits, out, scale=None)
  • attention_mha_fp16_masked(q, k, v, *, logits, out, scale=None)
  • attention_mha_bf16_masked(q, k, v, *, logits, out, qkv_token_stride=None, scale=None)
  • allocate_workspace(q, k)
  • forward_seqused_static(q, k, v, valid_k, *, logits, out, scale=None)

Inputs use (sequence, heads, head_dim). forward_static is the CUDA Graph hot-path API: allocate logits and out once and reuse their addresses. Rows wider than 1024 keys use a deterministic multi-pass softmax.

This package contains the native masked-MHA execution path used by FlashRT's GROOT runtimes on SM110 and SM120. It is separate from FlashAttention-4.

Downloads last month
23
kernel
cuda
attention
inference
cuda-graphs
apache-2.0
Supported hardwares new
CUDA
12.011.0a
DGX Spark
GB10
128GB
GPU
RTX PRO 6000 WS
96GB
GPU
RTX PRO 6000 Max-Q
96GB
GPU
RTX PRO 5000
48GB
GPU
RTX PRO 4500 WS
32GB
GPU
RTX PRO 4000
24GB
GPU
RTX PRO 4000 SFF
24GB
GPU
RTX PRO 2000
16GB
RTX
RTX 5090
32GB
RTX
RTX 5090 D
32GB
RTX
RTX 5090 Mobile
24GB
RTX
RTX 5080
16GB
RTX
RTX 5080 Mobile
16GB
RTX
RTX 5070
12GB
RTX
RTX 5070 Mobile
8GB
RTX
RTX 5070 Ti
16GB
RTX
RTX 5070 Ti Mobile
12GB
RTX
RTX 5060 Ti
16GB
RTX
RTX 5060
8GB
RTX
RTX 5060 Mobile
8GB
RTX
RTX 5050
8GB
RTX
RTX 5050 Mobile
8GB
OS
linux
Arch
x86_64
Kernel Builder
2c40e10