Skip to content

Small optimizations to local_join kernels - #2751

Open
sherylll wants to merge 2 commits into
NVIDIA:mainfrom
sherylll:update-nnd-bbq
Open

sherylll wants to merge 2 commits into
NVIDIA:mainfrom
sherylll:update-nnd-bbq

Conversation

@sherylll

@sherylll sherylll commented Oct 6, 2026

Copy link
Copy Markdown
Contributor
  • Since u4 MMA only exists in a limited range of architectures, switch to native u8 MMA for better compatibility. Perf is not negatively affected because the kernel is latency bound.
  • In MMA kernels, add warp guard on MMA instructions for elements past the useful range (new_size/old_size)

Skip warps whose output block is out of range;
Remove redundant syncthreads in local_join_kernel_wmma
@copy-pr-bot

copy-pr-bot Bot commented Oct 6, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

1 participant