Skip to content

Run the NN-descent per-iteration host update on the calling thread - #2748

Closed
bdice wants to merge 1 commit into
NVIDIA:mainfrom
bdice:nn-descent-host-update-inline
Closed

bdice wants to merge 1 commit into
NVIDIA:mainfrom
bdice:nn-descent-host-update-inline

Conversation

@bdice

@bdice bdice commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

GNND::build (both the dense and the BBQ overload) started and joined a std::thread on every iteration to run update_and_sample (update_graph + sample_graph) while the GPU ran add_reverse_edges + local_join. Each of those threads is a new OpenMP root thread, so its parallel regions ran on a second OpenMP team next to the calling thread's team. That puts two full teams on the same cores, and the idle workers of one team spin (libomp's default KMP_BLOCKTIME is 200 ms) while the other team works. For small and medium graphs (CAGRA and HNSW-ACE partition builds, HNSW upper layers, batched all-neighbors), each iteration was dominated by thread start-up and oversubscription rather than GPU work: in nsys traces of the tests, 2.3-5.6 ms per iteration against 0.3-1.1 ms of GPU time.

This PR changes only cpp/src/neighbors/detail/nn_descent.cuh. Both build overloads now enqueue the iteration's GPU work first and then call update_and_sample on the calling thread.

  • The enqueue sequence is fully asynchronous, so the host work still overlaps the GPU work.
  • The host and device halves of an iteration touch disjoint buffers, and the synchronization points are unchanged. The output is the same; run-to-run variation from local_join's lock-ordered inserts is unchanged.
  • An exception thrown by local_join (for example a launch error or the __CUDA_ARCH__ < 700 check) while the helper thread was joinable used to end in std::terminate. It now propagates to the caller.

On the code before this change, fewer OpenMP threads did not help wall time (OMP_NUM_THREADS=8 was slightly slower than 36, 4 was 60% slower), but KMP_BLOCKTIME=0 was 12% faster with a quarter of the CPU time. That pointed at the two competing thread teams rather than the size of the host work.

Measurements

Single-process wall time of each test executable on an RTX 6000 Ada with a 36-thread host, before vs. after this change. Builds were run alternately, 2 or more repetitions each; mean of the runs.

executable before after change
NEIGHBORS_ANN_NN_DESCENT_TEST 52.6 s 33.7 s -36.0%
NEIGHBORS_ALL_NEIGHBORS_TEST 28.3 s 20.8 s -26.4%
NEIGHBORS_ANN_CAGRA_FLOAT_UINT32_TEST 77.5 s 66.8 s -13.8%
NEIGHBORS_ANN_CAGRA_HALF_UINT32_TEST 44.9 s 37.8 s -15.7%
NEIGHBORS_ANN_CAGRA_INT8_UINT32_TEST 67.8 s 60.4 s -10.9%
NEIGHBORS_ANN_CAGRA_UINT8_UINT32_TEST 70.8 s 63.0 s -11.1%
NEIGHBORS_ANN_CAGRA_BBQ_UINT32_TEST 3.8 s 1.9 s -51.2%
NEIGHBORS_ANN_CAGRA_FILTER_UDF_TEST 3.2 s 1.8 s -41.8%
NEIGHBORS_ANN_CAGRA_MERGE_TEST 1.6 s 1.1 s -33.4%
NEIGHBORS_ANN_CAGRA_TEST_BUGS 2.7 s 1.9 s -28.2%
NEIGHBORS_ANN_HNSW_ACE_FLOAT_UINT32_TEST 6.7 s 4.3 s -36.1%
NEIGHBORS_ANN_HNSW_ACE_HALF_UINT32_TEST 9.4 s 6.3 s -32.6%
NEIGHBORS_ANN_HNSW_ACE_INT8_UINT32_TEST 7.3 s 3.9 s -46.4%
NEIGHBORS_ANN_HNSW_ACE_UINT8_UINT32_TEST 7.1 s 4.2 s -40.9%
total 383.9 s 308.1 s -19.7%

All of these tests passed in every run, and the full ctest suite passes.

GNND::build started and joined a std::thread on every iteration to run update_graph and
sample_graph while the GPU ran add_reverse_edges and local_join. Each such thread brought up a
second OpenMP team that competed with the calling thread's team for the same cores. Enqueue the
iteration's (asynchronous) GPU work first and then run the host update on the calling thread; the
two halves touch disjoint buffers, so the overlap and all synchronization points are unchanged,
and exceptions from local_join now propagate instead of terminating the process.
@bdice bdice added improvement Improves an existing functionality non-breaking Introduces a non-breaking change labels Oct 6, 2026
@copy-pr-bot

copy-pr-bot Bot commented Oct 6, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@bdice

bdice commented Oct 6, 2026

Copy link
Copy Markdown
Contributor Author

Closing in favor of #2754.

@bdice bdice closed this Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

improvement Improves an existing functionality non-breaking Introduces a non-breaking change

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

1 participant