Repository navigation
Conversation
GNND::build started and joined a std::thread on every iteration to run update_graph and sample_graph while the GPU ran add_reverse_edges and local_join. Each such thread brought up a second OpenMP team that competed with the calling thread's team for the same cores. Enqueue the iteration's (asynchronous) GPU work first and then run the host update on the calling thread; the two halves touch disjoint buffers, so the overlap and all synchronization points are unchanged, and exceptions from local_join now propagate instead of terminating the process.
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
Contributor
Author
|
Closing in favor of #2754. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
GNND::build(both the dense and the BBQ overload) started and joined astd::threadon every iteration to runupdate_and_sample(update_graph+sample_graph) while the GPU ranadd_reverse_edges+local_join. Each of those threads is a new OpenMP root thread, so its parallel regions ran on a second OpenMP team next to the calling thread's team. That puts two full teams on the same cores, and the idle workers of one team spin (libomp's defaultKMP_BLOCKTIMEis 200 ms) while the other team works. For small and medium graphs (CAGRA and HNSW-ACE partition builds, HNSW upper layers, batched all-neighbors), each iteration was dominated by thread start-up and oversubscription rather than GPU work: in nsys traces of the tests, 2.3-5.6 ms per iteration against 0.3-1.1 ms of GPU time.This PR changes only
cpp/src/neighbors/detail/nn_descent.cuh. Bothbuildoverloads now enqueue the iteration's GPU work first and then callupdate_and_sampleon the calling thread.local_join's lock-ordered inserts is unchanged.local_join(for example a launch error or the__CUDA_ARCH__ < 700check) while the helper thread was joinable used to end instd::terminate. It now propagates to the caller.On the code before this change, fewer OpenMP threads did not help wall time (
OMP_NUM_THREADS=8was slightly slower than 36,4was 60% slower), butKMP_BLOCKTIME=0was 12% faster with a quarter of the CPU time. That pointed at the two competing thread teams rather than the size of the host work.Measurements
Single-process wall time of each test executable on an RTX 6000 Ada with a 36-thread host, before vs. after this change. Builds were run alternately, 2 or more repetitions each; mean of the runs.
NEIGHBORS_ANN_NN_DESCENT_TESTNEIGHBORS_ALL_NEIGHBORS_TESTNEIGHBORS_ANN_CAGRA_FLOAT_UINT32_TESTNEIGHBORS_ANN_CAGRA_HALF_UINT32_TESTNEIGHBORS_ANN_CAGRA_INT8_UINT32_TESTNEIGHBORS_ANN_CAGRA_UINT8_UINT32_TESTNEIGHBORS_ANN_CAGRA_BBQ_UINT32_TESTNEIGHBORS_ANN_CAGRA_FILTER_UDF_TESTNEIGHBORS_ANN_CAGRA_MERGE_TESTNEIGHBORS_ANN_CAGRA_TEST_BUGSNEIGHBORS_ANN_HNSW_ACE_FLOAT_UINT32_TESTNEIGHBORS_ANN_HNSW_ACE_HALF_UINT32_TESTNEIGHBORS_ANN_HNSW_ACE_INT8_UINT32_TESTNEIGHBORS_ANN_HNSW_ACE_UINT8_UINT32_TESTAll of these tests passed in every run, and the full
ctestsuite passes.