Repository navigation
fix: concurrent cache misses share one backing request - #1711
Draft
lucasmcdonald3 wants to merge 15 commits into
Draft
lucasmcdonald3 wants to merge 15 commits into
lucasmcdonald3 wants to merge 15 commits into
Conversation
When many encrypt/decrypt operations run concurrently against a cold cache for the same branch key, each cache miss independently queried the keystore, firing N DynamoDB GetItem and N KMS Decrypt calls instead of one. getBranchKeyMaterials now shares a single in-flight request per cache entry id, evicting it on settle so the cryptographic materials cache keeps ownership of caching and TTL. A rejected request is evicted too, so the next call retries rather than sharing the failure. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Track in-flight branch key requests per cache, so keyrings sharing a cache share requests. Coalesce concurrent caching CMM misses the same way; waiting encrypts read through the cache so each use counts against the entry's limits. Co-authored-by: Yosef Bensimchon <yosef.bensimchon@xero.com>
Copy branch key materials before sharing them with waiting callers, so a concurrent eviction cannot zero them first. In the caching CMM, let waiters request in parallel when the response cannot serve them.
Port the MPL's StormTracker to the hierarchical keyring: refresh entries in their grace period, let another caller fetch after graceInterval, fail waiters after inFlightTTL, and limit concurrent fetches to fanOut. In the caching CMM, callers join an in-flight request only while it has room in its limits, so a burst runs its requests in parallel. Waiters time out after inFlightTTL and retry a failed request once.
added 7 commits
October 9, 2026 13:49
Early refresh of cached branch keys is now an opt-in gracePeriod keyring option (seconds, default 0). With the default, a cached branch key is used until it expires and the keystore is not called, matching the behavior before storm tracking. When a keystore fetch fails, callers waiting on it now fail right away with the keystore's error instead of timing out after inFlightTTL with a generic error. Failed fetches no longer count toward fanOut, so failing branch keys cannot block fetches for other branch keys.
Remove a cache key's in-flight list once it is empty, and drop stuck requests nobody can join, so the map does not grow with every cache key ever used. Move SHARED_REQUESTS and the in-flight map to an internal module that the package index does not export.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Issue #, if available: #1663, #1665. Supersedes #1664.
Description of changes:
Concurrent cache misses share one backing request.
Hierarchical keyring
Also more changes to align behavior with the MPL:
Caching CMM
Also add the "Waiting callers time out after 10 seconds" behavior.
By submitting this pull request, I confirm that my contribution is made under the terms of the Apache 2.0 license.
Check any applicable: