Skip to content

Fix #3953: Dict Layout to short-circuit fetching codes for list_contains - #10100

Closed
pamod-madubashana wants to merge 4 commits into
vortex-data:developfrom
pamod-madubashana:pamod-madubashana/issue-3953-36407257745
Closed

pamod-madubashana wants to merge 4 commits into
vortex-data:developfrom
pamod-madubashana:pamod-madubashana/issue-3953-36407257745

Conversation

@pamod-madubashana

Copy link
Copy Markdown

Fixes #3953

What changed

DictReader::pruning_evaluation in vortex-layout/src/layouts/dict/reader.rs now handles list_contains(root, needle) filters.

When the needle side of the expression does not reference the root (typically a literal), the predicate is evaluated against just the small dictionary values array via the existing values_eval cache. If none of the dictionary values contain the needle (nulls treated as false, matching filter semantics), pruning returns an all-false mask for the full row range. Otherwise it returns the input mask unchanged.

Because pruning never touches the codes child, a fully pruned range short-circuits the scan before filter_evaluation ever fetches codes — exactly the expensive IO this avoids. Non-list_contains expressions keep the previous behavior (no values read during pruning), and list_contains expressions that do reference the root on the needle side are also left alone since they cannot be evaluated on the values array alone.

Why

A dict-encoded list column can answer list_contains membership from its dictionary alone when the answer is negative: if no dictionary entry contains the needle, no row can match. Previously pruning was a no-op for the dict layout, so every list_contains filter paid for a full codes read even when the dictionaries already proved the result empty.

Verification

Attempted cargo clippy -p vortex-layout --all-targets -- -D warnings, but it cannot build in this environment: the vortex-array build script requires the FlatBuffers compiler (failed to run flatc: No such file or directory). No test could be run as a result; the change was verified by careful reading against the existing filter_evaluation values-eval/mask logic instead.

@robert3005

Copy link
Copy Markdown
Contributor

The reason we haven't done this kind of operations is that we shouldn't select on expression id. There has to be a mechanism that makes it generic

pamod-madubashana and others added 2 commits September 30, 2026 09:39
Replace the ListContains-specific branch with a generic rule: any
boolean filter referencing the root whose whole tree is infallible is
evaluated against the cached dictionary values, pruning the range when
all values miss. Covers list_contains, eq, like, etc. without selecting
on expression id.
@pamod-madubashana

Copy link
Copy Markdown
Author

Addressed: replaced the ListContains-specific branch with a generic rule. pruning_evaluation now handles any boolean filter that references the root and whose whole tree is infallible: it evaluates the filter against the cached dictionary values and prunes the range when all values miss (nulls as false, matching filter semantics). This covers list_contains as well as eq/like/etc. without matching on expression id, and reuses the values_eval cache so a non-pruned range pays no second values read.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Dict Layout to short-circuit fetching codes for list_contains

2 participants