Skip to content

[Enhancement](file scanner) Support row-id fetch in FileScannerV2 - #67906

Merged
Gabriel39 merged 4 commits into
apache:masterfrom
Gabriel39:dev/file-scanner-v2-rowid-fetch
Sep 14, 2026
Merged

Gabriel39 merged 4 commits into
apache:masterfrom
Gabriel39:dev/file-scanner-v2-rowid-fetch

Conversation

@Gabriel39

@Gabriel39 Gabriel39 commented Sep 13, 2026 •

Copy link
Copy Markdown
Contributor

What problem does this PR solve?

Problem Summary:

FileScannerV2 did not support selective reads by absolute file row position. As a result, TopN two-phase materialization still had to use the legacy FileScanner for its second-phase Parquet/ORC fetch.

This change carries the selected row positions through TableReader and FileScanRequest, builds selective Parquet row ranges, seeks requested ORC rows, and routes supported second-phase fetches through FileScannerV2. Table-format variants that V2 does not support retain the existing V1 fallback. Phase two reuses the phase-one scanner selection policy, including the option presence bit, so disabling or omitting enable_file_scanner_v2 retains V1. Explicit row-ID requests bypass whole-chunk Parquet cache prefetch and range merging. Rebuilt fetch projections preserve Thrift slot presence bits and the pinned full schema's non-regular column categories, including slots pruned from phase one. Hive partition columns and synthesized metadata columns do not consume physical column indexes. Fetches retain original Iceberg file paths and row lineage while removing delete files, without mutating shared file mappings.

Release note

Support row-id fetch for Parquet and ORC in FileScannerV2.

Check List (For Author)

  • Test
    • Unit Test
      • ParquetScanTest.ReadsOnlyRequestedAbsoluteFileRowsAcrossRowGroups
      • NewOrcReaderTest.ReadsOnlyRequestedAbsoluteFileRows
      • RowIdStorageReaderTest.ExternalScannerSelectionRespectsRolloutOption
      • RowIdStorageReaderTest.ExternalScannerSelectionKeepsUnsupportedFormatsOnV1
      • ParquetScanTest.SparseRowIdsAvoidCachedRemoteChunkPrefetch
      • RowIdStorageReaderTest.ExternalFetchPartitionSlotsPreserveHivePositionMapping
      • RowIdStorageReaderTest.ExternalFetchPreservesPrunedMetadataCategories
      • RowIdStorageReaderTest.ExternalFetchPreservesIcebergFileMetadata
      • FileQueryScanNodeTest.testRowIdFetchRetainsCategoriesOfPrunedColumns
      • Added lazy metadata comparisons to test_iceberg_file_metadata_columns for Parquet and ORC.
      • Validation: 283 focused BE unit tests passed under ASAN using run-be-ut.sh; 16 targeted FE tests passed after temporarily excluding an unrelated pre-existing IVM test compilation error. Clang-format 16, FE Checkstyle, and build hygiene passed. Clang-tidy was blocked by pre-existing unmatched NOLINTEND diagnostics. Full SQL integration tests were not run locally.
  • Behavior changed:
    • Yes. TopN two-phase materialization uses FileScannerV2 for supported Parquet and ORC scans.
  • Does this need documentation?
    • No.

Check List (For Reviewer who merge this PR)

  • Confirm the release note
  • Confirm test cases
  • Confirm document
  • Add branch pick label

@Gabriel39
Gabriel39 requested a review from yiguolei as a code owner September 13, 2026 04:25
@hello-stephen

Copy link
Copy Markdown
Contributor

Thank you for your contribution to Apache Doris.
Don't know what should be done next? See How to process your PR.

Please clearly describe your PR:

  1. What problem was fixed (it's best to include specific error reporting information). How it was fixed.
  2. Which behaviors were modified. What was the previous behavior, what is it now, why was it modified, and what possible impacts might there be.
  3. What features were added. Why was this function added?
  4. Which code was refactored and why was this part of the code refactored?
  5. Which functions were optimized and what is the difference before and after the optimization?

@Gabriel39

Copy link
Copy Markdown
Contributor Author

run buildall

@Gabriel39

Copy link
Copy Markdown
Contributor Author

/review

@hello-stephen

Copy link
Copy Markdown
Contributor
TPC-H: Total hot run time: 16571 ms
machine: 'aliyun_ecs.c7a.8xlarge_32C64G'
scripts: https://github.com/apache/doris/tree/master/tools/tpch-tools
Tpch sf100 test result on commit f80bae195f7ed0db89c9beb6950475440750209c, data reload: false

------ Round 1 ----------------------------------
============================================
q1	17603	2985	2963	2963
q2	2077	248	216	216
q3	10251	859	504	504
q4	4673	254	205	205
q5	7668	563	385	385
q6	136	113	92	92
q7	521	513	386	386
q8	9241	950	860	860
q9	3429	2389	2348	2348
q10	6523	849	700	700
q11	396	199	190	190
q12	610	262	197	197
q13	18217	1527	1153	1153
q14	159	154	131	131
q15	q16	431	397	364	364
q17	1284	906	819	819
q18	3054	2247	2201	2201
q19	1256	828	794	794
q20	368	287	202	202
q21	5546	1630	1812	1630
q22	330	272	231	231
Total cold run time: 93773 ms
Total hot run time: 16571 ms

----- Round 2, with runtime_filter_mode=off -----
============================================
q1	3326	3280	3265	3265
q2	496	392	361	361
q3	2210	2354	2190	2190
q4	1182	1171	873	873
q5	2148	2066	2079	2066
q6	161	117	85	85
q7	1037	941	847	847
q8	1577	1378	1381	1378
q9	3073	3062	3038	3038
q10	1854	1799	1615	1615
q11	355	264	249	249
q12	444	430	335	335
q13	1447	1536	1136	1136
q14	172	175	154	154
q15	q16	387	392	356	356
q17	3604	3332	3250	3250
q18	4769	4348	4736	4348
q19	880	774	794	774
q20	1107	964	827	827
q21	3847	3078	3288	3078
q22	383	340	337	337
Total cold run time: 34459 ms
Total hot run time: 30562 ms

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Static review of exact head f80bae195f7ed0db89c9beb6950475440750209c: 2 blocking findings, both attached inline.

Checkpoint conclusions:

  • Goal / proof: The format-neutral row-ID plumbing and the Parquet/ORC absolute-row selection are coherent by static inspection, but the phase-two implementation does not preserve the scanner rollout choice and does not deliver selective remote Parquet I/O in a reachable cache configuration.
  • Scope / reuse: The change is focused and reuses TableReader, native readers, and existing selected-range/seek abstractions.
  • Concurrency / lifecycle: Per-task scanner and reader state, completion publication, EOF/close, and failure ownership were traced; no separate race, deadlock, leak, use-after-free, or initialization defect was substantiated.
  • Configuration / compatibility / parallel paths: enable_file_scanner_v2=false is retained in the phase-two runtime state but ignored by the new dispatch, allowing a V1 phase one to cross into V2. The older absent-option case has the same BE behavior. Unsupported connector/table-format paths still fall back to V1.
  • Data correctness / special cases: Strict ID validation, Split/Row-Group/stripe coordinates, duplicate restoration, schema/default/partition materialization, and ORC one-row seek parity were checked; no additional wrong-row issue was confirmed. This is a read-only path with no transaction, persistence, storage-format, or data-write change.
  • Performance / observability: A sparse Parquet request can start full projected Column-Chunk FileCache prefetch, while that dry-run traffic bypasses query reader/cache counters.
  • Tests: The two added direct-reader tests are statically coherent but do not cover end-to-end dispatch, false/unset rollout behavior, cached-remote byte selectivity, partial Splits, or richer projections. Per the review contract I did not run builds or tests. Current CI shows compile/style checks passing while BE UT, external/regression, and performance jobs are still pending; those are CI evidence, not independent execution.
  • Focus / completeness: No additional user focus was supplied. Two complete review rounds converged with every Round 2 normal and risk-focused reviewer returning NO_NEW_VALUABLE_FINDINGS; every candidate was accepted, deduplicated, or dismissed before submission.

Comment thread be/src/exec/rowid_fetcher.cpp Outdated
Comment thread be/src/format_v2/parquet/parquet_scan.cpp
@hello-stephen

Copy link
Copy Markdown
Contributor
TPC-DS: Total hot run time: 80928 ms
machine: 'aliyun_ecs.c7a.8xlarge_32C64G'
scripts: https://github.com/apache/doris/tree/master/tools/tpcds-tools
TPC-DS sf100 test result on commit f80bae195f7ed0db89c9beb6950475440750209c, data reload: false

query5	4259	401	331	331
query6	373	130	123	123
query7	4987	416	218	218
query8	286	126	120	120
query9	8680	2847	2861	2847
query10	409	207	184	184
query11	5380	1015	898	898
query12	116	73	70	70
query13	1193	451	309	309
query14	6098	2164	2056	2056
query14_1	1956	1926	1909	1909
query15	175	121	110	110
query16	907	375	353	353
query17	811	448	370	370
query18	2321	323	240	240
query19	167	136	111	111
query20	77	69	77	69
query21	198	100	85	85
query22	5304	5274	5318	5274
query23	7026	6143	6011	6011
query23_1	6029	6006	5954	5954
query24	7259	1103	763	763
query24_1	794	783	772	772
query25	432	303	251	251
query26	1232	242	131	131
query27	2782	385	255	255
query28	4718	1485	1488	1485
query29	920	432	355	355
query30	253	142	132	132
query31	822	405	333	333
query32	127	72	74	72
query33	455	228	178	178
query34	988	822	464	464
query35	404	411	348	348
query36	554	556	525	525
query37	117	80	70	70
query38	996	835	789	789
query39	480	480	484	480
query39_1	443	445	446	445
query40	205	91	85	85
query41	58	57	55	55
query42	76	71	71	71
query43	242	235	212	212
query44	979	532	549	532
query45	111	110	98	98
query46	791	862	560	560
query47	746	731	702	702
query48	308	288	237	237
query49	549	270	185	185
query50	752	270	195	195
query51	8158	8044	7986	7986
query52	67	71	59	59
query53	188	201	151	151
query54	210	156	141	141
query55	78	63	55	55
query56	179	154	148	148
query57	688	632	642	632
query58	218	205	158	158
query59	1191	1194	1080	1080
query60	231	175	172	172
query61	105	123	116	116
query62	345	205	168	168
query63	171	142	138	138
query64	2761	688	588	588
query65	1616	1621	1593	1593
query66	1857	254	203	203
query67	9569	9546	9626	9546
query68	2867	1152	688	688
query69	340	222	192	192
query70	631	626	629	626
query71	246	181	162	162
query72	2262	1655	1472	1472
query73	631	655	342	342
query74	1838	1199	1112	1112
query75	1154	1086	958	958
query76	2298	712	497	497
query77	246	240	205	205
query78	3870	3554	3201	3201
query79	2765	814	579	579
query80	1559	324	259	259
query81	512	152	131	131
query82	627	128	91	91
query83	270	205	198	198
query84	287	111	88	88
query85	827	336	275	275
query86	470	178	172	172
query87	1026	959	882	882
query88	2950	2081	2104	2081
query89	287	193	175	175
query90	2025	135	123	123
query91	124	120	97	97
query92	85	66	68	66
query93	1884	1123	660	660
query94	631	231	204	204
query95	514	253	228	228
query96	787	590	266	266
query97	1050	1070	969	969
query98	177	140	131	131
query99	419	338	312	312
Total cold run time: 178075 ms
Total hot run time: 80928 ms

@hello-stephen

Copy link
Copy Markdown
Contributor
ClickBench: Total hot run time: 14.43 s
machine: 'aliyun_ecs.c7a.8xlarge_32C64G'
scripts: https://github.com/apache/doris/tree/master/tools/clickbench-tools
ClickBench test result on commit f80bae195f7ed0db89c9beb6950475440750209c, data reload: false

query1	0.01	0.00	0.01
query2	0.07	0.03	0.04
query3	0.25	0.11	0.11
query4	1.60	0.10	0.09
query5	0.18	0.16	0.15
query6	1.28	0.68	0.69
query7	0.03	0.01	0.00
query8	0.04	0.04	0.03
query9	0.29	0.21	0.21
query10	0.35	0.34	0.35
query11	0.16	0.11	0.12
query12	0.14	0.12	0.11
query13	0.31	0.30	0.31
query14	0.44	0.45	0.44
query15	0.37	0.32	0.37
query16	0.22	0.24	0.22
query17	0.71	0.68	0.70
query18	0.19	0.17	0.17
query19	1.18	1.17	1.17
query20	0.02	0.01	0.01
query21	15.44	0.16	0.11
query22	5.09	0.04	0.04
query23	16.22	0.27	0.10
query24	3.01	0.31	0.24
query25	0.11	0.04	0.03
query26	0.81	0.18	0.11
query27	0.02	0.04	0.02
query28	3.65	0.53	0.26
query29	12.46	3.18	2.55
query30	0.24	0.11	0.12
query31	2.75	0.37	0.16
query32	3.54	0.32	0.23
query33	1.43	1.38	1.34
query34	15.42	2.21	1.78
query35	1.70	1.72	1.67
query36	0.45	0.29	0.29
query37	0.06	0.04	0.04
query38	0.04	0.04	0.02
query39	0.04	0.02	0.02
query40	0.11	0.08	0.08
query41	0.08	0.03	0.02
query42	0.04	0.02	0.02
query43	0.04	0.02	0.02
Total cold run time: 90.59 s
Total hot run time: 14.43 s

…ication

### What problem does this PR solve?

Related PR: apache#67906

Problem Summary: External row-id fetches ignored the scanner rollout option, and sparse Parquet requests could prefetch complete projected chunks. Reuse the phase-one selection policy and bypass whole-chunk prefetch and range merging for explicit row selections.

### Release note

Honor enable_file_scanner_v2 during external row-id fetches and avoid whole-chunk Parquet cache prefetch for sparse row selections.

### Check List (For Author)

- Test: 280 focused BE unit tests passed under ASAN using run-be-ut.sh; clang-format 16 and build hygiene passed. Clang-tidy was blocked by pre-existing NOLINT marker errors.
- Behavior changed: Yes, disabled or absent V2 options retain V1 and row-id reads avoid whole-chunk warm-up.
- Does this need documentation: No
@Gabriel39

Copy link
Copy Markdown
Contributor Author

run buildall

@Gabriel39

Copy link
Copy Markdown
Contributor Author

/review

### What problem does this PR solve?

Related PR: apache#67906

Problem Summary: Hive ORC TopN queries failed when second-phase projections contained partition columns. Rebuilt is_file_slot values lacked Thrift presence bits, so Hive V2 treated partition columns as physical columns and exhausted column_idxs. Build the fetch projection with Thrift setters and cover its interaction with the Hive positional mapper.

### Release note

Fix Hive TopN row-id fetches with partition columns and positional column mapping.

### Check List (For Author)

- Test: New unit test reproduced the failure before the fix; 281 focused BE unit tests passed under ASAN after it. Clang-format 16 and build hygiene passed. Clang-tidy is blocked by pre-existing NOLINT marker errors. External SQL validation is pending CI.
- Behavior changed: Yes, partition columns no longer consume physical file indexes in second-phase scans.
- Does this need documentation: No
@Gabriel39

Copy link
Copy Markdown
Contributor Author

run buildall

@Gabriel39

Copy link
Copy Markdown
Contributor Author

/review

@github-actions

Copy link
Copy Markdown
Contributor

Codex automated review failed and did not complete.

Error: Codex completed, but no new pull request review was submitted for the current head SHA.
Workflow run: https://github.com/apache/doris/actions/runs/34743989299

Please inspect the workflow logs and rerun the review after the underlying issue is resolved.

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Static review of exact head df828379c4d947648202b01ef563a1b48227dd70: 1 blocking finding, attached inline.

Critical checkpoint conclusions:

  • Goal / proof: The PR adds FileScannerV2 phase-two row-ID fetching for supported Parquet/ORC paths, but the rebuilt projection loses connector metadata semantics. The direct reader, rollout, and cached-remote tests do not prove lazy synthesized-column behavior.
  • Scope / reuse: The change is focused and reuses TableReader, native readers, selected-range, and seek abstractions; the gap is at the cross-phase projection boundary.
  • Concurrency: Remote-scheduler tasks operate per file. Shared status/completion publication is synchronized, scan blocks/statistics are disjoint by index, heavy I/O stays outside the completion lock, there is no multi-lock order, and every accepted task is awaited. No distinct race or deadlock was found.
  • Lifecycle: Task memory-tracker attachment, RuntimeState ownership, error exits, and scanner/reader RAII were traced. No leak, cycle, use-after-free, or cross-TU static-initialization issue was substantiated.
  • Configuration: No configuration is added. The existing enable_file_scanner_v2 presence/value policy is now reused for phase two, including false/absent fallback.
  • Compatibility: There is no storage-format, persistence, or public-symbol change. Unsupported table/format paths retain V1 fallback; no separate rolling-upgrade issue was found.
  • Parallel paths: V1/V2, Parquet/ORC, and Hive/Iceberg/other external-table hooks were compared. The existing rollout and Parquet-prefetch threads are hard duplicate fences; no other distinct parity defect survived review.
  • Special conditions: Row-ID order/range/format validation and duplicate scatter are coherent. The synthesized-category condition is not, which is the inline finding.
  • Test coverage: Added unit coverage exercises direct Parquet/ORC reads, helper reconstruction, rollout, and cached-remote sparse Parquet I/O. It lacks an end-to-end TopN lazy Iceberg _file/_pos test and corresponding negative coverage.
  • Test results: Per the review contract, I ran no builds or tests. Visible CI/style results and the author's reported focused ASAN tests are external evidence only, not independent execution.
  • Observability: Existing status context and profile counters are adequate for this read path; no new log or metric requirement was identified.
  • Transactions / persistence: Not applicable; this is a read-only scan path with no EditLog or failover state.
  • Data writes: Not applicable; there is no write, atomicity, or crash-consistency change.
  • FE-BE transport: The phase-two slot projection is reconstructed across the RPC boundary, but it does not retain TColumnCategory; this is the blocking defect.
  • Performance: The sparse Parquet prefetch issue already raised inline is fixed on this head, and ORC's remaining merge behavior matches V1. No additional CPU, memory, or I/O regression was established.
  • Other correctness: Status propagation, strict row-count checks, memory ownership, and nullable handling were inspected. One distinct P1 remains after all other candidates were deduplicated or dismissed.

Two complete review rounds converged with all Round 2 tracks returning NO_NEW_VALUABLE_FINDINGS.

const auto& slot = scan_slots[slot_idx];
const auto column_idx = scan_column_idxs[slot_idx];
TFileScanSlotInfo slot_info;
slot_info.__set_slot_id(slot.id());

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Preserve synthesized slot categories in the phase-two request

An Iceberg metadata column such as _file or _pos can be selected as a TopN lazy slot (reduced path inferred from the code: Materialize(lazy=[_file]) -> TopN -> IcebergFileScan(id, GLOBAL_ROWID)). This rebuild creates every non-partition slot with only is_file_slot=true. On the V2 path, the missing SYNTHESIZED category makes _build_projected_columns() mark the slot as physical; TableColumnMapper then skips the virtual _file/_pos mapping and materializes a missing-column default instead of the metadata value. Please carry the original category through the phase-two request/rebuild, including slots pruned from phase one, derive column_idxs from it, and add a lazy metadata-column regression test.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 32d20cb. FE now retains the pinned full schema's non-regular column categories in the scan parameters, including lazy columns absent from phase one's tuple. BE restores those categories when rebuilding the fetch projection and includes only REGULAR/GENERATED slots in column_idxs. An explicitly empty category map preserves ordinary physical columns with metadata-like names.

The fetch also preserves original Iceberg file paths and row-lineage metadata when removing delete files, using a local range copy so shared FileMapping entries are not modified.

Both new BE regressions failed before their respective fixes. Validation: 283 focused ASAN tests and 16 targeted FE tests passed; the FE run required a temporary exclusion of an unrelated pre-existing IVM test compilation error. Added eager-versus-lazy _file/_pos comparisons for Parquet and ORC, with a materialization-plan assertion. External SQL validation is pending the newly requested CI run.

@hello-stephen

Copy link
Copy Markdown
Contributor
TPC-H: Total hot run time: 16816 ms
machine: 'aliyun_ecs.c7a.8xlarge_32C64G'
scripts: https://github.com/apache/doris/tree/master/tools/tpch-tools
Tpch sf100 test result on commit df828379c4d947648202b01ef563a1b48227dd70, data reload: false

------ Round 1 ----------------------------------
============================================
q1	17569	2991	2997	2991
q2	2104	254	225	225
q3	10228	862	525	525
q4	4672	242	198	198
q5	7681	572	387	387
q6	144	113	92	92
q7	537	510	382	382
q8	9235	905	949	905
q9	3463	2396	2383	2383
q10	6555	855	736	736
q11	397	204	179	179
q12	613	259	201	201
q13	18119	1565	1127	1127
q14	154	153	139	139
q15	q16	431	399	362	362
q17	1402	895	793	793
q18	3059	2265	2231	2231
q19	1265	856	767	767
q20	367	289	203	203
q21	5583	1768	1883	1768
q22	324	264	222	222
Total cold run time: 93902 ms
Total hot run time: 16816 ms

----- Round 2, with runtime_filter_mode=off -----
============================================
q1	3324	3289	3320	3289
q2	492	382	386	382
q3	2205	2260	2158	2158
q4	1178	1156	881	881
q5	2138	2097	2064	2064
q6	169	117	87	87
q7	1047	953	843	843
q8	1566	1366	1385	1366
q9	3076	3081	3031	3031
q10	1849	1809	1636	1636
q11	357	270	252	252
q12	450	429	339	339
q13	1484	1516	1159	1159
q14	164	163	158	158
q15	q16	393	393	350	350
q17	3568	3275	3193	3193
q18	4773	4370	4687	4370
q19	855	798	883	798
q20	1028	1000	855	855
q21	3848	3122	3320	3122
q22	395	339	324	324
Total cold run time: 34359 ms
Total hot run time: 30657 ms

@hello-stephen

Copy link
Copy Markdown
Contributor
TPC-DS: Total hot run time: 81513 ms
machine: 'aliyun_ecs.c7a.8xlarge_32C64G'
scripts: https://github.com/apache/doris/tree/master/tools/tpcds-tools
TPC-DS sf100 test result on commit df828379c4d947648202b01ef563a1b48227dd70, data reload: false

query5	4227	408	323	323
query6	402	130	120	120
query7	4961	418	231	231
query8	285	118	119	118
query9	8703	2898	2911	2898
query10	405	224	186	186
query11	5376	1032	913	913
query12	120	69	68	68
query13	1187	442	315	315
query14	6063	2184	2085	2085
query14_1	1964	1950	1946	1946
query15	178	128	110	110
query16	928	374	332	332
query17	785	450	376	376
query18	2326	322	238	238
query19	169	139	110	110
query20	75	73	74	73
query21	199	102	87	87
query22	5371	5235	5223	5223
query23	6621	6184	6064	6064
query23_1	5872	6264	6030	6030
query24	7233	1111	793	793
query24_1	787	792	800	792
query25	434	296	254	254
query26	1239	238	129	129
query27	2774	411	252	252
query28	4678	1491	1509	1491
query29	985	417	340	340
query30	248	151	127	127
query31	819	391	318	318
query32	122	71	71	71
query33	464	213	159	159
query34	990	838	478	478
query35	404	400	332	332
query36	583	572	525	525
query37	116	86	73	73
query38	1000	836	801	801
query39	483	466	462	462
query39_1	456	469	443	443
query40	202	89	75	75
query41	54	55	50	50
query42	77	70	70	70
query43	238	241	206	206
query44	1002	533	537	533
query45	114	108	102	102
query46	762	800	545	545
query47	749	765	710	710
query48	320	311	218	218
query49	564	235	187	187
query50	721	260	196	196
query51	7948	7971	7907	7907
query52	66	70	64	64
query53	201	193	144	144
query54	202	172	145	145
query55	77	58	56	56
query56	193	152	156	152
query57	677	666	644	644
query58	197	263	175	175
query59	1195	1223	1059	1059
query60	235	184	176	176
query61	127	115	115	115
query62	364	202	173	173
query63	168	141	155	141
query64	2693	711	630	630
query65	1625	1613	1641	1613
query66	1874	277	221	221
query67	9768	9714	9573	9573
query68	2749	1194	678	678
query69	351	229	216	216
query70	677	605	608	605
query71	256	176	164	164
query72	2392	1651	1517	1517
query73	655	629	346	346
query74	1574	1237	1129	1129
query75	1187	1108	956	956
query76	2292	713	506	506
query77	251	263	211	211
query78	3956	3740	3285	3285
query79	2473	805	591	591
query80	1613	320	287	287
query81	485	156	131	131
query82	622	135	95	95
query83	273	206	192	192
query84	296	108	86	86
query85	796	349	284	284
query86	410	176	174	174
query87	1004	958	895	895
query88	2776	2111	2112	2111
query89	277	199	170	170
query90	1965	119	125	119
query91	130	121	97	97
query92	80	71	70	70
query93	1433	1152	663	663
query94	611	205	214	205
query95	539	251	299	251
query96	780	598	280	280
query97	1071	1079	1055	1055
query98	164	131	149	131
query99	412	344	305	305
Total cold run time: 176407 ms
Total hot run time: 81513 ms

@hello-stephen

Copy link
Copy Markdown
Contributor
ClickBench: Total hot run time: 14.81 s
machine: 'aliyun_ecs.c7a.8xlarge_32C64G'
scripts: https://github.com/apache/doris/tree/master/tools/clickbench-tools
ClickBench test result on commit df828379c4d947648202b01ef563a1b48227dd70, data reload: false

query1	0.01	0.01	0.00
query2	0.08	0.03	0.04
query3	0.24	0.10	0.10
query4	1.60	0.10	0.10
query5	0.18	0.16	0.15
query6	1.27	0.68	0.68
query7	0.04	0.01	0.00
query8	0.05	0.04	0.03
query9	0.29	0.22	0.22
query10	0.36	0.36	0.35
query11	0.16	0.12	0.12
query12	0.14	0.13	0.12
query13	0.31	0.31	0.31
query14	0.44	0.45	0.45
query15	0.37	0.36	0.36
query16	0.22	0.21	0.21
query17	0.67	0.73	0.69
query18	0.18	0.17	0.16
query19	1.27	1.21	1.17
query20	0.01	0.01	0.01
query21	15.43	0.15	0.11
query22	5.07	0.05	0.04
query23	16.19	0.26	0.10
query24	2.99	0.30	0.26
query25	0.12	0.05	0.03
query26	0.86	0.16	0.12
query27	0.04	0.03	0.03
query28	3.71	0.57	0.27
query29	12.42	3.16	2.60
query30	0.27	0.12	0.11
query31	2.76	0.38	0.18
query32	3.52	0.32	0.23
query33	1.42	1.46	1.41
query34	15.39	2.15	1.80
query35	1.75	1.75	1.74
query36	0.45	0.29	0.30
query37	0.06	0.04	0.04
query38	0.04	0.03	0.03
query39	0.03	0.02	0.02
query40	0.12	0.08	0.08
query41	0.08	0.02	0.03
query42	0.03	0.02	0.02
query43	0.03	0.02	0.03
Total cold run time: 90.67 s
Total hot run time: 14.81 s

### What problem does this PR solve?

Related PR: apache#67906

Problem Summary: Rebuilding a second-phase projection discarded synthesized slot categories, so lazy Iceberg metadata columns were treated as missing physical fields. Preserve non-regular categories from the pinned full schema, including slots pruned from phase one, and derive physical column indexes from those categories. Retain original file paths and row-lineage metadata when removing delete files, using a local range copy to avoid modifying shared file mappings.

### Release note

Preserve Iceberg metadata values in TopN lazy row-id fetches.

### Check List (For Author)

- Test: 283 focused BE ASAN tests and 16 targeted FE tests passed. FE validation required a local exclusion of an unrelated pre-existing IVM test compilation error; the exclusion was restored. Added Parquet/ORC lazy metadata regression coverage; external SQL validation awaits CI. Clang-format 16, FE Checkstyle, and build hygiene passed. Clang-tidy is blocked by pre-existing NOLINT marker errors.
- Behavior changed: Yes, lazy metadata columns retain their virtual values and original file metadata.
- Does this need documentation: No
@Gabriel39

Copy link
Copy Markdown
Contributor Author

run buildall

@Gabriel39

Copy link
Copy Markdown
Contributor Author

/review

@github-actions github-actions Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One blocking correctness issue remains.

Critical checkpoint conclusions:

  • Goal and scope: The PR adds standalone FileScannerV2 row-id fetch for external Top-N lazy materialization. The FE/BE/format/test changes are appropriately focused, but the rolling-upgrade gap below prevents the feature from being correct across supported mixed versions.
  • Concurrency and lifecycle: The scheduled fetch tasks, semaphore, mutex/condition/status coordination, accepted-task wait, scanner state initialization, cleanup, and error propagation were traced; no additional actionable issue was found.
  • Configuration and compatibility: No new configuration is introduced. Reusing the existing V2 option is unsafe when the new optional field 38 is absent, as described inline.
  • Parallel paths and conditions: V1 fallback, Parquet/ORC, reachable Iceberg/Hive paths, and Paimon/Hudi/system-table capability gates were checked. Current sparse-prefetch and synthesized-column fixes address their same-version paths; the missing field-presence condition remains.
  • Tests and results: The new unit/FE/regression coverage exercises current-version behavior and deterministic eager-versus-lazy equality, but it does not cover an old-FE payload without field 38 on a new BE. Per the review contract, I did not build or run tests. The PR discussion reports 283 focused BE ASAN tests and 16 FE tests passing, with the external SQL regression pending.
  • Observability and performance: Phase-two I/O/time counters and sparse-read attribution are present; no further actionable issue was found.
  • Transactions, persistence, and writes: Not applicable; this is a read-only scan path with no EditLog, storage-format, or write-path change.
  • FE/BE contract: The current FE producer and BE consumer for field 38 are wired consistently, but absence of the optional field is not negotiated safely.

Focus: No additional user-provided review focus was supplied.

Completion: Static review converged in two rounds. All normal full-review and dedicated risk-review passes returned NO_NEW_VALUABLE_FINDINGS in the final round. Existing inline threads were treated as duplicate fences.

Comment thread be/src/exec/rowid_fetcher.cpp
@hello-stephen

Copy link
Copy Markdown
Contributor
TPC-H: Total hot run time: 16723 ms
machine: 'aliyun_ecs.c7a.8xlarge_32C64G'
scripts: https://github.com/apache/doris/tree/master/tools/tpch-tools
Tpch sf100 test result on commit 32d20cb0d5a1d3b471f110b927c628b2c384e68c, data reload: false

------ Round 1 ----------------------------------
============================================
q1	17448	3076	3043	3043
q2	2101	261	225	225
q3	10139	878	519	519
q4	4683	254	200	200
q5	7693	555	389	389
q6	146	111	96	96
q7	525	510	389	389
q8	9343	878	885	878
q9	4001	2434	2361	2361
q10	6584	850	691	691
q11	415	204	177	177
q12	664	258	196	196
q13	18196	1535	1154	1154
q14	227	146	140	140
q15	q16	444	399	366	366
q17	1778	882	768	768
q18	3097	2294	2241	2241
q19	1625	877	697	697
q20	374	293	197	197
q21	5420	1769	1862	1769
q22	326	265	227	227
Total cold run time: 95229 ms
Total hot run time: 16723 ms

----- Round 2, with runtime_filter_mode=off -----
============================================
q1	3529	3370	3370	3370
q2	492	396	376	376
q3	2255	2299	2178	2178
q4	1203	1168	897	897
q5	2194	2135	2120	2120
q6	171	117	88	88
q7	1040	907	850	850
q8	1602	1394	1387	1387
q9	3147	3123	3118	3118
q10	1881	1808	1636	1636
q11	355	271	256	256
q12	460	431	346	346
q13	1485	1527	1163	1163
q14	175	175	162	162
q15	q16	392	403	368	368
q17	3618	3390	3297	3297
q18	4861	4494	4943	4494
q19	925	884	840	840
q20	987	984	826	826
q21	3820	3005	3239	3005
q22	401	346	335	335
Total cold run time: 34993 ms
Total hot run time: 31112 ms

@hello-stephen

Copy link
Copy Markdown
Contributor
TPC-DS: Total hot run time: 81952 ms
machine: 'aliyun_ecs.c7a.8xlarge_32C64G'
scripts: https://github.com/apache/doris/tree/master/tools/tpcds-tools
TPC-DS sf100 test result on commit 32d20cb0d5a1d3b471f110b927c628b2c384e68c, data reload: false

query5	4280	407	333	333
query6	414	143	125	125
query7	4934	384	237	237
query8	309	127	122	122
query9	8687	2909	2885	2885
query10	379	225	190	190
query11	5391	1028	902	902
query12	114	72	71	71
query13	1198	452	328	328
query14	6123	2239	2110	2110
query14_1	1985	1983	1977	1977
query15	173	125	110	110
query16	923	372	367	367
query17	803	453	371	371
query18	2336	328	240	240
query19	172	147	107	107
query20	86	71	73	71
query21	207	104	88	88
query22	5515	5381	5307	5307
query23	6648	6330	6078	6078
query23_1	6251	6087	6097	6087
query24	7210	1087	774	774
query24_1	760	761	755	755
query25	435	294	253	253
query26	1229	222	133	133
query27	2794	432	246	246
query28	4656	1487	1482	1482
query29	947	445	387	387
query30	246	155	128	128
query31	854	411	346	346
query32	143	77	76	76
query33	451	221	174	174
query34	1052	818	479	479
query35	393	391	344	344
query36	570	596	521	521
query37	123	80	69	69
query38	1025	852	808	808
query39	480	480	495	480
query39_1	469	438	480	438
query40	201	99	85	85
query41	66	50	52	50
query42	74	74	74	74
query43	239	244	214	214
query44	1027	537	537	537
query45	109	107	100	100
query46	754	813	516	516
query47	741	778	719	719
query48	307	311	234	234
query49	549	246	195	195
query50	706	262	195	195
query51	8143	8072	8153	8072
query52	70	69	62	62
query53	196	189	142	142
query54	208	155	244	155
query55	80	56	54	54
query56	194	168	158	158
query57	673	663	650	650
query58	202	165	167	165
query59	1210	1239	1100	1100
query60	238	188	179	179
query61	131	104	109	104
query62	421	220	179	179
query63	173	144	135	135
query64	2701	722	585	585
query65	1648	1617	1652	1617
query66	1858	262	211	211
query67	10123	9638	9510	9510
query68	2780	1110	722	722
query69	341	222	182	182
query70	664	611	603	603
query71	252	171	171	171
query72	2298	1735	1539	1539
query73	649	558	335	335
query74	1558	1224	1144	1144
query75	1164	1104	953	953
query76	2251	722	526	526
query77	241	254	213	213
query78	3974	3740	3225	3225
query79	1184	812	578	578
query80	1141	313	267	267
query81	478	156	132	132
query82	580	131	97	97
query83	325	214	190	190
query84	296	115	87	87
query85	819	337	284	284
query86	379	180	175	175
query87	1023	969	897	897
query88	2740	2125	2091	2091
query89	293	196	175	175
query90	1841	130	127	127
query91	133	122	102	102
query92	81	74	72	72
query93	1215	1205	722	722
query94	596	241	220	220
query95	502	310	238	238
query96	849	566	278	278
query97	1055	1032	979	979
query98	137	134	137	134
query99	426	345	314	314
Total cold run time: 175549 ms
Total hot run time: 81952 ms

@hello-stephen

Copy link
Copy Markdown
Contributor
ClickBench: Total hot run time: 16.49 s
machine: 'aliyun_ecs.c7a.8xlarge_32C64G'
scripts: https://github.com/apache/doris/tree/master/tools/clickbench-tools
ClickBench test result on commit 32d20cb0d5a1d3b471f110b927c628b2c384e68c, data reload: false

query1	0.00	0.00	0.00
query2	0.23	0.08	0.07
query3	0.43	0.19	0.20
query4	2.18	0.19	0.18
query5	0.27	0.24	0.24
query6	1.46	0.39	0.40
query7	0.05	0.01	0.01
query8	0.07	0.05	0.06
query9	0.45	0.28	0.28
query10	0.40	0.41	0.38
query11	0.27	0.16	0.15
query12	0.28	0.16	0.16
query13	0.38	0.40	0.39
query14	0.47	0.47	0.47
query15	0.52	0.40	0.41
query16	0.28	0.27	0.29
query17	0.65	0.68	0.69
query18	0.24	0.24	0.24
query19	1.23	1.11	1.07
query20	0.01	0.01	0.01
query21	15.39	0.28	0.25
query22	4.88	0.11	0.10
query23	15.92	0.39	0.24
query24	2.58	0.48	0.32
query25	0.12	0.09	0.07
query26	0.73	0.22	0.17
query27	0.08	0.07	0.07
query28	3.49	0.62	0.36
query29	12.41	3.26	2.66
query30	0.36	0.21	0.21
query31	2.75	0.42	0.21
query32	3.43	0.35	0.26
query33	1.35	1.40	1.43
query34	15.39	2.40	1.92
query35	1.90	1.85	1.87
query36	0.51	0.37	0.36
query37	0.09	0.06	0.06
query38	0.06	0.06	0.06
query39	0.06	0.04	0.04
query40	0.14	0.10	0.11
query41	0.11	0.06	0.06
query42	0.07	0.05	0.04
query43	0.06	0.06	0.05
Total cold run time: 91.75 s
Total hot run time: 16.49 s

@hello-stephen

Copy link
Copy Markdown
Contributor

BE UT Coverage Report

Increment line coverage 50.00% (10/20) 🎉

Increment coverage report
Complete coverage report

Category Coverage
Function Coverage 63.26% (29632/46843)
Line Coverage 48.17% (309578/642714)
Region Coverage 43.72% (249373/570382)
Branch Coverage 45.31% (116104/256231)

@hello-stephen

Copy link
Copy Markdown
Contributor

BE Regression && UT Coverage Report

Increment line coverage 70.00% (14/20) 🎉

Increment coverage report
Complete coverage report

Category Coverage
Function Coverage 76.34% (34642/45376)
Line Coverage 61.30% (389503/635433)
Region Coverage 57.53% (326898/568184)
Branch Coverage 58.36% (149096/255495)

@hello-stephen

Copy link
Copy Markdown
Contributor

FE UT Coverage Report

Increment line coverage 100.00% (15/15) 🎉
Increment coverage report
Complete coverage report

@Gabriel39
Gabriel39 merged commit bda646e into apache:master Sep 14, 2026
44 of 47 checks passed
yiguolei pushed a commit that referenced this pull request Sep 15, 2026
…cannerV2 #67906 (#67922)

### What problem does this PR solve?

Backport #67906 to branch-4.1.

TopN two-phase materialization can fetch selected Parquet and ORC file
rows through FileScannerV2. The second phase follows the scanner rollout
policy, avoids whole-chunk Parquet prefetch for sparse row selections,
and preserves partition/generated/synthesized column categories and
Iceberg file metadata. Exact-row fetches reject short results before
reordering, fill output batches across sparse ranges, and use
demand-page reads for Parquet projections.

Compatibility adjustments for branch-4.1:
- Preserve the existing Lance dataset-level uint64 row-ID fetch path,
physical split scheduling, condition-cache state, and Variant
projections.
- Apply name-based column classification to the existing IcebergScanNode
and keep relation-snapshot schema categories separate from physical
column positions.
- Use the existing Iceberg v3 row-lineage regression suite for
eager/lazy comparisons; the master-only `_file`/`_pos` feature and
connector framework are not prerequisites for this backport.

### Release note

Support row-id fetch for Parquet and ORC in FileScannerV2.

### Check List (For Author)

- FE: `FileQueryScanNodeTest` passed (14 tests, zero failures/errors);
Maven validate passed with zero Checkstyle violations.
- BE: 274 related ASAN unit tests passed across RowIdStorageReader,
FileScanner, FileScannerV2, Parquet, and ORC. clang-format 16 passed for
all 20 affected C++ files.
- Regression evidence: all five focused cases failed before the fixes
and passed afterward. A 1,024-row sparse selection now needs 8 output
batches at a 128-row cap; a one-row fetch over 20 flat columns retains
about 1.53 MiB of stream buffers instead of 160 MiB with default
settings.
- Validation limitation: the external Iceberg regression suite was not
run locally.
- Behavior changed: Yes. Supported TopN second-phase fetches use
FileScannerV2 when enabled.
- Does this need documentation: No.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants