Skip to content

Add opt-in numeric Boolean parsing for CSV scans - #17

Merged
osipovartem merged 2 commits into
variant-get-nested-extensionfrom
arrow-csv-numeric-boolean-copy
Oct 8, 2026
Merged

osipovartem merged 2 commits into
variant-get-nested-extensionfrom
arrow-csv-numeric-boolean-copy

Conversation

@osipovartem

@osipovartem osipovartem commented Oct 8, 2026 •

Copy link
Copy Markdown

Summary

  • Add an opt-in Arrow CSV reader mode accepting exact 0 and 1 in Boolean columns alongside true/false.
  • Keep default parsing unchanged; dispatch once per Boolean column and preserve the text fast path.

Why

Snowflake COPY INTO accepts CSV text 0 and 1 for BOOLEAN columns, while the strict Arrow CSV reader rejected them. This option lets Snowflake-compatible loaders opt in without changing general CSV behavior.

Validation

  • cargo +1.95.0 test --offline -p arrow-csv --lib: 101 passed.
  • cargo +1.95.0 clippy --offline -p arrow-csv --all-targets: passed with existing unrelated warnings.
  • Local release microbenchmark, 64K single-column rows: strict text ~82.8M rows/s; opt-in text ~83.2M rows/s; opt-in numeric ~94.7M rows/s (one run; within normal benchmark noise).
  • Narrow live Snowflake COPY INTO probe loaded true,1,0,false as TRUE,TRUE,FALSE,FALSE.

Independent read-only review: APPROVED (correctness, regressions, performance, API, tests).

@osipovartem

Copy link
Copy Markdown
Author

Pre-merge CI audit: the failing jobs are outside this PR diff. cargo fmt reports only parquet/src/encodings/decoding.rs; Linux/Windows tests fail in parquet::bad_data::non_standard_delta_blocks; clippy reports arrow-csv/src/reader/records.rs:672; rustdoc fails in arrow-avro. This PR changes only arrow-csv/src/reader/mod.rs. Focused Arrow CSV tests (101/101), local release benchmark, Snowflake probe, and independent read-only review passed. Merging under the repository's local-validation/bypass policy; the base-branch CI issues remain separate follow-up work.

@osipovartem
osipovartem merged commit 4282541 into variant-get-nested-extension Oct 8, 2026
21 of 28 checks passed
@osipovartem

Copy link
Copy Markdown
Author

Post-merge repeat of the same local release microbenchmark (3 runs, single Boolean CSV column, no parallel build): at 64K rows, strict text was 85.4-85.8M rows/s, opt-in text 85.3-85.5M rows/s, and opt-in 0/1 98.1-99.8M rows/s. At 4K rows, strict and opt-in text were both ~73-74M rows/s. The opt-in text difference is within this benchmark's run-to-run noise; the default strict parser remains on its original compile-time branch.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant