Skip to content

Add opt-in ordered Snowflake CSV schema inference - #15

Merged
osipovartem merged 2 commits into
variant-get-nested-extensionfrom
arrow-csv-snowflake-ordered
Oct 8, 2026
Merged

osipovartem merged 2 commits into
variant-get-nested-extensionfrom
arrow-csv-snowflake-ordered

Conversation

@osipovartem

Copy link
Copy Markdown

Summary

  • Add an opt-in Snowflake CSV inference mode. Default Arrow CSV inference and the existing order-independent exact-decimal mode are unchanged.
  • Preserve the first observed type and compact type/precision state so callers can merge streamed chunks in order without reparsing rows.
  • Match observed asymmetric transitions: NUMBER then REAL becomes TEXT, REAL then NUMBER stays REAL; DATE then TIMESTAMP stays DATE, the reverse becomes TEXT.
  • Clip combined fixed-point scale to fit precision 38 while leaving individually overflowing numeric values in the REAL family.

Evidence

  • Narrow Snowflake 10.36.101 probes checked both row orders for scientific notation, 39-digit integers, scale-38 fractions, and DATE/TIMESTAMP. Separate probes confirmed 37-digit integer + scale 2 -> NUMBER(38,1) and 38-digit integer + scale 1 -> NUMBER(38,0).
  • cargo +1.95.0 test -p arrow-csv --lib: 99/99 passed.
  • The ordinary path uses compile-time dispatch; no extra per-row classifier or materialization is added to default CSV reads.
  • Independent read-only review: approved; the reviewer also checked 1,000 three-chunk sequences under both merge groupings.

@github-actions github-actions Bot added the arrow label Oct 8, 2026
@osipovartem
osipovartem merged commit 41bbc10 into variant-get-nested-extension Oct 8, 2026
19 of 26 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant