Skip to content

Iceberg native scan fails queries on tables whose column names are case-distinct to Java but not to Rust #5540

Description

@dwsmith1983

Describe the bug

The native Iceberg scan builds its schema adapter with no JVM case tables and case sensitivity hardcoded to false (iceberg_scan.rs:226), so case-insensitive name matching falls back to Rust's str::to_lowercase. Rust ships newer Unicode data than the JDK, so codepoints that gained lowercase mappings in Unicode 14+ fold together in Rust while Java keeps them distinct. A table with two such columns is perfectly legal to Spark, but Comet's adapter folds them into a duplicate and the query dies with _LEGACY_ERROR_TEMP_2093 Found duplicate field(s) ... in case-insensitive mode, an error Spark itself would never raise.

The same divergence class was fixed for the parquet and Delta scan paths in #5365 by shipping the running JVM's case data to native (JvmCaseTables); the Iceberg path is the remaining consumer of the fallback (there's a code comment marking it).

Steps to reproduce

Columns are U+10570 (VITHKUQI CAPITAL LETTER A) and U+10597 (its lowercase), a Unicode-14 case pair that JDK 17 treats as distinct. They may render as boxes depending on your font; copy-paste still works, or use the constructed variant below.

CREATE TABLE hadoop_catalog.uni_case (`U+10570` INT, `U+10597` INT) USING iceberg;
INSERT INTO hadoop_catalog.uni_case VALUES (1, 2);
SELECT `U+10570`, `U+10597` FROM hadoop_catalog.db.uni_case;

Font-independent version (spark-shell):

val upper = new String(Character.toChars(0x10570))
val lower = new String(Character.toChars(0x10597))
spark.sql(s"CREATE TABLE hadoop_catalog.uni_case (`$upper` INT, `$lower` INT) USING iceberg")
spark.sql(s"INSERT INTO hadoop_catalog.uni_case VALUES (1, 2)")
spark.sql(s"SELECT `$upper`, `$lower` FROM hadoop_catalog.db.uni_case").collect()

Expected behavior

The query returns [1,2] under Comet exactly as under stock Spark. The adapter should also honor spark.sql.caseSensitive rather than hardcoding insensitive matching.

Additional context

Fix direction: thread the same case tables NativeScanCommon carries through the Iceberg scan config, and pass the real case-sensitivity flag.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:scanParquet scan / data readingbugSomething isn't workingpriority:mediumFunctional bugs, performance regressions, broken features

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions