Skip to content

fix: infer a Parquet column as nullable when some files lack it - #26104

Open
zhuqi-lucas wants to merge 3 commits into
apache:mainfrom
zhuqi-lucas:infer-nullable-partially-present-columns
Open

zhuqi-lucas wants to merge 3 commits into
apache:mainfrom
zhuqi-lucas:infer-nullable-partially-present-columns

Conversation

@zhuqi-lucas

Copy link
Copy Markdown
Contributor

Which issue does this PR close?

Rationale for this change

ListingTable schema inference over Parquet merges the per-file schemas with Schema::try_merge, which widens nullability only for fields that appear in more than one schema. A column that is required in some files and absent from the others therefore comes out NOT NULL, but reading the files without it fills it with nulls, and the scan fails:

Arrow error: Invalid argument error: Column 'c' is declared as non-nullable but contains null values

Pure SQL reproducer (on main):

COPY (SELECT 1 AS id, 10 AS c) TO '/tmp/evo/a.parquet';
COPY (SELECT 2 AS id)          TO '/tmp/evo/b.parquet';
CREATE EXTERNAL TABLE t STORED AS PARQUET LOCATION '/tmp/evo/';
SELECT * FROM t;  -- fails; DESCRIBE shows c Int64 NO

Before DataFusion 52 this was masked by SchemaAdapter::map_batch, which rebuilt batches under the table schema without validating nullability (removed in #18998). The strict check is right; the inferred schema is what is wrong.

What changes are included in this PR?

In ParquetFormat::infer_schema, count in how many files each top-level field appears and, after the merge, mark every field that is missing from at least one file as nullable. Fields present in every file keep whatever try_merge produced, so tables whose files all share a schema are unchanged.

Are these changes tested?

  • New schema_evolution.slt case: two COPY TO files, CREATE EXTERNAL TABLE without a schema, DESCRIBE shows the partial column as nullable, SELECT returns the null. Fails on main with the error above.
  • New schema_merge_marks_partially_present_columns_nullable in core/tests/parquet/schema.rs: three files, column in a different position, an already-nullable partial column, and both skip_metadata paths.

Are there any user-facing changes?

An inferred Parquet table schema now reports a column as nullable when some files do not contain it. Tables where every file has every column are unaffected. Explicitly declared schemas are not touched.

@github-actions github-actions Bot added core Core DataFusion crate sqllogictest SQL Logic Tests (.slt) datasource Changes to the datasource crate labels Oct 7, 2026
ListingTable schema inference merges per-file schemas with
Schema::try_merge, which only widens nullability for fields present in more
than one file. A column that exists, as required, in some files and not at
all in others therefore came out NOT NULL, and reading the files without it
failed with "declared as non-nullable but contains null values". Track
which files carry each field and mark fields that are missing from any file
nullable after the merge.
@zhuqi-lucas
zhuqi-lucas force-pushed the infer-nullable-partially-present-columns branch from a0d90a8 to 872bbd9 Compare October 7, 2026 08:29
@codecov-commenter

codecov-commenter commented Oct 7, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 82.76%. Comparing base (102a162) to head (cf0cf42).
⚠️ Report is 19 commits behind head on main.

Additional details and impacted files
@@            Coverage Diff             @@
##             main   #26104      +/-   ##
==========================================
+ Coverage   82.74%   82.76%   +0.02%     
==========================================
  Files        1147     1147              
  Lines      449767   450603     +836     
  Branches   449767   450603     +836     
==========================================
+ Hits       372157   372953     +796     
+ Misses      54942    54938       -4     
- Partials    22668    22712      +44     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@zhuqi-lucas
zhuqi-lucas marked this pull request as ready for review October 8, 2026 06:19
Copilot AI balanced review requested due to automatic review settings October 8, 2026 06:19

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Approval recommended

The implementation directly addresses the reported failure and includes comprehensive unit and end-to-end regression coverage.

0 open findings

What changed in this PR

Fixes Parquet schema inference so columns absent from some files are nullable, preventing scan failures during schema evolution.

Changes:

  • Tracks top-level field presence across Parquet files and adjusts merged nullability.
  • Adds unit coverage for reordered, required, nullable, and missing columns.
  • Adds an end-to-end SQL logic regression test.
File Description
datafusion/​datasource-parquet/​src/​file_format.rs Corrects inferred nullability for partially present fields.
datafusion/​core/​tests/​parquet/​schema.rs Tests schema inference and reads across evolving files.
datafusion/​sqllogictest/​test_files/​schema_evolution.slt Verifies SQL-visible schema and query results.

🧠 Review effort: Balanced


Give feedback about Copilot approvals in this survey to enter a drawing for a $150 gift card.

----
1 10
2 NULL

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could you please consider adding SELECT id FROM inferred_partial WHERE c IS NULL, expecting 2, here? This would also cover the optimizer behavior: the base incorrectly returns no rows, while this fix returns the expected row.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added in cf0cf42, thanks.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The struct child case is a separate hole, presence is only tracked at the top level here; please do open an issue with your reproducer, or i can help create the issue.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for adding the test and confirming the nested-field case! I’ll open a separate issue with the reproducer and investigate further.

@rgbuilds

rgbuilds commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

I also reproduced a related case where both files contain struct s, but only one contains its required child s.y. On both base and head, s.y is still inferred as non-nullable: selecting it fails when the missing value is filled with NULL, and WHERE s.y IS NULL incorrectly returns no rows.

Would it be useful to track this separately? I’d be happy to open an issue with the reproducer and investigate further.

@rgbuilds rgbuilds left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the fix! It works locally and Rust/SQL tests pass. LGTM.
I’ve included an optional test suggestion and a related follow-up observation for your consideration.

@adriangb adriangb left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @zhuqi-lucas !

@zhuqi-lucas

zhuqi-lucas commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor Author

Thanks @rgbuilds and @adriangb for review.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

core Core DataFusion crate datasource Changes to the datasource crate sqllogictest SQL Logic Tests (.slt)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Parquet schema inference marks a column NOT NULL when it is missing from some files, making the table unreadable

5 participants