Skip to content

fix: preserve outer join equivalences for null rows - #26042

Open
Prajwal-k-tech wants to merge 5 commits into
apache:mainfrom
Prajwal-k-tech:prajwal/26036-outer-join-equivalence
Open

Prajwal-k-tech wants to merge 5 commits into
apache:mainfrom
Prajwal-k-tech:prajwal/26036-outer-join-equivalence

Conversation

@Prajwal-k-tech

Copy link
Copy Markdown

Which issue does this PR close?

Rationale for this change

After a LEFT, RIGHT, or FULL OUTER JOIN, the nullable side can contain rows whose columns were extended with NULLs. An equivalence such as b = coalesce(a, 0) is true on the input rows but does not remain true for those null-extended rows. Retaining it can lead to incorrect ordering and LIMIT results.

What changes are included in this PR?

  • Filter nullable-side equivalence classes against a one-row all-NULL batch for the outer-join cases where that side may be null-extended.
  • Ignore volatile expressions during this check and discard classes that are no longer meaningful.
  • Add LEFT and FULL OUTER JOIN SQL regression cases, including the reported ORDER BY/LIMIT behavior.

What is the testing strategy for this PR?

  • Verified the new SQL cases fail on the unmodified code and pass with this change.
  • cargo test -p datafusion-sqllogictest --test sqllogictests -- joins
  • cargo test -p datafusion-physical-expr (1,685 passed, 2 ignored; 13 doctests passed)
  • RUST_BACKTRACE=1 cargo test -p datafusion (passed)
  • RUST_BACKTRACE=1 cargo test -p datafusion-cli (passed)
  • cargo clippy --all-targets --all-features -- -D warnings (passed)
  • cargo +1.99.0 fmt --all -- --check and git diff --check (passed)

This contribution was prepared with AI assistance.

Are there any user-facing changes?

This fixes incorrect query results for affected outer joins. There are no API changes.

Copilot AI balanced review requested due to automatic review settings October 4, 2026 22:20

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@github-actions github-actions Bot added physical-expr Changes to the physical-expr crates sqllogictest SQL Logic Tests (.slt) auto detected api change Auto detected API change labels Oct 4, 2026
@Prajwal-k-tech

Copy link
Copy Markdown
Author

I pushed a regression fix for the CI failure: zero-column batches now preserve their row count, with a focused Full-join test. GitHub marked the new fork workflow runs as action_required before starting jobs. Could a maintainer approve them so CI can verify the updated commit?

@Prajwal-k-tech

Copy link
Copy Markdown
Author

I also addressed the semver bot feedback by restoring the original public EquivalenceGroup::join signature and using a crate-internal schema-aware helper for join planning. The updated head is 5f1cabc; GitHub has created another set of action_required workflows for this commit. Could a maintainer approve the latest runs?

@kumarUjjawal
kumarUjjawal self-requested a review October 5, 2026 06:59
@github-actions github-actions Bot removed the auto detected api change Auto detected API change label Oct 5, 2026
@alamb

alamb commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Thanks @Prajwal-k-tech

Which of these Prs would you like us to review first?

I think it will be eaier to focus and get one done / in rather than trying to do a bunch in parallel

@Prajwal-k-tech

Copy link
Copy Markdown
Author

No preference on my end; going numerically (#26040, then #26041, then #26042) works well. Thanks!

@kumarUjjawal kumarUjjawal left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you @Prajwal-k-tech

Left few comments please take a look.

let expressions = class.into_iter().filter(|expr| {
!is_volatile(expr)
&& expr
.evaluate(null_batch)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Planning now runs every non-volatile expression in an equivalence class, including user and third-party ScalarUDFs, on a synthetic all-NULL batch; it does this every time join properties are recomputed, and a panic inside evaluate is not caught.

&& expr
.evaluate(null_batch)
.and_then(|value| value.into_array(1))
.is_ok_and(|array| array.null_count() == 1)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

array.null_count() == 1 reads the physical null buffer, so arrays with only logical nulls report 0 and their valid equivalences are dropped.

/// that may no longer hold after an outer join.
fn with_null_preserving_expressions(self, null_batch: &RecordBatch) -> Self {
let classes = self.classes.into_iter().filter_map(|class| {
let expressions = class.into_iter().filter(|expr| {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The filter keeps only members that evaluate to NULL, so it drops members that still agree on null-extended rows because they evaluate to the same non-NULL value.

)?);
}
JoinType::Full => {
let batch = null_batch(join_schema)?;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

null_batch allocates a new schema and one null array per output column on every outer-join property computation, even when the nullable side has no equivalence classes.**

schema
.fields()
.iter()
.map(|field| Field::new(field.name(), field.data_type().clone(), true))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

null_batch rebuilds each field with Field::new(name, type, true), which drops field metadata such as extension-type annotations.

statement ok
CREATE TABLE outer_join_sort_right (a BIGINT, b BIGINT) AS VALUES (2, 2), (3, 3);

query III

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The regression tests cover only LEFT and FULL OUTER JOIN; the new JoinType::Right branch (class.rs:874) and the cross-join path are never exercised.

Ok(())
}

use datafusion_expr::Operator;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The new test full_join_equivalences_support_empty_schema was inserted between the test module's use lines, leaving use datafusion_expr::Operator; stranded after a function.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The public EquivalenceGroup::join (no schema) keeps the old unsound outer-join behavior with no doc warning or deprecation, so external callers still get wrong equivalences.

@Prajwal-k-tech

Copy link
Copy Markdown
Author

Thanks for the detailed review. I pushed the follow-up to the existing fork branch in commits 16a78c3 and 63cc845.

  • Outer-join pruning now replaces nullable-side columns with NULL structurally and retains equivalence groups whose rewritten expressions match, including expressions that produce the same non-NULL result. It does not evaluate physical expressions or ScalarUDFs during planning.
  • The synthetic RecordBatch is removed, avoiding its allocations, physical-null-buffer assumptions, and reconstructed fields that lost metadata.
  • The public EquivalenceGroup::join now applies the same conservative outer-join handling without requiring a schema; its docs describe that behavior.
  • Added coverage for RIGHT JOIN, a no-key inner-join path, and expressions with matching non-NULL results after null extension. Moved the stranded import.

Validation: cargo fmt --all -- --check, workspace Clippy (cargo clippy --all-targets --all-features -- -D warnings), all 1,690 datafusion-physical-expr unit tests (1,688 passed, 2 ignored), and the focused joins.slt suite pass. The repository-wide extended tests could not complete because this checkout lacks datafusion/testing/data; a representative failure reports that ARROW_TEST_DATA is unset and asks for the test-data submodule. I have not counted that suite as passing.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

physical-expr Changes to the physical-expr crates sqllogictest SQL Logic Tests (.slt)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Wrong ORDER BY results after an outer join when the nullable side equates a column to coalesce/CASE

4 participants