Skip to content

fix: match functional dependencies by column position, not by name - #26150

Open
jayzhan211 wants to merge 1 commit into
apache:mainfrom
jayzhan211:fix/fd-match-columns-by-position
Open

jayzhan211 wants to merge 1 commit into
apache:mainfrom
jayzhan211:fix/fd-match-columns-by-position

Conversation

@jayzhan211

Copy link
Copy Markdown
Contributor

Which issue does this PR close?

  • None yet. The bug and its repro are below.

Rationale for this change

If x is a primary key, the optimizer treats CAST(x AS ...) as x. It then drops the other ORDER BY or GROUP BY keys that x determines, so the query returns wrong results:

CREATE TABLE t (x DOUBLE, y VARCHAR, PRIMARY KEY (x))
  AS VALUES (1.1, 'b'), (1.2, 'a'), (2.1, 'c');

SELECT x, y FROM t ORDER BY CAST(x AS INT), y;
-- main:     1.1 b / 1.2 a / 2.1 c   (`y` is dropped from the sort)
-- expected: 1.2 a / 1.1 b / 2.1 c

SELECT CAST(x AS INT) k, count(*) n FROM t GROUP BY CAST(x AS INT), y;
-- main:     (1, 2) / (2, 1)         (`y` is dropped from the grouping)
-- expected: (1, 1) / (1, 1) / (2, 1)

SELECT y, count(*) FROM t GROUP BY CAST(x AS INT);
-- main:     accepted
-- expected: Error: Column in SELECT must be in GROUP BY or an aggregate function

CAST(x AS INT) is named t.x, because casts are left out of expression names. The functional dependency helpers matched GROUP BY and ORDER BY expressions to key columns by comparing names, so the cast matched the key t.x.

The same happens with TRY_CAST, with a UNIQUE NOT NULL key, and with a key that comes from an inner GROUP BY. The ORDER BY case comes from the sort key pruning added in #21362 (54.0.0).

What changes are included in this PR?

  • Index-based helpers. The four helpers in functional_dependencies.rs now take, for each GROUP BY or ORDER BY expression, the index of the input field it references (Option<usize>) instead of its name. A computed expression has no index, so it never matches a key. The new datafusion_expr::utils::passthrough_field_index returns the index for a column reference, aliased or not.
  • One GROUP BY list. The dependencies of an Aggregate's output use the GROUP BY list that its schema is built from (grouping_set_to_exprlist). Before, they used a list de-duplicated by name, which merged CAST(x AS INT) and x.
  • Projections resolve columns. A projection now finds each input column with DFSchema::index_of_column_by_name instead of comparing "qualifier.name" strings. The unit test projection_duplicate_flattened_name_uses_first_input_index asserted the old string behaviour: a column named "orders.id" got the dependency of the different field orders.id. It is renamed and now asserts that each column keeps its own dependency.

What is the testing strategy for this PR?

  • New tests. Section 6 of functional_dependencies.slt has one query for each user of the helpers:

    • ORDER BY pruning;
    • GROUP BY pruning;
    • Aggregate output dependencies;
    • GROUP BY expansion.

    On main, each of these returns a wrong result or accepts an invalid query.

  • Existing tests are unchanged, apart from the rewritten unit test.

  • Planning time. Planning is not slower. The sql_planner TPC-H and TPC-DS benchmarks, whose tables have primary keys, are about 5% faster, because the new code no longer renders expression names.

Benchmark numbers (3 interleaved runs per side)
Benchmark main This PR
physical_plan_tpch_all 17.3 / 17.8 / 18.0 ms 16.1 / 17.0 / 17.2 ms
physical_plan_tpcds_all 291 / 300 / 301 ms 276 / 280 / 282 ms

Are there any user-facing changes?

  • Results. The queries above return correct results.

  • API change. Four pub functions in datafusion_common take &[Option<usize>] instead of &[String]:

    • aggregate_functional_dependencies
    • get_target_functional_dependencies
    • get_required_group_by_exprs_indices
    • get_required_sort_exprs_indices

    The upgrade guide has a migration note.

@jayzhan211 jayzhan211 added the api change Changes the API exposed to users of the crate label Oct 9, 2026
@github-actions github-actions Bot added documentation Improvements or additions to documentation logical-expr Logical plan and expressions optimizer Optimizer rules sqllogictest SQL Logic Tests (.slt) common Related to common crate labels Oct 9, 2026
@jayzhan211
jayzhan211 marked this pull request as ready for review October 9, 2026 05:39
@jayzhan211
jayzhan211 requested review from adriangb and kosiew October 9, 2026 05:39
@adriangb
adriangb requested a balanced review from Copilot October 9, 2026 05:51

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The new passthrough helper can incorrectly resolve ambiguous unqualified columns to the first matching field.

1 open finding
What changed in this PR

Fixes functional-dependency matching by using field positions rather than potentially colliding expression names.

Changes:

  • Adds index-based dependency helpers and passthrough-column resolution.
  • Updates aggregate, projection, GROUP BY, and ORDER BY handling.
  • Adds regression tests and migration documentation.
File Description
docs/​source/​library-user-guide/​upgrading/​56.0.0.md Documents the API migration.
datafusion/​sqllogictest/​test_files/​functional_dependencies.slt Adds CAST regression coverage.
datafusion/​optimizer/​src/​optimize_projections/​mod.rs Uses field indices for GROUP BY pruning.
datafusion/​optimizer/​src/​eliminate_duplicated_expr.rs Uses field indices for sort pruning.
datafusion/​expr/​src/​utils.rs Adds passthrough-field resolution.
datafusion/​expr/​src/​logical_plan/​plan.rs Updates aggregate and projection dependencies.
datafusion/​expr/​src/​logical_plan/​builder.rs Updates implicit GROUP BY expansion.
datafusion/​common/​src/​functional_dependencies.rs Converts dependency helpers to index-based APIs.

🧠 Review effort: Balanced


💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

/// expressions can have the same name, e.g. `CAST(t.a AS INT)` is named `t.a`.
pub fn passthrough_field_index(expr: &Expr, schema: &DFSchema) -> Option<usize> {
match expr {
Expr::Column(col) => schema.maybe_index_of_column(col),
@codecov-commenter

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 96.49123% with 2 lines in your changes missing coverage. Please review.
✅ Project coverage is 82.76%. Comparing base (791660c) to head (e0cbb0a).

Files with missing lines Patch % Lines
datafusion/expr/src/logical_plan/plan.rs 85.71% 0 Missing and 2 partials ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main   #26150      +/-   ##
==========================================
- Coverage   82.76%   82.76%   -0.01%     
==========================================
  Files        1147     1147              
  Lines      450580   450563      -17     
  Branches   450580   450563      -17     
==========================================
- Hits       372944   372921      -23     
- Misses      54929    54932       +3     
- Partials    22707    22710       +3     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@kosiew kosiew left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@jayzhan211,

Thanks for working on this. The functional dependency fix looks sound, and the new regression tests cover the reported CAST issue. I have one optional suggestion for additional test coverage, but it is not blocking. Approving this change.


# 6.4 `y` is not determined by the GROUP BY expression, so it can't be selected.
query error DataFusion error: Error during planning: Column in SELECT must be in GROUP BY or an aggregate function
SELECT y, count(*) FROM t_cast GROUP BY CAST(x AS INT);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we also add a regression test for TRY_CAST over a UNIQUE NOT NULL key? For example, ('bad-a', 'b') and ('bad-b', 'a') both produce NULL when cast to INT, so the test could verify that ORDER BY retains y, GROUP BY preserves both groups, and selecting y when grouping only by the cast fails planning. This is optional coverage since the new helper already handles TRY_CAST conservatively.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

api change Changes the API exposed to users of the crate common Related to common crate documentation Improvements or additions to documentation logical-expr Logical plan and expressions optimizer Optimizer rules sqllogictest SQL Logic Tests (.slt)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants