Skip to content

fix: give a multi-set aggregate the grouping set index Substrait defines - #25527

Open
namanjain24-sudo wants to merge 4 commits into
apache:mainfrom
namanjain24-sudo:fix-substrait-grouping-set-index
Open

namanjain24-sudo wants to merge 4 commits into
apache:mainfrom
namanjain24-sudo:fix-substrait-grouping-set-index

Conversation

@namanjain24-sudo

Copy link
Copy Markdown
Contributor

Which issue does this PR close?

Rationale for this change

Substrait ends an AggregateRel with more than one grouping set with an extra i32 whose value is "the zero-based index of the grouping set that yielded the record" (Aggregate Operation).

DataFusion ends the same aggregate with __grouping_id, which packs two things: a bitmask with a bit set for every grouping column the set leaves out, counting from the last column, and an ordinal that separates repeated sets. The consumer and the producer both treated that column as Substrait's index, so:

  • a consumed plan returned the bitmask where the spec asks for the index, typed UInt8 rather than Int32;
  • a produced plan wrote the bitmask into the column another engine, such as substrait-java, reads as the index.

The two coincide for some lists of sets, which is why this went unnoticed: for (a, b) then (a) both are 0 then 1. For (a) then (b) the bitmask is 1 then 2 while the index is 0 then 1.

What changes are included in this PR?

Both values identify the set a row came from, and every set has its own __grouping_id, so each side can be written as a map of the other. The maps live in a new logical_plan::grouping_set module shared by the two directions:

  • The consumer projects CASE WHEN __grouping_id = <id of set 0> THEN 0 ... ELSE <last index> END, so the plan ends with the index, as a required Int32.
  • The producer leaves the AggregateRel holding what the spec defines and projects __grouping_id back from the index above it, which is also where the reordering to DataFusion's [groups, grouping_id, measures] now happens. A plan DataFusion writes therefore reads correctly in another engine, and still round trips here.

The last set is the ELSE arm rather than a WHEN: the values are exhaustive, and an ELSE keeps the column non-nullable, which Substrait requires of the index and DataFusion of __grouping_id.

Two details the maps have to respect:

  • Set order. Substrait has no ROLLUP or CUBE, so the producer writes both as a list of sets, reversed for ROLLUP and as a powerset for CUBE. The index follows that list, so the expansion is now one function used both to write the groupings and to compute the ids.
  • Repeated sets. GROUPING SETS ((a), (a)) is two sets with two indexes, which DataFusion separates with the ordinal packed above the bitmask, so the map stays one-to-one.

A single grouping set has no index column and is untouched, as is SELECT DISTINCT.

What is the testing strategy for this PR?

Consumer, in aggregation_tests.rs:

  • multiple_grouping_sets_emit_the_set_index is the issue's plan: sets (a) then (b), where the index differs from the bitmask. It checks the column is a required Int32 holding 0 and 1, where main gives a UInt8 holding 1 and 2.
  • duplicate_grouping_sets_are_separate_indexes: the same set twice gets index 0 and 1, with its rows once per index.

Round trip, in roundtrip_logical_plan.rs:

  • aggregate_grouping_sets_keep_grouping_function: GROUPING(a) reads __grouping_id, so this only holds if the index is mapped back to it. The sets are (a), (c), (a, c), chosen so the index and the bitmask differ; with (a, c) first the test would pass either way.
  • aggregate_duplicate_grouping_sets: GROUPING SETS ((a), (a), ()) keeps each occurrence's rows.
  • aggregate_grouping_sets_wider_grouping_id: nine grouping columns, so __grouping_id is a UInt16 and the map has to carry the wider literal.
  • aggregate_grouping_sets now asserts that the AggregateRel no longer remaps its output and that the projection above it carries the map and the mapping.

Each half was reverted on its own to check the tests pin it:

  • without the consumer change, multiple_grouping_sets_emit_the_set_index, aggregate_duplicate_grouping_sets and aggregate_grouping_sets_keep_grouping_function fail;
  • without the producer change, the last two fail, together with aggregate_grouping_sets;
  • with the UInt16 arm of the literal narrowed to UInt8, only aggregate_grouping_sets_wider_grouping_id fails.

cargo test -p datafusion-substrait passes (58 unit, 215 integration with the 6 already ignored, 3 doc tests), as do cargo fmt --all -- --check, cargo clippy -p datafusion-substrait --all-targets --features physical -- -D warnings and cargo xtask ci step test substrait. A Substrait round trip of aggregate.slt, which is not in that job, reports the same 10 pre-existing failures as main.

Are there any user-facing changes?

Plans consumed from Substrait end a multi-set aggregate with the grouping set index, a required Int32, instead of DataFusion's __grouping_id. Plans produced for a multi-set aggregate now carry a projection above the AggregateRel that maps that index back to __grouping_id, so the AggregateRel itself holds the column the spec describes. No Rust API changes.

@github-actions github-actions Bot added the substrait Changes to the substrait crate label Sep 20, 2026
@codecov-commenter

codecov-commenter commented Sep 20, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 86.44068% with 40 lines in your changes missing coverage. Please review.
✅ Project coverage is 82.77%. Comparing base (21a3215) to head (473ecb8).
⚠️ Report is 210 commits behind head on main.

Files with missing lines Patch % Lines
...afusion/substrait/src/logical_plan/grouping_set.rs 87.81% 13 Missing and 11 partials ⚠️
...ait/src/logical_plan/producer/rel/aggregate_rel.rs 87.34% 2 Missing and 8 partials ⚠️
...ait/src/logical_plan/consumer/rel/aggregate_rel.rs 68.42% 0 Missing and 6 partials ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main   #25527      +/-   ##
==========================================
+ Coverage   82.50%   82.77%   +0.26%     
==========================================
  Files        1140     1148       +8     
  Lines      438050   450800   +12750     
  Branches   438050   450800   +12750     
==========================================
+ Hits       361410   373135   +11725     
- Misses      54833    54932      +99     
- Partials    21807    22733     +926     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

Substrait ends an AggregateRel with more than one grouping set with an i32
holding the zero-based index of the set that produced the row. DataFusion
ends the same aggregate with __grouping_id, which packs a bitmask of the
columns a set leaves out with an ordinal that separates repeated sets. The
consumer mapped one onto the other, so a consumed plan returned the bitmask
where the spec asks for the index, and the producer wrote the bitmask into
the column another engine reads as the index.

Both identify the set a row came from, so each is now written as a map of
the other: the consumer projects the index, and the producer projects
__grouping_id back from the index above an AggregateRel that carries what
the spec defines.
@namanjain24-sudo
namanjain24-sudo force-pushed the fix-substrait-grouping-set-index branch from 710f66e to 8d16378 Compare September 24, 2026 13:45
@namanjain24-sudo

Copy link
Copy Markdown
Contributor Author

FYI, cargo test datafusion-cli (amd64) is failing here on a MinIO image pull (quay.io 401 unauthorized on quay.io/minio/minio:RELEASE.2025-02-28T09-55-16Z), not on anything in this PR — same registry issue just hit #25529's and #25664's merge-queue runs today. Nothing else is failing.

@namanjain24-sudo

Copy link
Copy Markdown
Contributor Author

Bumping this — still open for review whenever someone has bandwidth (the earlier MinIO CI failure was unrelated infra flakiness, not this PR).

@kosiew kosiew left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@namanjain24-sudo,

Thanks for working on this. The grouping-set index mapping looks like the right direction, but I found three edge cases that can currently fail valid plans. I also left one optional test suggestion.

}
let ordinal = masks.iter().filter(|seen| **seen == mask).count() as u64;
masks.push(mask);
ids.push((ordinal << group_count) | mask);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This can shift by 64 when group_count == 64, which panics even when the ordinal is zero. Please handle the 64-column zero-ordinal case without shifting, reject nonzero ordinals that need more than 64 bits, and add producer and consumer boundary regressions.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed: at group_count==64 the zero-ordinal case now skips the shift, and a repeated set at 64 columns returns a clean error instead of panicking. Added producer/consumer regressions plus unit tests on grouping_set_ids directly.

Also found while testing: DataFusion's own native aggregate execution (physical-plan/src/aggregates/mod.rs, around line 3216) has the exact same unconditional ordinal << n at n == 64, independent of Substrait. Out of scope here, but flagging in case it's worth its own issue.

(0..grouping_id_index)
.chain(grouping_id_index + 1..schema.fields().len())
.map(|index| Arc::clone(schema.field(index)))
.chain(std::iter::once(Arc::new(index_field)))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

from_unqualified_fields drops qualifiers, so valid joined columns with the same base name can fail serialization as duplicate fields. Please preserve the qualified identities here, or use unique positional placeholders for this temporary schema, and add a joined-column regression.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed: builds the temporary schema from qualified field specs (new_with_metadata) instead of from_unqualified_fields, so joined columns keep their qualifiers. Added a joined-column regression.

}
let grouping_id = Expr::Column(grouping_id);
let grouping_id_type = grouping_id.get_type(schema)?;
let set_index = index_from_grouping_id(&grouping_id, &grouping_id_type, set_ids)?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The fixed grouping_set_index alias can collide with a real user column of the same name on both the consumer and producer paths. Please generate a collision-free synthetic name and use that same chosen name for the producer reference, with regressions for both boundaries.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed: added unique_grouping_set_index_name (shared helper in grouping_set.rs) that falls back to a free name when the schema already has one named grouping_set_index; used by both the producer and the consumer. Added a regression with a real column of that name.

/// With more than eight grouping columns `__grouping_id` is a `UInt16` rather
/// than a `UInt8`, so the map back to it has to carry the wider literal.
#[tokio::test]
async fn aggregate_grouping_sets_wider_grouping_id() -> Result<()> {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggestion: it would be useful to add a roundtrip with exactly eight distinct grouping expressions and a repeated set, so widening to UInt16 is caused only by the ordinal bit. This would complement the existing duplicate-set and wider-ID tests.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added: 8 distinct expressions with one set repeated, isolating that the UInt16 widening follows from the duplicate ordinal bit alone, not the column count.

- grouping_set_ids: at exactly 64 grouping columns the mask alone fills
  the u64, so `ordinal << 64` panicked even when ordinal was 0. Now the
  zero-ordinal case at 64 columns skips the shift, and a repeated set at
  64 columns (which would need a 65th bit) is a clean error instead.
- Producer: the grouping-set-index projection's temporary schema used
  DFSchema::from_unqualified_fields, dropping qualifiers and failing on
  two joined columns that only collide once reduced to a bare name.
  Built from qualified field specs instead.
- Consumer and producer: both used the fixed literal name
  "grouping_set_index" for their synthetic column, which could collide
  with a real column of that name. Added unique_grouping_set_index_name
  to pick a free name in the schema at hand, used by both sides.

Also adds the suggested coverage for 8 distinct grouping expressions
with one repeated, isolating that the UInt16 widening follows from the
duplicate ordinal bit alone, not the column count.

@kosiew kosiew left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@namanjain24-sudo,

Thanks for the follow-up changes. The 64-column guard, qualified schema handling, and collision-free index naming address the earlier concerns.

One correctness issue remains in grouping_set_ids: duplicate ordinals can still overflow when there are fewer than 64 grouping columns. Please add a representability check before shifting and cover the boundary cases with regression tests.

Once that is addressed, I'll be happy to take another look.

}
mask
} else {
(ordinal << group_count) | mask

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The new guard handles 64 columns, but ordinals can still overflow below that limit. With 63 columns and three identical sets, 2 << 63 wraps to 0, giving the first and third sets the same ID and breaking the reverse mapping. Could you check that the ordinal fits in the remaining 64 - group_count bits before shifting, while still allowing valid duplicates and the zero-ordinal 64-column case? Please add boundary tests and producer/consumer rejection tests for this case. checked_shl alone won't catch discarded high bits.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed: computed the representable range from the actual bits left above the mask (64 - group_count) and check it before shifting, for every group_count rather than only 64. Verified by reverting: got exactly the predicted collision, [0, 1<<63, 0] for 3x at 63 columns. Added boundary unit tests (last-ok vs first-overflow at 62 and 63 columns) plus a producer round-trip at the ok boundary and a rejection test at the overflow boundary.

…ping_set_ids

The group_count == 64 guard only caught the one case where the shift
amount itself is out of range. Below 64 columns the shift amount is
always valid, but the *value* can still overflow out of the u64 and
silently lose high bits: at 63 columns a third occurrence of the same
set needs ordinal 2 (0b10), and 2u64 << 63 == 0, colliding with the
first occurrence's id instead of erroring. checked_shl does not catch
this either, since the shift amount is in range.

Now the representable range is computed from the actual bits left
above the mask (64 - group_count) and checked before shifting, for
every group_count rather than only 64. Adds boundary unit tests at 62
and 63 columns (last representable ordinal vs. first that overflows)
and producer round-trip/rejection tests for the same boundary.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

substrait Changes to the substrait crate

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Substrait: the grouping set column holds __grouping_id, not the set's index

3 participants