Repository navigation
Conversation
- Enum descriptor counted once. - Partial retained allocations preserved. - Added None/full/partial capacity and lifecycle tests.
- Test coverage extended to cover longer sequences - Replacement key charges only new allocation, not old key
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #26074 +/- ##
==========================================
- Coverage 82.70% 82.70% -0.01%
==========================================
Files 1147 1147
Lines 447631 447680 +49
Branches 447631 447680 +49
==========================================
+ Hits 370222 370262 +40
+ Misses 54921 54919 -2
- Partials 22488 22499 +11 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
comphead
left a comment
There was a problem hiding this comment.
Thanks @kosiew. The fix is correct. size_of::<GroupOrdering>() already covers the inline variant, and both callers (grouped_hash_stream.rs and common_ordered.rs) only sum component sizes, so nothing is double counted at the call sites. size() runs once per batch over a constant-size state, so no benchmark is needed.
Inline comments cover test dedup and one pre-existing gap in heap_size. Smaller items:
State::sizeis heap-only too. Renaming it toheap_sizemakes the ownership rule uniform inpartial.rs.- Spare
order_indicescapacity cannot arise throughGroupOrdering::try_new(it clones the vec), so that part oftest_size_partial_retained_allocationsonly guardsVecAllocExt::allocated_size. If it stays,Vec::with_capacity(32)pluspush(0)replaces thevec![0],extendandtruncatesetup and makes thecapacity() > len()assert unnecessary. - Optional: a short comment on the
None | Full(_) => 0arm (they hold only inline state) would flag it for anyone who later adds a heap field toGroupOrderingFull, since the enum no longer delegates to it.
| #[test] | ||
| fn test_size_none() { | ||
| assert_eq!(GroupOrdering::None.size(), size_of::<GroupOrdering>()); | ||
| } | ||
|
|
||
| #[test] | ||
| fn test_size_full() -> Result<()> { | ||
| let mut ordering = GroupOrdering::try_new(&InputOrderMode::Sorted)?; | ||
| let expected = size_of::<GroupOrdering>(); | ||
| assert_eq!(ordering.size(), expected); | ||
|
|
||
| ordering.new_groups(&[], &[0, 1, 2], 3)?; | ||
| assert_eq!(ordering.size(), expected); | ||
| ordering.remove_groups(2); | ||
| assert_eq!(ordering.size(), expected); | ||
| ordering.input_done(); | ||
| assert_eq!(ordering.size(), expected); | ||
| ordering.reset(); | ||
| assert_eq!(ordering.size(), expected); | ||
| Ok(()) | ||
| } |
There was a problem hiding this comment.
One test is enough here: assert that None and a Sorted ordering both report exactly size_of::<GroupOrdering>(). None | Full(_) never reads the variant, so the new_groups, remove_groups, input_done and reset steps cannot change the result. The first Full assertion already catches the old code, and test_size_none passed before this PR too.
| ordering.remove_groups(1); | ||
| assert_eq!(ordering.size(), in_progress); | ||
|
|
||
| // Replacing the key charges only the new payload, not the previous key. | ||
| let replacement_key = "updated sort key with a longer payload"; | ||
| let replacement_group_values: Vec<ArrayRef> = | ||
| vec![Arc::new(StringArray::from(vec![replacement_key]))]; | ||
| ordering.new_groups(&replacement_group_values, &[1], 2)?; | ||
| let replaced = | ||
| expected + ScalarValue::Utf8(Some(replacement_key.to_owned())).size(); | ||
| assert!(replaced > in_progress); | ||
| assert_eq!(ordering.size(), replaced); | ||
|
|
||
| // Completing or resetting drops the key, but retains order-index capacity. | ||
| ordering.input_done(); | ||
| assert_eq!(ordering.size(), expected); | ||
| ordering.reset(); | ||
| assert_eq!(ordering.size(), expected); | ||
| ordering.new_groups(&batch_group_values, &[0, 1], 2)?; | ||
| assert_eq!(ordering.size(), in_progress); | ||
| ordering.reset(); | ||
| assert_eq!(ordering.size(), expected); |
There was a problem hiding this comment.
heap_size is a pure function of the current state, so these steps cannot fail independently of the earlier ones:
remove_groupsleavessort_keyuntouched.- The key replacement block hits the same
InProgressarm as the firstnew_groupscall. Replacement itself is already covered bytest_group_ordering_partialinpartial.rs. resetand the re-populate steps repeat theStartandInProgressassertions made above.
The spare-capacity assert, one new_groups call with the Utf8 key and the input_done assert cover every accounting term this PR touches at about half the length.
| /// Returns retained heap allocations, excluding the inline descriptor | ||
| /// already counted by [`super::GroupOrdering::size`]. | ||
| pub(crate) fn heap_size(&self) -> usize { | ||
| self.order_indices.allocated_size() + self.state.size() |
There was a problem hiding this comment.
Not introduced here, so a follow-up under #23393 is fine. order_indices is charged by capacity, but State::size charges sort_key by length. get_row_at_idx collects into Result<Vec<_>>, and a std-only stand-in with a 64-byte element ends at len = 1, capacity = 4. So a one-column key likely holds 3 spare ScalarValue slots (192 bytes) that heap_size does not report, and the in_progress expectation in the new test pins that length-based value. Vec<ScalarValue> already implements DFHeapSize (capacity based), which CountGroupsAccumulator::size uses the same way. ScalarValue::size_of_vec minus the Vec header, as in nth_value.rs, is the other existing option.
Which issue does this PR close?
sizefunctions #23393Rationale for this change
GroupOrdering::size()currently counts both the outer enum descriptor and the inline active variant descriptor. Because the active variant is stored insideGroupOrdering, this double-counts inline memory for partial and full ordering states.This change makes the
GroupOrderingenum the single owner of the descriptor charge. Partial ordering continues to include retainedorder_indicescapacity and state-owned allocations, while full ordering adds no additional descriptor charge.What changes are included in this PR?
GroupOrdering::size()chargesize_of::<GroupOrdering>()exactly once forNone,Partial, andFull.GroupOrderingPartial::size()toheap_size()and make it report only:order_indicesallocationGroupOrderingFull::size(), since full ordering has no additional heap allocation to add beyond the outer enum descriptor.GroupOrdering::size()andGroupOrderingPartial::heap_size().Are these changes tested?
Yes. This PR adds:
test_size_none, which verifies thatGroupOrdering::Nonereports exactlysize_of::<GroupOrdering>().test_size_full, which verifies that full ordering reports exactly oneGroupOrderingdescriptor before and afternew_groups,remove_groups,input_done, andreset.test_size_partial_retained_allocations, which verifies that:order_indicescapacity remains chargedScalarValue::Utf8sort key is included in the reported sizeinput_doneandresetdrop state-owned key allocations while retaining theorder_indicescapacityThe tests use deterministic
size_of, vector capacity, andScalarValue::size()expectations rather than allocator-observed byte counts.Are there any user-facing changes?
No public API changes are introduced.
This changes internal memory accounting reported by
GroupOrdering::size()so that inline descriptor memory is no longer double-counted. Grouping and aggregation behavior is otherwise unchanged by this patch.LLM-generated code disclosure
This PR includes LLM-generated code and comments. All LLM-generated content has been manually reviewed.