Skip to content

GC & Deduplicate String View on Spill - #23565

Merged
adriangb merged 4 commits into
apache:mainfrom
pydantic:spill-dedup-view-arrays
Oct 7, 2026
Merged

adriangb merged 4 commits into
apache:mainfrom
pydantic:spill-dedup-view-arrays

Conversation

@cetra3

@cetra3 cetra3 commented Jul 14, 2026 •

Copy link
Copy Markdown
Contributor

Which issue does this PR close?

Rationale for this change

Spill files inflate string views via GC. Arrow's gc() copies the bytes of every view separately, so a highly repeated string view column (for example a dictionary-encoded Parquet column) spills one copy per row, causing more memory and disk pressure.

What changes are included in this PR?

The spill path now compacts StringView / BinaryView arrays by copying each distinct value once:

  • Values are deduplicated with a hash table over the views, and null views are zeroed.
  • The first 256 non-inline values are sampled. If they contain almost no repeats, compaction falls back to plain gc(), so all-distinct data does not pay for hashing.

Are these changes tested?

Yes:

  • test_gc_copies_repeated_values_once checks that repeated values are written once, both when views share one buffer (as with Parquet dictionary columns) and when they reference separate copies.
  • test_gc_distinct_values checks the fallback to gc() for distinct values.
  • Existing spill tests cover the rest of the path.

Benchmarks (spill_views, sort_tpch with a memory limit) show 0.78–0.94x on repeated-value spills, and no change for all-distinct data or sort_tpch. See the results in the comments.

Are there any user-facing changes?

No, this is an internal method change.

@github-actions github-actions Bot added the physical-plan Changes to the physical-plan crate label Jul 14, 2026
@cetra3

cetra3 commented Jul 14, 2026

Copy link
Copy Markdown
Contributor Author

run benchmarks

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c4969050947-1047-9mz5q 6.12.85+ #1 SMP Mon May 11 08:17:35 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing spill-dedup-view-arrays (25cc81b) to f755cb4 (merge-base) diff using: clickbench_partitioned
Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c4969050947-1048-jrc8t 6.12.85+ #1 SMP Mon May 11 08:17:35 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing spill-dedup-view-arrays (25cc81b) to f755cb4 (merge-base) diff using: tpcds
Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c4969050947-1049-jv4c6 6.12.85+ #1 SMP Mon May 11 08:17:35 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing spill-dedup-view-arrays (25cc81b) to f755cb4 (merge-base) diff using: tpch
Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and spill-dedup-view-arrays
--------------------
Benchmark tpch_sf1.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃                           HEAD ┃        spill-dedup-view-arrays ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 1  │ 39.76 / 42.28 ±2.36 / 46.10 ms │ 38.31 / 39.36 ±1.13 / 41.55 ms │ +1.07x faster │
│ QQuery 2  │ 20.40 / 20.67 ±0.15 / 20.84 ms │ 19.38 / 19.75 ±0.21 / 19.97 ms │     no change │
│ QQuery 3  │ 32.53 / 33.62 ±1.17 / 35.85 ms │ 31.08 / 32.92 ±1.19 / 33.92 ms │     no change │
│ QQuery 4  │ 18.39 / 18.63 ±0.23 / 19.06 ms │ 17.77 / 18.19 ±0.69 / 19.57 ms │     no change │
│ QQuery 5  │ 41.46 / 43.24 ±1.42 / 45.74 ms │ 38.81 / 40.97 ±1.09 / 41.75 ms │ +1.06x faster │
│ QQuery 6  │ 16.97 / 18.09 ±1.77 / 21.61 ms │ 16.29 / 16.61 ±0.29 / 17.14 ms │ +1.09x faster │
│ QQuery 7  │ 50.62 / 51.41 ±0.78 / 52.72 ms │ 44.85 / 46.22 ±0.85 / 47.26 ms │ +1.11x faster │
│ QQuery 8  │ 46.31 / 46.75 ±0.27 / 47.05 ms │ 43.39 / 43.60 ±0.14 / 43.81 ms │ +1.07x faster │
│ QQuery 9  │ 53.80 / 56.20 ±1.64 / 58.08 ms │ 49.75 / 50.35 ±0.70 / 51.73 ms │ +1.12x faster │
│ QQuery 10 │ 46.12 / 47.10 ±1.29 / 49.53 ms │ 42.49 / 42.62 ±0.07 / 42.71 ms │ +1.11x faster │
│ QQuery 11 │ 14.76 / 15.23 ±0.65 / 16.51 ms │ 13.75 / 14.09 ±0.33 / 14.72 ms │ +1.08x faster │
│ QQuery 12 │ 25.58 / 27.38 ±2.28 / 31.51 ms │ 23.98 / 24.51 ±0.28 / 24.82 ms │ +1.12x faster │
│ QQuery 13 │ 34.22 / 35.84 ±1.47 / 38.41 ms │ 32.56 / 34.22 ±1.70 / 37.26 ms │     no change │
│ QQuery 14 │ 25.80 / 26.05 ±0.38 / 26.81 ms │ 23.89 / 24.05 ±0.12 / 24.22 ms │ +1.08x faster │
│ QQuery 15 │ 33.68 / 33.92 ±0.20 / 34.18 ms │ 31.49 / 31.64 ±0.22 / 32.08 ms │ +1.07x faster │
│ QQuery 16 │ 14.91 / 15.07 ±0.11 / 15.20 ms │ 13.92 / 14.27 ±0.20 / 14.54 ms │ +1.06x faster │
│ QQuery 17 │ 78.08 / 81.94 ±3.68 / 87.54 ms │ 73.34 / 74.14 ±0.46 / 74.62 ms │ +1.11x faster │
│ QQuery 18 │ 62.08 / 64.01 ±1.86 / 67.06 ms │ 60.27 / 61.21 ±0.60 / 62.16 ms │     no change │
│ QQuery 19 │ 34.21 / 34.48 ±0.19 / 34.70 ms │ 33.15 / 33.43 ±0.44 / 34.30 ms │     no change │
│ QQuery 20 │ 33.32 / 34.07 ±1.08 / 36.21 ms │ 32.50 / 32.78 ±0.24 / 33.16 ms │     no change │
│ QQuery 21 │ 56.18 / 58.90 ±1.57 / 61.01 ms │ 54.99 / 57.40 ±1.66 / 59.99 ms │     no change │
│ QQuery 22 │ 14.28 / 14.47 ±0.17 / 14.77 ms │ 14.15 / 14.29 ±0.08 / 14.39 ms │     no change │
└───────────┴────────────────────────────────┴────────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━┓
┃ Benchmark Summary                      ┃          ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━┩
│ Total Time (HEAD)                      │ 819.37ms │
│ Total Time (spill-dedup-view-arrays)   │ 766.60ms │
│ Average Time (HEAD)                    │  37.24ms │
│ Average Time (spill-dedup-view-arrays) │  34.85ms │
│ Queries Faster                         │       13 │
│ Queries Slower                         │        0 │
│ Queries with No Change                 │        9 │
│ Queries with Failure                   │        0 │
└────────────────────────────────────────┴──────────┘

Resource Usage

tpch — base (merge-base)

Metric Value
Wall time 5.0s
Peak memory 1.2 GiB
Avg memory 672.2 MiB
CPU user 23.5s
CPU sys 1.9s
Peak spill 0 B

tpch — branch

Metric Value
Wall time 5.0s
Peak memory 1.2 GiB
Avg memory 531.0 MiB
CPU user 22.2s
CPU sys 1.7s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and spill-dedup-view-arrays
--------------------
Benchmark tpcds_sf1.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃                                  HEAD ┃               spill-dedup-view-arrays ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 1  │           5.47 / 5.98 ±0.88 / 7.75 ms │           5.60 / 6.13 ±0.85 / 7.82 ms │     no change │
│ QQuery 2  │        80.96 / 81.51 ±0.44 / 82.19 ms │        80.62 / 81.03 ±0.28 / 81.36 ms │     no change │
│ QQuery 3  │        29.29 / 29.53 ±0.15 / 29.70 ms │        29.14 / 29.43 ±0.16 / 29.62 ms │     no change │
│ QQuery 4  │     475.66 / 480.38 ±5.78 / 489.39 ms │     485.60 / 488.82 ±4.76 / 498.17 ms │     no change │
│ QQuery 5  │        51.65 / 52.10 ±0.33 / 52.60 ms │        52.18 / 52.59 ±0.34 / 53.12 ms │     no change │
│ QQuery 6  │        36.81 / 37.28 ±0.46 / 38.17 ms │        37.05 / 37.20 ±0.17 / 37.45 ms │     no change │
│ QQuery 7  │        94.14 / 94.63 ±0.35 / 95.18 ms │        94.21 / 94.79 ±0.38 / 95.29 ms │     no change │
│ QQuery 8  │        36.90 / 37.29 ±0.34 / 37.90 ms │        37.08 / 37.57 ±0.55 / 38.62 ms │     no change │
│ QQuery 9  │        52.74 / 54.35 ±1.43 / 56.29 ms │        54.73 / 55.23 ±0.42 / 56.00 ms │     no change │
│ QQuery 10 │        62.62 / 62.97 ±0.32 / 63.38 ms │        63.03 / 63.39 ±0.23 / 63.71 ms │     no change │
│ QQuery 11 │     293.29 / 295.79 ±1.52 / 297.29 ms │     292.55 / 295.59 ±1.75 / 297.54 ms │     no change │
│ QQuery 12 │        28.56 / 28.93 ±0.33 / 29.53 ms │        28.40 / 28.77 ±0.50 / 29.72 ms │     no change │
│ QQuery 13 │     118.39 / 119.23 ±0.59 / 120.01 ms │     119.05 / 120.06 ±0.79 / 120.94 ms │     no change │
│ QQuery 14 │     414.53 / 419.74 ±4.20 / 425.27 ms │     418.00 / 420.22 ±2.50 / 424.67 ms │     no change │
│ QQuery 15 │        57.30 / 58.17 ±0.61 / 59.10 ms │        57.35 / 58.04 ±0.60 / 59.17 ms │     no change │
│ QQuery 16 │           6.78 / 6.89 ±0.15 / 7.17 ms │           6.83 / 6.96 ±0.12 / 7.17 ms │     no change │
│ QQuery 17 │        80.64 / 81.68 ±1.31 / 84.21 ms │        81.48 / 82.23 ±0.93 / 84.03 ms │     no change │
│ QQuery 18 │     124.33 / 125.52 ±1.47 / 128.35 ms │     124.33 / 124.98 ±0.61 / 125.96 ms │     no change │
│ QQuery 19 │        41.28 / 41.51 ±0.16 / 41.71 ms │        41.74 / 42.10 ±0.35 / 42.62 ms │     no change │
│ QQuery 20 │        35.46 / 35.86 ±0.28 / 36.17 ms │        35.54 / 35.98 ±0.29 / 36.34 ms │     no change │
│ QQuery 21 │        17.41 / 17.72 ±0.17 / 17.89 ms │        17.52 / 17.79 ±0.23 / 18.07 ms │     no change │
│ QQuery 22 │        62.37 / 63.71 ±1.48 / 66.54 ms │        62.46 / 63.76 ±2.02 / 67.75 ms │     no change │
│ QQuery 23 │     347.33 / 348.96 ±1.26 / 351.07 ms │     343.44 / 347.81 ±3.50 / 354.13 ms │     no change │
│ QQuery 24 │     225.96 / 229.60 ±5.94 / 241.39 ms │     226.43 / 229.68 ±2.35 / 232.42 ms │     no change │
│ QQuery 25 │     111.08 / 112.04 ±0.89 / 113.73 ms │     111.19 / 112.31 ±1.03 / 113.88 ms │     no change │
│ QQuery 26 │        57.66 / 59.07 ±2.17 / 63.39 ms │        57.63 / 59.87 ±3.34 / 66.36 ms │     no change │
│ QQuery 27 │           6.28 / 6.41 ±0.15 / 6.70 ms │           6.28 / 6.46 ±0.21 / 6.86 ms │     no change │
│ QQuery 28 │        61.46 / 62.21 ±0.71 / 63.42 ms │        61.97 / 62.63 ±0.86 / 64.32 ms │     no change │
│ QQuery 29 │       97.33 / 98.19 ±1.08 / 100.17 ms │       97.77 / 99.77 ±2.18 / 103.61 ms │     no change │
│ QQuery 30 │        32.25 / 34.38 ±2.99 / 40.23 ms │        32.51 / 33.45 ±0.75 / 34.65 ms │     no change │
│ QQuery 31 │     110.69 / 112.33 ±1.72 / 115.50 ms │     112.24 / 113.11 ±0.48 / 113.67 ms │     no change │
│ QQuery 32 │        20.50 / 20.66 ±0.27 / 21.20 ms │        20.33 / 20.68 ±0.33 / 21.29 ms │     no change │
│ QQuery 33 │        37.91 / 38.31 ±0.37 / 38.95 ms │        37.95 / 40.06 ±3.84 / 47.73 ms │     no change │
│ QQuery 34 │         9.85 / 10.01 ±0.18 / 10.35 ms │         9.82 / 10.01 ±0.21 / 10.40 ms │     no change │
│ QQuery 35 │        73.03 / 74.88 ±1.99 / 78.68 ms │        73.21 / 73.97 ±0.40 / 74.34 ms │     no change │
│ QQuery 36 │           5.94 / 6.09 ±0.15 / 6.37 ms │           5.88 / 6.04 ±0.18 / 6.40 ms │     no change │
│ QQuery 37 │           6.92 / 7.05 ±0.06 / 7.11 ms │           6.98 / 7.11 ±0.12 / 7.32 ms │     no change │
│ QQuery 38 │        62.21 / 62.80 ±0.30 / 63.08 ms │        62.92 / 63.11 ±0.15 / 63.34 ms │     no change │
│ QQuery 39 │        90.53 / 90.91 ±0.39 / 91.58 ms │        91.18 / 92.82 ±1.83 / 96.35 ms │     no change │
│ QQuery 40 │        23.21 / 24.50 ±0.70 / 25.06 ms │        23.40 / 23.80 ±0.27 / 24.13 ms │     no change │
│ QQuery 41 │        11.57 / 11.71 ±0.17 / 12.04 ms │        11.57 / 11.77 ±0.20 / 12.15 ms │     no change │
│ QQuery 42 │        23.72 / 24.08 ±0.29 / 24.60 ms │        23.92 / 24.11 ±0.15 / 24.36 ms │     no change │
│ QQuery 43 │           5.04 / 5.11 ±0.09 / 5.28 ms │           5.14 / 5.25 ±0.13 / 5.50 ms │     no change │
│ QQuery 44 │           9.23 / 9.42 ±0.11 / 9.57 ms │           9.70 / 9.76 ±0.06 / 9.84 ms │     no change │
│ QQuery 45 │        38.87 / 39.17 ±0.19 / 39.42 ms │        38.53 / 38.96 ±0.25 / 39.25 ms │     no change │
│ QQuery 46 │        11.69 / 11.97 ±0.20 / 12.14 ms │        11.94 / 12.14 ±0.11 / 12.25 ms │     no change │
│ QQuery 47 │     227.08 / 229.72 ±1.95 / 232.66 ms │     225.44 / 227.67 ±2.41 / 232.34 ms │     no change │
│ QQuery 48 │       96.61 / 98.03 ±1.53 / 100.87 ms │        96.65 / 97.55 ±0.87 / 99.21 ms │     no change │
│ QQuery 49 │        76.30 / 77.23 ±1.29 / 79.78 ms │        77.04 / 77.97 ±0.60 / 78.65 ms │     no change │
│ QQuery 50 │        58.36 / 59.50 ±0.70 / 60.25 ms │        58.98 / 62.57 ±3.87 / 69.80 ms │  1.05x slower │
│ QQuery 51 │        90.98 / 92.89 ±2.20 / 96.62 ms │        92.43 / 94.43 ±1.23 / 96.17 ms │     no change │
│ QQuery 52 │        23.48 / 23.94 ±0.25 / 24.25 ms │        24.16 / 24.42 ±0.29 / 24.97 ms │     no change │
│ QQuery 53 │        29.31 / 29.64 ±0.28 / 30.13 ms │        30.33 / 30.59 ±0.30 / 31.17 ms │     no change │
│ QQuery 54 │        55.08 / 57.50 ±3.44 / 64.32 ms │        56.63 / 59.05 ±2.98 / 64.81 ms │     no change │
│ QQuery 55 │        23.14 / 23.53 ±0.57 / 24.66 ms │        24.27 / 24.79 ±0.31 / 25.16 ms │  1.05x slower │
│ QQuery 56 │        39.19 / 39.89 ±0.58 / 40.90 ms │        39.72 / 40.17 ±0.50 / 41.13 ms │     no change │
│ QQuery 57 │     176.55 / 180.83 ±3.85 / 188.04 ms │     178.98 / 181.17 ±2.32 / 185.42 ms │     no change │
│ QQuery 58 │     115.91 / 117.99 ±1.62 / 120.84 ms │     116.98 / 119.13 ±2.41 / 123.58 ms │     no change │
│ QQuery 59 │     119.38 / 120.32 ±1.13 / 121.98 ms │     119.29 / 120.52 ±0.70 / 121.36 ms │     no change │
│ QQuery 60 │        39.38 / 40.29 ±0.65 / 41.21 ms │        40.08 / 40.57 ±0.65 / 41.81 ms │     no change │
│ QQuery 61 │        12.86 / 13.00 ±0.20 / 13.39 ms │        12.91 / 13.09 ±0.14 / 13.33 ms │     no change │
│ QQuery 62 │        46.65 / 48.11 ±2.37 / 52.82 ms │        47.21 / 48.35 ±1.90 / 52.14 ms │     no change │
│ QQuery 63 │        30.18 / 30.79 ±0.61 / 31.73 ms │        30.53 / 31.74 ±1.52 / 34.74 ms │     no change │
│ QQuery 64 │     408.48 / 415.90 ±7.68 / 429.92 ms │     419.87 / 421.64 ±1.59 / 424.32 ms │     no change │
│ QQuery 65 │     145.49 / 149.23 ±2.92 / 153.88 ms │     146.68 / 149.17 ±1.54 / 150.58 ms │     no change │
│ QQuery 66 │        79.88 / 82.67 ±3.89 / 90.37 ms │        81.41 / 82.44 ±0.97 / 84.26 ms │     no change │
│ QQuery 67 │     236.24 / 241.32 ±4.47 / 247.68 ms │     239.08 / 242.90 ±2.52 / 246.03 ms │     no change │
│ QQuery 68 │        11.87 / 12.07 ±0.26 / 12.58 ms │        12.13 / 12.27 ±0.19 / 12.64 ms │     no change │
│ QQuery 69 │        57.04 / 57.63 ±0.54 / 58.56 ms │        58.74 / 60.35 ±2.82 / 65.99 ms │     no change │
│ QQuery 70 │     106.38 / 111.64 ±7.32 / 125.95 ms │     106.27 / 110.01 ±4.45 / 117.87 ms │     no change │
│ QQuery 71 │        35.71 / 36.03 ±0.26 / 36.49 ms │        36.06 / 36.31 ±0.16 / 36.47 ms │     no change │
│ QQuery 72 │ 2090.46 / 2210.07 ±80.08 / 2310.63 ms │ 2163.70 / 2219.16 ±48.84 / 2298.38 ms │     no change │
│ QQuery 73 │          9.55 / 9.82 ±0.25 / 10.25 ms │         9.66 / 12.36 ±3.92 / 19.98 ms │  1.26x slower │
│ QQuery 74 │     168.61 / 172.01 ±3.12 / 177.72 ms │     168.68 / 171.05 ±2.22 / 173.97 ms │     no change │
│ QQuery 75 │     149.63 / 152.55 ±4.16 / 160.80 ms │     149.87 / 151.10 ±0.85 / 152.34 ms │     no change │
│ QQuery 76 │        35.55 / 36.18 ±0.52 / 36.88 ms │        35.31 / 35.82 ±0.26 / 36.01 ms │     no change │
│ QQuery 77 │        61.76 / 62.13 ±0.23 / 62.40 ms │        61.58 / 62.02 ±0.45 / 62.87 ms │     no change │
│ QQuery 78 │     197.32 / 200.01 ±2.10 / 202.70 ms │     195.17 / 200.98 ±5.61 / 209.00 ms │     no change │
│ QQuery 79 │        67.50 / 72.05 ±6.50 / 84.64 ms │        67.75 / 69.85 ±3.53 / 76.88 ms │     no change │
│ QQuery 80 │      99.99 / 101.59 ±1.63 / 103.82 ms │      99.49 / 100.93 ±1.24 / 102.96 ms │     no change │
│ QQuery 81 │        25.63 / 25.91 ±0.17 / 26.15 ms │        25.96 / 28.24 ±3.72 / 35.66 ms │  1.09x slower │
│ QQuery 82 │        16.29 / 16.76 ±0.25 / 17.03 ms │        17.27 / 18.93 ±1.85 / 21.72 ms │  1.13x slower │
│ QQuery 83 │        40.66 / 44.96 ±4.96 / 54.29 ms │        40.60 / 41.08 ±0.65 / 42.34 ms │ +1.09x faster │
│ QQuery 84 │        30.85 / 31.18 ±0.21 / 31.45 ms │        30.88 / 31.07 ±0.10 / 31.15 ms │     no change │
│ QQuery 85 │     107.28 / 108.31 ±0.80 / 109.31 ms │     107.16 / 108.96 ±1.49 / 111.69 ms │     no change │
│ QQuery 86 │        25.24 / 25.45 ±0.27 / 25.90 ms │        24.94 / 26.25 ±1.00 / 27.93 ms │     no change │
│ QQuery 87 │        63.38 / 65.58 ±2.74 / 70.82 ms │        63.16 / 63.95 ±0.82 / 65.48 ms │     no change │
│ QQuery 88 │        62.82 / 63.50 ±0.55 / 64.15 ms │        63.80 / 64.33 ±0.48 / 65.03 ms │     no change │
│ QQuery 89 │        35.59 / 35.96 ±0.22 / 36.19 ms │        35.57 / 37.12 ±1.92 / 40.88 ms │     no change │
│ QQuery 90 │        17.20 / 17.37 ±0.16 / 17.67 ms │        17.37 / 17.55 ±0.20 / 17.89 ms │     no change │
│ QQuery 91 │        45.89 / 47.22 ±1.15 / 49.24 ms │        46.77 / 47.58 ±0.66 / 48.67 ms │     no change │
│ QQuery 92 │        29.03 / 29.74 ±0.75 / 31.16 ms │        29.29 / 29.90 ±0.72 / 31.27 ms │     no change │
│ QQuery 93 │        49.75 / 50.47 ±0.79 / 51.93 ms │        50.03 / 50.64 ±0.47 / 51.46 ms │     no change │
│ QQuery 94 │        38.22 / 38.91 ±0.50 / 39.67 ms │        38.28 / 38.70 ±0.34 / 39.21 ms │     no change │
│ QQuery 95 │        81.96 / 83.58 ±1.69 / 86.66 ms │        82.66 / 84.66 ±1.75 / 86.85 ms │     no change │
│ QQuery 96 │        24.39 / 24.71 ±0.29 / 25.14 ms │        24.47 / 24.63 ±0.15 / 24.89 ms │     no change │
│ QQuery 97 │        46.64 / 47.40 ±0.62 / 48.39 ms │        47.06 / 47.37 ±0.25 / 47.61 ms │     no change │
│ QQuery 98 │        42.10 / 42.91 ±0.44 / 43.35 ms │        42.63 / 43.44 ±0.50 / 43.92 ms │     no change │
│ QQuery 99 │        70.09 / 70.46 ±0.32 / 71.05 ms │        69.93 / 71.81 ±2.11 / 75.67 ms │     no change │
└───────────┴───────────────────────────────────────┴───────────────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary                      ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)                      │ 10007.10ms │
│ Total Time (spill-dedup-view-arrays)   │ 10055.70ms │
│ Average Time (HEAD)                    │   101.08ms │
│ Average Time (spill-dedup-view-arrays) │   101.57ms │
│ Queries Faster                         │          1 │
│ Queries Slower                         │          5 │
│ Queries with No Change                 │         93 │
│ Queries with Failure                   │          0 │
└────────────────────────────────────────┴────────────┘

Resource Usage

tpcds — base (merge-base)

Metric Value
Wall time 55.0s
Peak memory 2.0 GiB
Avg memory 1.3 GiB
CPU user 228.6s
CPU sys 5.8s
Peak spill 0 B

tpcds — branch

Metric Value
Wall time 55.0s
Peak memory 2.2 GiB
Avg memory 1.5 GiB
CPU user 222.2s
CPU sys 5.9s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and spill-dedup-view-arrays
--------------------
Benchmark clickbench_partitioned.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃                                  HEAD ┃               spill-dedup-view-arrays ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 0  │          1.25 / 4.02 ±5.44 / 14.89 ms │          1.22 / 3.99 ±5.44 / 14.88 ms │     no change │
│ QQuery 1  │        12.70 / 12.84 ±0.09 / 12.95 ms │        12.40 / 12.77 ±0.27 / 13.23 ms │     no change │
│ QQuery 2  │        35.80 / 36.36 ±0.36 / 36.79 ms │        35.67 / 35.95 ±0.36 / 36.61 ms │     no change │
│ QQuery 3  │        30.82 / 31.49 ±0.68 / 32.75 ms │        30.59 / 30.80 ±0.19 / 31.15 ms │     no change │
│ QQuery 4  │     219.80 / 228.85 ±4.65 / 232.90 ms │     220.22 / 227.50 ±4.02 / 230.88 ms │     no change │
│ QQuery 5  │     272.53 / 275.50 ±2.40 / 279.69 ms │     268.71 / 274.43 ±4.67 / 282.29 ms │     no change │
│ QQuery 6  │           1.26 / 1.41 ±0.24 / 1.88 ms │           1.24 / 1.41 ±0.24 / 1.88 ms │     no change │
│ QQuery 7  │        13.53 / 13.67 ±0.16 / 13.91 ms │        13.40 / 13.61 ±0.15 / 13.78 ms │     no change │
│ QQuery 8  │     318.06 / 326.57 ±5.60 / 332.74 ms │     323.03 / 331.22 ±4.58 / 336.78 ms │     no change │
│ QQuery 9  │     457.93 / 463.57 ±3.65 / 468.97 ms │    455.97 / 468.33 ±10.04 / 484.29 ms │     no change │
│ QQuery 10 │        70.47 / 71.30 ±0.73 / 72.40 ms │        70.60 / 75.63 ±8.38 / 92.35 ms │  1.06x slower │
│ QQuery 11 │        81.92 / 82.76 ±0.50 / 83.44 ms │        80.64 / 82.29 ±2.61 / 87.47 ms │     no change │
│ QQuery 12 │     271.52 / 276.53 ±4.62 / 284.48 ms │     268.29 / 279.93 ±8.78 / 294.33 ms │     no change │
│ QQuery 13 │    378.37 / 393.95 ±14.04 / 418.05 ms │    375.01 / 383.14 ±10.56 / 403.19 ms │     no change │
│ QQuery 14 │     285.60 / 291.61 ±4.83 / 299.89 ms │     279.54 / 292.76 ±7.92 / 300.60 ms │     no change │
│ QQuery 15 │     273.96 / 280.09 ±3.87 / 284.58 ms │     274.98 / 285.61 ±7.05 / 292.92 ms │     no change │
│ QQuery 16 │    627.08 / 639.96 ±12.19 / 656.35 ms │    624.06 / 634.64 ±15.43 / 665.17 ms │     no change │
│ QQuery 17 │    628.06 / 637.35 ±12.65 / 661.47 ms │     627.76 / 636.06 ±5.87 / 643.34 ms │     no change │
│ QQuery 18 │ 1277.71 / 1297.36 ±16.85 / 1326.15 ms │ 1273.05 / 1309.98 ±25.78 / 1349.45 ms │     no change │
│ QQuery 19 │        28.87 / 34.42 ±9.45 / 53.27 ms │        27.47 / 27.74 ±0.21 / 28.02 ms │ +1.24x faster │
│ QQuery 20 │    515.66 / 527.27 ±11.49 / 543.78 ms │     519.96 / 525.18 ±5.83 / 536.00 ms │     no change │
│ QQuery 21 │     515.86 / 522.39 ±6.74 / 534.31 ms │     514.76 / 523.97 ±8.11 / 537.60 ms │     no change │
│ QQuery 22 │    993.42 / 999.55 ±6.77 / 1009.95 ms │   981.34 / 996.05 ±13.43 / 1014.97 ms │     no change │
│ QQuery 23 │ 3062.01 / 3098.23 ±38.93 / 3173.76 ms │ 3068.63 / 3120.26 ±37.76 / 3162.40 ms │     no change │
│ QQuery 24 │        41.41 / 42.73 ±1.37 / 45.20 ms │        40.81 / 42.03 ±0.89 / 43.45 ms │     no change │
│ QQuery 25 │     113.66 / 114.04 ±0.41 / 114.76 ms │     111.03 / 111.93 ±0.78 / 113.33 ms │     no change │
│ QQuery 26 │        41.52 / 42.20 ±0.47 / 43.00 ms │        41.58 / 43.34 ±2.05 / 46.74 ms │     no change │
│ QQuery 27 │    670.59 / 679.48 ±10.59 / 700.12 ms │     668.73 / 681.20 ±7.41 / 690.93 ms │     no change │
│ QQuery 28 │ 3019.28 / 3046.35 ±23.56 / 3084.36 ms │ 3060.72 / 3075.89 ±18.86 / 3112.02 ms │     no change │
│ QQuery 29 │        40.87 / 42.24 ±2.20 / 46.63 ms │       40.88 / 48.06 ±12.64 / 73.28 ms │  1.14x slower │
│ QQuery 30 │    307.57 / 318.44 ±18.14 / 354.52 ms │     299.62 / 307.09 ±7.17 / 319.99 ms │     no change │
│ QQuery 31 │     278.20 / 290.83 ±7.95 / 301.96 ms │    277.20 / 293.43 ±12.97 / 309.85 ms │     no change │
│ QQuery 32 │   923.31 / 957.06 ±30.85 / 1003.81 ms │    949.60 / 971.42 ±17.08 / 990.64 ms │     no change │
│ QQuery 33 │ 1477.83 / 1499.48 ±15.84 / 1515.48 ms │ 1434.93 / 1489.19 ±35.53 / 1530.22 ms │     no change │
│ QQuery 34 │ 1497.68 / 1529.92 ±17.91 / 1546.25 ms │ 1491.65 / 1528.99 ±25.64 / 1561.06 ms │     no change │
│ QQuery 35 │    278.59 / 306.51 ±29.30 / 358.79 ms │    283.22 / 335.27 ±74.30 / 482.55 ms │  1.09x slower │
│ QQuery 36 │        66.18 / 72.47 ±4.66 / 78.49 ms │       65.94 / 77.11 ±10.67 / 96.44 ms │  1.06x slower │
│ QQuery 37 │        35.48 / 37.67 ±2.20 / 41.91 ms │        35.49 / 37.25 ±1.92 / 40.79 ms │     no change │
│ QQuery 38 │        41.60 / 42.64 ±1.12 / 44.74 ms │       41.43 / 48.82 ±13.17 / 75.13 ms │  1.14x slower │
│ QQuery 39 │     143.52 / 151.58 ±8.57 / 167.87 ms │    138.54 / 151.49 ±10.32 / 163.38 ms │     no change │
│ QQuery 40 │        14.10 / 14.89 ±0.46 / 15.53 ms │        14.07 / 14.42 ±0.34 / 15.05 ms │     no change │
│ QQuery 41 │        13.71 / 13.99 ±0.32 / 14.58 ms │        13.55 / 14.05 ±0.31 / 14.52 ms │     no change │
│ QQuery 42 │        13.11 / 13.49 ±0.32 / 13.98 ms │        13.11 / 13.40 ±0.17 / 13.65 ms │     no change │
└───────────┴───────────────────────────────────────┴───────────────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary                      ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)                      │ 19763.05ms │
│ Total Time (spill-dedup-view-arrays)   │ 19857.62ms │
│ Average Time (HEAD)                    │   459.61ms │
│ Average Time (spill-dedup-view-arrays) │   461.81ms │
│ Queries Faster                         │          1 │
│ Queries Slower                         │          5 │
│ Queries with No Change                 │         37 │
│ Queries with Failure                   │          0 │
└────────────────────────────────────────┴────────────┘

Resource Usage

clickbench_partitioned — base (merge-base)

Metric Value
Wall time 100.0s
Peak memory 11.4 GiB
Avg memory 4.6 GiB
CPU user 1018.4s
CPU sys 68.9s
Peak spill 0 B

clickbench_partitioned — branch

Metric Value
Wall time 100.0s
Peak memory 10.3 GiB
Avg memory 4.1 GiB
CPU user 1014.1s
CPU sys 71.7s
Peak spill 0 B

File an issue against this benchmark runner

@cetra3
cetra3 force-pushed the spill-dedup-view-arrays branch from 25cc81b to d6079c7 Compare July 14, 2026 12:45
@cetra3

cetra3 commented Jul 14, 2026

Copy link
Copy Markdown
Contributor Author

Those benchmarks don't actually stress this path at all. In fact I don't think there are any standard benchmarks that will sweat this particular change. There is a sort_tpch benchmark but that has a high cardinality l_comment string view column, which, once again, doesn't hit this path or show any benefits here.

fn gc_dedup_view<T: ByteViewType>(
array: &GenericByteViewArray<T>,
) -> GenericByteViewArray<T> {
let mut builder = GenericByteViewBuilder::<T>::with_capacity(array.len())

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Every qualifying Utf8View and BinaryView now allocates a hash table sized for all rows and hashes every non-inline value. This is valuable for repeated values, but unique strings or large binary values get no disk reduction over gc() while paying extra CPU and peak memory during spilling.

@cetra3 cetra3 Aug 11, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes this is the trade off. However, I believe the size of this is actually dictated by the configured batch size so it's never going to be more than that.

So with 16384 as the default batch size this could use ~ 16384 * 64 = 1MiB <+ hash internals> RAM to write this out + CPU overhead of hash lookups which I think is ~ log(n). So yeah high cardinality string views will probably cause increased memory/cpu but not by much. Keeping in mind that we are spilling here to reduce RAM usage, and so when we consider the re-hydrated spilled file, the memory usage is much more reduced for low cardinality views.

In high cardinality scenarios this will add overhead, but it's hard to measure without doing a second pass to work out what the cardinality is. Maybe there is a better way to gauge the cardinality of the input view and use that to decide? Not sure there is a good method for this.

But, in production we have seen a good reduction in file usage/memory pressure with gc & dedup (i.e, we have applied this as a patch to our version of DF), so I feel that for real use cases the trade off is worth it.

@@ -1345,7 +1359,7 @@ mod tests {
let size_without_gc: usize = array_without_gc
.data_buffers()
.iter()
.map(|buffer| buffer.capacity())
.map(|buffer| buffer.len())

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Keep coverage for both length and capacity.

@kumarUjjawal

Copy link
Copy Markdown
Contributor

Those benchmarks don't actually stress this path at all. In fact I don't think there are any standard benchmarks that will sweat this particular change. There is a sort_tpch benchmark but that has a high cardinality l_comment string view column, which, once again, doesn't hit this path or show any benefits here.

we can compare the repeated-label reproduction with a forced-spill high-cardinality case such as sort-TPCH Q3.

@cetra3
cetra3 force-pushed the spill-dedup-view-arrays branch 2 times, most recently from d6079c7 to 20bc3e1 Compare August 11, 2026 03:42
@cetra3

cetra3 commented Aug 11, 2026

Copy link
Copy Markdown
Contributor Author

run benchmark sort_tpch

env:
DATAFUSION_RUNTIME_MEMORY_LIMIT: "256M"

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5248827939-1538-nxfvd 6.12.85+ #1 SMP Wed Jun 17 20:31:55 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing spill-dedup-view-arrays (20bc3e1) to a9b61ab (merge-base) diff

Run configuration
run benchmark sort_tpch
env:
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "256M"

Results will be posted here when complete


File an issue against this benchmark runner

@cetra3

cetra3 commented Aug 11, 2026

Copy link
Copy Markdown
Contributor Author

run benchmark sort_tpch

env:
DATAFUSION_RUNTIME_MEMORY_LIMIT: "512M"

@codecov-commenter

codecov-commenter commented Aug 11, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 88.95349% with 19 lines in your changes missing coverage. Please review.
✅ Project coverage is 82.73%. Comparing base (1c43063) to head (0335ffa).

Files with missing lines Patch % Lines
datafusion/physical-plan/src/spill/mod.rs 88.95% 1 Missing and 18 partials ⚠️
Additional details and impacted files
@@           Coverage Diff            @@
##             main   #23565    +/-   ##
========================================
  Coverage   82.73%   82.73%            
========================================
  Files        1147     1147            
  Lines      449389   449540   +151     
  Branches   449389   449540   +151     
========================================
+ Hits       371788   371924   +136     
+ Misses      54943    54940     -3     
- Partials    22658    22676    +18     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5248861843-1539-vqgpb 6.12.85+ #1 SMP Wed Jun 17 20:31:55 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing spill-dedup-view-arrays (20bc3e1) to a9b61ab (merge-base) diff

Run configuration
run benchmark sort_tpch
env:
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "512M"

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing spill-dedup-view-arrays (20bc3e1) to a9b61ab (merge-base) diff

Run configuration
run benchmark sort_tpch
env:
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "256M"
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and spill-dedup-view-arrays
--------------------
Benchmark sort_tpch1.json
--------------------
┏━━━━━━━┳━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━┓
┃ Query ┃      HEAD ┃ spill-dedup-view-arrays ┃       Change ┃
┡━━━━━━━╇━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━┩
│ Q1    │ 121.46 ms │               121.23 ms │    no change │
│ Q2    │ 117.93 ms │               116.89 ms │    no change │
│ Q3    │      FAIL │                    FAIL │ incomparable │
│ Q4    │ 179.42 ms │               176.85 ms │    no change │
│ Q5    │ 319.50 ms │               315.05 ms │    no change │
│ Q6    │ 332.11 ms │               335.62 ms │    no change │
│ Q7    │      FAIL │                    FAIL │ incomparable │
│ Q8    │      FAIL │                    FAIL │ incomparable │
│ Q9    │      FAIL │                    FAIL │ incomparable │
│ Q10   │      FAIL │                    FAIL │ incomparable │
│ Q11   │      FAIL │                    FAIL │ incomparable │
└───────┴───────────┴─────────────────────────┴──────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary                      ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)                      │ 1070.42ms │
│ Total Time (spill-dedup-view-arrays)   │ 1065.63ms │
│ Average Time (HEAD)                    │  214.08ms │
│ Average Time (spill-dedup-view-arrays) │  213.13ms │
│ Queries Faster                         │         0 │
│ Queries Slower                         │         0 │
│ Queries with No Change                 │         5 │
│ Queries with Failure                   │         6 │
└────────────────────────────────────────┴───────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and spill-dedup-view-arrays
--------------------
Benchmark sort_tpch1.json
--------------------
┏━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━┓
┃ Query ┃                              HEAD ┃           spill-dedup-view-arrays ┃       Change ┃
┡━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━┩
│ Q1    │ 121.46 / 122.26 ±0.89 / 123.92 ms │ 121.23 / 122.02 ±0.74 / 123.11 ms │    no change │
│ Q2    │ 117.93 / 120.64 ±2.59 / 125.51 ms │ 116.89 / 118.47 ±1.28 / 119.67 ms │    no change │
│ Q3    │                              FAIL │                              FAIL │ incomparable │
│ Q4    │ 179.42 / 183.07 ±3.03 / 187.73 ms │ 176.85 / 180.64 ±2.20 / 182.97 ms │    no change │
│ Q5    │ 319.50 / 332.54 ±6.73 / 338.12 ms │ 315.05 / 327.28 ±8.44 / 339.33 ms │    no change │
│ Q6    │ 332.11 / 334.52 ±2.43 / 338.76 ms │ 335.62 / 341.01 ±5.78 / 349.72 ms │    no change │
│ Q7    │                              FAIL │                              FAIL │ incomparable │
│ Q8    │                              FAIL │                              FAIL │ incomparable │
│ Q9    │                              FAIL │                              FAIL │ incomparable │
│ Q10   │                              FAIL │                              FAIL │ incomparable │
│ Q11   │                              FAIL │                              FAIL │ incomparable │
└───────┴───────────────────────────────────┴───────────────────────────────────┴──────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary                      ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)                      │ 1093.04ms │
│ Total Time (spill-dedup-view-arrays)   │ 1089.41ms │
│ Average Time (HEAD)                    │  218.61ms │
│ Average Time (spill-dedup-view-arrays) │  217.88ms │
│ Queries Faster                         │         0 │
│ Queries Slower                         │         0 │
│ Queries with No Change                 │         5 │
│ Queries with Failure                   │         6 │
└────────────────────────────────────────┴───────────┘

Memory Pool Peaks

Peak MemoryPool reservation per query — what DataFusion's accounting believes it reserved. Recorded only when the benchmark runs with DATAFUSION_RUNTIME_MEMORY_LIMIT set.

Base: a9b61ab (merge-base) | Changed: spill-dedup-view-arrays

sort_tpch — sort_tpch1

Query Base Changed Change
1 240.8 MiB 241.2 MiB +0.1%
2 248.5 MiB 250.5 MiB +0.8%
3 254.4 MiB 254.9 MiB +0.2%
4 250.9 MiB 249.7 MiB -0.5%
5 247.5 MiB 250.6 MiB +1.3%
6 249.5 MiB 245.1 MiB -1.8%
7 259.5 MiB 253.3 MiB -2.4%
8 255.7 MiB 255.5 MiB -0.1%
9 256.9 MiB 257.9 MiB +0.4%
10 255.1 MiB 255.8 MiB +0.3%
11 256.2 MiB 254.1 MiB -0.8%

Pool accounting vs. process RSS

Max pool peak is the largest reservation any single query in the run reached; peak RSS covers the whole invocation, including data loading and allocator retention, and the two high-water marks need not coincide in time. The gap is therefore an upper bound on what the pool did not account for, not a measurement of it.

Benchmark Side Max pool peak Peak RSS Gap RSS / pool
sort_tpch base (a9b61ab (merge-base)) 259.5 MiB 1.2 GiB 1011.7 MiB 4.9×
sort_tpch changed (spill-dedup-view-arrays) 257.9 MiB 1.3 GiB 1.0 GiB 5.0×
Resource Usage

sort_tpch — base (merge-base)

Metric Value
Wall time 10.0s
Peak memory 1.2 GiB
Avg memory 577.1 MiB
CPU user 28.4s
CPU sys 5.5s
Peak spill 0 B

sort_tpch — branch

Metric Value
Wall time 10.0s
Peak memory 1.3 GiB
Avg memory 569.4 MiB
CPU user 23.1s
CPU sys 3.3s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing spill-dedup-view-arrays (20bc3e1) to a9b61ab (merge-base) diff

Run configuration
run benchmark sort_tpch
env:
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "512M"
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and spill-dedup-view-arrays
--------------------
Benchmark sort_tpch1.json
--------------------
┏━━━━━━━┳━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━┓
┃ Query ┃      HEAD ┃ spill-dedup-view-arrays ┃       Change ┃
┡━━━━━━━╇━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━┩
│ Q1    │ 120.96 ms │               121.04 ms │    no change │
│ Q2    │ 105.45 ms │               104.87 ms │    no change │
│ Q3    │ 582.82 ms │               609.50 ms │    no change │
│ Q4    │ 197.22 ms │               196.43 ms │    no change │
│ Q5    │ 264.57 ms │               263.53 ms │    no change │
│ Q6    │ 276.51 ms │               276.76 ms │    no change │
│ Q7    │      FAIL │                    FAIL │ incomparable │
│ Q8    │ 382.54 ms │               387.57 ms │    no change │
│ Q9    │ 413.23 ms │               415.20 ms │    no change │
│ Q10   │      FAIL │                    FAIL │ incomparable │
│ Q11   │ 247.95 ms │               252.52 ms │    no change │
└───────┴───────────┴─────────────────────────┴──────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary                      ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)                      │ 2591.24ms │
│ Total Time (spill-dedup-view-arrays)   │ 2627.42ms │
│ Average Time (HEAD)                    │  287.92ms │
│ Average Time (spill-dedup-view-arrays) │  291.94ms │
│ Queries Faster                         │         0 │
│ Queries Slower                         │         0 │
│ Queries with No Change                 │         9 │
│ Queries with Failure                   │         2 │
└────────────────────────────────────────┴───────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and spill-dedup-view-arrays
--------------------
Benchmark sort_tpch1.json
--------------------
┏━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━┓
┃ Query ┃                              HEAD ┃           spill-dedup-view-arrays ┃       Change ┃
┡━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━┩
│ Q1    │ 120.96 / 121.64 ±0.87 / 123.23 ms │ 121.04 / 121.94 ±0.84 / 123.53 ms │    no change │
│ Q2    │ 105.45 / 106.33 ±1.05 / 107.90 ms │ 104.87 / 106.37 ±1.41 / 108.52 ms │    no change │
│ Q3    │ 582.82 / 586.85 ±3.16 / 591.54 ms │ 609.50 / 611.27 ±1.30 / 612.81 ms │    no change │
│ Q4    │ 197.22 / 199.23 ±2.16 / 202.18 ms │ 196.43 / 198.91 ±1.58 / 201.36 ms │    no change │
│ Q5    │ 264.57 / 265.77 ±1.12 / 267.35 ms │ 263.53 / 264.73 ±1.46 / 267.11 ms │    no change │
│ Q6    │ 276.51 / 278.07 ±1.64 / 281.14 ms │ 276.76 / 277.97 ±0.83 / 279.34 ms │    no change │
│ Q7    │                              FAIL │                              FAIL │ incomparable │
│ Q8    │ 382.54 / 388.93 ±4.68 / 395.35 ms │ 387.57 / 396.01 ±6.72 / 405.82 ms │    no change │
│ Q9    │ 413.23 / 418.91 ±8.37 / 435.57 ms │ 415.20 / 421.61 ±5.88 / 429.31 ms │    no change │
│ Q10   │                              FAIL │                              FAIL │ incomparable │
│ Q11   │ 247.95 / 256.02 ±7.97 / 270.72 ms │ 252.52 / 259.40 ±4.67 / 265.63 ms │    no change │
└───────┴───────────────────────────────────┴───────────────────────────────────┴──────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary                      ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)                      │ 2621.75ms │
│ Total Time (spill-dedup-view-arrays)   │ 2658.21ms │
│ Average Time (HEAD)                    │  291.31ms │
│ Average Time (spill-dedup-view-arrays) │  295.36ms │
│ Queries Faster                         │         0 │
│ Queries Slower                         │         0 │
│ Queries with No Change                 │         9 │
│ Queries with Failure                   │         2 │
└────────────────────────────────────────┴───────────┘

Memory Pool Peaks

Peak MemoryPool reservation per query — what DataFusion's accounting believes it reserved. Recorded only when the benchmark runs with DATAFUSION_RUNTIME_MEMORY_LIMIT set.

Base: a9b61ab (merge-base) | Changed: spill-dedup-view-arrays

sort_tpch — sort_tpch1

Query Base Changed Change
1 240.1 MiB 241.2 MiB +0.4%
2 279.9 MiB 282.1 MiB +0.8%
3 510.7 MiB 503.0 MiB -1.5%
4 319.6 MiB 322.1 MiB +0.8%
5 316.8 MiB 321.1 MiB +1.4%
6 405.0 MiB 407.1 MiB +0.5%
7 506.0 MiB 502.8 MiB -0.6%
8 502.1 MiB 503.2 MiB +0.2%
9 500.7 MiB 498.1 MiB -0.5%
10 503.6 MiB 509.8 MiB +1.2%
11 508.3 MiB 520.5 MiB +2.4%

Pool accounting vs. process RSS

Max pool peak is the largest reservation any single query in the run reached; peak RSS covers the whole invocation, including data loading and allocator retention, and the two high-water marks need not coincide in time. The gap is therefore an upper bound on what the pool did not account for, not a measurement of it.

Benchmark Side Max pool peak Peak RSS Gap RSS / pool
sort_tpch base (a9b61ab (merge-base)) 510.7 MiB 1.7 GiB 1.2 GiB 3.4×
sort_tpch changed (spill-dedup-view-arrays) 520.5 MiB 1.5 GiB 1.0 GiB 3.0×
Resource Usage

sort_tpch — base (merge-base)

Metric Value
Wall time 15.0s
Peak memory 1.7 GiB
Avg memory 923.5 MiB
CPU user 52.0s
CPU sys 7.8s
Peak spill 0 B

sort_tpch — branch

Metric Value
Wall time 15.0s
Peak memory 1.5 GiB
Avg memory 882.8 MiB
CPU user 53.5s
CPU sys 7.9s
Peak spill 0 B

File an issue against this benchmark runner

@cetra3

cetra3 commented Aug 11, 2026

Copy link
Copy Markdown
Contributor Author

run benchmark sort_tpch

env:
DATAFUSION_RUNTIME_MEMORY_LIMIT: "512M"
DATAFUSION_EXECUTION_TARGET_PARTITIONS: "4"

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5249003089-1540-tzvxj 6.12.85+ #1 SMP Wed Jun 17 20:31:55 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing spill-dedup-view-arrays (20bc3e1) to a9b61ab (merge-base) diff

Run configuration
run benchmark sort_tpch
env:
  DATAFUSION_EXECUTION_TARGET_PARTITIONS: "4"
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "512M"

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing spill-dedup-view-arrays (20bc3e1) to a9b61ab (merge-base) diff

Run configuration
run benchmark sort_tpch
env:
  DATAFUSION_EXECUTION_TARGET_PARTITIONS: "4"
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "512M"
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and spill-dedup-view-arrays
--------------------
Benchmark sort_tpch1.json
--------------------
┏━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━┓
┃ Query ┃       HEAD ┃ spill-dedup-view-arrays ┃       Change ┃
┡━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━┩
│ Q1    │  150.39 ms │               151.06 ms │    no change │
│ Q2    │  135.83 ms │               136.40 ms │    no change │
│ Q3    │  602.63 ms │               699.14 ms │ 1.16x slower │
│ Q4    │  249.09 ms │               252.54 ms │    no change │
│ Q5    │  264.93 ms │               267.72 ms │    no change │
│ Q6    │  290.46 ms │               292.75 ms │    no change │
│ Q7    │  725.17 ms │               734.68 ms │    no change │
│ Q8    │  503.43 ms │               528.52 ms │    no change │
│ Q9    │  540.82 ms │               565.44 ms │    no change │
│ Q10   │ 1227.66 ms │              1263.74 ms │    no change │
│ Q11   │  425.69 ms │               454.23 ms │ 1.07x slower │
└───────┴────────────┴─────────────────────────┴──────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary                      ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)                      │ 5116.11ms │
│ Total Time (spill-dedup-view-arrays)   │ 5346.22ms │
│ Average Time (HEAD)                    │  465.10ms │
│ Average Time (spill-dedup-view-arrays) │  486.02ms │
│ Queries Faster                         │         0 │
│ Queries Slower                         │         2 │
│ Queries with No Change                 │         9 │
│ Queries with Failure                   │         0 │
└────────────────────────────────────────┴───────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and spill-dedup-view-arrays
--------------------
Benchmark sort_tpch1.json
--------------------
┏━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━┓
┃ Query ┃                                 HEAD ┃               spill-dedup-view-arrays ┃       Change ┃
┡━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━┩
│ Q1    │    150.39 / 151.27 ±1.12 / 153.35 ms │     151.06 / 151.79 ±0.92 / 153.55 ms │    no change │
│ Q2    │    135.83 / 138.62 ±2.36 / 142.10 ms │     136.40 / 139.36 ±3.03 / 143.18 ms │    no change │
│ Q3    │    602.63 / 613.91 ±6.65 / 622.24 ms │     699.14 / 706.24 ±7.39 / 720.05 ms │ 1.15x slower │
│ Q4    │    249.09 / 256.85 ±5.68 / 263.50 ms │     252.54 / 257.31 ±4.92 / 266.73 ms │    no change │
│ Q5    │    264.93 / 268.12 ±2.14 / 270.99 ms │     267.72 / 271.01 ±1.86 / 273.09 ms │    no change │
│ Q6    │    290.46 / 295.84 ±3.90 / 301.12 ms │     292.75 / 296.51 ±2.50 / 300.19 ms │    no change │
│ Q7    │    725.17 / 735.03 ±5.19 / 738.83 ms │     734.68 / 742.74 ±5.68 / 750.06 ms │    no change │
│ Q8    │    503.43 / 506.24 ±2.00 / 508.36 ms │     528.52 / 532.73 ±3.08 / 537.99 ms │ 1.05x slower │
│ Q9    │    540.82 / 549.17 ±7.11 / 559.09 ms │     565.44 / 572.55 ±5.80 / 579.23 ms │    no change │
│ Q10   │ 1227.66 / 1237.65 ±8.35 / 1251.47 ms │ 1263.74 / 1276.78 ±12.75 / 1298.87 ms │    no change │
│ Q11   │    425.69 / 431.41 ±3.38 / 436.28 ms │     454.23 / 457.47 ±3.83 / 464.54 ms │ 1.06x slower │
└───────┴──────────────────────────────────────┴───────────────────────────────────────┴──────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary                      ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)                      │ 5184.10ms │
│ Total Time (spill-dedup-view-arrays)   │ 5404.49ms │
│ Average Time (HEAD)                    │  471.28ms │
│ Average Time (spill-dedup-view-arrays) │  491.32ms │
│ Queries Faster                         │         0 │
│ Queries Slower                         │         3 │
│ Queries with No Change                 │         8 │
│ Queries with Failure                   │         0 │
└────────────────────────────────────────┴───────────┘

Memory Pool Peaks

Peak MemoryPool reservation per query — what DataFusion's accounting believes it reserved. Recorded only when the benchmark runs with DATAFUSION_RUNTIME_MEMORY_LIMIT set.

Base: a9b61ab (merge-base) | Changed: spill-dedup-view-arrays

sort_tpch — sort_tpch1

Query Base Changed Change
1 174.9 MiB 175.5 MiB +0.3%
2 219.9 MiB 220.6 MiB +0.3%
3 505.0 MiB 503.0 MiB -0.4%
4 266.1 MiB 264.6 MiB -0.6%
5 295.0 MiB 295.0 MiB +0.0%
6 356.1 MiB 355.7 MiB -0.1%
7 509.0 MiB 506.4 MiB -0.5%
8 505.8 MiB 504.1 MiB -0.3%
9 505.1 MiB 505.1 MiB +0.0%
10 510.1 MiB 510.1 MiB +0.0%
11 508.9 MiB 508.9 MiB +0.0%

Pool accounting vs. process RSS

Max pool peak is the largest reservation any single query in the run reached; peak RSS covers the whole invocation, including data loading and allocator retention, and the two high-water marks need not coincide in time. The gap is therefore an upper bound on what the pool did not account for, not a measurement of it.

Benchmark Side Max pool peak Peak RSS Gap RSS / pool
sort_tpch base (a9b61ab (merge-base)) 510.1 MiB 1.8 GiB 1.3 GiB 3.7×
sort_tpch changed (spill-dedup-view-arrays) 510.1 MiB 1.5 GiB 990.1 MiB 2.9×
Resource Usage

sort_tpch — base (merge-base)

Metric Value
Wall time 30.0s
Peak memory 1.8 GiB
Avg memory 987.4 MiB
CPU user 79.9s
CPU sys 16.6s
Peak spill 0 B

sort_tpch — branch

Metric Value
Wall time 30.0s
Peak memory 1.5 GiB
Avg memory 1022.6 MiB
CPU user 83.0s
CPU sys 17.0s
Peak spill 0 B

File an issue against this benchmark runner

@adriangb
adriangb force-pushed the spill-dedup-view-arrays branch from 20bc3e1 to cad94ba Compare September 15, 2026 14:41
@adriangb
adriangb force-pushed the spill-dedup-view-arrays branch from cad94ba to ca22a41 Compare September 22, 2026 15:53
@adriangb

adriangb commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Local results with the spill_views SQL suite from #25625 (1M rows from Parquet, 64-byte strings, each query sets its own memory limit and spills on both sides). Both builds use the same main commit plus the suite.

Time: 3 batches × 10 interleaved rounds (random arm order; arms main, main, PR, PR), 100 timings per arm per batch, median. Spill bytes from EXPLAIN ANALYZE. M4 Pro laptop SSD, machine under load, so the bot results are better for absolute numbers.

Query Spilled main Spilled PR Time main Time PR PR / main (batch 1, 2, 3) A/A range
q01 sort, 1 distinct value 84.2 MB 23.2 MB 54–66 ms 54–66 ms 1.01x, 0.99x, 1.00x 0.99–1.03x
q02 sort, 1000 distinct values 84.2 MB 30.6 MB 52–64 ms 63–74 ms 1.16x, 1.17x, 1.22x 0.98–1.03x
q03 sort, BinaryView, 1000 distinct values 84.2 MB 30.6 MB 51–61 ms 63–74 ms 1.20x, 1.21x, 1.23x 1.00–1.06x
q04 sort, all distinct 84.2 MB 84.2 MB 65–75 ms 86–104 ms 1.37x, 1.32x, 1.33x 0.98–1.01x
q05 GROUP BY, all distinct 84.2 MB 84.2 MB 65–90 ms 71–101 ms 1.12x, 1.07x, 1.08x 0.95–1.03x
q05 failures at 96M 0 / 80 2 / 80
#23564 repro (shuffled key), 64M 83.2 MB 30.5 MB

Bot results with the suite will follow after #25625 merges and this PR is rebased.

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing spill-dedup-view-arrays (9c2f8f1) to ed43e69 (merge-base) diff

Run configuration
run benchmark sort_tpch
env:
  DATAFUSION_EXECUTION_TARGET_PARTITIONS: "4"
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "512M"
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and spill-dedup-view-arrays
--------------------
Benchmark sort_tpch1.json
--------------------
┏━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━┓
┃ Query ┃       HEAD ┃ spill-dedup-view-arrays ┃       Change ┃
┡━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━┩
│ Q1    │  135.77 ms │               135.43 ms │    no change │
│ Q2    │  136.39 ms │               137.98 ms │    no change │
│ Q3    │       FAIL │                    FAIL │ incomparable │
│ Q4    │  198.64 ms │               201.01 ms │    no change │
│ Q5    │  229.05 ms │               231.48 ms │    no change │
│ Q6    │  253.46 ms │               253.52 ms │    no change │
│ Q7    │  693.15 ms │               670.37 ms │    no change │
│ Q8    │  477.49 ms │               472.08 ms │    no change │
│ Q9    │  520.82 ms │               524.98 ms │    no change │
│ Q10   │ 1221.39 ms │              1209.75 ms │    no change │
│ Q11   │  374.48 ms │               370.50 ms │    no change │
└───────┴────────────┴─────────────────────────┴──────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary                      ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)                      │ 4240.66ms │
│ Total Time (spill-dedup-view-arrays)   │ 4207.08ms │
│ Average Time (HEAD)                    │  424.07ms │
│ Average Time (spill-dedup-view-arrays) │  420.71ms │
│ Queries Faster                         │         0 │
│ Queries Slower                         │         0 │
│ Queries with No Change                 │        10 │
│ Queries with Failure                   │         1 │
└────────────────────────────────────────┴───────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and spill-dedup-view-arrays
--------------------
Benchmark sort_tpch1.json
--------------------
┏━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━┓
┃ Query ┃                                  HEAD ┃               spill-dedup-view-arrays ┃       Change ┃
┡━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━┩
│ Q1    │     135.77 / 136.60 ±0.91 / 138.13 ms │     135.43 / 136.28 ±1.34 / 138.92 ms │    no change │
│ Q2    │     136.39 / 137.38 ±1.03 / 138.64 ms │     137.98 / 139.76 ±1.27 / 141.91 ms │    no change │
│ Q3    │                                  FAIL │                                  FAIL │ incomparable │
│ Q4    │     198.64 / 204.05 ±3.88 / 208.53 ms │     201.01 / 205.01 ±3.40 / 210.60 ms │    no change │
│ Q5    │     229.05 / 232.24 ±2.74 / 237.04 ms │     231.48 / 234.62 ±1.82 / 236.51 ms │    no change │
│ Q6    │     253.46 / 256.68 ±2.41 / 259.31 ms │     253.52 / 257.43 ±2.75 / 261.46 ms │    no change │
│ Q7    │     693.15 / 696.44 ±2.64 / 700.04 ms │     670.37 / 681.34 ±9.61 / 696.97 ms │    no change │
│ Q8    │     477.49 / 487.06 ±5.58 / 493.82 ms │     472.08 / 479.51 ±5.29 / 486.68 ms │    no change │
│ Q9    │     520.82 / 528.81 ±5.57 / 534.84 ms │     524.98 / 526.64 ±1.28 / 528.48 ms │    no change │
│ Q10   │ 1221.39 / 1234.24 ±10.21 / 1252.06 ms │ 1209.75 / 1230.46 ±14.64 / 1247.69 ms │    no change │
│ Q11   │     374.48 / 376.96 ±2.46 / 380.83 ms │     370.50 / 375.86 ±5.44 / 384.93 ms │    no change │
└───────┴───────────────────────────────────────┴───────────────────────────────────────┴──────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary                      ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)                      │ 4290.47ms │
│ Total Time (spill-dedup-view-arrays)   │ 4266.93ms │
│ Average Time (HEAD)                    │  429.05ms │
│ Average Time (spill-dedup-view-arrays) │  426.69ms │
│ Queries Faster                         │         0 │
│ Queries Slower                         │         0 │
│ Queries with No Change                 │        10 │
│ Queries with Failure                   │         1 │
└────────────────────────────────────────┴───────────┘

Memory Pool Peaks

Peak MemoryPool reservation per query — what DataFusion's accounting believes it reserved. Recorded only for benchmarks that write a results JSON (the dfbench suites, and the suites bench.sh runs through the Criterion SQL harness since #25644), and only under a memory limit: DATAFUSION_RUNTIME_MEMORY_LIMIT, or one the suite sets itself.

Base: ed43e69 (merge-base) | Changed: spill-dedup-view-arrays

sort_tpch — sort_tpch1

Query Base Changed Change
1 175.1 MiB 174.9 MiB -0.1%
2 220.4 MiB 219.6 MiB -0.3%
3 518.9 MiB 539.3 MiB +3.9%
4 265.2 MiB 264.9 MiB -0.1%
5 294.0 MiB 294.0 MiB +0.0%
6 355.7 MiB 355.2 MiB -0.1%
7 509.0 MiB 509.0 MiB +0.0%
8 508.7 MiB 508.7 MiB +0.0%
9 497.9 MiB 499.7 MiB +0.4%
10 510.1 MiB 510.1 MiB +0.0%
11 524.7 MiB 522.2 MiB -0.5%

Pool accounting vs. process RSS

Max pool peak is the largest reservation any single query in the run reached; peak RSS covers the whole invocation, including data loading and allocator retention, and the two high-water marks need not coincide in time. The gap is therefore an upper bound on what the pool did not account for, not a measurement of it.

Benchmark Side Max pool peak Peak RSS Gap RSS / pool
sort_tpch base (ed43e69 (merge-base)) 524.7 MiB 1.2 GiB 746.5 MiB 2.4×
sort_tpch changed (spill-dedup-view-arrays) 539.3 MiB 1.1 GiB 602.7 MiB 2.1×
Resource Usage

sort_tpch — base (merge-base)

Metric Value
Wall time 25.0s
Peak memory 1.2 GiB
Avg memory 838.2 MiB
CPU user 63.2s
CPU sys 15.1s
Peak spill 973.0 MiB

sort_tpch — branch

Metric Value
Wall time 25.0s
Peak memory 1.1 GiB
Avg memory 792.2 MiB
CPU user 64.0s
CPU sys 15.0s
Peak spill 1013.6 MiB

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing spill-dedup-view-arrays (9c2f8f1) to ed43e69 (merge-base) diff

Run configuration
run benchmark sort_tpch
env:
  DATAFUSION_EXECUTION_TARGET_PARTITIONS: "4"
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "512M"
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and spill-dedup-view-arrays
--------------------
Benchmark sort_tpch1.json
--------------------
┏━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━┓
┃ Query ┃       HEAD ┃ spill-dedup-view-arrays ┃       Change ┃
┡━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━┩
│ Q1    │  136.56 ms │               133.79 ms │    no change │
│ Q2    │  137.59 ms │               136.45 ms │    no change │
│ Q3    │       FAIL │                    FAIL │ incomparable │
│ Q4    │  201.39 ms │               197.86 ms │    no change │
│ Q5    │  230.92 ms │               226.47 ms │    no change │
│ Q6    │  253.33 ms │               251.33 ms │    no change │
│ Q7    │  687.93 ms │               658.99 ms │    no change │
│ Q8    │  483.50 ms │               479.94 ms │    no change │
│ Q9    │  521.77 ms │               528.19 ms │    no change │
│ Q10   │ 1209.55 ms │              1215.69 ms │    no change │
│ Q11   │  365.47 ms │               367.81 ms │    no change │
└───────┴────────────┴─────────────────────────┴──────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary                      ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)                      │ 4228.00ms │
│ Total Time (spill-dedup-view-arrays)   │ 4196.51ms │
│ Average Time (HEAD)                    │  422.80ms │
│ Average Time (spill-dedup-view-arrays) │  419.65ms │
│ Queries Faster                         │         0 │
│ Queries Slower                         │         0 │
│ Queries with No Change                 │        10 │
│ Queries with Failure                   │         1 │
└────────────────────────────────────────┴───────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and spill-dedup-view-arrays
--------------------
Benchmark sort_tpch1.json
--------------------
┏━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query ┃                                 HEAD ┃               spill-dedup-view-arrays ┃        Change ┃
┡━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ Q1    │    136.56 / 138.54 ±1.99 / 142.21 ms │     133.79 / 134.76 ±1.00 / 136.66 ms │     no change │
│ Q2    │    137.59 / 138.87 ±1.31 / 141.37 ms │     136.45 / 138.99 ±1.96 / 141.41 ms │     no change │
│ Q3    │                                 FAIL │                                  FAIL │  incomparable │
│ Q4    │    201.39 / 205.84 ±3.15 / 209.32 ms │     197.86 / 199.57 ±1.20 / 201.57 ms │     no change │
│ Q5    │    230.92 / 233.35 ±1.63 / 235.57 ms │     226.47 / 229.98 ±2.04 / 232.23 ms │     no change │
│ Q6    │    253.33 / 257.42 ±2.24 / 259.50 ms │     251.33 / 255.29 ±3.90 / 261.63 ms │     no change │
│ Q7    │   687.93 / 703.31 ±10.09 / 718.83 ms │     658.99 / 665.73 ±8.37 / 681.86 ms │ +1.06x faster │
│ Q8    │   483.50 / 508.69 ±17.56 / 532.85 ms │     479.94 / 488.06 ±5.86 / 496.35 ms │     no change │
│ Q9    │    521.77 / 535.65 ±9.39 / 543.85 ms │     528.19 / 540.41 ±9.54 / 556.68 ms │     no change │
│ Q10   │ 1209.55 / 1223.07 ±8.42 / 1231.72 ms │ 1215.69 / 1230.15 ±17.24 / 1262.27 ms │     no change │
│ Q11   │   365.47 / 379.06 ±10.53 / 395.51 ms │     367.81 / 370.66 ±2.51 / 374.99 ms │     no change │
└───────┴──────────────────────────────────────┴───────────────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary                      ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)                      │ 4323.80ms │
│ Total Time (spill-dedup-view-arrays)   │ 4253.61ms │
│ Average Time (HEAD)                    │  432.38ms │
│ Average Time (spill-dedup-view-arrays) │  425.36ms │
│ Queries Faster                         │         1 │
│ Queries Slower                         │         0 │
│ Queries with No Change                 │         9 │
│ Queries with Failure                   │         1 │
└────────────────────────────────────────┴───────────┘

Memory Pool Peaks

Peak MemoryPool reservation per query — what DataFusion's accounting believes it reserved. Recorded only for benchmarks that write a results JSON (the dfbench suites, and the suites bench.sh runs through the Criterion SQL harness since #25644), and only under a memory limit: DATAFUSION_RUNTIME_MEMORY_LIMIT, or one the suite sets itself.

Base: ed43e69 (merge-base) | Changed: spill-dedup-view-arrays

sort_tpch — sort_tpch1

Query Base Changed Change
1 175.1 MiB 175.5 MiB +0.2%
2 219.4 MiB 220.4 MiB +0.4%
3 510.9 MiB 539.3 MiB +5.6%
4 266.1 MiB 262.4 MiB -1.4%
5 294.0 MiB 294.0 MiB +0.0%
6 355.7 MiB 355.2 MiB -0.1%
7 509.0 MiB 509.0 MiB +0.0%
8 508.7 MiB 508.7 MiB +0.0%
9 497.9 MiB 505.3 MiB +1.5%
10 510.1 MiB 510.1 MiB +0.0%
11 523.1 MiB 523.6 MiB +0.1%

Pool accounting vs. process RSS

Max pool peak is the largest reservation any single query in the run reached; peak RSS covers the whole invocation, including data loading and allocator retention, and the two high-water marks need not coincide in time. The gap is therefore an upper bound on what the pool did not account for, not a measurement of it.

Benchmark Side Max pool peak Peak RSS Gap RSS / pool
sort_tpch base (ed43e69 (merge-base)) 523.1 MiB 1.1 GiB 581.4 MiB 2.1×
sort_tpch changed (spill-dedup-view-arrays) 539.3 MiB 1.1 GiB 606.2 MiB 2.1×
Resource Usage

sort_tpch — base (merge-base)

Metric Value
Wall time 25.0s
Peak memory 1.1 GiB
Avg memory 797.1 MiB
CPU user 66.1s
CPU sys 16.1s
Peak spill 1.1 GiB

sort_tpch — branch

Metric Value
Wall time 25.0s
Peak memory 1.1 GiB
Avg memory 772.1 MiB
CPU user 64.0s
CPU sys 15.1s
Peak spill 1.2 GiB

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing spill-dedup-view-arrays (9c2f8f1) to ed43e69 (merge-base) diff

Run configuration
run benchmark spill_views
env:
  DATAFUSION_EXECUTION_TARGET_PARTITIONS: "4"
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "512M"
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

group                                                    HEAD                                   spill-dedup-view-arrays
-----                                                    ----                                   -----------------------
spill_views/q01_sort_string_1_distinct_repeated          1.47    105.4±0.84ms        ? ?/sec    1.00     71.5±0.21ms        ? ?/sec
spill_views/q02_sort_string_1000_distinct_repeated       1.41    100.5±0.80ms        ? ?/sec    1.00     71.5±0.54ms        ? ?/sec
spill_views/q03_sort_binary_1000_distinct_repeated       1.39    100.2±2.10ms        ? ?/sec    1.00     71.8±0.45ms        ? ?/sec
spill_views/q04_sort_string_all_distinct_distinct        1.01    126.2±1.18ms        ? ?/sec    1.00    125.4±1.02ms        ? ?/sec
spill_views/q05_group_by_string_all_distinct_distinct    1.04    135.6±2.74ms        ? ?/sec    1.00    130.9±3.38ms        ? ?/sec

Memory Pool Peaks

Peak MemoryPool reservation per query — what DataFusion's accounting believes it reserved. Recorded only for benchmarks that write a results JSON (the dfbench suites, and the suites bench.sh runs through the Criterion SQL harness since #25644), and only under a memory limit: DATAFUSION_RUNTIME_MEMORY_LIMIT, or one the suite sets itself.

Base: ed43e69 (merge-base) | Changed: spill-dedup-view-arrays

spill_views

Query Base Changed Change
spill_views/q01_sort_string_1_distinct_repeated 40.6 MiB 40.6 MiB +0.0%
spill_views/q02_sort_string_1000_distinct_repeated 39.8 MiB 39.8 MiB +0.0%
spill_views/q03_sort_binary_1000_distinct_repeated 39.8 MiB 39.8 MiB +0.0%
spill_views/q04_sort_string_all_distinct_distinct 102.6 MiB 102.6 MiB +0.0%
spill_views/q05_group_by_string_all_distinct_distinct 96.0 MiB 96.0 MiB +0.0%

Pool accounting vs. process RSS

Max pool peak is the largest reservation any single query in the run reached; peak RSS covers the whole invocation, including data loading and allocator retention, and the two high-water marks need not coincide in time. The gap is therefore an upper bound on what the pool did not account for, not a measurement of it.

Benchmark Side Max pool peak Peak RSS Gap RSS / pool
spill_views base (ed43e69 (merge-base)) 102.6 MiB 788.7 MiB 686.1 MiB 7.7×
spill_views changed (spill-dedup-view-arrays) 102.6 MiB 750.4 MiB 647.8 MiB 7.3×
Resource Usage

spill_views — base (merge-base)

Metric Value
Wall time 365.2s
Peak memory 788.7 MiB
Avg memory 58.9 MiB
CPU user 62.6s
CPU sys 29.6s
Peak spill 84.2 MiB

spill_views — branch

Metric Value
Wall time 565.2s
Peak memory 750.4 MiB
Avg memory 41.5 MiB
CPU user 73.6s
CPU sys 22.6s
Peak spill 72.5 MiB

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing spill-dedup-view-arrays (9c2f8f1) to ed43e69 (merge-base) diff

Run configuration
run benchmark spill_views
env:
  DATAFUSION_EXECUTION_TARGET_PARTITIONS: "4"
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "512M"
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

group                                                    HEAD                                   spill-dedup-view-arrays
-----                                                    ----                                   -----------------------
spill_views/q01_sort_string_1_distinct_repeated          1.50    109.7±0.91ms        ? ?/sec    1.00     73.3±0.37ms        ? ?/sec
spill_views/q02_sort_string_1000_distinct_repeated       1.42    104.7±0.69ms        ? ?/sec    1.00     73.7±0.42ms        ? ?/sec
spill_views/q03_sort_binary_1000_distinct_repeated       1.40    103.3±1.03ms        ? ?/sec    1.00     74.0±0.53ms        ? ?/sec
spill_views/q04_sort_string_all_distinct_distinct        1.00    130.2±1.03ms        ? ?/sec    1.00    130.1±0.96ms        ? ?/sec
spill_views/q05_group_by_string_all_distinct_distinct    1.00    135.7±4.61ms        ? ?/sec    1.01    136.7±4.90ms        ? ?/sec

Memory Pool Peaks

Peak MemoryPool reservation per query — what DataFusion's accounting believes it reserved. Recorded only for benchmarks that write a results JSON (the dfbench suites, and the suites bench.sh runs through the Criterion SQL harness since #25644), and only under a memory limit: DATAFUSION_RUNTIME_MEMORY_LIMIT, or one the suite sets itself.

Base: ed43e69 (merge-base) | Changed: spill-dedup-view-arrays

spill_views

Query Base Changed Change
spill_views/q01_sort_string_1_distinct_repeated 40.6 MiB 40.6 MiB +0.0%
spill_views/q02_sort_string_1000_distinct_repeated 39.8 MiB 39.8 MiB +0.0%
spill_views/q03_sort_binary_1000_distinct_repeated 39.8 MiB 39.8 MiB +0.0%
spill_views/q04_sort_string_all_distinct_distinct 102.6 MiB 102.6 MiB +0.0%
spill_views/q05_group_by_string_all_distinct_distinct 96.0 MiB 96.0 MiB +0.0%

Pool accounting vs. process RSS

Max pool peak is the largest reservation any single query in the run reached; peak RSS covers the whole invocation, including data loading and allocator retention, and the two high-water marks need not coincide in time. The gap is therefore an upper bound on what the pool did not account for, not a measurement of it.

Benchmark Side Max pool peak Peak RSS Gap RSS / pool
spill_views base (ed43e69 (merge-base)) 102.6 MiB 732.2 MiB 629.6 MiB 7.1×
spill_views changed (spill-dedup-view-arrays) 102.6 MiB 774.8 MiB 672.2 MiB 7.6×
Resource Usage

spill_views — base (merge-base)

Metric Value
Wall time 380.2s
Peak memory 732.2 MiB
Avg memory 53.7 MiB
CPU user 63.0s
CPU sys 30.4s
Peak spill 84.2 MiB

spill_views — branch

Metric Value
Wall time 580.3s
Peak memory 774.8 MiB
Avg memory 41.6 MiB
CPU user 75.7s
CPU sys 24.3s
Peak spill 84.2 MiB

File an issue against this benchmark runner

@cetra3

cetra3 commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor Author

Results for 9c2f8f1 (keep shared buffers instead of deduplicating values)

9c2f8f1 replaces the value dedup with a check on the buffers: merge repeated entries of the same buffer, then run gc() only if it writes less than the merged buffers hold. No values are hashed or sampled. Details are in the commit message.

The issue's repro

1M rows, 1000 distinct labels, -m 64M, default partitions. ORDER BY id DESC, because main now skips the sort for ORDER BY id on this file. Three runs each, identical.

Build Spilled bytes
main 83.2 MB
previous version (7effeab) 83.2 MB
9c2f8f1 31.0 MB

The previous version did not reduce this at all: sorted by id, the labels come out in cycling order, so the sample saw 256 distinct values and fell back to gc() in every batch.

Multi-level merge: the repro at 10M rows, 1 partition

Median of 5 interleaved runs. Spilled rows above 10M are rows re-spilled by extra merge passes.

Limit Spilled: main → 9c2f8f1 Spilled rows Time
24M 1959 → 1152 MB 23.5M → 20.0M 832 → 751 ms
32M 1477 → 334 MB 17.8M → 10.0M (one pass) 715 → 490 ms
48M 832 → 353 MB 10.0M → 10.0M 571 → 501 ms
64M 832 → 371 MB 10.0M → 10.0M 574 → 494 ms
128M 832 → 444 MB 10.0M → 10.0M 614 → 578 ms

Smaller spilled batches lower the per-file reservation in the multi-level merge, so more files fit in each pass. At 32M the merge needs one pass instead of about 1.8.

spill_views and mixed data, local

datafusion-cli, 1M rows, 4 partitions, 15 rounds with the build order shuffled each round. Ratios are the median per-round time over main. "N% repeated" sets N% of the rows to one value, with the rest distinct.

Case Spilled: main / previous / 9c2f8f1 previous / main 9c2f8f1 / main
q01 sort, 1 distinct value 84.2 / 23.2 / 23.8 MB 0.97x 0.85x
q02 sort, 1000 distinct values 84.2 / 30.6 / 31.1 MB 1.05x 0.88x
q03 sort, BinaryView 84.2 / 30.6 / 31.1 MB 1.01x 0.85x
q04 sort, all distinct 84.2 / 84.2 / 84.2 MB 1.02x 1.00x
q05 GROUP BY, all distinct 84.2 / 84.2 / 84.2 MB 0.94x 1.01x
5% repeated 84.2 / 81.1 / 84.2 MB 1.53x 1.04x
20% repeated 84.2 / 72.0 / 84.2 MB 1.56x 0.99x
50% repeated 84.2 / 53.7 / 84.2 MB 1.48x 1.04x
80% repeated 84.2 / 35.4 / 84.2 MB 1.31x 1.00x
1000 labels in cycling order 84.2 / 84.2 / 31.1 MB 1.03x 0.79x
  • Shared-dictionary data spills 2.7–3.6x less and runs 12–21% faster.
  • Data without shared buffers takes the same gc() path as main: the same bytes, and times within noise (slower than main in 6–9 of 15 rounds).
  • The trade-off: repeated values stored as separate copies (the 50% and 80% cases) are no longer deduplicated. The previous version spilled 36–58% less there, at 1.31–1.48x the time.

Benchmark bot

Two A/B runs of 9c2f8f1 against ed43e69 (merge-base), DATAFUSION_RUNTIME_MEMORY_LIMIT=512M, 4 partitions. The A/A (main vs main) runs are from the previous round, on the same merge-base. Ratios are branch / base, mean times.

spill_views (each query sets its own limit: 40M for q01–q03, 96M for q04–q05)

Query A/B run 1 A/B run 2 A/A (previous round) Previous version (7effeab)
q01 sort, 1 distinct value 109.7 → 73.3 ms (0.67x) 105.4 → 71.5 ms (0.68x) 0.99x 0.78–0.79x
q02 sort, 1000 distinct values 104.7 → 73.7 ms (0.70x) 100.5 → 71.5 ms (0.71x) 0.99x 0.92x
q03 sort, BinaryView 103.3 → 74.0 ms (0.72x) 100.2 → 71.8 ms (0.72x) 0.98x 0.94x
q04 sort, all distinct 1.00x 0.99x 0.99x 1.00–1.02x
q05 GROUP BY, all distinct 1.01x 0.97x 0.98x 0.99–1.01x

q05 passed at the default 96M limit in both runs.

sort_tpch (SF1, 512M, 4 partitions): no query changes beyond A/A noise. Q3 fails on both sides, as before.

Query A/B run 1 A/B run 2 A/A (previous round)
Q8 (l_comment) 509 → 488 ms (0.96x) 487 → 480 ms (0.98x) 0.99x
Q9 (l_comment) 536 → 540 ms (1.01x) 529 → 527 ms (1.00x) 1.00x
Q11 (l_comment) 379 → 371 ms (0.98x) 377 → 376 ms (1.00x) 0.99x
Total 4324 → 4254 ms (0.98x) 4290 → 4267 ms (0.99x) 0.99x

@adriangb

Copy link
Copy Markdown
Contributor

run benchmarks

env:
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "512M"

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5870585310-2877-4nbmw 6.12.94+ #1 SMP Wed Aug 19 07:47:20 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing spill-dedup-view-arrays (9c2f8f1) to ed43e69 (merge-base) diff

Run configuration
run benchmark clickbench_partitioned
env:
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "512M"

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5870585310-2878-jgccc 6.12.94+ #1 SMP Wed Aug 19 07:47:20 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing spill-dedup-view-arrays (9c2f8f1) to ed43e69 (merge-base) diff

Run configuration
run benchmark tpcds
env:
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "512M"

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark running (GKE) | trigger
Instance: c4a-highmem-16 (12 vCPU / 65 GiB) | Linux bench-c5870585310-2879-wkv4j 6.12.94+ #1 SMP Wed Aug 19 07:47:20 UTC 2026 aarch64 GNU/Linux

CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected

Comparing spill-dedup-view-arrays (9c2f8f1) to ed43e69 (merge-base) diff

Run configuration
run benchmark tpch
env:
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "512M"

Results will be posted here when complete


File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing spill-dedup-view-arrays (9c2f8f1) to ed43e69 (merge-base) diff

Run configuration
run benchmark tpch
env:
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "512M"
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and spill-dedup-view-arrays
--------------------
Benchmark tpch_sf1.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Query     ┃     HEAD ┃ spill-dedup-view-arrays ┃    Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ QQuery 1  │ 41.44 ms │                41.93 ms │ no change │
│ QQuery 2  │ 19.08 ms │                19.47 ms │ no change │
│ QQuery 3  │ 29.60 ms │                29.67 ms │ no change │
│ QQuery 4  │ 18.06 ms │                18.00 ms │ no change │
│ QQuery 5  │ 36.98 ms │                37.11 ms │ no change │
│ QQuery 6  │ 16.94 ms │                17.00 ms │ no change │
│ QQuery 7  │ 44.46 ms │                43.60 ms │ no change │
│ QQuery 8  │ 42.24 ms │                42.70 ms │ no change │
│ QQuery 9  │ 52.11 ms │                52.79 ms │ no change │
│ QQuery 10 │ 42.62 ms │                42.64 ms │ no change │
│ QQuery 11 │ 14.59 ms │                14.66 ms │ no change │
│ QQuery 12 │ 22.16 ms │                22.33 ms │ no change │
│ QQuery 13 │ 43.77 ms │                45.14 ms │ no change │
│ QQuery 14 │ 25.50 ms │                26.01 ms │ no change │
│ QQuery 15 │ 32.19 ms │                32.94 ms │ no change │
│ QQuery 16 │ 14.50 ms │                14.82 ms │ no change │
│ QQuery 17 │ 80.76 ms │                82.46 ms │ no change │
│ QQuery 18 │ 65.86 ms │                67.84 ms │ no change │
│ QQuery 19 │ 32.71 ms │                33.91 ms │ no change │
│ QQuery 20 │ 32.97 ms │                33.49 ms │ no change │
│ QQuery 21 │ 58.62 ms │                59.34 ms │ no change │
│ QQuery 22 │ 14.86 ms │                15.04 ms │ no change │
└───────────┴──────────┴─────────────────────────┴───────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━┓
┃ Benchmark Summary                      ┃          ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━┩
│ Total Time (HEAD)                      │ 782.04ms │
│ Total Time (spill-dedup-view-arrays)   │ 792.89ms │
│ Average Time (HEAD)                    │  35.55ms │
│ Average Time (spill-dedup-view-arrays) │  36.04ms │
│ Queries Faster                         │        0 │
│ Queries Slower                         │        0 │
│ Queries with No Change                 │       22 │
│ Queries with Failure                   │        0 │
└────────────────────────────────────────┴──────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and spill-dedup-view-arrays
--------------------
Benchmark tpch_sf1.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃                           HEAD ┃        spill-dedup-view-arrays ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 1  │ 41.44 / 42.19 ±1.06 / 44.27 ms │ 41.93 / 43.00 ±1.24 / 44.78 ms │     no change │
│ QQuery 2  │ 19.08 / 19.50 ±0.40 / 20.16 ms │ 19.47 / 20.04 ±0.59 / 21.18 ms │     no change │
│ QQuery 3  │ 29.60 / 29.77 ±0.11 / 29.94 ms │ 29.67 / 30.08 ±0.48 / 31.00 ms │     no change │
│ QQuery 4  │ 18.06 / 18.66 ±0.80 / 20.12 ms │ 18.00 / 18.62 ±0.85 / 20.27 ms │     no change │
│ QQuery 5  │ 36.98 / 37.86 ±0.63 / 38.77 ms │ 37.11 / 37.37 ±0.19 / 37.62 ms │     no change │
│ QQuery 6  │ 16.94 / 17.21 ±0.22 / 17.60 ms │ 17.00 / 17.17 ±0.10 / 17.31 ms │     no change │
│ QQuery 7  │ 44.46 / 45.79 ±0.84 / 46.99 ms │ 43.60 / 45.24 ±1.18 / 46.92 ms │     no change │
│ QQuery 8  │ 42.24 / 43.02 ±0.63 / 44.06 ms │ 42.70 / 43.58 ±0.76 / 44.85 ms │     no change │
│ QQuery 9  │ 52.11 / 53.57 ±1.07 / 55.22 ms │ 52.79 / 54.47 ±0.98 / 55.61 ms │     no change │
│ QQuery 10 │ 42.62 / 44.01 ±1.27 / 45.66 ms │ 42.64 / 43.19 ±0.37 / 43.78 ms │     no change │
│ QQuery 11 │ 14.59 / 15.13 ±0.76 / 16.63 ms │ 14.66 / 14.86 ±0.16 / 15.10 ms │     no change │
│ QQuery 12 │ 22.16 / 22.53 ±0.33 / 23.14 ms │ 22.33 / 22.58 ±0.18 / 22.87 ms │     no change │
│ QQuery 13 │ 43.77 / 44.76 ±0.76 / 46.08 ms │ 45.14 / 46.82 ±1.30 / 48.39 ms │     no change │
│ QQuery 14 │ 25.50 / 26.81 ±1.37 / 28.78 ms │ 26.01 / 27.13 ±1.58 / 30.21 ms │     no change │
│ QQuery 15 │ 32.19 / 32.50 ±0.37 / 33.23 ms │ 32.94 / 34.32 ±0.82 / 35.29 ms │  1.06x slower │
│ QQuery 16 │ 14.50 / 14.76 ±0.28 / 15.25 ms │ 14.82 / 15.11 ±0.19 / 15.40 ms │     no change │
│ QQuery 17 │ 80.76 / 82.76 ±2.49 / 87.46 ms │ 82.46 / 83.75 ±1.24 / 85.82 ms │     no change │
│ QQuery 18 │ 65.86 / 68.00 ±1.62 / 70.39 ms │ 67.84 / 69.83 ±1.44 / 71.96 ms │     no change │
│ QQuery 19 │ 32.71 / 33.73 ±1.14 / 35.66 ms │ 33.91 / 34.57 ±0.76 / 35.92 ms │     no change │
│ QQuery 20 │ 32.97 / 33.34 ±0.26 / 33.64 ms │ 33.49 / 34.09 ±0.35 / 34.56 ms │     no change │
│ QQuery 21 │ 58.62 / 60.07 ±0.82 / 60.85 ms │ 59.34 / 60.32 ±0.76 / 61.63 ms │     no change │
│ QQuery 22 │ 14.86 / 17.20 ±4.11 / 25.41 ms │ 15.04 / 16.28 ±1.99 / 20.24 ms │ +1.06x faster │
└───────────┴────────────────────────────────┴────────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━┓
┃ Benchmark Summary                      ┃          ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━┩
│ Total Time (HEAD)                      │ 803.17ms │
│ Total Time (spill-dedup-view-arrays)   │ 812.41ms │
│ Average Time (HEAD)                    │  36.51ms │
│ Average Time (spill-dedup-view-arrays) │  36.93ms │
│ Queries Faster                         │        1 │
│ Queries Slower                         │        1 │
│ Queries with No Change                 │       20 │
│ Queries with Failure                   │        0 │
└────────────────────────────────────────┴──────────┘

Memory Pool Peaks

Peak MemoryPool reservation per query — what DataFusion's accounting believes it reserved. Recorded only for benchmarks that write a results JSON (the dfbench suites, and the suites bench.sh runs through the Criterion SQL harness since #25644), and only under a memory limit: DATAFUSION_RUNTIME_MEMORY_LIMIT, or one the suite sets itself.

Base: ed43e69 (merge-base) | Changed: spill-dedup-view-arrays

tpch — tpch_sf1

Query Base Changed Change
Query 1 21.4 MiB 21.4 MiB +0.0%
Query 2 15.7 MiB 15.7 MiB +0.0%
Query 3 11.5 MiB 11.5 MiB +0.0%
Query 4 23.0 MiB 23.2 MiB +0.7%
Query 5 52.6 MiB 43.0 MiB -18.1%
Query 6 832 B 832 B +0.0%
Query 7 107.4 MiB 107.4 MiB +0.0%
Query 8 33.8 MiB 33.1 MiB -2.1%
Query 9 153.9 MiB 153.9 MiB +0.0%
Query 10 36.7 MiB 34.9 MiB -5.0%
Query 11 72.0 MiB 51.5 MiB -28.5%
Query 12 20.0 MiB 20.0 MiB -0.0%
Query 13 52.3 MiB 62.3 MiB +19.1%
Query 14 7.7 MiB 7.7 MiB +0.0%
Query 15 12.2 MiB 12.1 MiB -0.8%
Query 16 44.3 MiB 44.4 MiB +0.2%
Query 17 142.2 MiB 141.2 MiB -0.7%
Query 18 358.3 MiB 358.3 MiB +0.0%
Query 19 746.1 KiB 746.1 KiB +0.0%
Query 20 61.7 MiB 61.7 MiB -0.1%
Query 21 41.7 MiB 41.7 MiB +0.0%
Query 22 32.4 MiB 32.1 MiB -0.9%

Pool accounting vs. process RSS

Max pool peak is the largest reservation any single query in the run reached; peak RSS covers the whole invocation, including data loading and allocator retention, and the two high-water marks need not coincide in time. The gap is therefore an upper bound on what the pool did not account for, not a measurement of it.

Benchmark Side Max pool peak Peak RSS Gap RSS / pool
tpch base (ed43e69 (merge-base)) 358.3 MiB 1.3 GiB 998.5 MiB 3.8×
tpch changed (spill-dedup-view-arrays) 358.3 MiB 1.1 GiB 796.7 MiB 3.2×
Resource Usage

tpch — base (merge-base)

Metric Value
Wall time 5.0s
Peak memory 1.3 GiB
Avg memory 717.5 MiB
CPU user 22.7s
CPU sys 1.9s
Peak spill 3.3 MiB

tpch — branch

Metric Value
Wall time 5.0s
Peak memory 1.1 GiB
Avg memory 647.4 MiB
CPU user 22.7s
CPU sys 1.9s
Peak spill 0 B

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing spill-dedup-view-arrays (9c2f8f1) to ed43e69 (merge-base) diff

Run configuration
run benchmark tpcds
env:
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "512M"
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and spill-dedup-view-arrays
--------------------
Benchmark tpcds_sf1.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃       HEAD ┃ spill-dedup-view-arrays ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 1  │    5.50 ms │                 6.02 ms │  1.09x slower │
│ QQuery 2  │   81.22 ms │                80.35 ms │     no change │
│ QQuery 3  │   29.10 ms │                29.25 ms │     no change │
│ QQuery 4  │  481.86 ms │               459.87 ms │     no change │
│ QQuery 5  │   51.43 ms │                51.54 ms │     no change │
│ QQuery 6  │   36.20 ms │                35.91 ms │     no change │
│ QQuery 7  │   73.66 ms │                74.97 ms │     no change │
│ QQuery 8  │   36.52 ms │                36.24 ms │     no change │
│ QQuery 9  │   52.32 ms │                54.10 ms │     no change │
│ QQuery 10 │   62.20 ms │                61.96 ms │     no change │
│ QQuery 11 │  284.47 ms │               296.00 ms │     no change │
│ QQuery 12 │   28.94 ms │                29.68 ms │     no change │
│ QQuery 13 │  117.28 ms │               117.14 ms │     no change │
│ QQuery 14 │  394.81 ms │               396.61 ms │     no change │
│ QQuery 15 │   55.51 ms │                54.74 ms │     no change │
│ QQuery 16 │    6.56 ms │                 7.06 ms │  1.08x slower │
│ QQuery 17 │   80.25 ms │                79.48 ms │     no change │
│ QQuery 18 │  106.41 ms │               104.68 ms │     no change │
│ QQuery 19 │   41.68 ms │                41.75 ms │     no change │
│ QQuery 20 │   36.39 ms │                36.12 ms │     no change │
│ QQuery 21 │   17.05 ms │                17.54 ms │     no change │
│ QQuery 22 │   65.88 ms │                67.34 ms │     no change │
│ QQuery 23 │  400.63 ms │               413.60 ms │     no change │
│ QQuery 24 │  195.96 ms │               195.02 ms │     no change │
│ QQuery 25 │  107.36 ms │               108.26 ms │     no change │
│ QQuery 26 │   48.45 ms │                48.16 ms │     no change │
│ QQuery 27 │    6.29 ms │                 6.49 ms │     no change │
│ QQuery 28 │   57.01 ms │                60.14 ms │  1.06x slower │
│ QQuery 29 │   94.79 ms │                94.44 ms │     no change │
│ QQuery 30 │   33.95 ms │                32.79 ms │     no change │
│ QQuery 31 │  111.14 ms │               109.41 ms │     no change │
│ QQuery 32 │   20.20 ms │                20.90 ms │     no change │
│ QQuery 33 │   37.60 ms │                37.80 ms │     no change │
│ QQuery 34 │   10.36 ms │                10.05 ms │     no change │
│ QQuery 35 │   74.25 ms │                71.19 ms │     no change │
│ QQuery 36 │    6.15 ms │                 5.69 ms │ +1.08x faster │
│ QQuery 37 │    6.76 ms │                 6.86 ms │     no change │
│ QQuery 38 │   60.27 ms │                61.55 ms │     no change │
│ QQuery 39 │   73.69 ms │                73.56 ms │     no change │
│ QQuery 40 │   23.66 ms │                24.07 ms │     no change │
│ QQuery 41 │   11.80 ms │                11.91 ms │     no change │
│ QQuery 42 │   23.95 ms │                23.54 ms │     no change │
│ QQuery 43 │    4.67 ms │                 4.97 ms │  1.06x slower │
│ QQuery 44 │    8.79 ms │                 8.64 ms │     no change │
│ QQuery 45 │   38.18 ms │                37.00 ms │     no change │
│ QQuery 46 │   11.48 ms │                11.81 ms │     no change │
│ QQuery 47 │  224.56 ms │               225.39 ms │     no change │
│ QQuery 48 │   94.26 ms │                94.82 ms │     no change │
│ QQuery 49 │   70.90 ms │                71.07 ms │     no change │
│ QQuery 50 │   58.24 ms │                57.83 ms │     no change │
│ QQuery 51 │   93.81 ms │                95.74 ms │     no change │
│ QQuery 52 │   23.97 ms │                24.00 ms │     no change │
│ QQuery 53 │   29.60 ms │                29.67 ms │     no change │
│ QQuery 54 │   53.73 ms │                54.85 ms │     no change │
│ QQuery 55 │   23.23 ms │                23.97 ms │     no change │
│ QQuery 56 │   39.15 ms │                38.84 ms │     no change │
│ QQuery 57 │  168.79 ms │               169.39 ms │     no change │
│ QQuery 58 │  108.74 ms │               108.63 ms │     no change │
│ QQuery 59 │  114.85 ms │               116.34 ms │     no change │
│ QQuery 60 │   38.22 ms │                39.32 ms │     no change │
│ QQuery 61 │   11.24 ms │                12.08 ms │  1.08x slower │
│ QQuery 62 │   44.37 ms │                43.96 ms │     no change │
│ QQuery 63 │   29.80 ms │                29.85 ms │     no change │
│ QQuery 64 │  361.01 ms │               360.95 ms │     no change │
│ QQuery 65 │  132.37 ms │               131.45 ms │     no change │
│ QQuery 66 │   77.89 ms │                77.38 ms │     no change │
│ QQuery 67 │  294.00 ms │               296.14 ms │     no change │
│ QQuery 68 │   12.05 ms │                12.19 ms │     no change │
│ QQuery 69 │   56.64 ms │                56.48 ms │     no change │
│ QQuery 70 │  111.11 ms │               109.22 ms │     no change │
│ QQuery 71 │   34.98 ms │                35.48 ms │     no change │
│ QQuery 72 │ 1696.97 ms │              1716.93 ms │     no change │
│ QQuery 73 │    9.76 ms │                10.09 ms │     no change │
│ QQuery 74 │  160.42 ms │               165.08 ms │     no change │
│ QQuery 75 │  142.08 ms │               142.08 ms │     no change │
│ QQuery 76 │   33.96 ms │                35.14 ms │     no change │
│ QQuery 77 │   61.57 ms │                59.98 ms │     no change │
│ QQuery 78 │  165.59 ms │               169.00 ms │     no change │
│ QQuery 79 │   67.53 ms │                66.81 ms │     no change │
│ QQuery 80 │   96.52 ms │                95.61 ms │     no change │
│ QQuery 81 │   25.89 ms │                26.91 ms │     no change │
│ QQuery 82 │   17.12 ms │                16.95 ms │     no change │
│ QQuery 83 │   33.53 ms │                33.32 ms │     no change │
│ QQuery 84 │   29.94 ms │                29.54 ms │     no change │
│ QQuery 85 │  102.46 ms │               103.10 ms │     no change │
│ QQuery 86 │   25.97 ms │                25.51 ms │     no change │
│ QQuery 87 │   61.61 ms │                60.43 ms │     no change │
│ QQuery 88 │   62.30 ms │                61.62 ms │     no change │
│ QQuery 89 │   35.51 ms │                34.88 ms │     no change │
│ QQuery 90 │   17.15 ms │                17.21 ms │     no change │
│ QQuery 91 │   45.64 ms │                44.49 ms │     no change │
│ QQuery 92 │   30.41 ms │                29.44 ms │     no change │
│ QQuery 93 │   49.53 ms │                49.29 ms │     no change │
│ QQuery 94 │   38.63 ms │                38.39 ms │     no change │
│ QQuery 95 │   81.58 ms │                80.01 ms │     no change │
│ QQuery 96 │   24.16 ms │                23.51 ms │     no change │
│ QQuery 97 │   51.00 ms │                49.94 ms │     no change │
│ QQuery 98 │   43.48 ms │                42.63 ms │     no change │
│ QQuery 99 │   66.89 ms │                65.59 ms │     no change │
└───────────┴────────────┴─────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary                      ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)                      │ 9197.36ms │
│ Total Time (spill-dedup-view-arrays)   │ 9224.69ms │
│ Average Time (HEAD)                    │   92.90ms │
│ Average Time (spill-dedup-view-arrays) │   93.18ms │
│ Queries Faster                         │         1 │
│ Queries Slower                         │         5 │
│ Queries with No Change                 │        93 │
│ Queries with Failure                   │         0 │
└────────────────────────────────────────┴───────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and spill-dedup-view-arrays
--------------------
Benchmark tpcds_sf1.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃                                  HEAD ┃               spill-dedup-view-arrays ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 1  │           5.50 / 6.04 ±1.02 / 8.08 ms │           6.02 / 6.56 ±0.92 / 8.40 ms │  1.09x slower │
│ QQuery 2  │        81.22 / 81.56 ±0.27 / 82.02 ms │        80.35 / 81.16 ±0.56 / 82.06 ms │     no change │
│ QQuery 3  │        29.10 / 29.38 ±0.17 / 29.64 ms │        29.25 / 29.49 ±0.15 / 29.69 ms │     no change │
│ QQuery 4  │     481.86 / 490.09 ±6.25 / 498.56 ms │    459.87 / 481.16 ±10.72 / 488.57 ms │     no change │
│ QQuery 5  │        51.43 / 52.26 ±0.70 / 53.56 ms │        51.54 / 54.48 ±4.16 / 62.66 ms │     no change │
│ QQuery 6  │        36.20 / 36.85 ±0.48 / 37.58 ms │        35.91 / 36.85 ±0.54 / 37.45 ms │     no change │
│ QQuery 7  │        73.66 / 74.82 ±0.92 / 75.89 ms │        74.97 / 75.35 ±0.39 / 76.07 ms │     no change │
│ QQuery 8  │        36.52 / 37.88 ±2.18 / 42.21 ms │        36.24 / 38.00 ±2.60 / 43.16 ms │     no change │
│ QQuery 9  │        52.32 / 54.65 ±1.83 / 57.78 ms │        54.10 / 57.07 ±2.60 / 60.85 ms │     no change │
│ QQuery 10 │        62.20 / 62.70 ±0.51 / 63.67 ms │        61.96 / 62.22 ±0.23 / 62.53 ms │     no change │
│ QQuery 11 │     284.47 / 294.31 ±7.52 / 304.12 ms │     296.00 / 305.55 ±4.83 / 309.18 ms │     no change │
│ QQuery 12 │        28.94 / 29.46 ±0.29 / 29.71 ms │        29.68 / 30.09 ±0.35 / 30.61 ms │     no change │
│ QQuery 13 │     117.28 / 117.83 ±0.56 / 118.64 ms │     117.14 / 118.56 ±0.80 / 119.61 ms │     no change │
│ QQuery 14 │     394.81 / 402.64 ±5.38 / 410.03 ms │     396.61 / 400.09 ±3.57 / 406.80 ms │     no change │
│ QQuery 15 │        55.51 / 57.11 ±1.61 / 59.55 ms │        54.74 / 57.99 ±2.94 / 63.24 ms │     no change │
│ QQuery 16 │           6.56 / 6.91 ±0.18 / 7.05 ms │           7.06 / 7.16 ±0.12 / 7.37 ms │     no change │
│ QQuery 17 │        80.25 / 80.87 ±0.78 / 82.37 ms │        79.48 / 81.06 ±1.16 / 82.41 ms │     no change │
│ QQuery 18 │     106.41 / 107.08 ±0.60 / 108.16 ms │     104.68 / 108.08 ±2.01 / 110.26 ms │     no change │
│ QQuery 19 │        41.68 / 41.88 ±0.15 / 42.08 ms │        41.75 / 42.19 ±0.35 / 42.80 ms │     no change │
│ QQuery 20 │        36.39 / 36.71 ±0.19 / 36.90 ms │        36.12 / 36.81 ±0.42 / 37.30 ms │     no change │
│ QQuery 21 │        17.05 / 17.44 ±0.36 / 18.00 ms │        17.54 / 17.64 ±0.08 / 17.74 ms │     no change │
│ QQuery 22 │        65.88 / 67.43 ±1.05 / 68.43 ms │        67.34 / 68.45 ±0.86 / 69.84 ms │     no change │
│ QQuery 23 │     400.63 / 410.87 ±7.56 / 420.58 ms │     413.60 / 418.06 ±2.91 / 421.42 ms │     no change │
│ QQuery 24 │     195.96 / 201.04 ±4.33 / 208.32 ms │     195.02 / 200.71 ±4.64 / 204.97 ms │     no change │
│ QQuery 25 │     107.36 / 109.59 ±1.41 / 111.57 ms │     108.26 / 108.70 ±0.33 / 109.09 ms │     no change │
│ QQuery 26 │        48.45 / 49.49 ±0.92 / 51.05 ms │        48.16 / 50.00 ±1.78 / 53.40 ms │     no change │
│ QQuery 27 │           6.29 / 6.51 ±0.21 / 6.89 ms │           6.49 / 6.65 ±0.16 / 6.87 ms │     no change │
│ QQuery 28 │        57.01 / 59.51 ±2.15 / 62.26 ms │        60.14 / 60.75 ±0.40 / 61.34 ms │     no change │
│ QQuery 29 │       94.79 / 98.04 ±2.44 / 101.86 ms │        94.44 / 96.61 ±1.73 / 98.59 ms │     no change │
│ QQuery 30 │        33.95 / 34.95 ±1.20 / 37.27 ms │        32.79 / 34.34 ±2.24 / 38.79 ms │     no change │
│ QQuery 31 │     111.14 / 112.47 ±1.07 / 114.07 ms │     109.41 / 111.11 ±1.10 / 112.42 ms │     no change │
│ QQuery 32 │        20.20 / 20.75 ±0.28 / 20.92 ms │        20.90 / 21.13 ±0.22 / 21.45 ms │     no change │
│ QQuery 33 │        37.60 / 37.97 ±0.23 / 38.21 ms │        37.80 / 38.89 ±0.71 / 39.70 ms │     no change │
│ QQuery 34 │        10.36 / 10.60 ±0.18 / 10.77 ms │        10.05 / 10.48 ±0.43 / 11.12 ms │     no change │
│ QQuery 35 │        74.25 / 76.35 ±1.72 / 78.87 ms │        71.19 / 72.99 ±1.76 / 75.44 ms │     no change │
│ QQuery 36 │           6.15 / 6.34 ±0.19 / 6.67 ms │           5.69 / 5.94 ±0.27 / 6.30 ms │ +1.07x faster │
│ QQuery 37 │           6.76 / 6.90 ±0.08 / 7.00 ms │           6.86 / 7.04 ±0.12 / 7.20 ms │     no change │
│ QQuery 38 │        60.27 / 62.66 ±1.93 / 65.12 ms │        61.55 / 62.87 ±0.71 / 63.70 ms │     no change │
│ QQuery 39 │        73.69 / 74.95 ±0.95 / 76.07 ms │        73.56 / 76.22 ±2.10 / 79.73 ms │     no change │
│ QQuery 40 │        23.66 / 24.22 ±0.49 / 25.11 ms │        24.07 / 24.63 ±0.58 / 25.74 ms │     no change │
│ QQuery 41 │        11.80 / 11.93 ±0.08 / 12.03 ms │        11.91 / 12.06 ±0.09 / 12.20 ms │     no change │
│ QQuery 42 │        23.95 / 24.39 ±0.38 / 25.05 ms │        23.54 / 23.77 ±0.30 / 24.36 ms │     no change │
│ QQuery 43 │           4.67 / 4.89 ±0.28 / 5.41 ms │           4.97 / 5.15 ±0.19 / 5.45 ms │  1.05x slower │
│ QQuery 44 │           8.79 / 8.92 ±0.13 / 9.16 ms │           8.64 / 9.10 ±0.30 / 9.51 ms │     no change │
│ QQuery 45 │        38.18 / 40.79 ±1.74 / 42.76 ms │        37.00 / 38.72 ±1.32 / 40.63 ms │ +1.05x faster │
│ QQuery 46 │        11.48 / 12.34 ±0.52 / 12.91 ms │        11.81 / 11.98 ±0.12 / 12.18 ms │     no change │
│ QQuery 47 │     224.56 / 227.90 ±2.78 / 232.84 ms │     225.39 / 228.21 ±2.75 / 232.24 ms │     no change │
│ QQuery 48 │        94.26 / 95.29 ±0.62 / 96.05 ms │        94.82 / 95.98 ±0.83 / 97.20 ms │     no change │
│ QQuery 49 │        70.90 / 72.50 ±1.83 / 75.98 ms │        71.07 / 71.78 ±0.73 / 73.16 ms │     no change │
│ QQuery 50 │        58.24 / 58.51 ±0.21 / 58.84 ms │        57.83 / 59.80 ±2.22 / 63.69 ms │     no change │
│ QQuery 51 │        93.81 / 96.61 ±1.95 / 99.43 ms │        95.74 / 97.13 ±1.38 / 98.83 ms │     no change │
│ QQuery 52 │        23.97 / 25.53 ±1.69 / 28.81 ms │        24.00 / 24.35 ±0.26 / 24.64 ms │     no change │
│ QQuery 53 │        29.60 / 29.90 ±0.23 / 30.23 ms │        29.67 / 29.73 ±0.05 / 29.81 ms │     no change │
│ QQuery 54 │        53.73 / 54.16 ±0.29 / 54.46 ms │        54.85 / 55.62 ±0.67 / 56.49 ms │     no change │
│ QQuery 55 │        23.23 / 23.68 ±0.35 / 24.31 ms │        23.97 / 24.09 ±0.12 / 24.31 ms │     no change │
│ QQuery 56 │        39.15 / 39.52 ±0.27 / 39.97 ms │        38.84 / 39.52 ±0.47 / 40.05 ms │     no change │
│ QQuery 57 │     168.79 / 170.07 ±1.80 / 173.45 ms │     169.39 / 171.96 ±2.20 / 175.46 ms │     no change │
│ QQuery 58 │     108.74 / 111.00 ±2.27 / 114.90 ms │     108.63 / 109.20 ±0.42 / 109.71 ms │     no change │
│ QQuery 59 │     114.85 / 116.06 ±0.83 / 117.39 ms │     116.34 / 117.53 ±1.38 / 120.12 ms │     no change │
│ QQuery 60 │        38.22 / 38.60 ±0.37 / 39.29 ms │        39.32 / 39.90 ±0.44 / 40.65 ms │     no change │
│ QQuery 61 │        11.24 / 12.47 ±1.71 / 15.83 ms │        12.08 / 12.15 ±0.08 / 12.30 ms │     no change │
│ QQuery 62 │        44.37 / 45.02 ±0.60 / 46.14 ms │        43.96 / 45.84 ±2.82 / 51.44 ms │     no change │
│ QQuery 63 │        29.80 / 30.13 ±0.37 / 30.85 ms │        29.85 / 30.15 ±0.18 / 30.35 ms │     no change │
│ QQuery 64 │     361.01 / 365.63 ±4.05 / 371.38 ms │     360.95 / 365.38 ±2.95 / 368.67 ms │     no change │
│ QQuery 65 │     132.37 / 132.88 ±0.47 / 133.69 ms │     131.45 / 133.74 ±1.66 / 135.73 ms │     no change │
│ QQuery 66 │        77.89 / 79.34 ±1.61 / 82.11 ms │        77.38 / 79.50 ±2.53 / 84.42 ms │     no change │
│ QQuery 67 │    294.00 / 312.01 ±15.44 / 335.17 ms │     296.14 / 302.63 ±5.24 / 310.53 ms │     no change │
│ QQuery 68 │        12.05 / 12.14 ±0.08 / 12.24 ms │        12.19 / 12.68 ±0.31 / 12.96 ms │     no change │
│ QQuery 69 │        56.64 / 57.48 ±0.95 / 59.32 ms │        56.48 / 57.51 ±0.71 / 58.64 ms │     no change │
│ QQuery 70 │     111.11 / 113.20 ±2.61 / 118.21 ms │     109.22 / 111.25 ±3.15 / 117.48 ms │     no change │
│ QQuery 71 │        34.98 / 35.46 ±0.41 / 36.21 ms │        35.48 / 36.19 ±0.55 / 36.88 ms │     no change │
│ QQuery 72 │ 1696.97 / 1799.98 ±62.74 / 1871.63 ms │ 1716.93 / 1733.42 ±15.81 / 1756.05 ms │     no change │
│ QQuery 73 │         9.76 / 10.20 ±0.39 / 10.87 ms │        10.09 / 10.51 ±0.24 / 10.78 ms │     no change │
│ QQuery 74 │     160.42 / 168.71 ±4.69 / 173.03 ms │     165.08 / 168.79 ±3.54 / 175.22 ms │     no change │
│ QQuery 75 │     142.08 / 142.87 ±0.81 / 144.06 ms │     142.08 / 142.92 ±0.92 / 144.69 ms │     no change │
│ QQuery 76 │        33.96 / 34.63 ±0.58 / 35.29 ms │        35.14 / 36.02 ±1.20 / 38.37 ms │     no change │
│ QQuery 77 │        61.57 / 63.13 ±2.11 / 67.32 ms │        59.98 / 61.18 ±0.66 / 61.93 ms │     no change │
│ QQuery 78 │     165.59 / 170.81 ±3.57 / 175.55 ms │     169.00 / 172.50 ±2.71 / 176.34 ms │     no change │
│ QQuery 79 │        67.53 / 68.31 ±1.01 / 70.22 ms │        66.81 / 67.75 ±0.82 / 69.27 ms │     no change │
│ QQuery 80 │        96.52 / 98.47 ±1.13 / 99.71 ms │       95.61 / 98.39 ±1.77 / 101.18 ms │     no change │
│ QQuery 81 │        25.89 / 26.57 ±0.36 / 26.88 ms │        26.91 / 27.22 ±0.29 / 27.63 ms │     no change │
│ QQuery 82 │        17.12 / 17.31 ±0.17 / 17.62 ms │        16.95 / 17.21 ±0.21 / 17.56 ms │     no change │
│ QQuery 83 │        33.53 / 34.91 ±1.37 / 37.31 ms │        33.32 / 33.98 ±0.65 / 34.90 ms │     no change │
│ QQuery 84 │        29.94 / 30.25 ±0.22 / 30.49 ms │        29.54 / 30.43 ±0.94 / 32.06 ms │     no change │
│ QQuery 85 │     102.46 / 103.42 ±0.60 / 104.10 ms │     103.10 / 103.87 ±0.70 / 104.97 ms │     no change │
│ QQuery 86 │        25.97 / 26.29 ±0.22 / 26.58 ms │        25.51 / 25.87 ±0.41 / 26.67 ms │     no change │
│ QQuery 87 │        61.61 / 64.87 ±2.09 / 67.64 ms │        60.43 / 61.49 ±0.71 / 62.56 ms │ +1.05x faster │
│ QQuery 88 │        62.30 / 63.08 ±0.80 / 64.52 ms │        61.62 / 62.22 ±0.41 / 62.84 ms │     no change │
│ QQuery 89 │        35.51 / 36.20 ±0.39 / 36.66 ms │        34.88 / 35.46 ±0.51 / 36.36 ms │     no change │
│ QQuery 90 │        17.15 / 17.30 ±0.11 / 17.42 ms │        17.21 / 17.51 ±0.15 / 17.63 ms │     no change │
│ QQuery 91 │        45.64 / 46.48 ±1.18 / 48.77 ms │        44.49 / 44.84 ±0.38 / 45.47 ms │     no change │
│ QQuery 92 │        30.41 / 30.86 ±0.59 / 31.93 ms │        29.44 / 30.28 ±0.48 / 30.87 ms │     no change │
│ QQuery 93 │        49.53 / 50.05 ±0.51 / 50.80 ms │        49.29 / 50.78 ±0.92 / 51.79 ms │     no change │
│ QQuery 94 │        38.63 / 40.18 ±1.11 / 41.37 ms │        38.39 / 38.89 ±0.35 / 39.45 ms │     no change │
│ QQuery 95 │        81.58 / 82.38 ±0.76 / 83.57 ms │        80.01 / 80.89 ±0.83 / 82.18 ms │     no change │
│ QQuery 96 │        24.16 / 24.38 ±0.18 / 24.61 ms │        23.51 / 23.87 ±0.31 / 24.35 ms │     no change │
│ QQuery 97 │        51.00 / 51.93 ±0.91 / 53.68 ms │        49.94 / 51.88 ±2.01 / 55.76 ms │     no change │
│ QQuery 98 │        43.48 / 45.81 ±1.73 / 48.05 ms │        42.63 / 42.91 ±0.34 / 43.55 ms │ +1.07x faster │
│ QQuery 99 │        66.89 / 67.16 ±0.26 / 67.60 ms │        65.59 / 66.19 ±0.44 / 66.87 ms │     no change │
└───────────┴───────────────────────────────────────┴───────────────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━┓
┃ Benchmark Summary                      ┃           ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━┩
│ Total Time (HEAD)                      │ 9464.03ms │
│ Total Time (spill-dedup-view-arrays)   │ 9393.03ms │
│ Average Time (HEAD)                    │   95.60ms │
│ Average Time (spill-dedup-view-arrays) │   94.88ms │
│ Queries Faster                         │         4 │
│ Queries Slower                         │         2 │
│ Queries with No Change                 │        93 │
│ Queries with Failure                   │         0 │
└────────────────────────────────────────┴───────────┘

Memory Pool Peaks

Peak MemoryPool reservation per query — what DataFusion's accounting believes it reserved. Recorded only for benchmarks that write a results JSON (the dfbench suites, and the suites bench.sh runs through the Criterion SQL harness since #25644), and only under a memory limit: DATAFUSION_RUNTIME_MEMORY_LIMIT, or one the suite sets itself.

Base: ed43e69 (merge-base) | Changed: spill-dedup-view-arrays

tpcds — tpcds_sf1

Query Base Changed Change
Query 1 192.6 KiB 192.6 KiB +0.0%
Query 2 91.5 MiB 99.5 MiB +8.8%
Query 3 3.4 MiB 3.4 MiB +0.0%
Query 4 120.5 MiB 120.5 MiB +0.0%
Query 5 13.0 MiB 13.0 MiB +0.2%
Query 6 7.6 MiB 7.6 MiB -0.0%
Query 7 9.1 MiB 8.3 MiB -8.3%
Query 8 2.5 MiB 2.4 MiB -2.3%
Query 9 32.1 KiB 32.1 KiB +0.0%
Query 10 2.9 MiB 2.9 MiB +0.0%
Query 11 83.7 MiB 83.7 MiB +0.0%
Query 12 130.3 MiB 131.1 MiB +0.7%
Query 13 15.4 MiB 16.1 MiB +4.1%
Query 14 170.3 MiB 171.4 MiB +0.7%
Query 15 6.5 MiB 6.7 MiB +3.9%
Query 16 161.6 KiB 161.6 KiB +0.0%
Query 17 40.4 MiB 42.8 MiB +5.7%
Query 18 19.3 MiB 19.7 MiB +2.3%
Query 19 5.0 MiB 5.0 MiB -0.0%
Query 20 131.3 MiB 131.3 MiB +0.0%
Query 21 5.3 MiB 5.3 MiB +0.9%
Query 22 91.6 MiB 92.3 MiB +0.7%
Query 23 205.6 MiB 204.3 MiB -0.6%
Query 24 48.1 MiB 48.1 MiB +0.0%
Query 25 99.1 MiB 99.1 MiB +0.0%
Query 26 9.7 MiB 9.7 MiB -0.0%
Query 27 1.7 MiB 1.7 MiB +0.0%
Query 28 1.3 MiB 1.2 MiB -9.9%
Query 29 104.4 MiB 104.4 MiB +0.0%
Query 30 6.7 MiB 6.7 MiB -0.0%
Query 31 17.8 MiB 17.8 MiB +0.0%
Query 32 2.6 MiB 2.6 MiB +0.0%
Query 33 4.1 MiB 3.9 MiB -6.7%
Query 34 14.3 MiB 14.3 MiB +0.0%
Query 35 25.3 MiB 25.3 MiB +0.0%
Query 36 10.6 MiB 10.6 MiB +0.0%
Query 37 3.8 MiB 3.8 MiB +0.0%
Query 38 20.1 MiB 20.1 MiB +0.0%
Query 39 42.9 MiB 73.3 MiB +71.0%
Query 40 13.5 MiB 13.5 MiB +0.0%
Query 41 2.0 MiB 1.9 MiB -7.7%
Query 42 1.6 MiB 2.0 MiB +19.0%
Query 43 192.4 KiB 192.4 KiB +0.0%
Query 44 192.4 KiB 192.4 KiB +0.0%
Query 45 8.1 MiB 8.5 MiB +4.7%
Query 46 10.1 MiB 10.1 MiB +0.0%
Query 47 155.9 MiB 146.8 MiB -5.9%
Query 48 16.2 MiB 15.8 MiB -2.9%
Query 49 52.0 MiB 51.6 MiB -0.8%
Query 50 15.5 MiB 15.2 MiB -2.3%
Query 51 178.6 MiB 178.6 MiB -0.0%
Query 52 3.3 MiB 3.3 MiB +0.0%
Query 53 30.6 MiB 30.6 MiB +0.0%
Query 54 4.3 MiB 4.2 MiB -3.6%
Query 55 2.9 MiB 3.2 MiB +9.8%
Query 56 4.0 MiB 4.0 MiB +0.9%
Query 57 141.9 MiB 141.9 MiB +0.0%
Query 58 6.2 MiB 5.9 MiB -4.2%
Query 59 15.0 MiB 15.0 MiB +0.0%
Query 60 5.6 MiB 5.9 MiB +3.8%
Query 61 1.0 MiB 880.8 KiB -17.9%
Query 62 7.0 MiB 7.7 MiB +9.9%
Query 63 20.6 MiB 20.6 MiB -0.0%
Query 64 100.1 MiB 107.8 MiB +7.7%
Query 65 28.3 MiB 28.3 MiB -0.0%
Query 66 12.2 MiB 15.9 MiB +30.5%
Query 67 304.8 MiB 362.0 MiB +18.8%
Query 68 10.1 MiB 10.1 MiB +0.0%
Query 69 46.6 MiB 46.6 MiB -0.0%
Query 70 72.5 MiB 68.4 MiB -5.7%
Query 71 40.4 MiB 60.4 MiB +49.5%
Query 72 96.9 MiB 96.9 MiB +0.0%
Query 73 14.3 MiB 14.3 MiB +0.0%
Query 74 41.6 MiB 41.6 MiB +0.0%
Query 75 20.3 MiB 20.3 MiB +0.0%
Query 76 7.7 MiB 7.6 MiB -0.3%
Query 77 3.4 MiB 3.4 MiB -0.0%
Query 78 109.9 MiB 110.7 MiB +0.8%
Query 79 11.6 MiB 10.8 MiB -6.8%
Query 80 40.0 MiB 39.6 MiB -1.0%
Query 81 10.7 MiB 10.7 MiB +0.0%
Query 82 3.8 MiB 3.8 MiB +0.0%
Query 83 11.2 MiB 11.2 MiB +0.0%
Query 84 11.9 MiB 11.9 MiB +0.0%
Query 85 19.3 MiB 19.6 MiB +1.5%
Query 86 105.2 MiB 105.2 MiB +0.0%
Query 87 20.3 MiB 20.3 MiB +0.0%
Query 88 1.5 MiB 1.7 MiB +11.1%
Query 89 130.0 MiB 130.0 MiB +0.0%
Query 90 562.1 KiB 562.1 KiB +0.0%
Query 91 22.7 MiB 22.7 MiB +0.0%
Query 92 2.9 MiB 2.6 MiB -7.7%
Query 93 15.4 MiB 15.8 MiB +2.4%
Query 94 3.4 MiB 3.7 MiB +8.4%
Query 95 18.3 MiB 17.1 MiB -6.1%
Query 96 224.7 KiB 224.7 KiB +0.0%
Query 97 32.3 MiB 32.3 MiB +0.0%
Query 98 131.5 MiB 131.5 MiB +0.0%
Query 99 7.8 MiB 7.1 MiB -8.9%

Pool accounting vs. process RSS

Max pool peak is the largest reservation any single query in the run reached; peak RSS covers the whole invocation, including data loading and allocator retention, and the two high-water marks need not coincide in time. The gap is therefore an upper bound on what the pool did not account for, not a measurement of it.

Benchmark Side Max pool peak Peak RSS Gap RSS / pool
tpcds base (ed43e69 (merge-base)) 304.8 MiB 3.1 GiB 2.8 GiB 10.4×
tpcds changed (spill-dedup-view-arrays) 362.0 MiB 3.6 GiB 3.2 GiB 10.1×
Resource Usage

tpcds — base (merge-base)

Metric Value
Wall time 50.0s
Peak memory 3.1 GiB
Avg memory 2.1 GiB
CPU user 207.7s
CPU sys 6.7s
Peak spill 86.9 MiB

tpcds — branch

Metric Value
Wall time 50.0s
Peak memory 3.6 GiB
Avg memory 2.1 GiB
CPU user 207.5s
CPU sys 6.7s
Peak spill 55.1 MiB

File an issue against this benchmark runner

@adriangbot

Copy link
Copy Markdown

🤖 Benchmark completed (GKE) | trigger

Instance: c4a-highmem-16 (12 vCPU / 65 GiB)

Comparing spill-dedup-view-arrays (9c2f8f1) to ed43e69 (merge-base) diff

Run configuration
run benchmark clickbench_partitioned
env:
  DATAFUSION_RUNTIME_MEMORY_LIMIT: "512M"
CPU Details (lscpu)
Architecture:                            aarch64
CPU op-mode(s):                          64-bit
Byte Order:                              Little Endian
CPU(s):                                  16
On-line CPU(s) list:                     0-15
Vendor ID:                               ARM
Model name:                              Neoverse-V2
Model:                                   1
Thread(s) per core:                      1
Core(s) per cluster:                     16
Socket(s):                               -
Cluster(s):                              1
Stepping:                                r0p1
BogoMIPS:                                2000.00
Flags:                                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit uscat ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
L1d cache:                               1 MiB (16 instances)
L1i cache:                               1 MiB (16 instances)
L2 cache:                                32 MiB (16 instances)
L3 cache:                                80 MiB (1 instance)
NUMA node(s):                            1
NUMA node0 CPU(s):                       0-15
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Not affected
Vulnerability Spec rstack overflow:      Not affected
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; CSV2, BHB
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Not affected
Details

Comparing HEAD and spill-dedup-view-arrays
--------------------
Benchmark clickbench_partitioned.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━┓
┃ Query     ┃       HEAD ┃ spill-dedup-view-arrays ┃       Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━┩
│ QQuery 0  │    1.22 ms │                 1.22 ms │    no change │
│ QQuery 1  │   11.78 ms │                11.70 ms │    no change │
│ QQuery 2  │   37.10 ms │                37.23 ms │    no change │
│ QQuery 3  │   31.31 ms │                31.36 ms │    no change │
│ QQuery 4  │  329.84 ms │               332.95 ms │    no change │
│ QQuery 5  │  515.26 ms │               530.39 ms │    no change │
│ QQuery 6  │    1.28 ms │                 1.28 ms │    no change │
│ QQuery 7  │   12.82 ms │                12.80 ms │    no change │
│ QQuery 8  │  691.14 ms │               698.62 ms │    no change │
│ QQuery 9  │  573.73 ms │               591.03 ms │    no change │
│ QQuery 10 │   65.37 ms │                65.49 ms │    no change │
│ QQuery 11 │   75.82 ms │                76.38 ms │    no change │
│ QQuery 12 │  558.18 ms │               564.17 ms │    no change │
│ QQuery 13 │ 1501.29 ms │              1533.93 ms │    no change │
│ QQuery 14 │  621.10 ms │               621.18 ms │    no change │
│ QQuery 15 │  374.69 ms │               370.71 ms │    no change │
│ QQuery 16 │ 1306.28 ms │              1315.75 ms │    no change │
│ QQuery 17 │  942.31 ms │               930.31 ms │    no change │
│ QQuery 18 │ 2731.40 ms │              2719.93 ms │    no change │
│ QQuery 19 │   27.72 ms │                27.31 ms │    no change │
│ QQuery 20 │  512.11 ms │               509.42 ms │    no change │
│ QQuery 21 │  499.84 ms │               497.95 ms │    no change │
│ QQuery 22 │  968.31 ms │               969.45 ms │    no change │
│ QQuery 23 │       FAIL │                    FAIL │ incomparable │
│ QQuery 24 │   40.41 ms │                40.19 ms │    no change │
│ QQuery 25 │  105.26 ms │               104.28 ms │    no change │
│ QQuery 26 │   40.82 ms │                42.63 ms │    no change │
│ QQuery 27 │  514.42 ms │               507.28 ms │    no change │
│ QQuery 28 │ 3373.94 ms │              3323.85 ms │    no change │
│ QQuery 29 │   41.95 ms │                41.92 ms │    no change │
│ QQuery 30 │  438.82 ms │               434.37 ms │    no change │
│ QQuery 31 │  531.96 ms │               524.00 ms │    no change │
│ QQuery 32 │ 3494.58 ms │              3505.65 ms │    no change │
│ QQuery 33 │ 4294.38 ms │              4346.00 ms │    no change │
│ QQuery 34 │ 4290.76 ms │              4394.65 ms │    no change │
│ QQuery 35 │  337.05 ms │               333.52 ms │    no change │
│ QQuery 36 │   64.25 ms │                65.55 ms │    no change │
│ QQuery 37 │   35.55 ms │                35.51 ms │    no change │
│ QQuery 38 │   40.81 ms │                42.08 ms │    no change │
│ QQuery 39 │  136.44 ms │               131.19 ms │    no change │
│ QQuery 40 │   14.18 ms │                13.88 ms │    no change │
│ QQuery 41 │   13.55 ms │                13.42 ms │    no change │
│ QQuery 42 │   12.98 ms │                12.92 ms │    no change │
└───────────┴────────────┴─────────────────────────┴──────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary                      ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)                      │ 30212.03ms │
│ Total Time (spill-dedup-view-arrays)   │ 30363.42ms │
│ Average Time (HEAD)                    │   719.33ms │
│ Average Time (spill-dedup-view-arrays) │   722.94ms │
│ Queries Faster                         │          0 │
│ Queries Slower                         │          0 │
│ Queries with No Change                 │         42 │
│ Queries with Failure                   │          1 │
└────────────────────────────────────────┴────────────┘

Distribution per query (min / mean ±stddev / max):

Comparing HEAD and spill-dedup-view-arrays
--------------------
Benchmark clickbench_partitioned.json
--------------------
┏━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━━━━┓
┃ Query     ┃                                   HEAD ┃               spill-dedup-view-arrays ┃        Change ┃
┡━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━━━━┩
│ QQuery 0  │           1.22 / 4.04 ±5.50 / 15.04 ms │          1.22 / 3.92 ±5.33 / 14.58 ms │     no change │
│ QQuery 1  │         11.78 / 11.83 ±0.06 / 11.96 ms │        11.70 / 11.85 ±0.08 / 11.94 ms │     no change │
│ QQuery 2  │         37.10 / 37.55 ±0.27 / 37.80 ms │        37.23 / 37.48 ±0.15 / 37.64 ms │     no change │
│ QQuery 3  │         31.31 / 32.23 ±0.82 / 33.58 ms │        31.36 / 31.66 ±0.24 / 31.94 ms │     no change │
│ QQuery 4  │      329.84 / 335.53 ±3.79 / 340.27 ms │     332.95 / 340.61 ±5.82 / 348.12 ms │     no change │
│ QQuery 5  │      515.26 / 526.19 ±7.12 / 534.58 ms │     530.39 / 535.19 ±3.35 / 540.90 ms │     no change │
│ QQuery 6  │            1.28 / 1.44 ±0.24 / 1.91 ms │           1.28 / 1.43 ±0.23 / 1.89 ms │     no change │
│ QQuery 7  │         12.82 / 13.17 ±0.18 / 13.33 ms │        12.80 / 12.87 ±0.07 / 12.97 ms │     no change │
│ QQuery 8  │     691.14 / 702.73 ±10.11 / 715.35 ms │     698.62 / 707.55 ±8.04 / 722.44 ms │     no change │
│ QQuery 9  │     573.73 / 585.79 ±11.82 / 605.36 ms │    591.03 / 604.89 ±10.45 / 618.20 ms │     no change │
│ QQuery 10 │         65.37 / 67.13 ±1.50 / 69.11 ms │        65.49 / 67.78 ±1.29 / 69.38 ms │     no change │
│ QQuery 11 │         75.82 / 78.57 ±1.79 / 80.86 ms │        76.38 / 78.76 ±1.77 / 80.53 ms │     no change │
│ QQuery 12 │     558.18 / 572.51 ±15.61 / 596.56 ms │    564.17 / 576.36 ±12.07 / 596.45 ms │     no change │
│ QQuery 13 │  1501.29 / 1533.91 ±27.79 / 1569.67 ms │ 1533.93 / 1550.20 ±21.76 / 1591.71 ms │     no change │
│ QQuery 14 │     621.10 / 637.22 ±14.28 / 660.83 ms │    621.18 / 649.29 ±14.86 / 663.91 ms │     no change │
│ QQuery 15 │      374.69 / 378.96 ±3.47 / 384.80 ms │     370.71 / 381.76 ±7.06 / 390.91 ms │     no change │
│ QQuery 16 │  1306.28 / 1321.61 ±11.28 / 1335.19 ms │  1315.75 / 1331.16 ±8.18 / 1338.42 ms │     no change │
│ QQuery 17 │     942.31 / 961.46 ±11.09 / 973.99 ms │    930.31 / 962.18 ±19.58 / 985.88 ms │     no change │
│ QQuery 18 │  2731.40 / 2762.65 ±28.15 / 2799.90 ms │ 2719.93 / 2767.96 ±26.20 / 2793.26 ms │     no change │
│ QQuery 19 │         27.72 / 28.14 ±0.61 / 29.34 ms │        27.31 / 27.74 ±0.34 / 28.28 ms │     no change │
│ QQuery 20 │      512.11 / 518.26 ±6.79 / 531.47 ms │     509.42 / 512.19 ±2.13 / 514.17 ms │     no change │
│ QQuery 21 │      499.84 / 503.09 ±3.41 / 509.08 ms │     497.95 / 510.94 ±9.42 / 521.84 ms │     no change │
│ QQuery 22 │      968.31 / 973.54 ±3.04 / 977.49 ms │     969.45 / 974.52 ±3.85 / 979.46 ms │     no change │
│ QQuery 23 │                                   FAIL │                                  FAIL │  incomparable │
│ QQuery 24 │         40.41 / 50.75 ±9.77 / 66.04 ms │        40.19 / 44.80 ±7.03 / 58.60 ms │ +1.13x faster │
│ QQuery 25 │     105.26 / 120.59 ±20.04 / 160.04 ms │     104.28 / 108.92 ±3.44 / 113.91 ms │ +1.11x faster │
│ QQuery 26 │         40.82 / 41.03 ±0.16 / 41.20 ms │       42.63 / 54.03 ±17.11 / 87.95 ms │  1.32x slower │
│ QQuery 27 │      514.42 / 521.56 ±4.63 / 528.11 ms │     507.28 / 514.03 ±4.25 / 520.46 ms │     no change │
│ QQuery 28 │   3373.94 / 3386.43 ±7.94 / 3395.49 ms │ 3323.85 / 3362.82 ±25.63 / 3400.67 ms │     no change │
│ QQuery 29 │         41.95 / 42.31 ±0.35 / 42.95 ms │        41.92 / 43.28 ±2.45 / 48.18 ms │     no change │
│ QQuery 30 │      438.82 / 446.74 ±8.46 / 459.60 ms │    434.37 / 451.14 ±13.79 / 468.17 ms │     no change │
│ QQuery 31 │      531.96 / 543.97 ±9.41 / 554.91 ms │    524.00 / 537.70 ±13.16 / 558.14 ms │     no change │
│ QQuery 32 │   3494.58 / 3504.73 ±7.80 / 3513.28 ms │  3505.65 / 3514.44 ±6.03 / 3523.08 ms │     no change │
│ QQuery 33 │ 4294.38 / 4487.46 ±153.94 / 4717.91 ms │ 4346.00 / 4401.22 ±46.65 / 4473.91 ms │     no change │
│ QQuery 34 │  4290.76 / 4410.62 ±82.46 / 4510.00 ms │ 4394.65 / 4445.95 ±47.29 / 4529.58 ms │     no change │
│ QQuery 35 │      337.05 / 346.48 ±7.90 / 358.15 ms │     333.52 / 343.70 ±6.97 / 351.32 ms │     no change │
│ QQuery 36 │         64.25 / 68.77 ±4.45 / 76.11 ms │        65.55 / 68.85 ±5.69 / 80.20 ms │     no change │
│ QQuery 37 │         35.55 / 36.45 ±0.57 / 37.35 ms │        35.51 / 36.39 ±1.00 / 38.30 ms │     no change │
│ QQuery 38 │         40.81 / 44.58 ±4.37 / 53.02 ms │        42.08 / 44.45 ±1.94 / 46.26 ms │     no change │
│ QQuery 39 │      136.44 / 138.19 ±1.84 / 141.69 ms │     131.19 / 140.04 ±9.07 / 156.08 ms │     no change │
│ QQuery 40 │         14.18 / 15.10 ±0.96 / 16.35 ms │        13.88 / 14.35 ±0.34 / 14.87 ms │     no change │
│ QQuery 41 │         13.55 / 13.84 ±0.35 / 14.48 ms │        13.42 / 14.45 ±1.13 / 15.84 ms │     no change │
│ QQuery 42 │         12.98 / 13.11 ±0.15 / 13.39 ms │        12.92 / 14.18 ±2.05 / 18.28 ms │  1.08x slower │
└───────────┴────────────────────────────────────────┴───────────────────────────────────────┴───────────────┘
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┳━━━━━━━━━━━━┓
┃ Benchmark Summary                      ┃            ┃
┡━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╇━━━━━━━━━━━━┩
│ Total Time (HEAD)                      │ 30820.29ms │
│ Total Time (spill-dedup-view-arrays)   │ 30833.03ms │
│ Average Time (HEAD)                    │   733.82ms │
│ Average Time (spill-dedup-view-arrays) │   734.12ms │
│ Queries Faster                         │          2 │
│ Queries Slower                         │          2 │
│ Queries with No Change                 │         38 │
│ Queries with Failure                   │          1 │
└────────────────────────────────────────┴────────────┘

Memory Pool Peaks

Peak MemoryPool reservation per query — what DataFusion's accounting believes it reserved. Recorded only for benchmarks that write a results JSON (the dfbench suites, and the suites bench.sh runs through the Criterion SQL harness since #25644), and only under a memory limit: DATAFUSION_RUNTIME_MEMORY_LIMIT, or one the suite sets itself.

Base: ed43e69 (merge-base) | Changed: spill-dedup-view-arrays

clickbench_partitioned

Query Base Changed Change
Query 0 0 B 0 B 0.0%
Query 1 104 B 104 B +0.0%
Query 2 936 B 936 B +0.0%
Query 3 312 B 312 B +0.0%
Query 4 274.3 MiB 272.3 MiB -0.7%
Query 5 342.7 MiB 353.7 MiB +3.2%
Query 6 0 B 0 B 0.0%
Query 7 50.2 MiB 40.3 MiB -19.7%
Query 8 129.8 MiB 130.9 MiB +0.8%
Query 9 333.6 MiB 333.7 MiB +0.0%
Query 10 108.7 MiB 104.8 MiB -3.7%
Query 11 121.5 MiB 115.3 MiB -5.1%
Query 12 325.9 MiB 331.0 MiB +1.6%
Query 13 303.3 MiB 304.6 MiB +0.4%
Query 14 389.5 MiB 397.4 MiB +2.0%
Query 15 281.6 MiB 280.2 MiB -0.5%
Query 16 429.8 MiB 382.3 MiB -11.1%
Query 17 414.8 MiB 407.3 MiB -1.8%
Query 18 379.0 MiB 383.8 MiB +1.3%
Query 19 0 B 0 B 0.0%
Query 20 104 B 104 B +0.0%
Query 21 3.3 MiB 3.7 MiB +9.1%
Query 22 3.1 MiB 3.6 MiB +16.8%
Query 23 0 B 0 B 0.0%
Query 24 59.4 MiB 60.7 MiB +2.1%
Query 25 157.6 MiB 174.1 MiB +10.5%
Query 26 57.9 MiB 57.9 MiB +0.0%
Query 27 2.4 MiB 2.4 MiB +0.0%
Query 28 315.8 MiB 314.7 MiB -0.4%
Query 29 624 B 624 B +0.0%
Query 30 378.9 MiB 392.2 MiB +3.5%
Query 31 146.7 MiB 148.2 MiB +1.0%
Query 32 278.0 MiB 288.8 MiB +3.9%
Query 33 383.6 MiB 359.6 MiB -6.2%
Query 34 372.9 MiB 370.4 MiB -0.7%
Query 35 292.0 MiB 294.2 MiB +0.7%
Query 36 115.2 MiB 107.7 MiB -6.5%
Query 37 6.9 MiB 6.9 MiB +0.0%
Query 38 5.7 MiB 5.6 MiB -0.6%
Query 39 298.0 MiB 283.4 MiB -4.9%
Query 40 2.0 MiB 2.0 MiB -1.2%
Query 41 3.1 MiB 3.1 MiB +0.0%
Query 42 1.7 MiB 1.6 MiB -6.0%

Pool accounting vs. process RSS

Max pool peak is the largest reservation any single query in the run reached; peak RSS covers the whole invocation, including data loading and allocator retention, and the two high-water marks need not coincide in time. The gap is therefore an upper bound on what the pool did not account for, not a measurement of it.

Benchmark Side Max pool peak Peak RSS Gap RSS / pool
clickbench_partitioned base (ed43e69 (merge-base)) 429.8 MiB 24.2 GiB 23.8 GiB 57.8×
clickbench_partitioned changed (spill-dedup-view-arrays) 407.3 MiB 22.1 GiB 21.7 GiB 55.5×
Resource Usage

clickbench_partitioned — base (merge-base)

Metric Value
Wall time 160.0s
Peak memory 24.2 GiB
Avg memory 7.3 GiB
CPU user 1463.4s
CPU sys 307.3s
Peak spill 5.2 GiB

clickbench_partitioned — branch

Metric Value
Wall time 160.0s
Peak memory 22.1 GiB
Avg memory 7.9 GiB
CPU user 1457.9s
CPU sys 309.3s
Peak spill 5.2 GiB

File an issue against this benchmark runner

@adriangb adriangb left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Benchmarks and implementation look good to me, I'll leave this open for @kumarUjjawal to weigh in

@kumarUjjawal kumarUjjawal left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good 👍

@adriangb
adriangb enabled auto-merge September 29, 2026 10:56
Jefffrey pushed a commit to apache/arrow-rs that referenced this pull request Oct 1, 2026
# Which issue does this PR close?

Fixes #11280

And also helps with: apache/datafusion#23565

# Rationale for this change

See the original issue, but, basically the interleave kernel can make
sub-optimal string views especially when buffers are shared.

# What changes are included in this PR?

We use a btreemap on the seen buffers (not individual values) and use
that to prevent us from cloning the same buffer multiple times.

# Are these changes tested?

Yes tested, and locally benchmarks have shown a moderate speed increase
in some of the interleave benchmarks, on my dev machine which has an
**AMD Ryzen 9 9950X3D** processor:

| benchmark | main | this PR | change |

|-----------------------------------------------------------------------|----------|----------|--------|
| interleave str_view(0.0) 100 [0..100, 100..230, 450..1000] | 284.5 ns
| 268.6 ns | -5.6% |
| interleave str_view(0.0) 400 [0..100, 100..230, 450..1000] | 706.9 ns
| 563.9 ns | -20.2% |
| interleave str_view(0.0) 1024 [0..100, 100..230, 450..1000] | 1.717 µs
| 1.250 µs | -27.2% |
| interleave str_view(0.0) 1024 [0..100, 100..230, 450..1000, 0..1000] |
1.783 µs | 1.274 µs | -28.6% |


# Are there any user-facing changes?

No user facing changes, just an optimization.
cetra3 added a commit to pydantic/arrow-rs that referenced this pull request Oct 2, 2026
# Which issue does this PR close?

Fixes apache#11280

And also helps with: apache/datafusion#23565

# Rationale for this change

See the original issue, but, basically the interleave kernel can make
sub-optimal string views especially when buffers are shared.

# What changes are included in this PR?

We use a btreemap on the seen buffers (not individual values) and use
that to prevent us from cloning the same buffer multiple times.

# Are these changes tested?

Yes tested, and locally benchmarks have shown a moderate speed increase
in some of the interleave benchmarks, on my dev machine which has an
**AMD Ryzen 9 9950X3D** processor:

| benchmark | main | this PR | change |

|-----------------------------------------------------------------------|----------|----------|--------|
| interleave str_view(0.0) 100 [0..100, 100..230, 450..1000] | 284.5 ns
| 268.6 ns | -5.6% |
| interleave str_view(0.0) 400 [0..100, 100..230, 450..1000] | 706.9 ns
| 563.9 ns | -20.2% |
| interleave str_view(0.0) 1024 [0..100, 100..230, 450..1000] | 1.717 µs
| 1.250 µs | -27.2% |
| interleave str_view(0.0) 1024 [0..100, 100..230, 450..1000, 0..1000] |
1.783 µs | 1.274 µs | -28.6% |

# Are there any user-facing changes?

No user facing changes, just an optimization.

(cherry picked from commit 007894e)
@adriangb
adriangb added this pull request to the merge queue Oct 7, 2026
Merged via the queue into apache:main with commit b350ea6 Oct 7, 2026
42 checks passed
@adriangb
adriangb deleted the spill-dedup-view-arrays branch October 7, 2026 19:39
Omega359 pushed a commit to Omega359/arrow-datafusion that referenced this pull request Oct 11, 2026
…iew columns (apache#25625)

## Which issue does this PR close?

- Refers to apache#23564.
- Precursor for apache#23565. This PR
must merge first, so the benchmark bot can compare `main` and that PR
with `run benchmark spill_views`.

## Rationale for this change

apache#23565 changes how spilled
`StringView` and `BinaryView` columns are compacted. No current
benchmark shows the effect:

- The `sort_tpch` string columns are either high-cardinality
(`l_comment`) or have a dictionary buffer that is too small to be
compacted (`l_shipinstruct`).
- Most suites spill only when the environment sets a low memory limit.

## What changes are included in this PR?

A new SQL benchmark suite `spill_views` (in
`benchmarks/sql_benchmarks/spill_views/`) and a `bench.sh` entry for it.
There is no change to the spill code.

Each query reads 1M rows from a Parquet file that the suite writes with
`COPY`. The data must come from Parquet, because then the views of a
dictionary-encoded column point into one shared buffer. Each query sets
its own memory limit and `target_partitions = 4`, so it spills in any
environment.

| Query | Subgroup | What it does | Memory limit |
|---|---|---|---|
| `q01_sort_string_1_distinct` | `repeated` | `ORDER BY` a shuffled key,
`StringView` payload with 1 distinct value (64 bytes) | 40M |
| `q02_sort_string_1000_distinct` | `repeated` | Same, with 1000
distinct values | 40M |
| `q03_sort_binary_1000_distinct` | `repeated` | Same as q02, with a
`BinaryView` payload | 40M |
| `q04_sort_string_all_distinct` | `distinct` | Same, with all-distinct
values (compaction cannot remove data) | 96M |
| `q05_group_by_string_all_distinct` | `distinct` | `GROUP BY` an
all-distinct string key (compaction cannot remove data) | 96M |

The asserts check the row count, the number of distinct values, that the
column is read as a view type, and that the memory limit and
`target_partitions` settings are applied.

## What is the testing strategy for this PR?

I ran the suite with `benchmark_runner` and with `./bench.sh run
spill_views` (the `cargo bench --bench sql` path that the bot uses). All
asserts pass and each query takes less than 150 ms per iteration.

I also ran the same queries with `EXPLAIN ANALYZE`, on `main` and with
this commit on top of apache#23565.
All queries spill on both sides. Spilled bytes are from the sort or
final aggregation operator. Times are the median of 40 iterations of
`benchmark_runner` on a laptop (Apple M4 Pro, local SSD), with runs of
both builds alternated:

| Query | Spilled (`main`) | Spilled (PR) | Time (`main`) | Time (PR) |
|---|---|---|---|---|
| q01 sort, 1 distinct | 84.2 MB | 23.2 MB | 63.5 ms | 57.0 ms |
| q02 sort, 1000 distinct | 84.2 MB | 30.6 MB | 57.0 ms | 65.3 ms |
| q03 sort, 1000 distinct binary | 84.2 MB | 30.6 MB | 54.0 ms | 65.9 ms
|
| q04 sort, all distinct | 84.2 MB | 84.2 MB | 71.0 ms | 95.7 ms |
| q05 group by, all distinct | 84.2 MB | 84.2 MB | 75.2 ms | 77.7 ms |

The machine was under load, so the times are only approximate. On a
local SSD, spill I/O is cheap. The benchmark bot will give better
numbers.

## Are there any user-facing changes?

No. This PR only adds a benchmark.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Omega359 pushed a commit to Omega359/arrow-datafusion that referenced this pull request Oct 11, 2026
…nment (apache#25634)

## Which issue does this PR close?

- Follow-up to apache#25625.

## Rationale for this change

Each `spill_views` query hardcodes its memory limit, and a benchmark-bot
trigger comment can only set environment variables. So the only way to
change a limit today is another DataFusion PR.

At the current defaults every query spills on `main` and passes (q05,
the GROUP BY, passed 30 of 30 executions at every limit from 96M to
192M). But apache#23565 fails q05 at
96M on every bot run (4 of 4), which aborts the whole suite run, so that
PR gets no numbers.

## What changes are included in this PR?

- `SPILL_VIEWS_LIMIT_REPEATED` (default `40M`) and
`SPILL_VIEWS_LIMIT_DISTINCT` (default `96M`) set the limit of each
subgroup, also as `--limit-repeated` / `--limit-distinct` on
`benchmark_runner`. The defaults do not change.
- The docs suggest `128M` as an override: q05 stops spilling at about
`192M`, so larger values no longer measure the spill path.

## Are these changes tested?

Ran the suite through `benchmark_runner` at the defaults and with
`SPILL_VIEWS_LIMIT_DISTINCT=128M`. The suite's asserts read back the
limit each query ran with, so a wrong override fails the run.

## Are there any user-facing changes?

No, benchmarks only.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Omega359 pushed a commit to Omega359/arrow-datafusion that referenced this pull request Oct 11, 2026
…ess (apache#25644)

## Which issue does this PR close?

- None. This is follow-up work to
apache#23985, which added per-query
`pool_peak_bytes` to the `dfbench` results JSON.

## Rationale for this change

The suites that `bench.sh` runs through the Criterion SQL harness
(`cargo bench --bench sql`: `spill_views`, `wide_schema`,
`predicate_eval`, and others) do not report a per-query memory pool
peak. The harness prints one number to stdout and writes no results
JSON. For example, this benchmark bot run of `spill_views` with
`DATAFUSION_RUNTIME_MEMORY_LIMIT: "512M"` has no pool peak section:
apache#23565 (comment)

A second problem: a suite that sets its own limit with `SET
datafusion.runtime.memory_limit` records nothing, even if the output
were written. `spill_views` does this in its `init` script. The `SET`
makes `SessionContext` build a new `RuntimeEnv` with a new pool, and
that drops the `PeakRecordingPool` that `CommonOpt::runtime_env_builder`
installed.

## What changes are included in this PR?

Two commits:

1. **Keep the recorder across a SQL `SET`.** After a benchmark's `load`
and `init` steps, `prepare_benchmark` puts a `PeakRecordingPool` in
front of the session's pool again if the pool has a finite limit and no
recorder. The new `RuntimeEnv` is installed the same way `SET
datafusion.runtime.*` does it
(`SessionStateBuilder::from(state).with_runtime_env(..)`), so tables and
config are kept. The simple runner (`benchmark_runner -o`) also goes
through `prepare_benchmark`, so it gets this fix too.
2. **Write a results JSON from the Criterion harness.** When
`BENCH_RESULTS_FILE` is set, the harness writes the same JSON format as
`dfbench -o`. Each case is the Criterion id (for example,
`spill_views/q01_sort_string_1_distinct_repeated`, the same as in
`critcmp`). `pool_peak_bytes` covers every execution Criterion makes of
the query, including warm-up. The `load`, `init` and `assert` steps are
not included. `iterations` is empty, because Criterion keeps the
timings. `bench.sh` passes `results/<name>/criterion/<file>.json` at
each SQL-harness call site. The subdirectory keeps the file out of
`bench.sh compare`, which reads `results/<name>/*.json` as timing
results. This means no second timing table.

The benchmark bot reads the new file in
adriangb/datafusion-benchmarking#38. The bot
takes `bench.sh` from `main`, so it shows these tables only after this
PR is merged. Until then, and for any side that is older than this PR,
the bot shows a note that the pool peaks are not available.

Not changed here: `SET datafusion.runtime.memory_limit` always builds a
`GreedyMemoryPool` and drops any wrapper, so `--mem-pool-type` has no
effect on suites that set their own limit. The fix in this PR stays in
the benchmarks crate. A core change that keeps a pool wrapper across
`SET` is possible, but it is out of scope for this PR.

## What is the testing strategy for this PR?

New unit tests in `benchmarks/src/sql_benchmark_runner.rs`:

- `sql_memory_limit_keeps_the_peak_recorder`: a SQL `SET` replaces the
harness pool, and the recorder is put back and records the next query.
- `no_memory_limit_gets_no_recorder`: without any limit, nothing
changes.
- `criterion_harness_writes_pool_peak_per_case`: an end-to-end Criterion
run of a suite whose `init` sets the limit writes a JSON with a non-zero
peak and empty `iterations`.

If the rewrap is disabled, the first and third tests fail.

`./bench.sh run spill_views`, the command that the bot runs
(`SQL_CARGO_COMMAND="cargo bench --bench sql -- --save-baseline ..."`),
writes `results/<name>/criterion/spill_views.json`. The peaks are the
same with `DATAFUSION_RUNTIME_MEMORY_LIMIT=512M` and with no env limit,
because the suite's own limits apply (40M for q01 to q03, 96M for q04
and q05):

| Query | `pool_peak_bytes` |
| --- | --- |
| `spill_views/q01_sort_string_1_distinct_repeated` | 42614784 (40.6
MiB) |
| `spill_views/q02_sort_string_1000_distinct_repeated` | 41709504 (39.8
MiB) |
| `spill_views/q03_sort_binary_1000_distinct_repeated` | 41709504 (39.8
MiB) |
| `spill_views/q04_sort_string_all_distinct_distinct` | 107544576 (102.6
MiB) |
| `spill_views/q05_group_by_string_all_distinct_distinct` | 100633968
(96.0 MiB) |

The q01 and q04 peaks are a little above their limits. The pool allows
this: infallible `grow` calls are not limited, and the recorder counts
what the pool grants.

`./bench.sh compare` with two such result sets reads no file.
`sort_tpch` (dfbench, queries 1 and 2,
`DATAFUSION_RUNTIME_MEMORY_LIMIT=512M`) is unchanged. It writes
`results/<name>/sort_tpch1.json` with 5 iterations and `pool_peak_bytes`
for each query, and it writes no `criterion/` directory.

## Are there any user-facing changes?

Only for benchmark tooling. There is a new optional `BENCH_RESULTS_FILE`
env var for `cargo bench --bench sql` (documented in
`benchmarks/sql_benchmarks/README.md`). `bench.sh run` now writes
`results/<name>/criterion/*.json` for SQL-harness suites. A SQL-set
memory limit now reports a pool peak. There are no public API changes
outside the benchmarks crate.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Omega359 pushed a commit to Omega359/arrow-datafusion that referenced this pull request Oct 11, 2026
## Which issue does this PR close?

- Closes apache#23564

## Rationale for this change

Spill files inflate string views via GC. Arrow's `gc()` copies the bytes
of every view separately, so a highly repeated string view column (for
example a dictionary-encoded Parquet column) spills one copy per row,
causing more memory and disk pressure.

## What changes are included in this PR?

The spill path now compacts `StringView` / `BinaryView` arrays by
copying each distinct value once:

- Values are deduplicated with a hash table over the views, and null
views are zeroed.
- The first 256 non-inline values are sampled. If they contain almost no
repeats, compaction falls back to plain `gc()`, so all-distinct data
does not pay for hashing.

## Are these changes tested?

Yes:

- `test_gc_copies_repeated_values_once` checks that repeated values are
written once, both when views share one buffer (as with Parquet
dictionary columns) and when they reference separate copies.
- `test_gc_distinct_values` checks the fallback to `gc()` for distinct
values.
- Existing spill tests cover the rest of the path.

Benchmarks (`spill_views`, `sort_tpch` with a memory limit) show
0.78–0.94x on repeated-value spills, and no change for all-distinct data
or `sort_tpch`. See the results in the comments.

## Are there any user-facing changes?

No, this is an internal method change.

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Adrian Garcia Badaracco <1755071+adriangb@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

physical-plan Changes to the physical-plan crate v56.0.0

Projects

None yet

Development

Successfully merging this pull request may close these issues.

External sort spill GC inflates deduplicated string views

5 participants