Repository navigation
Show decimal add bug - #25974
Draft
kamilk-pal wants to merge 2 commits into
Draft
Show decimal add bug#25974kamilk-pal wants to merge 2 commits into
kamilk-pal wants to merge 2 commits into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Which issue does this PR close?
No linked issue; this PR is a failing regression reproducer.
Rationale for this change
Adding
12345678901234567890123456789012asDECIMAL(32, 0)to1.00000000000000000000asDECIMAL(38, 20)overflows in DataFusion. The inferred result keeps scale 20, so aligning the integer's scale multiplies its underlying value by10^20and exceeds Decimal128's capacity.The expected type and sums in this test follow Spark with
spark.sql.decimalOperations.allowPrecisionLoss=true: the result isdecimal(38,6)with value12345678901234567890123456789013.000000.PostgreSQL uses arbitrary-precision
numericarithmetic and preserves scale 20, returning12345678901234567890123456789013.00000000000000000000. The expected(38,6)type therefore describes Spark's scale-reduction policy rather than a type shared by all three engines.What changes are included in this PR?
Regression cases in
datafusion/sqllogictest/test_files/decimal.sltcheck the inferred result type and addition with both literal and column inputs. Comments identify the Spark configuration behind the expected output and PostgreSQL's different representation.The assertions intentionally expect the desired successful result and are expected to fail until the DataFusion behavior is changed.
What is the testing strategy for this PR?
Run the DataFusion SLT from this PR's checkout, then compare the same operands on PostgreSQL and Spark using the commands below. The PostgreSQL and Spark examples exercise both literals and typed columns.
Observed results for both literal and column inputs:
Decimal128(38, 20)numeric, scale 2012345678901234567890123456789013.00000000000000000000allowPrecisionLoss=truedecimal(38,6)12345678901234567890123456789013.000000The DataFusion result above was measured with the released Python binding, not a build of this PR's exact revision. The pinned Rust toolchain was unavailable in the validation workspace. The command below runs the actual SLT against the source checkout when that toolchain is installed.
DataFusion
From the repository root, with the repository's Rust toolchain and test dependencies installed:
cargo test --profile ci --test sqllogictests -- decimal.sltThis executes DataFusion itself. The new type assertion expects
Decimal128(38, 6)and is intentionally expected to fail while the engine infersDecimal128(38, 20); the runner can stop at that first failure.PostgreSQL
Prerequisite: Docker. This starts a temporary database without publishing a host port, runs both forms of the query, and removes the container on exit:
Both queries should report
numeric, scale20, and12345678901234567890123456789013.00000000000000000000.PG_COMPAT=truedoes not run this particulardecimal.sltfile: the PostgreSQL compatibility runner selectspg_compat_*files. Also,arrow_typeofis specific to DataFusion, so this example uses PostgreSQL'spg_typeofandscalefunctions.Spark
Prerequisites: Java 17 or newer and
uv. This provisions Python 3.12 and PySpark 4.2.0, then runs a local Apache Spark session:Both queries should report
decimal(38,6)and12345678901234567890123456789013.000000.The precision-loss setting matters: changing
spark.sql.decimalOperations.allowPrecisionLosstofalsegivesdecimal(38,20)and an overflow with ANSI mode enabled. With ANSI mode disabled, the result isNULL. The ordinary DataFusion SLT runner, including its Spark-compatibility suite, does not run Apache Spark; the PySpark command above does.Are there any user-facing changes?
No implementation change. This PR documents and reproduces the decimal-addition behavior with an intentionally failing test.