Executable cases comparing Substrait output schemas and relation semantics across libraries and engines. Expectations are written from the specification text they cite, independently of the implementation answers. The authored corpora target Substrait v0.102.0 and support the relation conformance proposal.
| Corpus | Scope | Participants |
|---|---|---|
derived-schema/ |
108 generated plans testing schema derivation | substrait-java, substrait-python, substrait-go, substrait-validator, Isthmus/Calcite, DataFusion, DuckDB, Spark, Acero, Gluten |
tests/relations/ |
71 hand-written cases testing relation schemas and rows | substrait-java, substrait-go, DuckDB, DataFusion |
producers/ |
SQL-produced plans: declared types, bindings and consumer schemas | DuckDB, Isthmus, Spark and DataFusion producers; schema consumers |
The results on GitHub Pages show the expectation and saved answer for each cell in the authored corpora. Schema differences groups differing cases by participant; FINDINGS.md links reproducers to upstream reports.
Producer plans are checked against the deriver's reading of v0.102.0, without a second oracle. Their results include a consumer matrix; schema comparisons establish neither row correctness nor general plan validity.
The schema columns were taken 2026-10-06 against the versions in
probe/versions.env. They answer the 98 cases that carry an expectation.
Gluten is measured separately over the virtual-table variant and is outside the table below.
| matched | differed | unsupported | |
|---|---|---|---|
| substrait-java | 92 | 3 | 3 |
| substrait-python | 95 | 1 | 2 |
| substrait-go | 69 | 2 | 27 |
| substrait-validator | 64 | 14 | 20 |
| Isthmus/Calcite | 57 | 4 | 37 |
| DataFusion | 59 | 7 | 32 |
| DuckDB | 36 | 21 | 41 |
| Spark | 31 | 6 | 61 |
| Acero | 4 | 16 | 78 |
A difference is a disagreement with this repository's reading of the specification, not proof of a defect. The columns measure different consumer APIs and type systems; totals do not rank engines. A schema match alone establishes neither correct rows nor independent type derivation. The Java row is calibration; the expand cases illustrate its limits. METHOD.md explains these boundaries and the declaration swap.
The relation matrix compares schemas and, for executing participants, rows. Its measurement notes describe the saved results and each participant's comparison limits.
From the repository root:
bash probe/selfcheck.sh # committed-file checks; python3, no network
bash probe/replay_column.sh DUCKDB # pinned schema column
bash probe/relations/replay.sh DUCKDB # pinned relation columnEach replay downloads and builds the named participant, then requires the saved answers back. The schema probe guide and relation probe guide list participant names and prerequisites. Gluten has a separate run path. The self-check executes no participant and does not validate the specification reading.
- Method: expectation sources, calibration, coverage and measurement limits.
- Findings: specification questions, implementation reports and reproducers.
- Relation test vectors: protobuf bundle contract and case authoring.
- Generators: building and extending the schema corpus.
- Independent deriver: a second encoding of the specification rules.
- Producer corpus: saved SQL-produced plans, declaration checks and consumer results.
- Producer shapes: plans from SQL passed between implementations.
Expectation corrections should name the case and the specification rule. Outstanding review questions are listed in METHOD.md.